Reference
Glossary
Shared definitions for terms used across these documents. Defined words are underlined wherever they appear. Hover over one, or tap it, to read its definition without leaving the page.
- ADR
An architecture decision record: a short, permanent note of one significant decision, covering its context, the options considered and the consequences. ADRs are superseded by new records rather than edited.
- Back-reference
In LZ77 compression, an instruction to copy L bytes starting D bytes back in the output already produced, instead of storing them again. DEFLATE allows distances up to 32,768 bytes.
- Backoff
Waiting longer after each failed attempt before retrying, usually doubling the delay each time. Without jitter, clients that failed together retry together.
- Barrier
A GPU instruction that makes every thread in a threadgroup wait until all of them arrive, and makes their earlier memory writes visible to each other. There is no safe barrier across threadgroups.
- Burst
A short spike of requests above the average rate. In a token bucket, the largest burst allowed is the bucket's capacity b.
- Cache miss
A lookup that finds nothing in the cache, so the request falls through to the slower source of truth and usually back-fills the cache.
- Compute kernel
A function that runs on the GPU once per thread of a dispatch. Each thread reads its own index to decide which part of the work it does.
- CRC-32
A 32-bit checksum computed by polynomial division over the bits of the data. ZIP and gzip store one per file so a decoder can detect corruption. Checksums of adjacent pieces can be combined, which allows computing it in parallel.
- DEFLATE
The compression format inside ZIP, gzip and PNG: LZ77 back-references plus Huffman coding, organised in blocks. Its bit stream has no index, so decoding is serial.
- Dispatch
One launch of a compute kernel over a grid of threads, divided into threadgroups. Dispatches in one command buffer run in order, which is the safe way to order work between threadgroups.
- Divergence
When lanes of one SIMD group take different sides of a branch. The hardware runs each side in turn with the other lanes switched off, so the branch costs the sum of both sides.
- GCRA
The generic cell rate algorithm. It is equivalent to a token bucket, but stores one timestamp per key: the theoretical arrival time of the next conforming request.
- Huffman code
A variable-length binary code that gives frequent symbols short bit patterns and rare ones long patterns. No code is a prefix of another, so a decoder can read them back to back without separators.
- Jitter
Randomness added to retry or polling delays, so many clients that start together don't keep acting in lock-step and overload a recovering service.
- Lane
One thread's slot within a SIMD group. Each lane has its own registers, but all lanes execute the same instruction at the same time.
- lcov
A plain-text coverage format that lists, per source file, how many times each line ran. Most coverage tools can write it, which makes it the common currency between them and tools like patchcov.
- Leaky bucket
A queue that drains at a fixed rate. It smooths output to a constant pace instead of allowing bursts; compare with a token bucket.
- LZ77
The compression idea, from Lempel and Ziv in 1977, of replacing repeated data with back-references to an earlier occurrence.
- Marker
In rapidgzip-style decoding, a 16-bit placeholder meaning "the byte n positions before this chunk started". It lets a decoder start mid-stream without the history, and is replaced by the real byte once the previous chunk is decoded.
- Merge base
The newest commit that a branch and its target (usually
main) share: where the branch forked. Comparing against it, rather than the tip ofmain, shows only what the branch itself changed.- Monotonic clock
A clock that only moves forwards. Use it to measure elapsed time, because the wall clock can jump backwards when NTP corrects it.
- Occupancy
How many SIMD groups a GPU core keeps resident at once. More resident groups give the core more work to switch to while others wait for memory. Registers and threadgroup memory per group limit it.
- p99
The 99th-percentile latency: 99% of requests complete at least this fast. Tail latency like this, rather than the average, is what users notice.
- Patch coverage
The share of the executable lines a change added that the tests ran. Unlike the project-wide total, it shows the gaps in new code even when they are too small to move the total.
Announcing patchcov lets you try it.
- Rate limiting
Capping how many requests a client may make in a period of time, to protect a service and share its capacity fairly.
Token bucket rate limiting explains the most common algorithm.
- Read replica
A copy of a database that serves reads and follows the primary asynchronously, so it can lag slightly behind.
- Retry-After
An HTTP response header that tells the client how long to wait before retrying, sent with
429or503.- Retry budget
A cap on retries as a fraction of normal traffic (for example 10%). It stops retries from multiplying load during an outage.
- Shuffle
A GPU instruction in which every lane reads a register value from another lane of its SIMD group, without going through memory. In Metal it is
simd_shuffle; in HLSL,WaveReadLaneAt.- SIMD group
The set of threads a GPU executes in lock-step: 32 on Apple GPUs (NVIDIA calls it a warp, Vulkan a subgroup, HLSL a wave). One instruction is fetched once and executed by every lane.
- SLO
A service level objective: a reliability target, such as 99.9% of requests succeeding over 30 days. It is set below 100% on purpose, leaving an error budget.
- SPIR-V
The Khronos binary intermediate language for GPU programs. Vulkan drivers consume it directly, and SPIRV-Cross can translate it into Metal Shading Language.
- Threadgroup
A group of GPU threads that runs on one core and can share threadgroup memory and barriers. Vulkan calls it a workgroup and CUDA a thread block.
- Threadgroup memory
Small, fast on-chip memory shared by the threads of one threadgroup: at most 32 KiB per threadgroup on Apple GPUs. HLSL calls it
groupshared.- TLS
Transport Layer Security: the protocol that encrypts and authenticates a connection, as in HTTPS. A gateway that "terminates TLS" decrypts traffic there.
- Token bucket
A rate limiter that refills tokens at rate r up to a capacity b, and admits each request only by spending a token. It allows bursts up to b but enforces r on average.
- Unified memory
A design, used by Apple Silicon, in which the CPU and GPU share the same physical DRAM. A buffer can be visible to both without being copied, though first GPU use of a buffer still has a cost.