Reference

Glossary

Shared definitions for terms used across these documents. Defined words are underlined wherever they appear. Hover over one, or tap it, to read its definition without leaving the page.

ADR

An architecture decision record: a short, permanent note of one significant decision, covering its context, the options considered and the consequences. ADRs are superseded by new records rather than edited.

Back-reference

In LZ77 compression, an instruction to copy L bytes starting D bytes back in the output already produced, instead of storing them again. DEFLATE allows distances up to 32,768 bytes.

Backoff

Waiting longer after each failed attempt before retrying, usually doubling the delay each time. Without jitter, clients that failed together retry together.

Barrier

A GPU instruction that makes every thread in a threadgroup wait until all of them arrive, and makes their earlier memory writes visible to each other. There is no safe barrier across threadgroups.

Burst

A short spike of requests above the average rate. In a token bucket, the largest burst allowed is the bucket's capacity b.

Cache miss

A lookup that finds nothing in the cache, so the request falls through to the slower source of truth and usually back-fills the cache.

Compute kernel

A function that runs on the GPU once per thread of a dispatch. Each thread reads its own index to decide which part of the work it does.

CRC-32

A 32-bit checksum computed by polynomial division over the bits of the data. ZIP and gzip store one per file so a decoder can detect corruption. Checksums of adjacent pieces can be combined, which allows computing it in parallel.

Cyclic redundancy check (Wikipedia)

DEFLATE

The compression format inside ZIP, gzip and PNG: LZ77 back-references plus Huffman coding, organised in blocks. Its bit stream has no index, so decoding is serial.

RFC 1951

Dispatch

One launch of a compute kernel over a grid of threads, divided into threadgroups. Dispatches in one command buffer run in order, which is the safe way to order work between threadgroups.

Divergence

When lanes of one SIMD group take different sides of a branch. The hardware runs each side in turn with the other lanes switched off, so the branch costs the sum of both sides.

GCRA

The generic cell rate algorithm. It is equivalent to a token bucket, but stores one timestamp per key: the theoretical arrival time of the next conforming request.

Generic cell rate algorithm (Wikipedia)

Huffman code

A variable-length binary code that gives frequent symbols short bit patterns and rare ones long patterns. No code is a prefix of another, so a decoder can read them back to back without separators.

Huffman coding (Wikipedia)

Jitter

Randomness added to retry or polling delays, so many clients that start together don't keep acting in lock-step and overload a recovering service.

Lane

One thread's slot within a SIMD group. Each lane has its own registers, but all lanes execute the same instruction at the same time.

lcov

A plain-text coverage format that lists, per source file, how many times each line ran. Most coverage tools can write it, which makes it the common currency between them and tools like patchcov.

Leaky bucket

A queue that drains at a fixed rate. It smooths output to a constant pace instead of allowing bursts; compare with a token bucket.

LZ77

The compression idea, from Lempel and Ziv in 1977, of replacing repeated data with back-references to an earlier occurrence.

LZ77 and LZ78 (Wikipedia)

Marker

In rapidgzip-style decoding, a 16-bit placeholder meaning "the byte n positions before this chunk started". It lets a decoder start mid-stream without the history, and is replaced by the real byte once the previous chunk is decoded.

Merge base

The newest commit that a branch and its target (usually main) share: where the branch forked. Comparing against it, rather than the tip of main, shows only what the branch itself changed.

Monotonic clock

A clock that only moves forwards. Use it to measure elapsed time, because the wall clock can jump backwards when NTP corrects it.

Occupancy

How many SIMD groups a GPU core keeps resident at once. More resident groups give the core more work to switch to while others wait for memory. Registers and threadgroup memory per group limit it.

p99

The 99th-percentile latency: 99% of requests complete at least this fast. Tail latency like this, rather than the average, is what users notice.

Patch coverage

The share of the executable lines a change added that the tests ran. Unlike the project-wide total, it shows the gaps in new code even when they are too small to move the total.

Announcing patchcov lets you try it.

Rate limiting

Capping how many requests a client may make in a period of time, to protect a service and share its capacity fairly.

Token bucket rate limiting explains the most common algorithm.

Read replica

A copy of a database that serves reads and follows the primary asynchronously, so it can lag slightly behind.

Retry-After

An HTTP response header that tells the client how long to wait before retrying, sent with 429 or 503.

RFC 9110 §10.2.3

Retry budget

A cap on retries as a fraction of normal traffic (for example 10%). It stops retries from multiplying load during an outage.

Shuffle

A GPU instruction in which every lane reads a register value from another lane of its SIMD group, without going through memory. In Metal it is simd_shuffle; in HLSL, WaveReadLaneAt.

SIMD group

The set of threads a GPU executes in lock-step: 32 on Apple GPUs (NVIDIA calls it a warp, Vulkan a subgroup, HLSL a wave). One instruction is fetched once and executed by every lane.

SLO

A service level objective: a reliability target, such as 99.9% of requests succeeding over 30 days. It is set below 100% on purpose, leaving an error budget.

SPIR-V

The Khronos binary intermediate language for GPU programs. Vulkan drivers consume it directly, and SPIRV-Cross can translate it into Metal Shading Language.

SPIR-V registry

Threadgroup

A group of GPU threads that runs on one core and can share threadgroup memory and barriers. Vulkan calls it a workgroup and CUDA a thread block.

Threadgroup memory

Small, fast on-chip memory shared by the threads of one threadgroup: at most 32 KiB per threadgroup on Apple GPUs. HLSL calls it groupshared.

TLS

Transport Layer Security: the protocol that encrypts and authenticates a connection, as in HTTPS. A gateway that "terminates TLS" decrypts traffic there.

Token bucket

A rate limiter that refills tokens at rate r up to a capacity b, and admits each request only by spending a token. It allows bursts up to b but enforces r on average.

Token bucket rate limiting

Unified memory

A design, used by Apple Silicon, in which the CPU and GPU share the same physical DRAM. A buffer can be visible to both without being copied, though first GPU use of a buffer still has a cost.