Generated by Codex with GPT 5.6 Sol XHigh
At Internet scale, data representation becomes infrastructure. Cloudflare’s DNS platform keeps more than 250 billion cache entries in memory, so even one unnecessary byte per entry expands into more than 250 gigabytes across the fleet. The challenge was not simply to squeeze the cache into less RAM. Every saved byte had to preserve the behavior of a latency-sensitive resolver, and any extra parsing, allocation, or pointer chasing could erase the gain on the request path.
The official Cloudflare Blog published this post on August 27, 2026. It describes five successive changes to Big Pineapple, the Rust platform behind 1.1.1.1, Gateway DNS, DNS Firewall, and other Cloudflare DNS services. Together, the changes reduced the benchmarked memory footprint of a cache entry from 953 bytes to 420 bytes, lowered the fleet’s working set by roughly 100 terabytes, raised insert throughput by 43%, and cut lookup latency by 19%.
The important story is not the headline number alone. Cloudflare treated memory layout, workload frequency, allocator behavior, protocol encoding, and CPU locality as one systems problem. The result shows why optimization at scale often begins below the algorithmic level, in the exact shape and lifetime of data.
Remove mutability after the write path
A cached DNS response is built once and never modified. Its original Rust representation did not encode that fact. The cache key and entry contained eight Vec and String fields, each of which stored a pointer, a length, and a capacity. Capacity is useful while a collection may grow, but after insertion it becomes permanent overhead. A Vec may also retain unused heap space left by its growth strategy.
Cloudflare converted those fields to Box<[T]> and Box<str>. These fixed-size forms retain a pointer and length but drop capacity and cannot grow. Saving eight bytes on each of eight fields removed 64 bytes from every entry, while trimming unused backing allocations as well. Across the live cache, this first change alone accounted for more than 15 terabytes.
The same principle applied to the answer, authority, and additional sections of a DNS response. Instead of three independently allocated lists, Big Pineapple now stores one record sequence plus two u16 offsets marking the section boundaries. Two pointers and two lengths became two small integers, saving another 28 bytes per entry. Packing several Boolean values into bitflags also removed alignment padding, illustrating that field size is not the same as structure cost: changing a few small fields can let the compiler shrink the surrounding layout by more than their nominal byte count.
These changes make an architectural property explicit in the type system. Data that becomes immutable after construction should not continue paying for mutation, growth, or separately owned containers on every read.
Store common facts once
Each DNS record normally carries an owner name, the domain to which the record belongs. In the cache’s common case, that owner is identical to the queried domain already stored in the cache key. Repeating it in every record preserved a convenient self-contained representation, but multiplied heap allocations for information already available during every lookup.
Big Pineapple changed the owner field to an optional boxed name. A missing value means “use the name from the cache key”; only records whose owner actually differs, such as addresses reached through a CNAME, keep a full name. Response construction restores the common owner without allocating. The representation is a little less locally convenient, because a record can no longer always be interpreted in isolation, but it matches the actual dependency of the read path and avoids storing the same fact many times.
This is a useful normalization rule for hot in-memory systems: repetition that looks harmless in an object model can dominate cost when the repeated value has billions of instances. The right boundary is often the unit that is already fetched together, not the most self-contained abstraction.
Optimize for the traffic distribution, then recover locality
Rust enums reserve enough inline space for their largest variant. Big Pineapple’s parsed RecordData enum therefore occupied 144 bytes because a rare NAPTR record needed 136 bytes, even though A and AAAA records require only 4 and 16 bytes and together represent more than 80% of traffic. Most records were paying over 120 bytes for variants they did not contain.
The first response was to keep common small variants inline and box large variants. That compressed the enum to 24 bytes and made the usual case far cheaper, at the cost of a slightly more expensive rare case. But boxing exposed the next bottleneck. Each large value required its own allocation, jemalloc rounded requests up to size classes, and pointers scattered related records around the heap. A layout that saved inline space could still waste allocator space and CPU cache lines.
The final design moved record data into a single contiguous Box<[u8]>. Each record is stored as a two-byte length followed by its DNS wire-format bytes. This middle ground avoids caching an entire ready-made DNS message, which would complicate conditional DNSSEC handling, while also avoiding a tree of parsed enum values and per-record allocations.
The wire representation improves both space and the lookup path. Most record types, including A, AAAA, TXT, and DNSSEC records, can be copied directly into the outgoing response rather than serialized field by field. Records containing names, such as CNAME, NS, MX, and SOA, are still parsed when name compression must be applied. Sequential access replaces random indexing, but entries contain few records, so the cost is negligible. Better locality and less serialization reduced lookup latency by 5% for this change alone.
Insertion also uses a reusable scratch buffer. It accumulates the encoded records without repeatedly growing new vectors, then performs one allocation and one copy into the exact-size boxed slice. That removed the many boxed allocations and avoided retaining the unused tail of a shrunken vector, increasing insert throughput by 13% for the wire-format change.
Measure the representation, then verify the system
Cloudflare did not infer fleet savings by multiplying size_of values. Its benchmark populated the cache with a production-shaped mix of A, AAAA, and variable-length TXT records, while a custom allocator counted both allocation size and frequency per entry. The same harness measured complete insert and lookup operations, preventing a memory improvement from hiding a latency regression.
The team then checked the result in production because resident memory also depends on traffic mix, cache occupancy, allocator state, and everything else in the process. As the changes rolled out between May 18 and July 6, steady-state p99 memory per instance fell from 9.3 GB to 5.3 GB; p90 fell from 6.5 GB to 3.8 GB. The benchmark showed a 56% smaller net entry footprint and 58% fewer allocated bytes, while production showed about a 42–43% reduction in total resident memory across the measured percentiles.
The distinction matters. A microbenchmark explains where bytes and nanoseconds went; a staged rollout shows whether those gains survive real allocation patterns and workloads. Both are needed when a local representation change is expected to remove 100 terabytes from a distributed system.
The broader engineering lesson
The optimizations worked because they aligned representation with invariants and frequency. Immutable data stopped paying for spare capacity. Adjacent sections stopped paying for separate ownership. Common owner names became an implicit reference to an existing key. Rare large variants stopped setting the size of common small ones. Finally, protocol bytes became the native cached form where parsing added no value.
This is more than a collection of Rust tricks. It is a method for high-scale performance work: measure allocations on realistic data, make workload skew an explicit design input, examine padding and allocator bins rather than only logical field sizes, and judge memory and latency together. Fewer bytes can mean fewer allocations, better cache locality, less serialization, and faster operations—not merely a smaller bill.
Cloudflare plans to spend the freed memory on a larger cache rather than simply shrinking the fleet. That closes the systems loop: a denser representation increases effective capacity, a larger cache can raise hit rates, and higher hit rates reduce queries to upstream authoritative servers. At sufficient scale, choosing the right in-memory shape does not just optimize one process; it changes the behavior and economics of the network around it.