· via Hacker News – Front Page (native)
Zero-copy Rust parsing: why reading a u64 shouldn't touch the heap
A post by Sebastian Sastre, widely shared via Hacker News, shows how a casual byte copy in a Rust u64 parser forces heap allocation, and how zero-copy decoding reaches hundreds of millions of ops per second.

The allocation hiding in a tiny parser
A blog post by Sebastian Sastre that reached the Hacker News front page dissects a small Rust function that converts eight bytes of a buffer into a u64. The starting version looks harmless: it checks the length, copies the first eight bytes into a Vec, and passes that to u64::from_le_bytes. The code is short, returns proper errors, and appears cheap.
The problem is the Vec. In Rust a Vec always allocates on the heap, and whether that allocation is cheap depends on the allocator having a free block ready. Under load it may not, and the allocator then has to ask the operating system for more memory, turning a supposedly trivial parse into a potential syscall. Sastre's argument is that the copy serves no purpose: the eight bytes sitting in the buffer already are the number, so duplicating them just to read them is accidental complexity on the path to the result.
Reading the value where it sits
His remedy is to read in place, the approach generally known as zero-copy parsing. Because the function already receives a &[u8], it can check that eight readable bytes exist via bytes.get(..8), convert them into a fixed-size [u8; 8] on the stack, and let from_le_bytes finish the job. The bytes still move onto the stack, but the heap is never involved. He ties this to Rust's broader strength: when types and every error mode are modelled properly, the compiler's guarantees translate into runtime behaviour that costs nothing extra.
Growing the idea into a full codec
Sastre then extends the technique beyond a single integer. For a matching-engine project he is building a command codec where each command is a variant of an enum: limit orders carrying account, instrument, price and quantity fields, cancels identified by a single order id, and market orders with no price field at all. Every field type is deliberately Copy, nothing in the enum owns a buffer, and the enum as a whole is therefore Copy — a deliberate design choice, he explains, so that parsing and downstream processing can happen in place.
On the wire, a kind byte selects the layout, and each layout has a fixed length: 57 bytes for a limit order, 16 for a cancel-by-order. Decoders read fields at known offsets and wrap them in typed constructors. Encoding mirrors decoding: values are written at those same offsets into a buffer the caller already owns, with no intermediate allocations, and the journal frame and datagram paths share a single write.
What it buys in practice
The post reports benchmark results of roughly 4.1 nanoseconds per decoded element across command types, which works out to between about 238 and 246 million commands per second: cancel-by-order at about 243 million, new limit orders at about 238 million, and new market orders at about 245 million. Sastre also invites readers to verify the syscall behaviour themselves with a small Rust program run under Docker, noting that glibc serves allocations above 128 KiB through mmap, which makes the cost observable in a way smaller allocations hide.
Why it matters
Parsing sits on the hot path of almost every data pipeline, feed handler and message-driven service, yet it is rarely scrutinised. The post's core lesson generalises well beyond Rust trading systems: convenience copies hide inside innocent-looking code and become expensive precisely when a system is under load, because that is when allocators turn to the operating system. Designing data types so they can be read where they already sit — and written into buffers the caller owns — removes an entire class of avoidable cost, whether the workload is market data, log ingest, metrics or any serialization boundary. As the author's measurements show, the payoff is throughput in the hundreds of millions of elements per second rather than marginal gains.
- #rust
- #zero-copy
- #performance
- #parsing
- #systems-programming