deniz.in

Markets

Weather

Loading weather

· via dev.to (home feed)

Independent sub-$1 benchmark measures X25519MLKEM768 TLS handshake overhead on Graviton3

An independent benchmark on two budget Graviton3 instances found X25519MLKEM768 adds no visible cost to median TLS handshakes but raises p99 latency from 5 ms to 8 ms, reproducible for under $0.30.

Independent sub-$1 benchmark measures X25519MLKEM768 TLS handshake overhead on Graviton3

What the benchmark measures

Both AWS and Cloudflare have publicly described the overhead of post-quantum TLS as minimal, but those figures are vendor-published and come from infrastructure most teams will never touch. A new benchmark series by Asif Naeem on dev.to sets out to check the claim independently, on hardware anyone can rent, with every script and raw result file published for reruns.

The first post measures X25519MLKEM768 — the hybrid that pairs classical X25519 with ML-KEM-768, the lattice-based KEM standardized as FIPS 203 — against plain X25519. The rig is two c7g.large spot instances (the "$0.04/hour Graviton3" of the post's title) placed in the same us-east-1 availability zone, so cross-AZ latency cannot swamp the effect. One instance serves Ubuntu 24.04 with nginx compiled from source against OpenSSL 3.5.0; the other drives load with Gatling 3.15.1 on JDK 21. Traffic is TLS 1.3 only, the response body is three bytes so the test isolates handshakes rather than transfers, session resumption is disabled, and each arm runs three five-minute trials with a 30-second warmup discarded. A full reproduction, the author estimates, takes about 50 minutes and somewhere between roughly $0.15 and $0.30 in AWS charges.

One deliberate choice: the server presents a self-signed ECDSA P-256 certificate rather than an ML-DSA-65 one, because Gatling's underlying BoringSSL does not advertise ML-DSA-65 support. According to the post, this mirrors how PQ TLS is actually deployed today, with no public CA issuing ML-DSA-65 certificates — current deployments serve classical certificates alongside a hybrid key exchange, which is exactly the pattern measured here.

The run that almost got published

The first serious attempt pushed 1,000 requests per second and produced absurd-looking numbers: one-millisecond minimums but means of roughly 1.2–1.3 seconds in both arms. The post diagnoses this as client-side CPU saturation — two vCPUs on the load generator cannot complete that many fresh handshakes per second, so requests queue and the queuing noise swallows any real PQ-versus-classical difference. The PQ arm's per-trial p99 spread from about 5.9 s to 7.6 s was itself larger than any signal being measured. The lesson the author draws: validate the load generator before trusting its output, because client-side cost is easy to forget when the discussion is framed around server overhead.

Results at 300 requests per second

Backed off to 300 requests per second, the load generator sits within its capacity and the handshake itself becomes measurable. Median-of-three-trial latencies:

Metric X25519 X25519MLKEM768 Overhead
Mean 2 ms 2 ms —
p50 1 ms 2 ms +1 ms
p95 4 ms 4 ms —
p99 5 ms 8 ms +3 ms (about 60%)
Max 14 ms 41 ms +27 ms

The post is candid about variance: per-trial PQ p99 values were 5, 8 and 39 ms, with the last an outlier the author suspects came from spot-instance CPU steal or a JVM garbage-collection pause, though the cause has not been isolated.

Three conclusions follow. In the common case, PQ overhead is essentially invisible at this scale — mean handshakes are identical at millisecond resolution. The cost concentrates in the tail, where p99 grows and worst-case scheduling amplifies the variance of KEM operations. And the figures are a lower bound rather than an upper bound: Gatling's integer-millisecond timing can hide sub-millisecond deltas, possibly around half a millisecond, so a reading of "no difference" may really mean "small difference."

What it cannot answer yet

The run covers one request rate, one CPU architecture, one payload size and fresh handshakes only. Session resumption skips the KEM entirely, so what was measured is effectively the worst case, while real traffic mixes both. Announced follow-ups include payload sweeps from 100 bytes to 1 MB, resumption mixes including 0-RTT, the same benchmark on Intel (c7i) and AMD (c6a) instances, and concurrency sweeps to find where saturation actually lands. The post also discloses drafting assistance from Claude, while stating that the measurements, setup and interpretation are the author's own.

Why it matters

The industry is defaulting toward hybrid post-quantum key exchange, and capacity planning rests on overhead figures that are usually taken on vendor trust. This benchmark makes the claim-check cheap and repeatable — pocket-change cost, a public repository, raw per-trial data — and its early findings cut both ways: reassurance for the common case, where the hybrid KEM adds nothing visible at the median, and a caution for tail-latency SLOs, where p99 grew about 60% even at modest load. Just as valuable is the methodology. Publishing a failed saturated run, showing trial-by-trial variance, and inviting corrections and independent reruns is how infrastructure benchmarks earn credibility.

  • #post-quantum-cryptography
  • #tls
  • #benchmark
  • #aws
  • #cloud

Related posts