· via dev.to (home feed)
ONNX Runtime's 6 MB WASM file outlasts a 44 MB model in browser AI cold start
Cold-start timings from ImgIng show the 5.95 MB onnxruntime-web WASM binary arriving about nine seconds after a 44 MB model, because ORT only fetches it once a session is created.

A small file became the long pole
A developer responsible for the model loading path at ImgIng, a browser-based image processing service, has published cold-start timings for its in-browser background removal feature. Writing on dev.to, they report that on a first visit the 5.95 MB onnxruntime-web WebAssembly binary finished arriving roughly nine seconds after the 44 MB model it is supposed to execute, and that inference could not begin until it landed.
How the test was run
The measurement was taken on the evening of 28 September 2026, in a fresh browser context on an M4 Mac with 16 GB of RAM running Chromium 149. Rather than clicking through the interface, the author invoked the same internal segmentation function the start button calls, logging every network request and progress event. The test image was a synthetic bottle; its content does not affect timing because the model resizes all inputs to 1024×1024. Two configurations were measured: WebGPU on Metal, and a Chromium instance with no usable GPU adapter, the state produced by disabling hardware acceleration.
Where the time went
With WebGPU available, the runtime JavaScript was ready at 32 ms. The 44,279,201-byte model from ModelScope's CDN finished downloading at 4,188 ms, and its SHA-256 verification completed by 4,207 ms. Session setup began at 4,250 ms, and that is exactly when onnxruntime-web requested ort-wasm-simd-threaded.asyncify.wasm — 5,955,745 bytes gzipped — from ImgIng's own origin. That download finished around 13.2 seconds in; inference started at 13,725 ms and the result arrived at 14,594 ms.
Without a GPU adapter, the shape was the same: the WebGPU attempt failed, ORT logged a switch-to-WASM message at 13,551 ms, and the run completed at 16,015 ms.
Two structural details stand out. The WASM file is requested only when a session is created, so it begins after the model has already downloaded and the two transfers never overlap. And the WebGPU path needs the file too, because ORT's WebGPU execution provider is compiled into that same WASM binary — hardware acceleration does not let a site skip it.
File size was not the problem
Per-file curl measurements at the same time of evening put the runtime download at 9.19 and 10.78 seconds across two tries, between 552 and 648 KB/s. The 44 MB model from ModelScope took 4.20 seconds, about 10.5 MB/s — roughly 16 to 19 times faster per byte. The author notes this was an ordinary home connection and may not hold in other regions or at other hours. A later cold run saw the runtime land at +34.2 seconds against the model's +4.3, but the machine was shared with other work, so only the ordering is treated as reliable.
The pattern repeats on the heavier professional tier (BEN2 FP16, 219,121,675 bytes): with WebGPU, the model finished at +20.5 seconds, the runtime at +29.8, and the whole call took 34.8 seconds.
Warm caches make it disappear
The asymmetry only bites on the first visit. The runtime is served with Cache-Control: max-age=31536000, immutable and is never requested again, while the model comes back from Cache Storage. On a warm cache the entire call completed in 930 ms with WebGPU and 2,643 ms without an adapter.
Why it matters
The infrastructure effort is inverted relative to the cost. The model gets a CDN, a backup domain and a hash check; the runtime is a single request to the site's own origin, and it starts late. The author's advice for anyone shipping in-browser inference is concrete: open the Network waterfall on a cold load, look past the largest file, compare when the .wasm request fires against the model download, and measure the throughput your own origin actually delivers for it. The numbers point at two obvious levers — fetching the runtime before the model finishes, and giving it CDN treatment comparable to the weights — which would collapse two sequential downloads into one and remove most of the first-visit penalty.
- #onnx
- #webassembly
- #webgpu
- #in-browser-ai
- #performance