· via dev.to (home feed)
Ollama's llama.cpp runner eliminates measured 90x display lag in local-LLM streaming
Benchmarks posted on dev.to show Ollama's old engine caused 90x worse display lag on all-core inference; the llama.cpp-based runner in 0.40.2 fixed it, but thread caps still matter.

A developer building a local-LLM coding IDE noticed something counterintuitive: giving the model every CPU core made the application feel slower even though tokens kept arriving at the same rate. Writing on dev.to, the author measured the discrepancy and found that on Ollama 0.21.0, inference across all six logical cores of a Ryzen 5 4500U delivered the same generation speed as a four-core run but roughly 90 times worse display lag at the 95th percentile — 181 ms versus 2 ms between a token arriving on the wire and appearing in the DOM. UI click response during generation was about six times slower, at 929 ms versus 158 ms.
How the lag was measured
The benchmark, described on dev.to, deliberately separated generation speed from perceived responsiveness. A thin proxy between the IDE and Ollama timestamped every NDJSON chunk as it arrived, while a MutationObserver inside the app recorded when streaming text actually reached the DOM. Event-loop drift, effective frame rate and a probe that clicked a real UI button mid-generation rounded out the measurements. Conditions were held fixed — temperature 0, a fixed seed, 600 predicted tokens — with medians taken over 60-second windows.
Generation speed came out essentially identical on the old engine, 4.6 versus 4.7 characters per second, because token production during decode is limited by memory bandwidth rather than available compute. The display pipeline was another story: on all cores, tokens waited up to half a second before being painted, and process inspection showed the runner occupying about 5.7 cores continuously. The subjective impression that the capped run felt snappier was grounded in measurable rendering delay, not imagination.
What changed in Ollama 0.40.2
Mid-verification, the author upgraded to Ollama 0.40.2 and the premise shifted. The gap nearly vanished: P95 display lag of 13 ms with six threads versus 3 ms with four, effectively equal event-loop delay, frame rates pinned at 59.9, and click responses of 123 ms versus 114 ms. A larger 9.6 GB model behaved the same way.
Inspecting the runner process explained why. Version 0.21.0 used Ollama's older engine, which by default took every core. Version 0.40.2 runs llama-server from llama.cpp, and an unspecified thread count now means three threads, so the runner never seizes the whole machine to begin with. Even an explicit num_thread of 6 left roughly 4.3 to 4.6 cores in use, according to the post, because the newer engine's worker threads idle rather than busy-spin. That leaves about 1.5 to 2 spare cores at all times so the renderer, the OS and other applications can preempt, and UI responsiveness reportedly held between 90 and 180 ms even during prompt prefill.
Why the thread cap still ships
Given the fix, a thread-limit setting might look redundant, but the author kept it for two reasons. Engine internals shift between versions — the default moved from all cores to three within one release — whereas an explicit num_thread acts as a stable ceiling independent of version, model or whatever else is running on the machine. Just as importantly, the option is only reachable through Ollama's native /api/chat endpoint: the OpenAI-compatible /v1 API has no channel for it, nor for the think flag, which suppresses the long reasoning trace that thinking models such as qwen3.5:4b emit before answering even a plain greeting.
That API limitation became a product decision. The author shipped CPU Threads and Thinking settings in Teaspoon IDE v1.2.0, a free Electron-based AI coding IDE, and moved its Ollama integration to /api/chat so the native options can actually be passed through. Most VS Code extensions built on OpenAI-compatible clients have no way to send these parameters at all, the post notes.
Why it matters
The story is a useful corrective for anyone running local models. Tokens per second describe only half the experience: identical generation speed can coexist with a 90x difference in perceived responsiveness when the renderer is starved. It also shows that allocating all cores is not automatically optimal, and that the right allocation depends on engine internals that can change without notice between releases. Finally, it argues for splitting wire-arrival time from DOM render time in any local-LLM benchmark, converting a subjective sense of sluggishness into a concrete figure. The author cautions that the data reflects two to three runs on a single Windows machine, so the exact numbers are one-machine results; the durable contribution is the measurement approach itself.
- #local-llm
- #ollama
- #llama-cpp
- #streaming
- #cpu-inference
- #ui-performance