· via dev.to (home feed)
Ollama's small default context window silently truncates prompts; llm-doctor diagnoses local setups
A dev.to post details how Ollama 0.34.4 served qwen3:1.7b with a 4,096-token window instead of the supported 40,960, silently dropping most of a long prompt. The author built llm-doctor to catch such issues.

What happened
A developer reports on dev.to that Ollama 0.34.4, running qwen3:1.7b on a MacBook Air, was serving the model with a 4,096-token context window even though the model supports 40,960 tokens. When the author sent a prompt of roughly 7,000 tokens, the API answered with a normal HTTP 200 response — but the model had only processed the final 2,050 tokens. No error was raised and no warning was printed.
That silent failure prompted the author to build llm-doctor, an open-source diagnostic utility that inspects local LLM installations and reports configuration problems that would otherwise go unnoticed.
Why the failure is easy to miss
Context truncation at the runtime level is difficult to detect from the outside. The request succeeds, the model responds coherently, and nothing in the output signals that most of the input was discarded. According to the dev.to post, Ollama applies a small default context window, and the OpenAI-compatible endpoint that coding agents typically target has no way to raise it on a per-request basis. An agent that feeds several files into a local model can therefore lose the majority of its prompt while everything looks healthy on the wire.
What the scanner checks
Beyond context size, the post identifies several other problems that accumulate in local model setups:
- Duplicated model weights. Each runtime stores its own copy of the same GGUF file, and Ollama names its copy with a sha256 hash rather than the model name, so there is no obvious way to see that LM Studio is holding an identical multi-gigabyte file.
- Outdated chat templates. Quantizers sometimes repair broken chat templates after a release, but files downloaded earlier keep the flawed version. Tool calls can then fail in ways that resemble a weak model rather than a fixable packaging bug.
- Context window defaults, as described above.
- Storage lock-in. Because Ollama's blobs carry no readable names, moving to another runtime generally means re-downloading every model.
The llm-doctor command scans all detected model stores, llm-doctor fix previews deduplication and cleanup, and llm-doctor fix --yes applies the changes. In the author's demo fixture, where Ollama and LM Studio both held the same GGUF alongside one orphaned blob, the fix replaced the LM Studio copy with a symlink to the Ollama blob and removed the orphan, reclaiming 140 MB. Every change is recorded in a JSONL log file under the user's home directory. An unbundle command can also expose Ollama models to llama.cpp.
The endpoint probe
The feature the author says they rely on most is the endpoint probe, invoked with llm-doctor endpoint against a running server. It verifies context length, tool calls, parallel tool calls, streaming tool calls, JSON schema output and think tags, then prints a concrete recommendation for each failing check. Pointed at the misconfigured Ollama setup that started the project, it reported 4,096 effective tokens against the 40,960 the model advertises and marked the context check as failed.
Caveats
The tool requires Python 3.10 or newer and is not yet published on PyPI, so it must be installed directly from its GitHub repository. The author's testing has been mostly on macOS with Ollama, and they explicitly ask for reports from Linux and Windows users, as well as for details on model stores the scanner does not yet recognize. The project is released under the MIT license.
Why it matters
Local LLM stacks are usually assembled from several independent pieces — runtimes, quantized weights, chat templates and agent clients — and the interfaces between them fail quietly. A context mismatch does not crash anything; it just quietly degrades answers, which users naturally blame on the model. For anyone wiring coding agents to a local OpenAI-compatible endpoint, a deterministic probe that reports the effective context length and actual tool-call support replaces guesswork with a checkable result. The story also illustrates a broader gap in the local ecosystem: where a hosted API would reject an oversized request, local runtimes tend to trim it and say nothing.
- #ollama
- #local-llm
- #context-window
- #developer-tools
- #open-source