deniz.in

Markets

Weather

Loading weather

· via dev.to (home feed)

Jylus benchmark: 98.77% less LLM input context, 528/528 strict accuracy

A dev.to post by Jylus founder Josh Hodgetts reports that a compact, conflict-aware evidence pack beat both full context and BM25-only RAG on a frozen 528-question benchmark, using far fewer input tokens.

Jylus benchmark: 98.77% less LLM input context, 528/528 strict accuracy

What the benchmark measured

A dev.to post by Josh Hodgetts, founder of Jylus, describes a frozen benchmark of 528 questions across four data domains, scored by a deterministic scorer, with every question put to the same model (Gemini 3.1 Flash Lite) under identical settings. The only variable was how evidence reached the model.

Three configurations were compared. Supplying the model with the full context averaged 205,129 input tokens per question and produced 78.79% accuracy under strict scoring. A BM25-only retrieval setup cut average input to 8,128 tokens, but accuracy slipped to 76.14%. The third configuration, Jylus's own Context Pack, averaged 2,532 tokens — 98.77% less input than full context — and answered all 528 questions correctly under the same scorer.

The BM25 row is the one Hodgetts singles out. Shrinking the prompt by itself made things slightly worse, not better: the model was still handed overlapping, partly conflicting records and left to work out which of them counted.

The underlying problem: records that revise each other

The post illustrates the failure mode with two records. One, logged minutes after the fact, notes a temperature warning on Wednesday at 09:00. A second record, entered on Friday, explains that the warning came from a faulty sensor. Ask what caused Wednesday's warning given everything known now, and the Friday correction is fair game. Ask what was known about the warning on Wednesday, and the correction must stay out of the answer, because it did not exist yet.

Returning both records is technically complete, but it offloads a judgment call — deciding which evidence is admissible for the question asked — onto the model. According to the post, that is where confidently wrong answers originate, and the pattern generalises to support histories, subscription changes, incident investigations and any dataset where a later entry reframes an earlier event.

What the tool actually does

Jylus sits between the data source and the model. It retrieves candidate evidence, resolves state and relationships, and compiles a bounded pack that carries source references while deliberately preserving relevant conflicts and gaps rather than smoothing them away. The model then reasons over that prepared evidence.

The developer-facing output is intended to make provenance inspectable: which source supports a given fact, whether the fact applies at the time the question concerns, whether a newer record has superseded it, whether sources disagree, and whether enough evidence exists to answer at all. Hodgetts frames the token budget as a constraint on evidence volume that should not hide the uncertainty needed to interpret it.

The limits of the result

The post is direct about what the numbers do not establish. The benchmark is Jylus's own adversarial design across four data domains and has not been independently reproduced. The 100% figure means 528 out of 528 under this test's strict scoring, not universal accuracy. BM25-only retrieval stands in for just one slice of possible RAG architectures. And the experiment does not isolate which part of the evidence preparation — retrieval, state resolution, conflict retention or the token bound — actually drove the improvement. The methodology is public, and a playground accepts synthetic records without an account for anyone who wants to probe the behaviour directly.

Lessons for RAG and prompt design

Even taken cautiously, the result suggests testable ideas for anyone building retrieval systems over mutable data:

  • More context is not monotonically better. Conflicting versions of a fact can subtract accuracy while adding information.
  • Relevance is not the same as admissibility. A retrieved record also has to be valid for the point in time the question concerns.
  • Provenance metadata — source, supersession, disagreement, gaps — surfaces correctness problems before generation instead of after.
  • "Not enough evidence" needs to be an answer the system can actually return.

Why it matters

If the pattern holds beyond this one benchmark, it points at a rare combination: a large cost reduction and a correctness gain from the same change. The failure mode it targets — a model quietly blending stale and current records into one confident answer — is common in production systems over data that changes. That makes this less a vendor showcase than a concrete design hypothesis: prepare evidence for time and conflict before prompting, then measure whether accuracy holds. Reproducing the experiment on your own data is the obvious next step.

  • #rag
  • #llm
  • #context-engineering
  • #retrieval
  • #benchmark

Related posts