deniz.in

Markets

Weather

Loading weather

· via dev.to (home feed)

Testing LLM features with golden sets, deterministic checks and eval gates in CI

A dev.to field guide lays out how to test LLM features with property-based evals, version-controlled golden sets, calibrated LLM-as-judge scoring and CI gates on pass-rate regression.

Testing LLM features with golden sets, deterministic checks and eval gates in CI

A field guide published on dev.to by Ahmed Mahmoud lays out a testing approach for features built on large language models, arguing that conventional unit tests cannot see the failures that matter. In their place he proposes property-based evals, a version-controlled golden set, deterministic assertions and a regression gate wired into CI.

The article opens with the failure that motivated it: a route converting free-text support messages into structured tickets worked fine until a colleague tweaked the wording of the tone instructions in the system prompt. Extraction began inventing order IDs for messages that contained none, and every test in the repository stayed green, because the tests covered the code around the model, not what the model said.

Why exact-match tests fail

According to Mahmoud, an LLM is a non-deterministic function, so asserting the output equals a fixed string fails for reasons unrelated to correctness. Setting temperature to zero makes output more stable but not byte-identical: providers batch requests and floating-point accumulation differs between batches, and no major provider guarantees reproducible tokens. Exact-string suites flip between red and green without code changes, and teams learn to ignore them.

The recommended fix is to split the code at the model boundary. Prompt assembly, retry logic, output parsing and every branch consuming the parsed result are ordinary deterministic functions and deserve ordinary unit tests. Only the call that crosses into the model gets an eval, and that eval scores properties true of any correct answer: does the output parse against the schema, stay in the requested language, cite only identifiers present in the input, and refuse when the context contains no answer.

Golden sets grown from production bugs

A golden set, in the author's definition, is a version-controlled file of real inputs paired with the properties their outputs must satisfy. His started at roughly a dozen cases and grew every time production surprised him; every bug he fixed in an AI feature became another row in the file. He prefers real cases over synthetic ones, because the awkward ones catch regressions: the empty message, the message written in Arabic, the message that is only an order number, the prompt-injection attempt pasted out of an email.

Two storage rules have paid off, he writes: keep fixtures as JSON or JSONL next to the prompt they test so the case list is reviewable in a pull request, and keep the prompt in its own file rather than a template literal inside a route handler, so a prompt change and its eval results appear in the same diff. His sample code uses the Vercel AI SDK's generateObject with a Zod schema, so a shape regression throws before results flow into assertions.

Deterministic checks before any judge call

The guide's rule of thumb: use a deterministic check whenever the failure can be expressed as a predicate over the string, and reach for a judge only for properties that require reading comprehension. Deterministic checks — schema validation, required substrings, forbidden strings, numeric ranges — catch most regressions, run in milliseconds and cost nothing.

The author maps properties to methods: schema validation for output shape, substring or set checks for grounding (hallucinated identifiers are string-detectable, and the most damaging error class), regex denylists for forbidden content such as prompt leakage, and simple computation for language, length and format. Whether the answer actually addressed the question is left to the judge.

Keeping LLM-as-judge honest

A judge is a second model call scoring the first model's output against a written rubric, with its own error rate and its own bill. Mahmoud treats the judge as production code with a version: pin the exact model ID rather than a floating alias, because a silent provider-side model change would read as a quality regression in your product. Ask binary questions instead of one-to-ten scores, since asking whether the answer states a refund deadline is reproducible while helpfulness ratings drift. Keep the rubric in a committed, reviewed file. Calibrate by hand-labelling twenty to thirty outputs and comparing the judge's verdicts — until a judge has been checked against human labels, its score carries no demonstrated meaning. Where budget allows, use a different model family for judging than for generation, since models tend to rate their own phrasing favourably.

Gating CI on regression, not perfection

The final recommendation: fail the build when the pass rate drops below a baseline committed to the repository, and treat raising that baseline as a reviewed commit. The gate targets regression rather than a perfect score.

Why it matters

Most teams shipping AI features already have CI that verifies everything except the behaviour customers actually see. This guide offers an incremental path — property assertions, a growing golden set, a calibrated judge and a regression gate — that fits existing test runners rather than requiring a new platform. It also pulls the riskiest artifacts in an AI feature, namely prompts, fixtures, rubrics and baselines, into version control where they become reviewable in diffs.

  • #llm-evals
  • #testing
  • #ci-cd
  • #llm-as-judge
  • #ai-engineering

Related posts