deniz.in

Markets

Weather

Loading weather

· via dev.to (home feed)

TypeScript contract tests catch silent breaks when you upgrade LLM model strings

A dev.to post argues that swapping an LLM model string is a breaking dependency change, then builds an offline TypeScript test gate that mocks two model versions and diffs their behavior.

TypeScript contract tests catch silent breaks when you upgrade LLM model strings

A post on dev.to makes a deceptively simple argument: the model identifier sitting inside your API call is not configuration. It is one half of a live API contract, and changing it should be treated the same way as bumping a major dependency — with a test suite that runs before the change ships. The author then builds that suite in TypeScript, entirely offline.

A month of contract changes, as told by the post

To justify the effort, the post catalogs four provider releases from late September, each of which altered documented behavior:

On September 17, Google shipped Antigravity Agent 09-2026. According to the post, its built-in tools moved to PascalCase parameters, and write_file(path, content) was replaced by write_to_file or replace_file_content. The older antigravity-preview-05-2026 is retired on October 5, 2026.

On September 22, Anthropic launched Claude Opus 5.5. The cited release notes say thinking payloads of type disabled and enabled now produce 400 errors, as do tool_choice types any and tool.

On September 28 came Claude Sonnet 5.5, whose release notes list five ways existing Sonnet 5 code can break: between_tools replaces disabled for turning off up-front thinking at high effort or below; forced tool use returns a 400; thinking blocks are now tied to the model and conversation; the earlier computer_20251124 computer-use tool is no longer accepted on the Claude API or Google Cloud; and the advisor tool rejects Claude Opus 4.8, Opus 4.7 and Sonnet 5 in the advisor role. A separate "what's new" page adds a sixth change: text between tool calls now arrives inside thinking blocks — a response-shape change that fails no request at all.

On September 29, OpenAI released GPT-6.1 Sol at DevDay, and per its model page the none and minimal reasoning efforts are unsupported.

Three companies, the post notes, one repeating pattern.

Loud failures versus quiet ones

The core distinction the article draws is between the two failure modes. A 400 error is loud — it fails fast and gets noticed. But a change like progress text migrating into thinking blocks, where it renders as empty at default display settings, is quiet. Users simply see nothing before a tool call fires, and any test that only checks status codes will pass. That asymmetry is why the author considers a diff between model versions necessary, not just assertion-based tests.

Neutral types, then two mocked versions

The walkthrough starts with a Node 18+ project and three dev dependencies: typescript, tsx and @types/node. Step one defines neutral request and response types that describe the application's own shape — prompt, token budget, thinking mode, tool choice — rather than mirroring any vendor SDK. The author's stated principle: contracts should encode what your app needs, not what the provider exposes.

Step two builds two mocks of the same provider, claude-sonnet-5 and claude-sonnet-5-5. The newer mock reproduces the documented rejections for disabled thinking and forced tool use, plus the shape shift that moves the inter-tool progress note into a thinking block. Error wording and reply payloads are invented, and nothing touches the network or requires an API key. The post is blunt about the alternative: a mock that accepts every request would prove nothing.

Seven contracts, each returning a verdict

Step three encodes seven promises as plain functions: the output JSON has a numeric total and a string currency; tool calls carry a well-formed input with a tool_use stop reason; progress text is visible before a tool call; forced tool use still produces a call; the thinking-off fast path returns a 200; a low token budget stops with max_tokens; and a declined request returns a 200 with a refusal stop reason, matching the documented behavior. Each check returns null on success or a human-readable reason on failure — no assertion framework required.

The remaining steps, outlined in the post's table of contents, run the same contracts against both mocked versions and diff the results, add checks derived directly from the release notes' known breaks, and assemble everything into a gate that blocks an upgrade before it reaches production. The framing is that one check reads the release notes while the other compares two model versions directly.

Why it matters

Most teams change a model string the way they change a log level. This post reframes that string as a versioned dependency whose upgrades break things in two ways: loudly, with rejections, and quietly, with reshaped responses. Contract tests against mocked providers close both gaps cheaply — they run in CI, need no credentials, and catch shape drift that status-code checks miss. The four September releases the author cites, spread across three vendors, suggest this is not an occasional hazard but the normal cost of living on fast-moving model APIs. For anyone maintaining production LLM integrations, the post offers both the argument and a working starting point.

  • #ai
  • #typescript
  • #testing
  • #llm
  • #api-design

Related posts