· via Hacker News – Front Page (native)
Jev-driven pipeline passes 76% of SREGym-Lite SRE diagnoses without an LLM agent
A programmatic collector turns Kubernetes state into evidence, and the Jev decision model picks root causes, passing 76.2% of SREGym-Lite fault diagnoses at a median of 14.6 seconds.
A pipeline with no agent in the loop
The team behind SREGym has published results from an automated incident-diagnosis pipeline that contains no LLM agent at all. Instead, a programmatic collector gathers and organizes Kubernetes evidence, and Jev — a model that answers questions by selecting from supplied options — makes the judgment calls. According to the sregym.com blog post, which surfaced on Hacker News's front page, the pipeline passed 80 of 105 diagnoses (76.2%) across 21 SREGym-Lite faults, with a median diagnosis time of 14.6 seconds.
The post follows an earlier experiment in which Jev acted as a decision aid for an LLM agent, helping rank the agent's proposed tests and reviewing evidence before submission. The new work drops the agent and hands the diagnostic decisions to Jev alone.
How the pipeline works
Because Jev can only choose among provided answers, the pipeline must supply both the evidence and the candidate answers. The collector reads Kubernetes objects, events, recent pod logs and resource usage, then groups observations by component — a Deployment, for instance — and summarizes the failure signals. Jev reviews those summaries and picks the component most worth inspecting.
The collector then gathers finer detail on that component and prepares numbered evidence items. Jev decides whether the component caused the fault, was affected by it, or is irrelevant, and selects the evidence that best supports its answer. From those choices the pipeline assembles and submits a diagnosis. If the evidence cannot back the hypothesis, it moves to the next candidate.
The current version examines one candidate at a time. Jev never generates commands or writes the report text; the authors have published the implementation on GitHub.
The webhook fault, step by step
The post walks through the Social Network application's mutating_webhook_resource_limits fault. Pods created for nginx-thrift kept being OOMKilled: the Deployment template specified a 256Mi memory limit, but live Pods arrived with 16Mi. A mutating admission webhook was rewriting limits at pod creation, and four other webhook configurations existed in the cluster, so searching by name alone could not isolate the culprit.
The collector summarized 27 Deployments in the social-network namespace and, for nginx-thrift, flagged both the template-versus-pod memory mismatch and a matching webhook, gatekeeper-mutating-webhook-configuration. Jev identified nginx-thrift as the likely origin and admission_webhook as the fault-carrying object type. Given 26 evidence items in a second call, it classified the component as the origin, categorized the cause as an admission or namespace policy issue, and singled out the evidence item showing the 16Mi limit alongside the matching webhook. The pipeline then inferred and named the mechanism in its submission. All five attempts on this fault passed the rubric — though, as the authors note, the collector did substantial work here, spotting the mismatch and narrowing the webhook candidates before Jev made any choice.
Results and consistency
The evaluation covered the 21 fault scenarios in the September 4 SREGym-Lite cohort, run five times with jev-1.13.0, with each run submitting one diagnosis. Scoring used gpt-6-astra at high reasoning effort with SREGym's nine-question rubric and a 0.70 pass threshold; the authors emphasize this is the historical cohort, not the current leaderboard set.
Beyond the 76.2% pass rate: 16 of 21 faults passed all five attempts and 5 failed all five; Jev made 252 calls in total (2.4 per attempt) with a median summed API latency of 0.53 seconds per attempt; the runs consumed 3.48 million input tokens at an estimated cost of $0.15, using TypeSafe's published pricing. The outcomes clustered completely: every fault was either five-for-five or zero-for-five, and 18 of 21 faults received the same diagnosis score in every run — though identical scores do not prove identical investigation paths. The five unrecoverable faults involved edge request filter CPU saturation, a Kafka poison pill, a search-rate retry collapse, a service DNS resolution failure, and Valkey auth disruption.
Cost versus LLM agents
Jev's 76.2% sits just under GPT-5.6 Sol (medium)'s 77.8% on the same faults, while running about 7× faster and costing roughly 200× less per diagnosis. The post also charts results against GPT-6 Astra, GPT-6.1 Sol, Claude Opus and Sonnet variants, Codex and Claude Code, noting these are separate experiments on the same 21 fault IDs.
The authors' central open question is granularity: a coarse summary may omit the detail that explains a fault, while forwarding every line of YAML can drown the useful signal. They argue that choosing what to show Jev may matter as much as the model's ability to judge it.
Why it matters
The results sketch a plausible two-tier architecture for automated incident response: a fast, cheap, constrained judge handles first-line triage, escalating to an LLM agent only when its supplied evidence and options run out. Jev is clearly less flexible than an agent — it cannot investigate freely, and its five zero-for-five faults mark exactly where the collector's evidence fell short. But the consistency across runs, the sub-15-second median, and a near-agent pass rate at a tiny fraction of the cost make this a concrete data point for teams building automated SRE tooling on Kubernetes.
- #sre
- #kubernetes
- #ai-agents
- #incident-response
- #diagnostics