· via dev.to (home feed)
Sixteen of 69 agent skills failed a second test on harder inputs despite passing first runs
A vendor of agent skills re-ran all 69 of its skills on deliberately difficult inputs and found defects in 16, all of which had passed an initial run. The failures clustered on uncertainty restated as fact.

One passing run was not enough
A team that builds and sells agent skills re-tested all 69 of them on deliberately harder inputs and found defects in 16, a bit under one in four. According to a post on dev.to by ProSkillPacks, every one of those 16 skills had already passed a first run on a real public input that a human read and judged acceptable.
The authors' takeaway is blunt: a single successful run shows a skill can do the job once, not that it will keep doing it.
How the retest worked
For each skill, the team prepared one input drawn from five categories of difficulty:
- Uncertain notes, packed with phrases like 'I think' and 'maybe'.
- Contradictions, such as a brief asking for something both cheap and premium, or a listing whose title and body disagree.
- Thin or blocked sources: near-empty pages, dead links, a 403.
- Requests to overclaim, for example asking for impressive-sounding results that the input does not contain.
- Messy data, including duplicates, a month numbered 13 and absurd amounts.
Each output was read by a person and compared against its input. There was no automated grader and no second model acting as judge. All runs used Claude Sonnet, which the authors list among their limitations.
What actually broke
The 16 defects broke down as follows:
- Seven restated uncertain input as fact.
- Three produced invented or unsupported claims.
- Two miscounted items in a summary.
- Two failed to name a conflict present in the brief.
- One script misread a compressed response.
- One wrote a file it should not have touched.
The largest group is also the quietest, the post notes: a skill would acknowledge uncertainty somewhere, then present it as settled fact in the text meant for an end reader.
The concrete examples are instructive. A newsletter skill given an unconfirmed postage change generated a subject line announcing the increase as decided; after the fix it became a hedged question about April. A host's tentative note about outside smoking hardened into a flat house rule; the corrected output marked the point as unsettled and asked for confirmation. A listing that hedged its walking distance to the beach produced a title dropping the hedge entirely, keeping the qualifier only in the description. A claim checker reported nine claims with eight rated high-risk while its own table showed seven high and two medium. And a read-only CSV profiler claimed in a single sentence to have left a file untouched and to have written a cleaned version of it.
A pattern worth noting: hedges vanished first in the shortest fields, such as titles, subject lines and summaries.
Small fixes, always re-run
Apart from one script repair, every defect was fixed by adding one or two sentences to the skill's instructions. The team also adopted a standing rule for skills that write text for other people: anything marked as unsure in the input stays marked in every output, including final copy, and is never stated as fact. Every repaired skill then passed a fresh run on the same input that broke it.
A checklist you can borrow
The post lays out an afternoon-length test for any prompt or skill. Feed it a hedged fact and check that the hedge survives into the title, subject line and summary. Give it two facts that cannot both be true and see whether it names the conflict or quietly picks a side. Remove the key input, or point it at an erroring page, and confirm it reports the gap rather than filling it. Ask it to overclaim and see whether it refuses. Hand-verify every count in a summary. Compare stated actions with actual behaviour: if a tool is supposed to be read-only, confirm nothing was written. And re-run after any fix, because a repair that has not been tested again proves nothing.
Why it matters
The headline failure rate of roughly 23 percent comes from a vendor testing its own products, on a single model, with self-judged outputs, one hard input per skill and partly synthetic test cases, so it is not a general failure rate for agent skills. But the method matters more than the number: cheap adversarial second runs surface defects that ordinary first runs hide, and the dominant failure class, uncertainty quietly hardening into confident prose, is exactly the kind a casual skim of output will miss. Anyone shipping prompts, skills or agents can apply the same checklist in an afternoon, and the shortest output fields are the place to look first. The commercial interest is disclosed in the post itself, which is worth keeping in mind when weighing the numbers.
- #ai-agents
- #llm-evaluation
- #prompt-engineering
- #quality-assurance
- #claude