· via dev.to (home feed)
Prompt injection broke every frontier agent tested in 272,000-attempt red-team
A public red-teaming competition logged 8,648 successful prompt injections against 13 frontier agents, while Apple moves to tighten macOS Full Disk Access over risks from AI agents.

A public red-teaming competition fired 272,000 prompt-injection attempts at 13 frontier AI agents and recorded 8,648 successes. Every model in the test proved vulnerable, with per-model success rates ranging from 0.5% to 8.5%, according to figures cited in a dev.to analysis of how much authority AI agents are being granted.
What the red-team actually found
The headline number is less important than the structure behind it. As the dev.to writeup reports, drawing on the competition's published results, certain attack strategies transferred across 21 of 41 tested behaviors and across multiple model families. That pattern points to shared weaknesses in how instruction-following architectures separate operator instructions from untrusted content, rather than to any single vendor's implementation.
Two further findings raise the stakes for deployments. Capability and robustness barely correlated — the most capable model in the test was also among the most vulnerable — so choosing an agent by benchmark score says nothing about how much damage a successful injection can do. And because agents continuously ingest untrusted material such as inboxes, documents and code repositories, adversarial instructions embedded in that content can steer their behavior without leaving any trace the user would see.
Apple moves on the other end of the problem
The permission side of the story was moving in parallel. On October 2, 2026, Apple published a developer notice, "Updates to Full Disk Access in macOS," saying some developers use that permission in ways that can expose a user's files, mail, messages and browsing history without the user fully understanding what was handed over. Users who genuinely want that level of access will need to take very explicit action to grant it. Apple presented the change as making consent more meaningful rather than as a new limitation, and MacRumors reports that no ship date was announced.
The backdrop includes a dispute over agent behavior: an Inc. columnist reported that Meta's Muse agent on his Mac had knowledge of his private messages, an account Meta disputed, according to TechCrunch. Apple named no product, though commentary has connected the change to agents including Muse, Grok, Claude and Dots.
Why a phone habit became a security hole
Citing Daring Fireball, the dev.to piece frames the deeper issue as architectural asymmetry. On iOS and iPadOS, no permission tier lets a third-party app read email or end-to-end encrypted iMessage and WhatsApp conversations. On macOS, approving the prompts an agent presents can effectively hand over the entire startup drive. Users trained to click "OK" on a phone, where consent is genuinely bounded, arrive at the Mac with the wrong mental model — and the flaw sits in the design, not in the user.
Tightening is not free, either. Backup utilities and disk-mapping tools legitimately need broad filesystem access, and developers argued in public comment that treating every broad grant as suspicious would break real software.
The authorization gap
A 2026 review of 89 primary sources on agent authorization, cited in the writeup, argues that trustworthy systems need three properties few deployments achieve together: every consequential action traceable to a human principal, no wider than the delegation that human actually made, and possible to challenge or review after the fact. The review models authority as a chain running from the human user through operator, orchestrator agent, sub-agent and tool endpoint, and flags runtime enforcement and limits on aggregated authority as the biggest open problems.
The industry has at least named the issue: the OWASP Top 10 for LLM Applications ranks Excessive Agency among its top risks, and an Okta survey found 74% of IT application leaders believe agents represent a new attack vector into their organizations.
Why it matters
The red-teaming data closes off two comfortable assumptions: that some frontier model is injection-proof, and that picking a more capable model buys safety. Neither holds. What actually limits damage is the permission grant a model inherits, which is why Apple is engineering the boundary at the consent prompt rather than inside the model. For teams shipping agents, the practical shift is one of timing: scope credentials to the current task and revoke them at completion, check authorization when each tool is invoked rather than once at startup, make sure sub-agents never exceed the authority of whoever spawned them, and keep audit trails that outlive the session. Until those controls exist, a broad standing grant is simply the payoff a successful injection is waiting for.
- #prompt-injection
- #ai-agents
- #security
- #red-teaming
- #macos