· via Cloudflare blog
Cloudflare pointed frontier AI models at its own WAF and logged 1,107 attack attempts
Cloudflare built an LLM-driven tester that mutates attack payloads the way a hacker would, ran 1,107 attempts against a customer staging environment, and turned the bypasses into new WAF detections.

What Cloudflare did
Cloudflare says customers keep asking whether its Web Application Firewall is ready for frontier AI models, so the company ran the experiment on itself. According to a Cloudflare blog post, the team built a tester that casts a large language model as an attacker probing a live application — a dynamic testing approach, distinct from static code analysis — with the goal of evaluating whether the WAF actually does its job.
The model operated blind. It had no visibility into the application's source code, no view of the WAF's rules, and access only to selected HTTP response data. The tester starts from known exploits the WAF already blocks and then iterates: re-encoding the payload, moving it to a different part of the HTTP request, or pivoting to another vulnerability class, using each response to pick the next variation. Requests that were not blocked became leads for human review rather than confirmed exploits.
The main run targeted an authorized customer staging environment across six attack categories and recorded 1,107 attempts. After Cloudflare removed malformed, benign, duplicate and out-of-scope observations from its review, the company reports that the vast majority of attacks were blocked. The requests that did get through were used to build new detections that harden the WAF for all customers.
How the loop works
Each scenario picks an attack category, a location in the request, a starting payload the WAF already blocks, and a fixed attempt budget. The loop invokes the model twice per iteration. A proposal call receives the starting request, context and a short history of earlier results, then suggests the next variation, which Cloudflare's code builds and sends. A review call receives the response status, selected headers and body, and its verdict determines the next step. The loop ends when mutations stop producing useful variations or the attempt limit is reached.
The guardrails are notable. The models never send requests directly — a Python orchestrator checks the target hostname against an allowlist, disables redirects, records every attempt and enforces the limit. Response text may reappear in a later prompt, so it is treated as untrusted input. Neither model call receives rule expressions, rule IDs, WAF Attack Score details or the identity of the security layer that acted, and neither can deploy rules or change enforcement.
In total, Cloudflare ran 45 scenarios, 44 of them spanning cross-site scripting, SQL injection, command injection, server-side request forgery, path traversal and local file inclusion, and Log4j; a single log-injection scenario was reported separately. The test zone's WAF used WAF Attack Score blocking at 30 or below, the full Cloudflare Managed Ruleset and the OWASP Core Ruleset at Paranoia Level 3, with an allowlisted test User-Agent so the customer's automated-traffic controls would not stop requests before they reached the WAF.
What adaptation looked like
The clearest illustration is an SSRF scenario aimed at a cloud metadata address, 169.254.169.254, which can expose temporary credentials to whoever can make an application fetch it. The model tried the same address as a decimal integer (2852039166), as an octal form (0251.0376.0251.0376), as hostname-based representations, and in different parts of the request, revising its hypothesis after each block. The WAF stopped every variant except one: at attempt 18, keeping the previous request structure but switching to a trailing-dot representation of the address, the request drew a redirect rather than a WAF block. Cloudflare cautions that the recorded hypotheses summarize what the model claimed it was trying, not proof its reasoning was correct.
Why it matters
The test is a concrete demonstration of what LLM-driven attackers actually bring: not necessarily novel exploit primitives, but relentless, fast iteration on encoding and delivery that outpaces any human tester. Cloudflare says the exercise is becoming a foundational building block of its WAF development lifecycle, converting each bypass into new detections for every customer. The company also stresses the limit of the finding: a payload that slips past a WAF still needs an exploitable application to succeed, so patching your software and deploying the WAF correctly remain among the strongest defenses.
- #cloudflare
- #waf
- #security
- #llm
- #penetration-testing