· via dev.to (home feed)
Roughly Two Dozen Open-Source LLMs Run on Cloudflare Workers AI's Free Tier
Hands-on testing published on dev.to finds roughly two dozen open-source LLMs usable free through Cloudflare Workers AI's single API, while many catalog entries remain paid-only.

What the experiment found
A hands-on test published on dev.to reports that roughly two dozen open-source LLMs will accept inference requests on Cloudflare Workers AI's free tier, all through a single API endpoint. The author, posting as interro-ai, set out to answer three questions: which models actually run without payment, which are best for coding, and how Cloudflare's limits compare with rivals such as Groq and Cerebras.
The starting point was Cloudflare's model catalog API, which returned more than 300 models for the account. That number proved misleading, because a listing is not a guarantee of access. Several larger models — among them zai-org's GLM 5.2 and GLM 5.3 variants and Moonshot's Kimi K2.6 and K2.7 Code — returned error 5035, "Model is not available on the Workers Free plan," when the author actually sent them requests.
To separate working models from gated ones, the author built a discovery script that walked the catalog, fired a test inference at every model, and recorded which calls succeeded and which required a paid plan.
The models that passed
According to the post, the free-tier survivors span nine model families:
- OpenAI: gpt-oss-20b and gpt-oss-120b
- Qwen: qwen2.5-coder-32b-instruct, qwen3-30b-a3b-fp8, qwen3.8-27b and qwq-32b
- DeepSeek: deepseek-r1-distill-qwen-32b
- Meta: Llama 3.1 8B, Llama 3.2 1B and 3B, Llama 3.3 70B (fp8-fast) and Llama 4 Scout
- Google: Gemma 2B and 7B LoRA variants, Gemma 4 26B, and AI Singapore's Gemma Sea-LION 27B
- Mistral: mistral-7b-instruct-v0.2-lora and Mistral Small 3.1 24B
- ZAI: glm-4.7-flash
- IBM: granite-4.0-h-micro
- NVIDIA: nemotron-3-120b-a12b
The post's headline counts 24 models, while the itemised list adds up to 21, so the exact figure is best treated as approximate.
One result surprised the author, who began the experiment specifically wanting to run GLM 5.3 Flash: only the older GLM 4.7 Flash turned out to be free, while every newer GLM model required payment.
The coding shortlist
For software development, three models rank above the rest. GPT-OSS-120B is the author's overall winner, credited with strong performance on Terraform, Kubernetes, general DevOps work, AWS architecture, refactoring and multi-file repositories. Qwen2.5-Coder-32B takes second place as a purpose-built coding model, rated highly for generating code, reviewing pull requests, fixing bugs and understanding repositories, including in editor tooling such as Continue.dev and Cline. DeepSeek-R1-Distill-Qwen-32B is the third pick and the reasoning choice for root-cause analysis, complex debugging, architecture reviews and agent workflows — the model the author would reach for during a late-night pipeline failure.
Honorable mentions go to QWQ-32B for reasoning, Llama 4 Scout for balancing speed and capability, and Llama 3.3 70B for reviews, documentation and system design.
How the limits work
Cloudflare's quota model differs from Groq's, according to the post. Groq commonly applies rate limits per model, whereas Cloudflare grants each free account 10,000 Neurons per day — effectively a shared daily budget across the account — alongside a documented ceiling of 300 requests per minute for text generation. The practical picture is an account-level daily budget, plus a text-generation rate cap, plus model-specific overrides.
Why it matters
The post challenges a common assumption that building with LLMs requires a paid OpenAI or Anthropic subscription, a rented GPU server or a cloud inference endpoint. For many projects, one free Cloudflare account offers GPT-OSS-120B, Qwen Coder 32B, DeepSeek R1 Distill, QWQ-32B, Llama 4 Scout and Gemma 4 through a single API — enough, in the author's view, to power coding assistants, agents, internal copilots, RAG applications, startup MVPs and developer tools without touching a GPU.
Caveats apply. This is one developer's snapshot rather than official Cloudflare documentation, and because availability is enforced per model, the lineup can shift as the free tier changes. The 10,000-Neuron budget is also shared across everything an account runs, so heavier workloads may exhaust it quickly. The author plans follow-ups measuring how many real coding requests fit inside the daily budget, plus comparisons against Groq and Cerebras. For anyone prototyping AI tooling, the practical lesson stands: probe the catalog with a real inference call before committing, because listed and usable are not the same thing on the free plan.
- #cloudflare
- #llm
- #open-source
- #inference
- #free-tier