· via Cloudflare blog
Cloudflare AI Search reaches general availability with image embeddings and PDF OCR
Cloudflare's managed AI search and retrieval pipeline is now generally available, adding native image embeddings, OCR for scanned PDFs, larger files and usage-based billing from November 1, 2026.

What Cloudflare announced
Cloudflare has moved AI Search, its managed indexing and retrieval service, into general availability, according to a post on the Cloudflare blog dated October 1, 2026. The product has existed in preview for more than a year and combines several of the company's existing building blocks — Workers AI for embeddings, Vectorize for vector storage and R2 for object storage — into one pipeline that handles parsing, chunking, embedding and retrieval. Cloudflare says it runs search on its own blog and developer docs, and that preview customers have used it for everything from internal documentation search to public site search.
Alongside GA, the company expanded support for content beyond plain text: native image embeddings, OCR for scanned PDFs and larger file limits.
Native image embeddings replace caption-only search
Cloudflare is candid that its earlier image support was crude: the service detected objects in a picture, generated a caption and then embedded that text. Anything the caption failed to mention — texture, layout, composition, fine visual detail — was effectively unsearchable.
The GA version embeds image pixels directly while keeping captions for textual understanding. To stop these richer representations from bloating storage and slowing queries, Cloudflare applies Matryoshka Representation Learning, which lets smaller embeddings retain useful signal. Native multimodal retrieval ships today with the Qwen3-VL-Embedding model.
At query time the service checks whether an instance's embedding model handles images. If it does, a query image is embedded straight into the same vector space as the indexed content. Instances running text-only models are not left out: the query image is converted to a text description using a component Cloudflare calls ToMarkdown, and that caption is used for the search instead.
The practical effect is that visual queries become possible — describing an image you want to find, supplying a similar image, or mixing both, such as asking for a bird with matching markings. Cloudflare points to product discovery, screenshot matching, charts, diagrams and scanned documents as collections where color, texture and spatial relationships matter.
OCR and larger files
AI Search now accepts text formats such as Markdown, HTML, CSV and JSON, plus PDFs, up to 10 MiB per file, raised from the previous 4 MiB cap. Because many PDFs are really scanned images with no extractable text layer, an optional OCR step reads the text off each page before chunking and embedding. OCR is available to every account and is billed under the new pricing as image-processing ingestion tokens.
What a query looks like
According to the Cloudflare blog, a query is optionally rewritten, then embedded — with image queries going directly to multimodal models or through captioning first on text-only setups. Vector and keyword search run in parallel, results are fused and optionally reranked, and the top chunks are either returned directly or handed to a generation model to write an answer.
Billing starts November 1, 2026
General availability brings paid billing, live from November 1, 2026, with a free allotment on all Workers plans. Costs fall into three buckets — ingestion, storage and queries — while parsing, chunking, embedding, keyword indexing and reranking are included. There are no instance hours, capacity units or monthly minimums to size in advance.
The paid rates are $0.75 per million ingestion tokens, with image processing as a $0.50 per million token add-on; $2.00 per GB-month for stored data; $0.75 per thousand semantic (vector and hybrid) queries; and $0.10 per thousand full-text queries. The free tier includes a shared pool of 5 million ingestion tokens per month across supported file types, 10 GB of storage, and — one change from the preview pricing announced during August 2026's Agents Week — 1,000 semantic queries plus 1,000 full-text queries instead of a shared pool of 2,000. Embedding and reranking are free with select Workers AI models; third-party models are billed separately.
What's next
Cloudflare says an ingestion pipeline for full video and audio processing is in the works so customers can search rich media assets. The keyword search engine is being reworked to scale better for large data stores, where the current implementation hits limits, and simpler setup paths are coming for websites already running on Cloudflare so that AI agents can discover and consume their content more easily.
Why it matters
Retrieval-augmented generation has become standard plumbing for AI applications, but assembling it usually means wiring together parsers, embedding APIs, a vector database, hybrid search and a reranker from several vendors. Cloudflare is packaging that entire stack as one managed, usage-priced service with a free tier large enough for small projects, which lowers the barrier for developers already in its ecosystem and adds competitive pressure on standalone vector database and search providers.
The multimodal upgrade is the more interesting signal. As AI systems increasingly work across images and documents rather than plain text, retrieval that operates on pixels instead of lossy captions becomes a meaningful capability — and Cloudflare is shipping it as a default part of the pipeline rather than something developers must assemble themselves. As with any vendor announcement, the capability and pricing claims are Cloudflare's own.
- #cloudflare
- #ai-search
- #vector-databases
- #rag
- #multimodal-ai