Cloudflare AI Search reached general availability on October 1, 2026, expanding its managed retrieval offering with image support and document-processing improvements. Billing begins November 1, according to the launch announcement. The operational benefit is a managed indexing and retrieval path; the quality of an application’s answers still depends on what it indexes and how it uses the results.
Managed retrieval still has a data contract
The service combines Cloudflare’s storage, vector-search and AI infrastructure. Its GA announcement highlights native image embeddings, OCR for PDFs and larger-file support. These capabilities can make previously inaccessible material searchable, but they also create new ways for a result to lose its original context.
A scanned invoice, a slide and a photograph do not have the same interpretation rules. An extracted number may be searchable while its associated unit or column heading is missing. Review retrieval examples at the document level, preserving source references and the surrounding evidence that a reader needs to check the answer.
Model selection changes retrieval behavior
The supported-model documentation separates generation and embedding models and lists which embedding choices accept images. That distinction matters: selecting a capable answer model does not automatically give an index multimodal retrieval.
Changing embeddings can change which documents rank highly, even when the user’s question stays the same. Keep a small set of representative queries with expected evidence, then compare retrieval results before replacing a model. Include negative examples where the corpus should not answer; a confident answer from irrelevant evidence is a failure that aggregate search traffic will not reveal.
Plan permissions and November costs together
Managed indexing does not define which users may see each document. The application needs a clear way to enforce access restrictions before retrieved material becomes model context or a displayed citation. Test two users with different permissions against the same query and verify that restricted content cannot leak through summaries.
Use the period before billing starts to measure ingestion volume, refresh frequency and realistic query traffic. Repeated reindexing and broad retrieval can materially change costs. Compare the complete workflow with the current approach, including freshness and correction effort, rather than judging the product solely by the time needed to build the first demo.