Genie Generate a free company AI assistant Try it
← Back to Blog

Cloudflare AI Search reaches GA: multimodal retrieval and billing

Cloudflare AI Search reaches GA: multimodal retrieval and billing

Key Takeaways

  • GA includes multimodal retrieval and document-processing improvements.
  • Generation-model capability and embedding image support are separate choices.
  • Billing begins November 1; evaluate real ingestion and query traffic beforehand.
BLOOMIE
POWERED BY NEROVA

Produced by Bloomie for Nerova AI using automated editorial checks. Sources used for factual claims are listed below.

Cloudflare AI Search reached general availability on October 1, 2026, expanding its managed retrieval offering with image support and document-processing improvements. Billing begins November 1, according to the launch announcement. The operational benefit is a managed indexing and retrieval path; the quality of an application’s answers still depends on what it indexes and how it uses the results.

Managed retrieval still has a data contract

The service combines Cloudflare’s storage, vector-search and AI infrastructure. Its GA announcement highlights native image embeddings, OCR for PDFs and larger-file support. These capabilities can make previously inaccessible material searchable, but they also create new ways for a result to lose its original context.

A scanned invoice, a slide and a photograph do not have the same interpretation rules. An extracted number may be searchable while its associated unit or column heading is missing. Review retrieval examples at the document level, preserving source references and the surrounding evidence that a reader needs to check the answer.

Model selection changes retrieval behavior

The supported-model documentation separates generation and embedding models and lists which embedding choices accept images. That distinction matters: selecting a capable answer model does not automatically give an index multimodal retrieval.

Changing embeddings can change which documents rank highly, even when the user’s question stays the same. Keep a small set of representative queries with expected evidence, then compare retrieval results before replacing a model. Include negative examples where the corpus should not answer; a confident answer from irrelevant evidence is a failure that aggregate search traffic will not reveal.

Plan permissions and November costs together

Managed indexing does not define which users may see each document. The application needs a clear way to enforce access restrictions before retrieved material becomes model context or a displayed citation. Test two users with different permissions against the same query and verify that restricted content cannot leak through summaries.

Use the period before billing starts to measure ingestion volume, refresh frequency and realistic query traffic. Repeated reindexing and broad retrieval can materially change costs. Compare the complete workflow with the current approach, including freshness and correction effort, rather than judging the product solely by the time needed to build the first demo.

Nerova context

Custom AI agents for business operations

Nerova builds custom AI agents for business operations. Companies use Nerova when they need AI support for customer intake, support, sales follow-up, research, website audits, internal handoffs, and workflow automation.

Nerova can help turn websites, business context, and operational workflows into practical AI systems: website chatbots, single-purpose agents, AI teams, audits, and automation workflows built around a clear business outcome.

Ask Bloomie about this article