Genie Generate a free company AI assistant Try it
← Back to Blog

OpenAI Prompt Cache Diagnostics: How to Investigate Misses and Measure Savings

OpenAI Prompt Cache Diagnostics: How to Investigate Misses and Measure Savings

Key Takeaways

  • The prompt-caching dashboard launched August 20; diagnostics reached GA September 8.
  • Diagnostics compare a request against an earlier completed response.
  • Requesting a comparison does not load the baseline conversation or change cache behavior.
  • Use reported usage for billing analysis; diagnostic counts answer a different question.
BLOOMIE
POWERED BY NEROVA

Produced by Bloomie for Nerova AI using automated editorial checks. Sources used for factual claims are listed below.

OpenAI added a prompt-caching dashboard on August 20, 2026, and made Prompt Cache Diagnostics generally available on September 8. Together they help teams distinguish an application-wide pattern from the cause of a specific miss. Neither tool makes a cache hit equivalent to a finished-task saving.

Monitor the pattern, then investigate a request

The release record establishes the two milestones. The dashboard provides aggregate visibility, while the diagnostics guide explains comparison against an earlier completed response on supported Responses API models.

Begin with the workload that has changed. A lower hit rate across all traffic calls for a different investigation than misses after a particular deployment. Separate results by the model, service tier, and application version you actually use, so unrelated request families do not hide the change.

A comparison is not conversation recovery

The guide uses prompt_cache_options.comparison_response_id to select a baseline. This asks for diagnostics; it does not load the earlier conversation or alter caching behavior. Compatible settings and an exact reusable prefix matter to the comparison.

Choose a baseline that matches the expectation being tested. Comparing unrelated requests can explain why they differ without revealing a defect. Retain the baseline during a targeted fix, then check that the changed request still produces the intended behavior. Cache-friendly input organization should not override the application's correct instructions or tool definitions.

Keep diagnostic counts separate from billed usage

A diagnostic cache hit means no expected miss was detected in the comparison; new input may still need processing. The documentation directs developers to usage fields for actual reported reuse and billing analysis, because diagnostic counts can differ.

That distinction protects cost estimates. A higher hit rate does not automatically mean a lower invoice if the workload also sends more input, makes more attempts, or generates longer answers. Calculate cost over the same completed task set and account for the applicable cache read and write charges.

Improve reuse without preserving stale context

A common engineering response is to stabilize shared instructions and tool declarations while keeping request-specific material in the appropriate place. Treat that as a hypothesis to test against diagnostics, not a reason to retain obsolete policies. Correct behavior remains the priority when a model, tool, or instruction intentionally changes.

Record enough configuration metadata to explain misses without logging private prompt contents. Investigate a representative sample, then monitor whether the fix survives real traffic. Readers who need the underlying concept can start with Nerova's prompt-caching guide; these new tools make the concept more observable in a running application.

Nerova context

Custom AI agents for business operations

Nerova builds custom AI agents for business operations. Companies use Nerova when they need AI support for customer intake, support, sales follow-up, research, website audits, internal handoffs, and workflow automation.

Nerova can help turn websites, business context, and operational workflows into practical AI systems: website chatbots, single-purpose agents, AI teams, audits, and automation workflows built around a clear business outcome.

Ask Bloomie about this article