GPT-Live 1 reached general availability in the OpenAI API on September 10, 2026, according to the current API changelog. It handles a spoken conversation while a separate backend model or agent performs reasoning and tool work. That architecture is the important change for developers building voice agents.
The dated release entry documents the API availability date. OpenAI’s GPT-Live guide explains full-duplex conversation and backend delegation. The GA milestone is distinct from earlier ChatGPT voice coverage.
The voice model and task backend have different jobs
GPT-Live manages listening, speaking, and conversational delegation. The backend carries detailed business rules and tool workflows. Developers can use Responses delegation or connect an existing service with client delegation, as described in the guide.
Separating those roles can make a voice application easier to reason about. A conversational response can acknowledge the caller while a backend retrieves an order. The application still needs to distinguish an acknowledgment from a confirmed result, so it does not tell the caller that an action succeeded before the underlying system accepts it.
Full duplex makes interruption handling more important
A caller can add information while the assistant speaks. That creates an opportunity for a more fluid interaction and a challenge for workflow state. A pending lookup might remain useful after an interruption, while a pending mutation might require cancellation or a new confirmation.
Define those behaviors in application code. A model should not decide implicitly whether a changed request supersedes an earlier transaction. Give each task an identity, preserve the current confirmed inputs, and make the final result attributable to the action actually executed.
Session duration is only part of the bill
The guide says GPT-Live voice sessions are billed by duration, per second, while backend model and tool usage are billed separately. A price comparison that counts only the voice session can therefore miss a material part of the application’s cost.
Measure complete calls: time spent waiting, backend requests, retries, and any external tools. Closing abandoned sessions matters for cost and resource cleanup. Also inspect whether additional reasoning saves enough later work to justify its latency and cost for the particular task.
What to validate before switching a production agent
Reuse a representative task set from the current voice system. Include corrections, overlapping speech, silence, slow tools, and disconnections. Check both what the assistant says and what the backend persists. A polite response after a failed write is still a failed workflow.
OpenAI explicitly leaves permissions, required confirmations, function execution, and saved progress to the application. Keep sensitive credentials on a trusted server and scope each operation. Nerova’s view is that GPT-Live offers a useful conversation layer, while dependable delegation depends on the surrounding system making authority and state explicit.