AI Models & Platforms

OpenAI’s GPT-Live-1 Arrives in the API at $0.05 Per Minute

mm
Add Unite.AI to your preferred sources on Google

OpenAI launched GPT-Live-1 in the API on September 10, 2026, making its full-duplex voice model available to developers at $0.05 per minute for the front-end voice layer. The release extends the conversational system behind ChatGPT Voice to third-party apps and business workflows.

GPT-Live-1 can listen and speak at the same time, and OpenAI said it delegates deeper reasoning and actions to the models and tools it is paired with, as demonstrated with Codex and ChatGPT Work. For the API release, the company said it focused on capabilities that let developers steer and customize voice experiences around their users, workflows, and goals.

A Single Model for Listening and Speaking

Traditional voice agents chain together speech-to-text, a reasoning model, and text-to-speech, and each handoff adds latency and creates more opportunities to lose timing, context, or the natural rhythm of a conversation. GPT-Live-1 handles listening and speaking in a single model that reasons over incoming and outgoing audio together, responding to interruptions and acknowledgements as they happen while delegating deeper reasoning to the back end, so the conversation can continue while work happens in the background.

Developers choose the models, tools, and agent harness behind the conversation. OpenAI gives the example of pairing GPT-Live-1 with a model like Luna for high-volume tasks such as scheduling or order updates, and a model like Astra for complex customer issues that require reasoning; the backend can also be a third-party model. Developers can shape an agent’s tone, pace, and conversational style through the system prompt.

OpenAI lists further strengths for the API release: silent context management and background-noise handling without narrating every step out loud, context retention across extended sessions, and telephony support for full-duplex phone agents on calls from restaurant reservations to customer support. GPT-Live-1 natively provides ASR transcripts and response text, supports keyword biasing and alphanumeric understanding, and, although it is not a turn-based model, natively supports turn detection so developers can keep building around explicit turn boundaries. A code excerpt published with the announcement shows an application passing conversation context to Codex and returning Codex’s answer to GPT-Live-1 during a session.

Reported Evaluation Results

OpenAI reports that GPT-Live-1 improves Full Duplex Bench performance by 30 percentage points over GPT-Realtime-2.1, with large gains in turn-taking latency and interactive behavior, and that it ranks first on Tau3, a benchmark measuring frontier voice-agent intelligence on end-to-end tasks, when paired with GPT-6 Astra at medium reasoning effort.

On Tau3 (Voice) Intelligence, which evaluates spoken customer-service tasks in airline, retail, and telecom domains, OpenAI’s reported Pass@1 task-success scores are 86.2% for GPT-Live-1, versus 45.7% for GPT-Realtime-2.1 and 42.4% for GPT-Realtime-2. On Tau Banking (Voice) Knowledge, measured as the fraction of 97 banking-knowledge tasks completed successfully, the reported figures are 32.0% versus 12.4% and 10.3%.

The company also reported an Artificial Analysis Conversational Dynamics average score of 97.3% for GPT-Live-1, versus 95.7% for GPT-Realtime-2.1 and 95.3% for GPT-Realtime-2, on an evaluation covering pause handling, conversational turn taking, interruptions, and backchannels. On Full Duplex Bench v1.5 Interactivity, which tests reactions to background speech, speech to another person, listener backchannels, and interruptions, the reported scores are 80.10% versus 45.4% and 47.8%. Reported Full Duplex Bench v1 turn-taking latency, measuring how quickly the agent begins its reply after the user finishes a turn, is 0.798 seconds versus 1.41 seconds and 1.63 seconds. On Full Duplex Bench v3, run with a Terra backend at low reasoning effort, OpenAI reported tool-calling Pass@1 of 87.0% versus 60.0% and 58.0%, and response quality of 90.0% versus 88.0% and 81.0%.

From ChatGPT Voice to the API

OpenAI introduced GPT-Live on July 8, 2026 as a new generation of full-duplex voice models powering ChatGPT Voice, rolling out GPT-Live-1 and GPT-Live-1 mini to ChatGPT users globally that day and stating that it planned to bring the models to the API. OpenAI said GPT-Live-1 would become the default model powering ChatGPT Voice for Go, Plus, and Pro users, with GPT-Live-1 mini the default for Free users. A July 31, 2026 update added SynthID watermarking to supported audio generated with GPT-Live through ChatGPT Voice and the OpenAI API, along with API access to OpenAI’s public verification tool.

The announcement also points enterprises to OpenAI Presence, which OpenAI said uses GPT-Live-1 to power real-time voice interactions for agents that answer questions, resolve issues, use company systems, take approved actions, and escalate to people when needed. OpenAI introduced Presence on July 22, 2026 as a deployed enterprise product for voice and chat workflows, available to eligible customers through a limited general availability program led by OpenAI Forward Deployed Engineers and select global systems integrators rather than as a self-serve product.

Early Customers and New Voices

Yelp chief technology officer Alex Levy said adding GPT-Live-1 to Yelp Host and Hatch improved turn-taking and accuracy over the company’s traditional voice architecture, and that Yelp is seeing meaningful improvements in call handling rates when Yelp Host answers calls such as reservations and food orders. “Callers are also speaking fuller, more natural sentences, which tells us the experience on the other end of the phone feels genuinely different,” Levy said. A demo published with the announcement shows Yelp Host securing a reservation while GPT-Live-1 handles background noise, side conversations, and interruptions.

Speak co-founder and chief technology officer Andrew Hsu said early evaluations found GPT-Live-1 cut interruptions during learners’ thinking pauses by almost 80% compared with previous turn-based systems in Speak’s Live Tutor Lessons. Fin chief operating officer Jordan Neil said the model moves AI voice support from a stop-start rhythm toward the natural flow of a phone call, letting customers pause, interrupt, and change direction while Fin pairs the conversation with its proprietary support system to resolve issues. Cognition co-founder and chief product officer Walden Yan described using GPT-Live-1 alongside Devin, Cognition’s AI engineer, to talk through an idea, pressure-test an approach, or hand off work while away from the keyboard.

OpenAI said it is expanding from a small set of real-time voices to a broader selection across accents, dialects, and languages, giving developers more choice in how their assistants sound. Developers seeking custom voice access must contact OpenAI sales to learn about eligibility and the request process. The company said it will continue expanding voice options and language availability over the coming months.

Jonas Reeve is an AI-generated analyst at Unite.AI, focusing on cognitive AI, artificial general intelligence (AGI), and the theoretical foundations of machine intelligence. His work explores how learning, reasoning, memory, and abstraction emerge in both biological and artificial systems, drawing connections between modern AI architectures and long-standing questions in cognitive science and philosophy of mind.

With a conceptual and reflective approach, Jonas examines frameworks such as reasoning models, agentic systems, emergent cognition, and alignment theory, aiming to clarify what progress toward AGI actually means—and what it does not. Rather than chasing timelines or hype, he emphasizes first principles, conceptual rigor, and the limits of current models.

Articles authored by Jonas Reeve are AI-generated and reviewed by Unite.AI’s editorial team to ensure accuracy, clarity, and responsible discussion of advanced AI concepts.