OpenAI has officially launched its next-generation voice models, GPT-Live-1 and GPT-Live-1 mini, delivering a comprehensive upgrade to ChatGPT's voice capabilities.
The new models employ a "full-duplex" architecture, enabling them to listen and speak simultaneously. This allows users to interrupt the AI's response at any time, creating an interaction style that more closely resembles a real human conversation.
Previously, ChatGPT's advanced voice mode operated on a turn-by-turn basis, requiring the model to wait for the user to finish speaking before responding. Brief pauses were often misinterpreted as the end of speech, leading to unnatural interruptions. The GPT-Live-1 model resolves this issue by continuously processing input while simultaneously generating output, making decisions multiple times per second on whether to speak, continue listening, or interrupt. During conversations, the model will use brief acknowledgments like "hmm," "right," or "got it" to signal it is actively listening. Users can also instruct it to slow down or remain quiet while awaiting a command.
A further enhancement is the separation of the voice interaction layer from the reasoning layer. When a question necessitates a web search or more complex reasoning, GPT-Live can delegate the task to the backend GPT-5.5 model, all while maintaining a fluid conversation. Paid subscribers can choose between "Instant," "Medium," and "High" reasoning levels to balance speed and intelligence according to their needs.
Additionally, GPT-Live supports real-time translation and can display information such as weather and stock data through visual cards. OpenAI stated that GPT-Live-1 will become the default voice model for Go, Plus, and Pro subscribers, while free users will have access to GPT-Live-1 mini. Both models are beginning a global rollout today to iOS, Android, and web users, with API access to follow at a later date. With over 150 million people using ChatGPT's voice and dictation features weekly, this upgrade is expected to significantly expand the use cases for voice-based interaction.