OpenAI has finished rolling out GPT-Live, the model now powering ChatGPT's voice mode, to every paying tier worldwide, with free users catching up as we speak. The pitch is simple: the old voice mode, built on a GPT-4o era model with a stale 2024 knowledge cutoff, was stiff. You'd talk, it would process, then it would answer, and the whole thing felt like talking into a walkie-talkie. GPT-Live runs full-duplex, meaning it listens and speaks at the same time, deciding multiple times a second whether to stay quiet, jump in, interrupt, or hand off to a tool. It throws in small acknowledgments like "mhmm" while you're still talking, lets you interrupt mid-sentence, and will slow down or pause if you ask it to. Nine voices got remastered for the launch.
GPT-Live itself is not the model doing the heavy thinking. When a question needs web search, deeper reasoning, or real work, it quietly delegates to OpenAI's frontier model, currently GPT-5.5, keeps the conversation going while that happens, and folds the answer back in when it's ready. That is a meaningful shift from "voice mode" as a separate, dumber product bolted onto the real model, to voice as just another interface into the same intelligence everyone is already using in text. Independent testing, including a widely read hands-on from developer Simon Willison, backs up that this is a real jump, not a marketing refresh, though he also flagged an early bug where the model would interrupt to laugh at things that weren't jokes, which OpenAI patched down after he reported it.
There's also a provenance angle worth noting. As of late July, audio generated through GPT-Live carries SynthID watermarking, with a public tool available to verify whether a given clip came from OpenAI's system. That matters more than it sounds like it should. As AI voice tools get good enough to actually replace a phone call or a brainstorming partner, being able to tell "was this a real person or a model" becomes a basic trust question, not a nerdy technical footnote.
Full-duplex is the real concept here. Old voice mode listened, stopped, thought, then spoke, one turn at a time. Full-duplex means the model is doing all of that at once, deciding dozens of times a second whether to talk, stay quiet, interrupt, or quietly hand a hard question to a bigger model mid-sentence without breaking the conversation. That's the shift that makes "voice assistant" start meaning something closer to a person on the other end of a call instead of a slow chatbot you talk to. Try it out.