After reading this, the reader knows Google launched two Gemini 3.8 voice models, and it matters because voice agents can now reason while speaking.

Google launches Gemini 3.8 Live voice models

Gemini 3.8 Live targets fast conversations at scale. Gemini 3.8 Live Extended Thinking adds deeper reasoning for multi-step voice tasks.

Google released Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking on September 15, bringing two new real-time audio models to its developer tools and consumer products.

In plain terms, Google made one model for quick voice interaction and another for harder work. Both can listen, see visual input, speak and use software tools. The launch matters because developers can build an agent that keeps a conversation moving while an API call or a longer reasoning step runs in the background.

Why it matters

Many voice systems still behave like telephone menus with better speech recognition. They wait for a complete request, process it and then answer. Google’s new models are designed for a continuous exchange: a user can interrupt, change direction or keep talking while the system performs another task.

That shifts the design problem for voice agents. The interface no longer needs to hide every tool call behind silence or filler. A model can acknowledge a request, continue the conversation and report progress while a booking, search or business-system action completes. Google demonstrated the extended model coordinating bookings and asynchronous function calls, although a vendor demonstration does not establish reliability in an independent production setting.

Gemini 3.8 Live is the lower-cost option intended for fast dialogue and visual grounding. Gemini 3.8 Live Extended Thinking is intended for tasks that require more reasoning. Google says the extended model can reason and speak at the same time, including giving short acknowledgements and progress updates during multi-step work.

Both models accept continuous audio, images, video and text. Google’s model card lists a context window of up to 128,000 tokens and audio or text output of up to 64,000 tokens. The company also says the regular Live model can switch automatically among 97 supported languages during a conversation.

The benchmark lead is real but narrow

Artificial Analysis, an independent model-testing company, gives Gemini 3.8 Live Extended Thinking an aggregate score of 82.6 on its Speech-to-Speech Index. The index gives equal weight to speech reasoning, agentic task completion, human preference and task success. Among models with complete data on that index, the extended Gemini model ranks first in the table published at launch.

Its strongest result is agentic work. The model completed 68.6% of the replica customer-service scenarios in the testing company’s τ-Voice evaluation. Each scenario has one valid final database state, so success measures whether the agent completed the task rather than whether its speech merely sounded convincing.

The result does not mean Gemini leads every voice measure. Artificial Analysis reports that Gemini 3.8 Live Extended Thinking scored 97.7% on its 1,000-question Big Bench Audio reasoning test, while StepAudio 3 Realtime scored 99.7%. The regular Gemini 3.8 Live model also received a higher human-preference rating than the extended model in the published table. Buyers therefore need to match the model to the job instead of treating one aggregate rank as a universal verdict.

Google’s own model card supplies another limit. It says both models may hallucinate, may respond slowly or time out, and have a January 2025 knowledge cutoff. Tool access or search can add current information, but that does not remove the need to test factual accuracy, recovery from interrupted calls and failure handling.

Availability comes in different stages

Developers can access both models through the Gemini API and Google AI Studio. Gemini 3.8 Live is also rolling out in Search Live, while enterprise access begins in private preview. The extended model is rolling out through Gemini Live and selected Google Workspace products, with access varying by subscription and product.

The Gemini Live API uses a stateful WebSocket connection for streaming audio, images and text. Google recommends ephemeral tokens instead of standard API keys when a client connects directly to the service in production. That detail matters because a low-latency design often moves the connection toward the user’s device, where a long-lived credential would be easier to expose.

Google says audio generated by its AI products carries its imperceptible SynthID watermark. The measure can help identify generated audio, but it addresses provenance rather than whether an agent’s action was correct or authorized.

The next evidence should come from deployments outside launch demonstrations: interrupted conversations, noisy audio, long tool calls and recovery after a failed transaction. The release makes simultaneous conversation and task execution available to developers now; production reliability remains the test that will determine whether it changes customer service and workplace software.

Verification

  1. VERIFIED — Google launched Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking on September 15, 2026. Primary source: https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-8-live-gemini-3-8-live-extended-thinking/
  2. VERIFIED — The regular model targets scale and cost efficiency; the extended model targets complex, multi-step reasoning. Primary source: https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-8-live-gemini-3-8-live-extended-thinking/
  3. VERIFIED — The models accept continuous audio, images, video and text; the model card lists a 128K-token context window and 64K-token output. Primary source: https://deepmind.google/models/model-cards/gemini-3-8-audio/
  4. VERIFIED — Google says the regular model switches among 97 languages and can continue dialogue during background tool calls; it demonstrated bookings and asynchronous function calls. Primary source: https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-8-live-gemini-3-8-live-extended-thinking/
  5. VERIFIED — Artificial Analysis reports an 82.6 aggregate index score and 68.6% τ-Voice task completion for the extended model; the benchmark uses valid database end states. Primary source: https://artificialanalysis.ai/speech-to-speech?api-benchmarks=agentic-performance-vs-cost-to-run
  6. VERIFIED — Big Bench Audio contains 1,000 questions; the extended model scored 97.7%, below StepAudio 3 Realtime at 99.7%, while the regular Live model had the higher preference rating in the table. Primary source: https://artificialanalysis.ai/speech-to-speech?api-benchmarks=agentic-performance-vs-cost-to-run
  7. VERIFIED — Google lists hallucinations, occasional slowness or timeouts, and a January 2025 knowledge cutoff as known limitations. Primary source: https://deepmind.google/models/model-cards/gemini-3-8-audio/
  8. VERIFIED — Access is rolling out through the Gemini API, AI Studio, Search Live, Gemini Live, Workspace and enterprise previews, with availability differing by product. Primary sources: https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-8-live-gemini-3-8-live-extended-thinking/ and https://deepmind.google/models/model-cards/gemini-3-8-audio/
  9. VERIFIED — The Live API uses stateful WebSockets, and Google recommends ephemeral tokens for direct client-to-server production connections. Primary source: https://ai.google.dev/gemini-api/docs/live-api
  10. VERIFIED — Google says audio from its AI products is watermarked with SynthID. Primary source: https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-8-live-gemini-3-8-live-extended-thinking/
  11. PARTIALLY VERIFIED — Simultaneous dialogue and tool execution can reduce silent waiting and broaden voice-agent interface designs. The capabilities are documented, but the operational consequence is an inference that requires production evidence. Primary source: https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-8-live-gemini-3-8-live-extended-thinking/

Glossary candidates

  • Full-duplex voice model
  • Visual grounding
  • Function calling
  • Stateful WebSocket
  • Ephemeral token
  • Speech-to-Speech Index
  • SynthID

Cold-reader sentence: Google released two real-time Gemini voice models that can keep talking while reasoning or using tools, but independent tests and Google’s model card show clear limits.