All articlesFrontier

Google launched Gemini 3.8 Live at $1.38/hour — here's what the price cut means for voice agent economics

Google released Gemini 3.8 Live and 3.8 Live Extended Thinking yesterday, undercutting OpenAI's GPT-Live-1 by ~70% on per-hour pricing. The models top Artificial Analysis's speech-to-speech leaderboard but lack full duplex.

Sep 16, 2026 3 min read
geminispeech-to-speechvoice-agentspricing

Google DeepMind released Gemini 3.8 Live and 3.8 Live Extended Thinking yesterday — two new speech-to-speech models that compete directly with OpenAI's GPT-Live family. The headline number: $1.38 per hour of voice conversation, roughly 70% cheaper than GPT-Live-1's reported $4.50/hour rate.

Both models topped Artificial Analysis's speech-to-speech leaderboard on launch day. Gemini 3.8 Live handles standard voice interactions; 3.8 Live Extended Thinking adds a longer reasoning window for complex queries. Simon Willison pointed GPT-6 Astra Extra High at the documentation and shipped a working demo in one session.

The pricing structure is different from OpenAI's. Google charges per audio hour; OpenAI charges per audio token with separate input/output rates. For a typical 20-minute customer service call, Gemini 3.8 Live costs $0.46. GPT-Live-1 costs approximately $1.50 for the same duration based on current published token rates. The difference compounds at scale: a contact center running 10,000 calls per month pays $4,600 on Gemini vs. $15,000 on GPT-Live-1.

The tradeoff is full duplex. GPT-Live-1 supports true simultaneous audio in both directions — the agent can interrupt itself mid-sentence when the caller speaks. Gemini 3.8 Live uses turn-based audio, closer to the Retell or Bland pattern where the agent waits for silence before responding. That latency delta matters for interruption-heavy use cases like appointment scheduling or troubleshooting calls.

What this changes for production deployments

We've deployed voice agents on Retell, ElevenLabs Conversational AI, and Vapi. Cost per call has been a blocker for mid-volume deployments — anything between 1,000 and 10,000 calls per month where per-minute pricing stacks up but commitment-tier discounts don't apply yet. Gemini 3.8 Live at $1.38/hour opens that mid-tier.

The Extended Thinking variant is interesting for complex phone workflows. A prior auth call that requires pulling patient history, checking formulary rules, and citing policy sections could justify the longer reasoning window. We haven't tested it yet, but the pattern fits.

Full duplex is still the better experience for high-churn interactions. A caller who interrupts three times in 90 seconds will notice the lag on turn-based audio. But most inbound service calls don't interrupt — they follow a script, ask clarifying questions, wait for the agent to finish. For those, Gemini 3.8 Live's price advantage is real.

The leaderboard caveat

Artificial Analysis ranked Gemini 3.8 Live first on launch day, but the leaderboard measures latency, accuracy, and naturalness on a fixed eval set. Production performance depends on accent coverage, background noise handling, and how the model behaves under real interruption patterns. OpenAI's GPT-Live-1 has been in the wild for six months; Gemini 3.8 Live shipped yesterday.

Google's documentation includes a transcription accuracy benchmark across 12 languages. Gemini 3.8 Live hit 94.2% word error rate (WER) on English, 91.8% on Spanish, 89.3% on Mandarin. GPT-Live-1's published WER is 93.7% English, 90.1% Spanish, 88.9% Mandarin. Close enough that real-world differences will come down to deployment tuning.

What we're testing next

We're spinning up a parallel branch of Goldie (Best Limo NY's inbound booking agent) on Gemini 3.8 Live this week. The current production version runs on ElevenLabs Conversational AI; we'll run 200 test calls through the Gemini variant and compare transcription accuracy, interruption handling, and cost per successful booking.

If Gemini 3.8 Live handles the interruption patterns without noticeable lag, the cost delta is enough to justify a migration for mid-volume clients. If the turn-based lag breaks the experience, it stays on the ElevenLabs stack.

Google shipping a $1.38/hour speech-to-speech model two weeks after OpenAI's latest GPT-Live release means the voice agent cost curve is compressing faster than expected. That's good for production deployments. The question is whether the turn-based architecture holds up under real caller behavior.

/ 06 — Start hereOne business day response

Tell us what you'd like built.

Send us a paragraph about the workflow, phone line, or tool you want built. We'll reply within one business day with a one-page plan, a fixed price, and a delivery date you can put on a calendar.

  • 30-min scoping call, free
  • Written proposal within 48 hours
  • Fixed price before we start
  • Most builds delivered in 2–8 weeks