Gemini 3.8 Live: Building Real-Time Voice AI Agents in 2026

Gemini 3.8 Live is the streaming model built for voice and video agents: about $0.005 per minute of audio in and $0.018 out, no caching, no structured outputs. Pricing, limits and a workable architecture.

Published 23 September 2026 by Jake Hissitt, Founder of Stob.AI.

Most model comparisons treat every model as a request-response API.

Gemini 3.8 Live is not that.

It is a streaming model built for continuous voice and video conversation, stable since 15 September 2026, and it is the only model in our comparison tool designed for that shape of work.

If you are building a phone agent, a voice concierge or a screen-sharing assistant, this is the relevant page.

Pricing, and why it is different Modality Price Practical rate Text input $0.75 per 1M tokens — Text output $4.50 per 1M tokens — Audio input $3.00 per 1M tokens about $0.005 per minute Audio output $12.00 per 1M tokens about $0.018 per minute Image / video input $1.00 per 1M tokens — The per-minute figures are the ones to budget with.

A ten-minute call with roughly balanced speaking time costs in the region of 12 cents in model spend.

Telephony, transcription fallbacks and your own infrastructure sit on top.

There is no prompt caching on Live.

A long system prompt is paid for on every session, so keep instructions tight and push reference material into tools rather than the prompt.

What it supports, and what it does not Supported: async function calling, search grounding, 131,072 token context, up to 65,536 tokens of output, streaming audio and video input.

Not supported: code execution, structured outputs, prompt caching.