夜雨聆风学习资料网

ARTICLE · 1024715

语音AI不再等回答,工具调用挪到后台|Voice AI That Acts While Talking

语音AI不再等回答,工具调用挪到后台|Voice AI That Acts While Talking

科技趋势 / TECHNOLOGY / 2026.09.17

把语音拆成快对话与慢推理两条线,计费单位变成通话时长。|Voice models split into fast talk and slow thinking; billing is per minute.

这次发布了什么 / What Shipped

Google 在九月十五日发布 Gemini 3.8 Live 与 Gemini 3.8 Live Extended Thinking 两个实时语音模型。前者的定位是低成本、可规模化:它近乎实时地处理画面输入,在对话中途自动切换九十七种语言,并把工具与接口调用放到后台执行;后者在说话的同时进行多步推理,用「我来查一下」这类口头衔接填住等待的空白。本文以北京时间九月十七日八点二十分为信息截点。

On September 15, Google released Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, two real-time voice models. The first is positioned for scale and cost efficiency: it processes visual input in near real time, switches between 97 languages mid-conversation, and runs tool and API calls in the background. The second reasons through multi-step tasks while it speaks, using spoken bridges such as "let me check that" to cover the pause. This article uses 08:20 Beijing time on September 17 as its information cut-off.

与旧做法相比,差异不在回答质量,而在架构。过去的语音助手通常把一句话拆成三步:先做语音识别转成文字,再交给语言模型生成回答,最后用语音合成读出来。三步串联会累积延迟,转写还容易丢掉语气和数字精度。原生语音到语音把这三步压进一个模型,并把耗时的工具调用挪到后台,让对话不必停下来等结果。

The difference from the old approach is architectural rather than a matter of answer quality. Earlier voice assistants typically split one utterance into three steps: speech recognition into text, a language model to draft a reply, and speech synthesis to read it back. Chained steps accumulate latency, and transcription tends to lose tone and numeric precision. Native speech-to-speech folds those steps into a single model and pushes the slow tool call into the background so the conversation does not have to stop and wait.

为什么拆成两个模型 / Why Two Models

拆成两个型号,是成本与能力之间的取舍。快对话模型按通话时长算更划算,适合客服热线、车载助手这类高频、低复杂度的场景;带思考的型号适合要多步操作的任务,例如改签机票、核对订单号、跨系统调数据。Google 在发布稿里没有公布具体单价,只强调成本效率,并给出一条可对照的成绩:带思考的型号在 Artificial Analysis 的语音到语音质量指数上得 82.6 分,位列第一。

Splitting the launch into two models is a trade-off between cost and capability. The fast model is the cheaper option per minute of conversation, which suits high-volume, low-complexity settings such as support lines and in-car assistants. The thinking model fits tasks that need several steps: rebooking a flight, reading back an order number, pulling data from another system. Google's announcement does not disclose per-minute or per-token pricing; it only stresses cost efficiency and offers one comparison point, with the thinking model scoring 82.6 on Artificial Analysis' Speech-to-Speech quality index, ranking first overall.

投资者更该留意的是计费单位。语音 Agent 按通话时长而不是按 token 计价,单位成本就直接与时长相乘,而最贵的一项往往不是模型,是「没有人说话的那段时间」。因此第一波受益方大概率是提供实时音视频管道的平台,发布稿中点名的 Agora、LiveKit、Pipecat、Vercel 都属于这一类:它们卖的是传输与并发,不与模型厂商争模型能力。

What investors should notice is the billing unit. Voice agents are priced by talk time rather than by token, so unit cost multiplies directly by duration — and the most expensive line item is usually the dead air, not the model. That makes the first wave of beneficiaries the platforms that supply real-time audio and video plumbing, among them Agora, LiveKit, Pipecat and Vercel, all named in the announcement. They sell transport and concurrency rather than competing with model makers on model quality.

谁的账会变 / Whose Economics Change

沿着产业链往下看:如果原生语音模型的延迟与成本继续下降,最先承压的是外包客服和简单的语音导航,这两类服务目前按坐席或按分钟收费;受益方是能把通话直接变成成交、或者变成可核算成本节省的企业。对模型厂商自己来说,用量扩张会先于收入出现,真正的分水岭是客户愿不愿意把语音放进收费环节,而不只是放进咨询环节。

Follow the chain one step down. If latency and cost keep falling, the first pressure point is outsourced call centres and simple voice menus, both of which currently charge by seat or by minute. The beneficiaries are companies that turn calls directly into revenue, or into cost savings they can audit. For model vendors themselves, volume growth will show up before revenue does. The real dividing line is whether customers put voice into a paid workflow rather than only a question-and-answer workflow.

一个可执行的动作:不要用「发布了更强的语音模型」当作判断依据。改看两个数字,一是企业侧语音 Agent 的实际通话时长与自动化完成率,二是同样这些时长里人工介入的次数。通话时长上升、人工介入下降,说明语音已经进入收费环节;两者同向上升,说明成本只是换了个位置。

One practical next step: do not judge this release by "a stronger voice model shipped". Watch two numbers instead. The first is actual talk time and automation completion rate on the enterprise side; the second is how often a human has to step in during those same minutes. Rising talk time with falling human intervention means voice has entered a paid workflow. Both rising together means the cost simply moved.

Risk notice / 风险提示风险提示:发布信息来自 Google 官方博客,官方未公布单价与规模数据;能力评测由厂商引用第三方指数,实际表现仍需独立验证。语音生成内容的合规与隐私要求因辖区而异。本文不构成投资建议。 / Risk: Details come from Google's official blog, which discloses no pricing or volume data; benchmark scores are vendor-cited third-party indices and still need independent validation. Compliance and privacy rules for synthetic speech vary by jurisdiction. Not investment advice.

资料来源 / Sources

  1. Google 官方博客:Introducing Gemini 3.8 Live and 3.8 Live Extended Thinking(2026-09-15,官方一手来源) / Google official blog - https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-8-live-gemini-3-8-live-extended-thinking/
  2. Unite.AI:Google Launches Gemini 3.8 Live and Extended Thinking Voice Models(2026-09-15,独立来源) / Independent coverage - https://www.unite.ai/google-launches-gemini-3-8-live-and-extended-thinking-voice-models
  3. ExplainX:Gemini 3.8 Live 与 Extended Thinking 能力与基准梳理(2026-09,独立来源) / Independent analysis - https://www.explainx.ai/blog/google-gemini-3-8-live-extended-thinking-2026

关注EthanFang,一起看清模型发布背后的成本账。 / Follow EthanFang to read the cost math behind each model launch.

相关学习资料

返回首页浏览学习资料