自然的语音 AI 是系统工程:Alex Smola 谈延迟、token 与可负担性
导语
语音界面一直在进步,也一直在破功——回复前多停的那半拍、被打断时的笨拙、不合时宜的语气。当 Sam Charrington 问 Boson AI 联合创始人兼 CEO Alex Smola 到底怎样才能补上这块短板时,贯穿整场对话的答案不是"更大的模型",而是一整套系统:一份从生物学借来的延迟预算、一门 token 经济学、一条以人类寿命为单位计量的数据流水线,以及一类全新的"智能体该如何在对话中举止"的基准测试。
错觉问题
Smola 把业界争相发布的产品降了一级: "语音只是一块中间垫脚石。" Alex Smola 00:48 “voice is an intermediate stepping stone” 原声音频直达锚点 播放原声 00:48 他说,这一切收敛的终点,是"你将和一个看起来、感觉起来都像人的视听智能体对话"——而他把那个终点放在离"自然"非常远的地方。 Alex Smola 01:00 “you'll be talking to an AV agent that looks and feels like a human” 原声音频直达锚点 播放原声 01:00 Charrington 则补充了用户侧的挫败感:ChatGPT 的高级语音模式一直在变好,却依然令人恼火,因为 "它真的只在完美条件下可用:安静的房间、没有背景噪音、打断时非常小心。" Sam Charrington 02:15 “it really only works in perfect conditions, meaning a silent room, no background noise, being very conscientious about interrupting” 原声音频直达锚点 播放原声 02:15
济州岛的一间嘈杂酒吧
Smola 的反例来自韩国济州岛 KDD 会议期间的一间酒吧,他一时兴起在那里演示了自家系统。 他说,模型"相当好地接住了有人切换到斯洛文尼亚语、另一个人又换到印地语,而酒吧非常吵"的场面。 Alex Smola 03:10 “our model was able to handle you know people switching to Slovenian and then somebody else to Hindi fairly well even though the bar was very noisy” 原声音频直达锚点 播放原声 03:10 他随即将功劳分了出去:"不小的一份功劳要记在 iOS 团队头上,他们在手机上做了非常好的降噪和音频分离," Alex Smola 03:23 “a non-trivial amount of the credit definitely goes to the iOS team doing really good noise cancellation and audio separation on the mobile phone” 原声音频直达锚点 播放原声 03:23 而单支麦克风大概永远不够——麦克风阵列才是已经出货的成熟解法。他的时间表很具体: "我想大概一年之后,音频就会基本做到万无一失," Alex Smola 05:14 “I think probably in a year audio will become pretty much bulletproof” 原声音频直达锚点 播放原声 05:14 虚拟人大约一年半后登场,带动画面孔的机器人更慢一些,因为硬件难做。
生物学定的时钟
工程的紧迫感来自一个生物学数字。 Smola 把预算定在"约 150 毫秒"——从"光子打到视网膜,到皮层开始对它做处理"的时间,这个数字稳定到可以当诊断指标用;从耳朵进来还要更短一点。 Alex Smola 06:39 “it takes about 150 milliseconds for you know basically a photon hitting your retina to your cortex actually doing something with it” 原声音频直达锚点 播放原声 06:39 他说人类在视听感知上大约以六到十赫兹的节奏运作,Boson 于是把模型调到同一个窗口: 中断处理在"大约 150 毫秒上下"完成。 Alex Smola 07:50 “it does interruption handling with about within about 150ish milliseconds or so” 原声音频直达锚点 播放原声 07:50
token 速率的两难
更深一层的约束是算术。在当前范式里,音频变成 token,token 过一遍 LLM 式骨干,再变回音频——而 "每秒 token 一多,模型就不能有太多参数", Alex Smola 09:23 “many tokens per second means your model cannot have too many parameters” 原声音频直达锚点 播放原声 09:23 token 速率慢下来,才腾得出预算装更大的模型。 文本在这点上占尽便宜:人类想要的语速合每秒"三到五个 token",而音频动辄十几个——大约每 100 毫秒就要吐一个 token。 Alex Smola 09:38 “it's about you know three to five tokens per second that humans want for audio you can easily have you know 10 plus” 原声音频直达锚点 播放原声 09:38 虚拟人视频则自带一层补偿: "背景里不会开过赛车", Alex Smola 11:53 “there are not going to be race cars driving in the background” 原声音频直达锚点 播放原声 11:53 所以虚拟人的画面大部分时间无聊、可高效压缩。
先定价,再定模型
Boson 从账单出发做设计。 流式传输一场对话,不该占用"一整台 Blackwell 服务器 GPU"——那样当演示很惊艳,当服务没人付得起。 Alex Smola 13:16 “if you need to use let's say you know a full Blackwell server GPU just for a single conversation then that may not be the most economically viable model” 原声音频直达锚点 播放原声 13:16 Smola 说,他们的做法是"先定价格,再倒回去", Alex Smola 13:45 “we went in with a price first and then work backwards” 原声音频直达锚点 播放原声 13:45 造客户真正付得起的东西。他给这套设计配的性能主张相当强—— 在基准测试上,"我们比——比如说——GPT、Gemini、还有 Grok 都要好,而成本只是几分之一。所以我们比 OpenAI 的模型好,成本是十分之一" Alex Smola 36:39 “we are better than let's say GPT and Gemini and and and Grok at a fraction of the cost. So we're better than OpenAI's models at one-tenth the cost” 原声音频直达锚点 播放原声 36:39 ——而且脚注是他自己补上的: "这些是对方关闭思考模式下的成绩," Alex Smola 37:19 “this is with thinking turned off in these models” 原声音频直达锚点 播放原声 37:19 因为每次回答前先停一大截去"思考",在对话里感觉并不自然。
给 LLM 教音频,别把老本教丢了
加上一个新模态,可能顺手抹掉它本来依托的智能。Smola 的警世故事是朋友的养女:只对她讲德语,几个月之内就 "把西班牙语的每一个词都忘光了" Alex Smola 16:05 “the girl had forgotten every single word of Spanish” 原声音频直达锚点 播放原声 16:05 ——他说,只被按在音频上的语言模型,同样会丢掉推理和语言能力。解法是围绕音频工作维持完整的训练流水线,它产出了 他称之为相当不错的 TTS 模型 Higgs Audio V2(去年发布),今年又加快节奏接连发布 TTS、ASR 和音频理解模型。 Alex Smola 17:12 “a fairly competent uh TTS model that was Higgs Audio V2 last year” 原声音频直达锚点 播放原声 17:12 Higgs Audio V2 这次发布在 Boson AI 官网上以 Higgs TTS 2 之名留有记录。 Boson AI 背景报道 Higgs TTS 2 独立媒体核实报道 阅读原文报道 在他看来这些类目必须分开: 语音识别是音频进、文本出,分辨不出"recognize speech"到底是"识别语音"还是"毁掉一片好海滩"(wreck a nice beach), Alex Smola 18:03 “whether to recognize speech means to recognize speech or whether it means to wreck a nice beach” 原声音频直达锚点 播放原声 18:03 音频理解模型才能真正对音频做推理。这也不是轻量微调: 一个已发布的模型"建立在 Llama 之上",并在模型卡里注明, Alex Smola 30:12 “one model that we released last year was built on top of Llama and we acknowledge them appropriately in our model card” 原声音频直达锚点 播放原声 30:12 但免费原料的逻辑依然成立——"如果有人送你免费的钢材,你不会去建钢厂,你会去造车" Alex Smola 31:17 “if somebody gives you free steel, you don't build a steel mill; you build a car” 原声音频直达锚点 播放原声 31:17 —— 而他估算,音频侧的工作量约为语言模型的一半或三分之一,并强调"不是十分之一"。 Alex Smola 31:53 “Maybe it's half or 1/3 uh but it's not one-tenth” 原声音频直达锚点 播放原声 31:53
一亿小时的飞轮
Smola 最先点名的差异化因素是数据: "我们持有的音频在一亿小时这个量级",他折算成"约等于 200 个人类一生。" Alex Smola 20:43 “we have in the order of 100 million hours of audio. Um and it's about 200 human lifetimes” 原声音频直达锚点 播放原声 20:43 靠的不是买标注,而是抓取之后按规模做提取、打标、归一化和转写。 Boson AI 自己的工程文章收窄了这个数字:文中表述为处理 1 亿+原始小时、得到超过 1000 万模型可用小时——因此访谈里的头条数字计的是原始音频,不是可直接训练的数据。 Boson AI 补充前提 Higgs Realtime 独立媒体核实报道 阅读原文报道 自建数据中心是护城河的一部分—— 放到 NeoCloud 上,"存储账单会把你吃穷" Alex Smola 21:30 “If you were to store this data on a NeoCloud, um the storage bill would eat you alive” 原声音频直达锚点 播放原声 21:30 ——而标注厂商的量级差了四个数量级: 有公司报价比一万小时标注,Smola 回答说 Boson 的规模"是它的一万倍"。 Alex Smola 22:48 “We are at between 10 we're at 10,000 times that scale” 原声音频直达锚点 播放原声 22:48 让旁人害怕的带噪标签,他们用经典统计学消化—— 他指出,中世纪的脚尺法字面上就是"脚尺和截尾均值估计量的出处" Alex Smola 25:28 “that's literally where the foot rule and trimmed mean estimators come from” 原声音频直达锚点 播放原声 25:28 —— 而闭环会自己转起来:素材够多,你就能"造出更好的模型,然后飞轮就转起来了", Alex Smola 27:10 “having a lot of this stuff allows you to build better models and then you get that flywheel” 原声音频直达锚点 播放原声 27:10 其中还包括利用"一期播客只有两个人"这类上下文做语音分离的巧办法。
前台交谈,后台推理
架构跟着经济学走。Boson 没有做单一的巨型端到端模型——快,但用他的话说"笨"——而是跑 "一种更两段式的架构:理解与推理在后台进行",前台由一个更小的模型撑住对话。 Alex Smola 34:27 “a little bit more of a two-stage architecture where the understanding and reasoning happens in one end with appropriate tool calls being fired off in the back end” 原声音频直达锚点 播放原声 34:27 开发者拿到的界面很熟悉:提示词写得和指挥普通 LLM"感觉一模一样","只不过我们的模型还会说话"。 Alex Smola 35:17 “the prompts that you write for our model are look look and feel the same as if you were just you know instructing a a regular LLM just that our model also talks” 原声音频直达锚点 播放原声 35:17 在 Higgs Live 演示里,模型"在后台做网页搜索",视结果返回快慢,要么直接回答,要么说一句"嘿,我去找找"。 Alex Smola 38:31 “This performs web search in the background and depending on how quickly gets the result back, it will just answer or it will actually tell you, hey, let me look for that” 原声音频直达锚点 播放原声 38:31 现成 MCP 服务器能用,但有上限: "能同时启用的服务器数量"目前"还比较有限", Alex Smola 39:35 “the quantity of servers at enabled at the same time right now is a little bit limited” 原声音频直达锚点 播放原声 39:35 因为一次大规模预填充会留下"一个大得你必须拖着走的 KV 缓存",让每个生成的 token 都更贵。 Alex Smola 40:08 “you have a large KV cache that you need to lug around with you that of course you know makes the token generation more expensive” 原声音频直达锚点 播放原声 40:08 慢工具只要提前打招呼就能活——"人类对延迟是非常宽容的",只要"被告知有延迟"。 Alex Smola 43:38 “humans are very tolerant to delays is if they are told that there's a delay” 原声音频直达锚点 播放原声 43:38 他的例证是苹果的开机进度条,收尾顺滑得可疑:在他的讲法里,说谎的是苹果,"微软的开机进度条才是实话"——一条按上次开机时长掐好时间的假进度条,妙就妙在没解决那个不可能的技术难题,却解决了体验。 Alex Smola 44:37 “No actually they lie to you. uh the Microsoft boot screen is the truth” 原声音频直达锚点 播放原声 44:37
为"举止"立基准
量什么,本身就是个研究问题。 Smola 主张,可打断性"必须随场景而定":日语听众那声短短的附和意思是"我收到了",不该让智能体停下;而"我没听懂"必须让它停。 Alex Smola 46:05 “know in Japanese it's very common to say hey. So basically you're back channeling the other person and saying hey I got it. Yeah. Oh that right. >> But if the agent stops for every one of those then that's going to be infuriating. >> Exactly. uh on the other hand uh if that if that person were to say well I don't understand you want the model to stop right so what that means is you you need to make sure that the interruptibility is really scene dependent so for” 原声音频直达锚点 播放原声 46:05 主动性基准已经公开,收录于 arXiv 论文 "ProactBench: Beyond What The User Asked For"。 arXiv 独立证实 ProactBench: Beyond What The User Asked For 独立媒体核实报道 阅读原文报道 它的同伴 IHBench 也有 arXiv 论文 "Evaluating Post-Interruption Recovery in Voice Agents with Structured Workflows",覆盖可打断性、响应速度、声音是否与内容相配等指标。 arXiv 独立证实 IHBench: Evaluating Post-Interruption Recovery in Voice Agents with Structured Workflows 独立媒体核实报道 阅读原文报道 优化目标则和代码生成不同: "在我们的场景里,我们操心的是人类的幸福感。" Alex Smola 47:46 “In our case, we worry about human happiness” 原声音频直达锚点 播放原声 47:46 任务完成度仍然算数,但"用得开心"是硬要求,不是点缀。他心中失败的样子,是《银河系漫游指南》BBC 剧集里那台欢快的飞船电脑,宣布 "两分钟后坠毁。我们都会死。" Alex Smola 49:18 “You will crash in 2 minutes. We will all die.” 原声音频直达锚点 播放原声 49:18 这个故事的寓意在于:"智商和情商是两回事",而语音两者都要。 Alex Smola 49:54 “there is a difference between IQ and EQ” 原声音频直达锚点 播放原声 49:54
学会讨人喜欢
社交行为的训练数据是个陷阱。爱情喜剧把尾随拍成浪漫,动作片以握手言和收场,而 Smola 说,"我衷心希望没有人拿《菲尔博士》去训练模型"还当作正常人类行为。 Alex Smola 52:09 “I sincerely hope that nobody will train a model on Dr. Phil and assume that this is normal human behavior” 原声音频直达锚点 播放原声 52:09 便宜又合乎伦理的练习来自模拟器—— "英伟达在发布一批数字人物和场景这件事上做得很棒", Alex Smola 55:14 “Nvidia did a great job at releasing some digital personas and some scenarios” 原声音频直达锚点 播放原声 55:14 Boson 用它们训练模型应对不配合的、粗声粗气的、存心搞坏系统的用户。另一半是记忆: 一份个人的交互历史,相当于"一个加强版 CRM,只不过现在人人都有一个", Alex Smola 57:14 “think of this as a glorified CRM but now for everybody” 原声音频直达锚点 播放原声 57:14 再与跨全体交互的通用改进相平衡—— 甚至细到礼仪层面:"在印度用左手吃饭是绝对犯忌讳的。" Alex Smola 58:29 “eating with your left hand in India is seriously frowned upon” 原声音频直达锚点 播放原声 58:29 不可复制的是天才: "才华很好,但才华不可重复、不可自动化。" Alex Smola 01:00:02 “Brilliance is good, but brilliance is not repeatable and automatable” 原声音频直达锚点 播放原声 01:00:02 替代品是体量—— "你可以把每一次交互都看成又一个新实验、一个新的数据点" Alex Smola 01:02:05 “you can think of each interaction as another new experiment, a new data point” 原声音频直达锚点 播放原声 01:02:05 ——再加上可以直接教的技巧,从管理培训到 "千禧世代三明治":按"表扬、批评、再表扬"的顺序把话说出口。 Alex Smola 01:02:40 “the millennial sandwich, right, where you have praise, critique, and praise” 原声音频直达锚点 播放原声 01:02:40 他的预测给这条弧线收尾:未来大概两三年会是一场"激动人心的革命",而且会推进得很快。 Alex Smola 01:03:46 “really going to be quite the an exciting revolution for the next maybe two to three years” 原声音频直达锚点 播放原声 01:03:46
访谈精粹15 组核心对谈
主持人与特邀嘉宾原声对谈精选,支持精准时间戳跳转
语音 AI 能否在嘈杂环境中自然交流,而不只是像我使用 ChatGPT 语音模式时那样依赖理想条件?
我们在济州岛一家嘈杂的酒吧演示过系统,现场有人切换到斯洛文尼亚语,也有人换成印地语,模型都处理得相当不错。这说明语音系统不一定只能在安静房间里工作,但功劳不能全算在模型头上:手机上的降噪和音频分离做得很好,iOS 团队贡献了相当一部分效果。我不认为随便拿一支麦克风挥动一下也能得到同样的结果。尤其同时存在多个声源时,单支麦克风恐怕无法彻底解决问题,真正做好降噪需要麦克风阵列。这些硬件能力已经有产品在用,我觉得我们正在取得进展。
语音 AI 面临的挑战中,哪些主要是工程问题,哪些主要是研究问题?
研究和工程必须一起做。我们希望打断处理大致落在人类感知所涉及的 150 毫秒量级,但音频还要经过编码成 token、模型处理、再解码成声音的过程。提高 token 频率有助于保真度,却会同时推高预填充和生成成本:每秒 token 越多,能负担的模型规模就越受限;频率降低,才有余地增加参数。每秒 10 个 token 也意味着每个覆盖约 100 毫秒。生理感知、交互设计、表示学习和服务成本,就在这里形成同一个取舍。缓冲有利于流式传输,但又不能妨碍及时打断。视频还带来帧率差异,以及维持一小时画面一致性的要求,不能只会生成短片。虚拟人画面通常大部分背景变化不大,可以利用这一特点提高压缩和实时运行效率;快速动作则会拉高码率。目标是用户负担得起的交互,而不只是大模型带来的漂亮演示。
你们为语音模型选择不同的设计取舍,主要是出于可负担性的考虑吗?
我们优先考虑可负担性。模型做大可以更聪明,但持续进行的对话也得在经济上成立。如果一场对话就要独占整块 Blackwell 服务器 GPU,演示或许很好看,客户却可能付不起钱。因此,我们先定价格目标,再倒推在这一成本内能做什么,而不是先追求模型能力、之后才考虑费用。