Natural Voice AI Is A Systems Problem: Alex Smola On Latency, Tokens, And Affordability
Introduction
Voice interfaces keep improving and keep breaking the spell — a beat too long before replying, a clumsy interruption, a tone that misses the room. When Sam Charrington asked Boson AI co-founder and CEO Alex Smola what it actually takes to fix this, the answer that ran through the conversation was less a smarter model than a whole system: a latency budget borrowed from biology, an economics of tokens, a data pipeline measured in human lifetimes, and a new class of benchmarks for how an agent should behave in conversation.
The Illusion Problem
Smola inverts the usual roadmap by demoting the product everyone is racing to ship: "voice is an intermediate stepping stone." Alex Smola 00:48 “voice is an intermediate stepping stone” Direct Audio Anchor Listen from 00:48 Where this converges, he said, is that "you'll be talking to an AV agent that looks and feels like a human" — and he put that destination very far from natural. Alex Smola 01:00 “you'll be talking to an AV agent that looks and feels like a human” Direct Audio Anchor Listen from 01:00 Charrington supplied the user-side frustration with today's frontier: ChatGPT's advanced voice mode keeps getting better yet stays infuriating, because "it really only works in perfect conditions, meaning a silent room, no background noise, being very conscientious about interrupting." Sam Charrington 02:15 “it really only works in perfect conditions, meaning a silent room, no background noise, being very conscientious about interrupting” Direct Audio Anchor Listen from 02:15
A Noisy Bar In Jeju
Smola's counterexample came from drinks at a bar in Jeju, South Korea, during the KDD conference, where he demoed the system on a dare to himself. The model, he said, handled "people switching to Slovenian and then somebody else to Hindi fairly well even though the bar was very noisy." Alex Smola 03:10 “our model was able to handle you know people switching to Slovenian and then somebody else to Hindi fairly well even though the bar was very noisy” Direct Audio Anchor Listen from 03:10 He then shared the credit rather than claiming it: "a non-trivial amount of the credit definitely goes to the iOS team doing really good noise cancellation and audio separation on the mobile phone," Alex Smola 03:23 “a non-trivial amount of the credit definitely goes to the iOS team doing really good noise cancellation and audio separation on the mobile phone” Direct Audio Anchor Listen from 03:23 and single microphones will probably never suffice — arrays are the solved, shipping path. His timeline is specific: "I think probably in a year audio will become pretty much bulletproof," Alex Smola 05:14 “I think probably in a year audio will become pretty much bulletproof” Direct Audio Anchor Listen from 05:14 avatars arrive in roughly a year and a half, and animated robot faces follow more slowly because hardware is hard.
Biology Sets The Clock
The engineering urgency comes from a biological number. Smola put the budget at "about 150 milliseconds" — the time from "a photon hitting your retina to your cortex actually doing something with it," stable enough to double as a diagnostic, and slightly shorter from the ears. Alex Smola 06:39 “it takes about 150 milliseconds for you know basically a photon hitting your retina to your cortex actually doing something with it” Direct Audio Anchor Listen from 06:39 Humans, he said, operate at around six to ten hertz for audiovisual perception, and Boson tuned the model to the same window: interruption handling completes "within about 150ish milliseconds or so." Alex Smola 07:50 “it does interruption handling with about within about 150ish milliseconds or so” Direct Audio Anchor Listen from 07:50
The Token-Rate Dilemma
The deeper constraint is arithmetic. In the current paradigm, audio becomes tokens, tokens pass through an LLM-style backbone, and tokens become audio again — and "many tokens per second means your model cannot have too many parameters," Alex Smola 09:23 “many tokens per second means your model cannot have too many parameters” Direct Audio Anchor Listen from 09:23 while a slower token rate frees budget for a bigger model. The asymmetry favors text, which runs at "three to five tokens per second that humans want," while audio can easily demand ten-plus — roughly one token per 100 milliseconds. Alex Smola 09:38 “it's about you know three to five tokens per second that humans want for audio you can easily have you know 10 plus” Direct Audio Anchor Listen from 09:38 Avatar video carries a compensating mercy: "there are not going to be race cars driving in the background," Alex Smola 11:53 “there are not going to be race cars driving in the background” Direct Audio Anchor Listen from 11:53 so most of an avatar feed is boring, effectively compressible video.
Price First, Model Second
Boson starts from the invoice. Streaming one conversation should not require "a full Blackwell server GPU just for a single conversation," Alex Smola 13:16 “if you need to use let's say you know a full Blackwell server GPU just for a single conversation then that may not be the most economically viable model” Direct Audio Anchor Listen from 13:16 which would be gorgeous as a demo and unaffordable as a service. Instead, Smola said, "we went in with a price first and then work backwards" Alex Smola 13:45 “we went in with a price first and then work backwards” Direct Audio Anchor Listen from 13:45 to something customers can actually pay for. The performance claim he attaches to that design is strong — on benchmarks, "we are better than let's say GPT and Gemini and and and Grok at a fraction of the cost. So we're better than OpenAI's models at one-tenth the cost" Alex Smola 36:39 “we are better than let's say GPT and Gemini and and and Grok at a fraction of the cost. So we're better than OpenAI's models at one-tenth the cost” Direct Audio Anchor Listen from 36:39 — and he supplied its footnote himself: "this is with thinking turned off in these models," Alex Smola 37:19 “this is with thinking turned off in these models” Direct Audio Anchor Listen from 37:19 because a long thinking pause before each reply would feel unnatural in conversation.
Teaching Audio To An LLM Without Forgetting
Adding a modality can erase the intelligence it was added to. Smola's cautionary tale is a friend's adopted daughter who, spoken to only in German, within months had "forgotten every single word of Spanish" Alex Smola 16:05 “the girl had forgotten every single word of Spanish” Direct Audio Anchor Listen from 16:05 — and a language model pushed onto audio alone, he said, loses its reasoning and language the same way. The remedy is a full training pipeline around the audio work, which produced what he called a fairly competent TTS model, Higgs Audio V2, released last year, followed this year by an accelerated release cadence of TTS, ASR, and audio-understanding models. Alex Smola 17:12 “a fairly competent uh TTS model that was Higgs Audio V2 last year” Direct Audio Anchor Listen from 17:12 The Higgs Audio V2 release is documented on Boson AI's site as Higgs TTS 2. Boson AI Context Higgs TTS 2 Verified Media Report Read Source Article The categories matter to him: speech recognition is audio-in, text-out and cannot resolve whether "recognize speech" means understanding speech or "wreck a nice beach," Alex Smola 18:03 “whether to recognize speech means to recognize speech or whether it means to wreck a nice beach” Direct Audio Anchor Listen from 18:03 while an audio-understanding model reasons over the audio itself. Nor is the adaptation light fine-tuning: one released model "was built on top of Llama," acknowledged in its model card, Alex Smola 30:12 “one model that we released last year was built on top of Llama and we acknowledge them appropriately in our model card” Direct Audio Anchor Listen from 30:12 but the logic of free inputs still holds — "if somebody gives you free steel, you don't build a steel mill; you build a car" Alex Smola 31:17 “if somebody gives you free steel, you don't build a steel mill; you build a car” Direct Audio Anchor Listen from 31:17 — and the audio-side effort, he estimated, is maybe half or a third of the language model's, while insisting "it's not one-tenth." Alex Smola 31:53 “Maybe it's half or 1/3 uh but it's not one-tenth” Direct Audio Anchor Listen from 31:53
The 100-Million-Hour Flywheel
The differentiator Smola names first is data: "we have in the order of 100 million hours of audio," which he sized up as "about 200 human lifetimes." Alex Smola 20:43 “we have in the order of 100 million hours of audio. Um and it's about 200 human lifetimes” Direct Audio Anchor Listen from 20:43 gathered not by buying annotations but by scraping and then extracting, tagging, normalizing, and transcribing at scale. Boson AI's own engineering post narrows that figure: it describes 100M-plus raw hours processed into over 10 million model-ready hours, so the interview's headline number counts raw audio rather than training-ready data. Boson AI With Caveats Higgs Realtime Verified Media Report Read Source Article Owning a data center is part of the moat — on a NeoCloud, "the storage bill would eat you alive" Alex Smola 21:30 “If you were to store this data on a NeoCloud, um the storage bill would eat you alive” Direct Audio Anchor Listen from 21:30 — and annotation vendors miss the scale by four orders of magnitude: offered 10,000 hours of annotation, Smola replied that Boson is "at 10,000 times that scale." Alex Smola 22:48 “We are at between 10 we're at 10,000 times that scale” Direct Audio Anchor Listen from 22:48 The noisy labels that scare everyone else are handled with classical statistics — the medieval foot rule, he noted, is literally "where the foot rule and trimmed mean estimators come from" Alex Smola 25:28 “that's literally where the foot rule and trimmed mean estimators come from” Direct Audio Anchor Listen from 25:28 — and the loop closes on itself: enough of this material lets you "build better models and then you get that flywheel," Alex Smola 27:10 “having a lot of this stuff allows you to build better models and then you get that flywheel” Direct Audio Anchor Listen from 27:10 including context tricks like separating the two voices on a podcast.
Foreground Talk, Background Reasoning
The architecture follows the economics. Rather than one giant end-to-end model — responsive but, in his word, dumb — Boson runs "a little bit more of a two-stage architecture where the understanding and reasoning happens" in the background while a smaller front end holds the conversation. Alex Smola 34:27 “a little bit more of a two-stage architecture where the understanding and reasoning happens in one end with appropriate tool calls being fired off in the back end” Direct Audio Anchor Listen from 34:27 Developers get a familiar surface: prompts "feel the same as if you were just" instructing a regular LLM, "just that our model also talks." Alex Smola 35:17 “the prompts that you write for our model are look look and feel the same as if you were just you know instructing a a regular LLM just that our model also talks” Direct Audio Anchor Listen from 35:17 On the Higgs Live demo the model "performs web search in the background" and, depending on how fast results return, either answers or says "hey, let me look for that." Alex Smola 38:31 “This performs web search in the background and depending on how quickly gets the result back, it will just answer or it will actually tell you, hey, let me look for that” Direct Audio Anchor Listen from 38:31 Off-the-shelf MCP servers work, with limits: "the quantity of servers" enabled simultaneously "is a little bit limited" today, Alex Smola 39:35 “the quantity of servers at enabled at the same time right now is a little bit limited” Direct Audio Anchor Listen from 39:35 because a massive prefill leaves "a large KV cache that you need to lug around," making every generated token more expensive. Alex Smola 40:08 “you have a large KV cache that you need to lug around with you that of course you know makes the token generation more expensive” Direct Audio Anchor Listen from 40:08 Slow tools are survivable when announced — "humans are very tolerant to delays," so long as "they are told that there's a delay." Alex Smola 43:38 “humans are very tolerant to delays is if they are told that there's a delay” Direct Audio Anchor Listen from 43:38 His exhibit is the Apple boot bar, which finishes suspiciously smoothly: in his telling, it is Apple that lies, and "the Microsoft boot screen is the truth" — a fake progress bar timed to last boot's duration, brilliant exactly because it solves the experience without solving an impossible technical problem. Alex Smola 44:37 “No actually they lie to you. uh the Microsoft boot screen is the truth” Direct Audio Anchor Listen from 44:37
Benchmarks For Behavior
What to measure is itself a research question. Interruptibility, Smola argued, "is really scene dependent": a Japanese listener's short backchannel means "I got it" and must not stop the agent, while "I don't understand" must. Alex Smola 46:05 “know in Japanese it's very common to say hey. So basically you're back channeling the other person and saying hey I got it. Yeah. Oh that right. >> But if the agent stops for every one of those then that's going to be infuriating. >> Exactly. uh on the other hand uh if that if that person were to say well I don't understand you want the model to stop right so what that means is you you need to make sure that the interruptibility is really scene dependent so for” Direct Audio Anchor Listen from 46:05 The proactivity benchmark is public, documented in the arXiv paper "ProactBench: Beyond What The User Asked For." arXiv Corroborates ProactBench: Beyond What The User Asked For Verified Media Report Read Source Article Its companion, IHBench, is documented in the arXiv paper "Evaluating Post-Interruption Recovery in Voice Agents with Structured Workflows," covering the metrics for interruptibility, responsiveness, and whether the audio matches what is said. arXiv Corroborates IHBench: Evaluating Post-Interruption Recovery in Voice Agents with Structured Workflows Verified Media Report Read Source Article The optimization target differs from code generation's: "In our case, we worry about human happiness." Alex Smola 47:46 “In our case, we worry about human happiness” Direct Audio Anchor Listen from 47:46 Task completion still counts, but enjoyment is a requirement, not a garnish. His model of failure is the cheerfully doomed ship computer from the Hitchhiker's Guide BBC series, announcing "You will crash in 2 minutes. We will all die." Alex Smola 49:18 “You will crash in 2 minutes. We will all die.” Direct Audio Anchor Listen from 49:18 The point of the story: "there is a difference between IQ and EQ," and voice needs both. Alex Smola 49:54 “there is a difference between IQ and EQ” Direct Audio Anchor Listen from 49:54
Learning To Be Pleasant
Training data for social behavior is a trap. Rom-coms reward stalking and movies end brawls with friendship, and "I sincerely hope that nobody will train a model on Dr. Phil" and call it normal behavior, Smola said. Alex Smola 52:09 “I sincerely hope that nobody will train a model on Dr. Phil and assume that this is normal human behavior” Direct Audio Anchor Listen from 52:09 Cheap, ethical practice comes from simulators — "Nvidia did a great job at releasing some digital personas and some scenarios," Alex Smola 55:14 “Nvidia did a great job at releasing some digital personas and some scenarios” Direct Audio Anchor Listen from 55:14 which Boson uses to train against uncooperative, abrasive, and system-breaking users. The other half is memory: a personal history of interactions amounts to "a glorified CRM but now for everybody," Alex Smola 57:14 “think of this as a glorified CRM but now for everybody” Direct Audio Anchor Listen from 57:14 balanced against improvements that generalize across everyone — including etiquette like the fact that "eating with your left hand in India is seriously frowned upon." Alex Smola 58:29 “eating with your left hand in India is seriously frowned upon” Direct Audio Anchor Listen from 58:29 What will not scale is individual genius: "Brilliance is good, but brilliance is not repeatable and automatable." Alex Smola 01:00:02 “Brilliance is good, but brilliance is not repeatable and automatable” Direct Audio Anchor Listen from 01:00:02 The substitute is volume — "you can think of each interaction as another new experiment, a new data point" Alex Smola 01:02:05 “you can think of each interaction as another new experiment, a new data point” Direct Audio Anchor Listen from 01:02:05 — plus technique that can simply be taught, from management training to the "millennial sandwich," where you deliver critique as "praise, critique, and praise." Alex Smola 01:02:40 “the millennial sandwich, right, where you have praise, critique, and praise” Direct Audio Anchor Listen from 01:02:40 His forecast closes the arc: an "exciting revolution for the next maybe two to three years," and one that is going to move fast. Alex Smola 01:03:46 “really going to be quite the an exciting revolution for the next maybe two to three years” Direct Audio Anchor Listen from 01:03:46
Interview Highlights15 exchanges
Direct dialogue & timestamps from the recording
Can voice AI work naturally in noisy environments, rather than only under the ideal conditions I have experienced with ChatGPT's voice mode?
Our system handled people switching to Slovenian and Hindi during a demonstration in a noisy bar in Jeju. That is an example of doing better than a voice system that works only in a quiet room, but the model should not receive all the credit. A non-trivial share belongs to the iOS team's noise cancellation and audio separation on the phone. I do not think simply waving a microphone around would have produced the same result. A single microphone probably cannot fully solve the problem, especially with several sound sources; microphone arrays are needed for proper noise cancellation. Those hardware capabilities already ship, and I think we are making progress.
How do you break down voice AI challenges into primarily engineering problems and primarily research problems?
Research and engineering have to be addressed together. We aim for interruption handling on roughly the 150-millisecond timescale discussed for human perception, while audio must be encoded into tokens, processed by a model, and decoded back into sound. A higher token frequency can improve fidelity, but it increases both prefill and generation costs. More tokens per second therefore constrain the model size we can afford; fewer tokens permit more parameters. Ten tokens per second also imply about 100 milliseconds per token. Biology, interaction design, representation learning, and serving costs all meet in that trade-off. Buffering adds another constraint: long buffers can help streaming, but the system still needs to be interruptible. Video brings different frame rates and the need for visual consistency over an hour, not just a short generated clip. An avatar feed has useful regularity because most of its background changes little. Exploiting that structure can make compression and real-time operation more economical, though rapid movements increase the bitrate. The objective is an affordable interaction, not simply the most impressive large-model demonstration.
Is affordability the main reason for choosing a different operating point for voice models?
Affordability comes first. A larger model can be smarter, but streaming conversations have to be economical. Using an entire Blackwell server GPU for a single conversation might deliver a gorgeous demonstration that customers cannot afford. We started with a price target and worked backward to determine what we could build within it, rather than optimizing model capability first and considering the cost later.