Natural Voice AI Is A Systems Problem: Alex Smola On Latency, Tokens, And Affordability

Introduction

Voice interfaces keep improving and keep breaking the spell — a beat too long before replying, a clumsy interruption, a tone that misses the room. When Sam Charrington asked Boson AI co-founder and CEO Alex Smola what it actually takes to fix this, the answer that ran through the conversation was less a smarter model than a whole system: a latency budget borrowed from biology, an economics of tokens, a data pipeline measured in human lifetimes, and a new class of benchmarks for how an agent should behave in conversation.

The Illusion Problem

Smola inverts the usual roadmap by demoting the product everyone is racing to ship: "voice is an intermediate stepping stone." Where this converges, he said, is that "you'll be talking to an AV agent that looks and feels like a human" — and he put that destination very far from natural. Charrington supplied the user-side frustration with today's frontier: ChatGPT's advanced voice mode keeps getting better yet stays infuriating, because "it really only works in perfect conditions, meaning a silent room, no background noise, being very conscientious about interrupting."

A Noisy Bar In Jeju

Smola's counterexample came from drinks at a bar in Jeju, South Korea, during the KDD conference, where he demoed the system on a dare to himself. The model, he said, handled "people switching to Slovenian and then somebody else to Hindi fairly well even though the bar was very noisy." He then shared the credit rather than claiming it: "a non-trivial amount of the credit definitely goes to the iOS team doing really good noise cancellation and audio separation on the mobile phone," and single microphones will probably never suffice — arrays are the solved, shipping path. His timeline is specific: "I think probably in a year audio will become pretty much bulletproof," avatars arrive in roughly a year and a half, and animated robot faces follow more slowly because hardware is hard.

Biology Sets The Clock

The engineering urgency comes from a biological number. Smola put the budget at "about 150 milliseconds" — the time from "a photon hitting your retina to your cortex actually doing something with it," stable enough to double as a diagnostic, and slightly shorter from the ears. Humans, he said, operate at around six to ten hertz for audiovisual perception, and Boson tuned the model to the same window: interruption handling completes "within about 150ish milliseconds or so."

The Token-Rate Dilemma

The deeper constraint is arithmetic. In the current paradigm, audio becomes tokens, tokens pass through an LLM-style backbone, and tokens become audio again — and "many tokens per second means your model cannot have too many parameters," while a slower token rate frees budget for a bigger model. The asymmetry favors text, which runs at "three to five tokens per second that humans want," while audio can easily demand ten-plus — roughly one token per 100 milliseconds. Avatar video carries a compensating mercy: "there are not going to be race cars driving in the background," so most of an avatar feed is boring, effectively compressible video.

Price First, Model Second

Boson starts from the invoice. Streaming one conversation should not require "a full Blackwell server GPU just for a single conversation," which would be gorgeous as a demo and unaffordable as a service. Instead, Smola said, "we went in with a price first and then work backwards" to something customers can actually pay for. The performance claim he attaches to that design is strong — on benchmarks, "we are better than let's say GPT and Gemini and and and Grok at a fraction of the cost. So we're better than OpenAI's models at one-tenth the cost" — and he supplied its footnote himself: "this is with thinking turned off in these models," because a long thinking pause before each reply would feel unnatural in conversation.

Teaching Audio To An LLM Without Forgetting

Adding a modality can erase the intelligence it was added to. Smola's cautionary tale is a friend's adopted daughter who, spoken to only in German, within months had "forgotten every single word of Spanish" — and a language model pushed onto audio alone, he said, loses its reasoning and language the same way. The remedy is a full training pipeline around the audio work, which produced what he called a fairly competent TTS model, Higgs Audio V2, released last year, followed this year by an accelerated release cadence of TTS, ASR, and audio-understanding models. The Higgs Audio V2 release is documented on Boson AI's site as Higgs TTS 2. The categories matter to him: speech recognition is audio-in, text-out and cannot resolve whether "recognize speech" means understanding speech or "wreck a nice beach," while an audio-understanding model reasons over the audio itself. Nor is the adaptation light fine-tuning: one released model "was built on top of Llama," acknowledged in its model card, but the logic of free inputs still holds — "if somebody gives you free steel, you don't build a steel mill; you build a car" — and the audio-side effort, he estimated, is maybe half or a third of the language model's, while insisting "it's not one-tenth."

The 100-Million-Hour Flywheel

The differentiator Smola names first is data: "we have in the order of 100 million hours of audio," which he sized up as "about 200 human lifetimes." gathered not by buying annotations but by scraping and then extracting, tagging, normalizing, and transcribing at scale. Boson AI's own engineering post narrows that figure: it describes 100M-plus raw hours processed into over 10 million model-ready hours, so the interview's headline number counts raw audio rather than training-ready data. Owning a data center is part of the moat — on a NeoCloud, "the storage bill would eat you alive" — and annotation vendors miss the scale by four orders of magnitude: offered 10,000 hours of annotation, Smola replied that Boson is "at 10,000 times that scale." The noisy labels that scare everyone else are handled with classical statistics — the medieval foot rule, he noted, is literally "where the foot rule and trimmed mean estimators come from" — and the loop closes on itself: enough of this material lets you "build better models and then you get that flywheel," including context tricks like separating the two voices on a podcast.

Foreground Talk, Background Reasoning

The architecture follows the economics. Rather than one giant end-to-end model — responsive but, in his word, dumb — Boson runs "a little bit more of a two-stage architecture where the understanding and reasoning happens" in the background while a smaller front end holds the conversation. Developers get a familiar surface: prompts "feel the same as if you were just" instructing a regular LLM, "just that our model also talks." On the Higgs Live demo the model "performs web search in the background" and, depending on how fast results return, either answers or says "hey, let me look for that." Off-the-shelf MCP servers work, with limits: "the quantity of servers" enabled simultaneously "is a little bit limited" today, because a massive prefill leaves "a large KV cache that you need to lug around," making every generated token more expensive. Slow tools are survivable when announced — "humans are very tolerant to delays," so long as "they are told that there's a delay." His exhibit is the Apple boot bar, which finishes suspiciously smoothly: in his telling, it is Apple that lies, and "the Microsoft boot screen is the truth" — a fake progress bar timed to last boot's duration, brilliant exactly because it solves the experience without solving an impossible technical problem.

Benchmarks For Behavior

What to measure is itself a research question. Interruptibility, Smola argued, "is really scene dependent": a Japanese listener's short backchannel means "I got it" and must not stop the agent, while "I don't understand" must. The proactivity benchmark is public, documented in the arXiv paper "ProactBench: Beyond What The User Asked For." Its companion, IHBench, is documented in the arXiv paper "Evaluating Post-Interruption Recovery in Voice Agents with Structured Workflows," covering the metrics for interruptibility, responsiveness, and whether the audio matches what is said. The optimization target differs from code generation's: "In our case, we worry about human happiness." Task completion still counts, but enjoyment is a requirement, not a garnish. His model of failure is the cheerfully doomed ship computer from the Hitchhiker's Guide BBC series, announcing "You will crash in 2 minutes. We will all die." The point of the story: "there is a difference between IQ and EQ," and voice needs both.

Learning To Be Pleasant

Training data for social behavior is a trap. Rom-coms reward stalking and movies end brawls with friendship, and "I sincerely hope that nobody will train a model on Dr. Phil" and call it normal behavior, Smola said. Cheap, ethical practice comes from simulators — "Nvidia did a great job at releasing some digital personas and some scenarios," which Boson uses to train against uncooperative, abrasive, and system-breaking users. The other half is memory: a personal history of interactions amounts to "a glorified CRM but now for everybody," balanced against improvements that generalize across everyone — including etiquette like the fact that "eating with your left hand in India is seriously frowned upon." What will not scale is individual genius: "Brilliance is good, but brilliance is not repeatable and automatable." The substitute is volume — "you can think of each interaction as another new experiment, a new data point" — plus technique that can simply be taught, from management training to the "millennial sandwich," where you deliver critique as "praise, critique, and praise." His forecast closes the arc: an "exciting revolution for the next maybe two to three years," and one that is going to move fast.

Interview Highlights15 exchanges

Direct dialogue & timestamps from the recording

QSam Charrington
01:39

Can voice AI work naturally in noisy environments, rather than only under the ideal conditions I have experienced with ChatGPT's voice mode?

AAlex Smola

Our system handled people switching to Slovenian and Hindi during a demonstration in a noisy bar in Jeju. That is an example of doing better than a voice system that works only in a quiet room, but the model should not receive all the credit. A non-trivial share belongs to the iOS team's noise cancellation and audio separation on the phone. I do not think simply waving a microphone around would have produced the same result. A single microphone probably cannot fully solve the problem, especially with several sound sources; microphone arrays are needed for proper noise cancellation. Those hardware capabilities already ship, and I think we are making progress.

QSam Charrington
06:07

How do you break down voice AI challenges into primarily engineering problems and primarily research problems?

AAlex Smola

Research and engineering have to be addressed together. We aim for interruption handling on roughly the 150-millisecond timescale discussed for human perception, while audio must be encoded into tokens, processed by a model, and decoded back into sound. A higher token frequency can improve fidelity, but it increases both prefill and generation costs. More tokens per second therefore constrain the model size we can afford; fewer tokens permit more parameters. Ten tokens per second also imply about 100 milliseconds per token. Biology, interaction design, representation learning, and serving costs all meet in that trade-off. Buffering adds another constraint: long buffers can help streaming, but the system still needs to be interruptible. Video brings different frame rates and the need for visual consistency over an hour, not just a short generated clip. An avatar feed has useful regularity because most of its background changes little. Exploiting that structure can make compression and real-time operation more economical, though rapid movements increase the bitrate. The objective is an affordable interaction, not simply the most impressive large-model demonstration.

QSam Charrington
12:50

Is affordability the main reason for choosing a different operating point for voice models?

AAlex Smola

Affordability comes first. A larger model can be smarter, but streaming conversations have to be economical. Using an entire Blackwell server GPU for a single conversation might deliver a gorgeous demonstration that customers cannot afford. We started with a price target and worked backward to determine what we could build within it, rather than optimizing model capability first and considering the cost later.