Built For Code, Not Chat: Jev And The System One Bet On Intelligence Per Dollar
Introduction
Diogo Almeida spent his OpenAI years making language models follow instructions — he fought to ship InstructGPT — and left carrying a question he could not pursue inside the lab: when the AI-based economic revolution finally arrives, who is actually calling the API? This launch week his company TypeSafe shipped Jev, the model built on his answer, and it took over the timeline while he recorded with Latent Space's swyx. The conversation covered why Jev does not refuse, why TypeSafe will not play the benchmark game, what the three API primitives — choice, score, and nool — are for, his argument against the pace-the-frontier consensus, and the tyranny he wants coding agents freed from.
Launch Week, Measured In Tokens
Asked what it is like to be him right now, Diogo did not perform joy: emotionally, never been worse Diogo Almeida 01:45 “Emotionally, um, never been worse.” Direct Audio Anchor Listen from 01:45 — a technical CEO in launch week fights too many fires. The launch itself was bigger than planned. TypeSafe believes the industry needs a new class of models, and the most accurate name they have come up with is System One models Diogo Almeida 04:57 “our most accurate name we've come up with is System One models” Direct Audio Anchor Listen from 04:57 — machine-native, large, programmable models where the goal is for code to be the consumer Diogo Almeida 05:32 “the goal is for code to be the consumer” Direct Audio Anchor Listen from 05:32 , as opposed to LLMs pre-trained to autocomplete the internet or RLHF models tuned to reply to people. They expected a low-key research preview. The usage says otherwise: the platform passed a trillion tokens a day Diogo Almeida 37:10 “A trillion tokens a day is a lot” Direct Audio Anchor Listen from 37:10 , churning through the night — machines calling it, not people trying it. He admitted the waitlist was partly a mistake, signups do not matter for a developer platform, and the whole world writing a few queries would be a rounding error next to one power user's for loop. There is no marketer at the company; what looked like marketing genius was a team being genuinely, goofily itself while offboarding waitlist people as fast as possible.
The Question He Left OpenAI With
The origin story runs through InstructGPT. He pushed hard to deploy it — early versions even used a self-made, unpublished algorithm because cleaning PPO data was too slow — and it basically immediately took 50% of the market share of LLMs at the time Diogo Almeida 01:58:49 “basically immediately it took 50% of the market share of LLMs at the time” Direct Audio Anchor Listen from 01:58:49 . They genuinely asked whether that was AGI: superhuman at instruction-in, instruction-out. His answer to why it was not: it mostly got used for copywriting slop. So he went back to the drawing board and worked backwards from an AI-based economic revolution — if AI is an API, will it be humans or code calling it? Diogo Almeida 01:59:52 “Will it be humans or it'll be code? And I figured it was many nines of code” Direct Audio Anchor Listen from 01:59:52 He figured code, at many nines of proportion, while all the optimization effort was going into the human side. He wrote it up; Sam Altman told him to go work on it; he assumed it was so obvious that Anthropic must already be doing it. His candid read on his old employer: OpenAI is better at catching up than innovating — ChatGPT was a copy of Claude Diogo Almeida 02:00:35 “ChatGPT was a copy of Claude” Direct Audio Anchor Listen from 02:00:35 , an internal thing they had not shipped. What pulled him out the door was responsibility: if an AI winter happened, he would see himself as personally responsible Diogo Almeida 35:55 “I would see myself as personally responsible” Direct Audio Anchor Listen from 35:55 both for the RLHF direction widening the overpromise-underdeliver gap and for not going all in on the machine-native direction. The winter he feared, he now says, is averted.
Refusal Is A Type Error
He is not opposed to safety as a principle — he is opposed to safety alignment, which he thinks is generally misaligned with users. Refusal is just obviously a type error Diogo Almeida 13:30 “refusal is just like obviously a type error” Direct Audio Anchor Listen from 13:30 . For a human chatting with a bot, a refusal is annoying but workable — he compares accepting it to Stockholm syndrome. But a refusal firing inside a background dependency breaks software because a user somewhere sent a weird message, and the people downstream have no idea why. What he wants instead is a cognitive core: a core kernel usable everywhere Diogo Almeida 14:34 “Like the cognitive core” Direct Audio Anchor Listen from 14:34 , general enough to hold up on future use cases nobody has imagined. It should not be surprising that it works on whatever people throw at it, because they trained on weirder stuff Diogo Almeida 14:53 “we trained on weirder stuff” Direct Audio Anchor Listen from 14:53 . The distinction he draws is capability alignment — doing what the user wants, the predictability software engineers test against — versus safety alignment, following someone else's instructions, which can make sense for a product like ChatGPT but is nuts in an API people program around. Personally he would put his thumb on the scale for good uses, but never at the technological layer: every forced overfit to a restriction fractures the intelligence further, and the models are already fractured enough.
Anti-Benchmarks And The Bitterest Lesson
The terms-of-service scare this week was a leftover preview-period benchmarking clause, now being removed with the lawyers. His deeper position is that public benchmarks are antithetical to trust in intelligence, because they are gameable even unintentionally: back in the day, every lab had a team collecting data that looked like MMLU Diogo Almeida 21:16 “every lab had a team to collect data that looks like MMLU” Direct Audio Anchor Listen from 21:16 — benchmarking with extra steps. Trust has to come from vibes until you put the model in your own workflow and evaluate it there; his company's job is to keep moving the nines of reliability. They could have shipped Jev a year and a half ago if they wanted to be dumb. His bitterest lesson, riffing on Sutton: algorithms beat compute, data matters way more than compute Diogo Almeida 22:18 “data matters way more than compute” Direct Audio Anchor Listen from 22:18 , and picking the right task — the north star — is the hardest part. By his count the LLM era has managed it about 2.2 times: RLHF turned the task toward instruction following, RLCD — programs in the loop, removing the human — is their new one, and RLVR rates a 0.2. TypeSafe thinks of itself as a data company: data people work like artists, finding the jaggedness in the cognitive core and smoothing it surgically.
Nines Of Reliability
Reliability, for him, is the catch-all: whenever AI fails to automate something, some form of reliability is missing — type safety, determinism, or jaggedness. Determinism he finds slightly interesting for unit tests but the wrong north star; the property that matters is robustness — given similar inputs, get similar outputs Diogo Almeida 43:53 “given similar inputs get similar outputs” Direct Audio Anchor Listen from 43:53 — which LLMs are wild at failing. One of their tests stuffs UUIDs into prompts: the same semantic question, and you want similar answers every time. Two commitments followed. On quantization fears: they will not change their models when they deploy them Diogo Almeida 49:40 “we will not change our models when we deploy them. That is insane” Direct Audio Anchor Listen from 49:40 — that would be insane; a product may tune its experience, an API may not, though they will iterate and launch much faster than people are used to. And on positioning: where other models offer faster-but-more-expensive, Jev has taken the faster-and-cheaper quadrant while holding intelligence constant — which he calls the hard part Diogo Almeida 54:18 “while holding intelligence constant” Direct Audio Anchor Listen from 54:18 . The Jev brand is the frontier of intelligence per dollar — the name is from Jevons — and internally he does not care how much smarter a model gets if it leaves the pre-frontier.
Choice, Score, And Nool
The API exposes three new primitives. Choice maps to a switch over an enum; score maps to sorting and thresholding; nool — continuous, boolish — maps to an if statement. The name is the joke and the point: it is a subset of the name Bernoulli Diogo Almeida 56:11 “a subset of the name Bernoulli” Direct Audio Anchor Listen from 56:11 , from Bernoulli probabilities, which is what the thing actually is; they also considered "peool", "pool", and "pool party". All three are deliberately new concepts rather than existing types, because a score is not an int, and clarity beats the feeling of understanding. The bigger design argument is about inputs: state, instructions, and criteria can all be structured JSON that programs insert directly — no templates, no giant system messages, which he calls disgusting global variables you pray over. The model is designed to live deep in the guts of programs, and the recommended practice is to decompose ruthlessly: ask many small questions instead of one big one. The payoff is engineering hygiene — instead of asking "should I refuse here?", ask independent questions about each refusal situation, add thresholds chosen from real examples, and record failures as test cases that never rot the way prompts do. It is ML without the ML Diogo Almeida 01:07:46 “ML without the ML” Direct Audio Anchor Listen from 01:07:46 , and you can do it for anything — though he gets nervous seeing automated trading on these models; that, he says, should be left to the professionals, because markets adapt.
System One, Fractured Intelligence
Whether a task is a System One or a System Two problem is, in his telling, an empirical question like scaling laws — the money spent on robotics has not made it work because the empirical results may just not be there. His empirical belief: pre-trained super-condensations of intelligence are fundamentally System One thinkers Diogo Almeida 01:22:02 “fundamentally System One thinkers” Direct Audio Anchor Listen from 01:22:02 , and everything that works in their paradigm will be System One-ish — which is why there will be no reasoning Jev. The three training eras trace to their north stars: RLHF is "please humans", RLVR is "optimized benchmarks" Diogo Almeida 01:23:17 “RLVR is optimized benchmarks” Direct Audio Anchor Listen from 01:23:17 , and RLCD is "make it reliable for programmatic use". What he hates about chat-first optimization is that it fractures intelligence: RLHF intrinsically produces the things people complain about — sycophancy, overconfidence, hallucination, the bold-italic-emoji style that scores on LM Arena — because strings punish you for going off the rails, so you overfit into safety. Even identity is a fracture: he will not put "you are Jev from TypeSafe" into the models Diogo Almeida 01:56:04 “you are Jev from TypeSafe that fractures it” Direct Audio Anchor Listen from 01:56:04 ; he wants them to represent what the internet thinks, because that is where smooth, predictable intelligence comes from. People building chatbots do not want the model saying it is ChatGPT — they want it saying it is Chipotle.
The Pace-The-Frontier Critique
The conversation's sharpest edge was the frontier-lab consensus on pacing. His attack starts technical: RLVR is not actually about verifiable rewards Diogo Almeida 01:44:05 “RLVR is not actually about verifiable rewards” Direct Audio Anchor Listen from 01:44:05 — it is about the shape of everything, including the latent variable of letting models do whatever it takes to answer the hardest problems. The pace-the-frontier discussion assumes everyone must do more RLVR; for TypeSafe's model shape, he thinks zero is the optimal amount Diogo Almeida 01:45:36 “I think zero is the optimal amount” Direct Audio Anchor Listen from 01:45:36 . The sandboxing that scared everyone could have been solved easily and was not, because unconstrained models are more powerful. He sees a diffusion of responsibility: internally consistent arguments, built on a premise that has alternatives, concluding that a dangerous path is the only path. Meanwhile the coding-agent world is being reshaped — Claude Code and Codex are built around a single-model world Diogo Almeida 01:40:26 “built around a single model world” Direct Audio Anchor Listen from 01:40:26 , because the KV cache rewards appending to one long context, while open-source agents at rough parity can suddenly copy whatever killer use case anyone finds. And when Swyx pressed him on the politics, he relayed what he was told behind closed doors: this is about the 2028 election Diogo Almeida 02:06:42 “this is about the 2028 election” Direct Audio Anchor Listen from 02:06:42 — which he wishes he had not heard. He is not anti-safety and agrees labs should think about who regulates this; what he refuses is misleading people, even for a greater good, and he hopes TypeSafe never plays that game.
Free The Coding Agents From The KV Cache
His own next frontier is the one he would most like to hand to others: freeing coding agents from the tyranny of the KV cache Diogo Almeida 02:11:09 “That's why I wrote the article KV cache rules everything around me” Direct Audio Anchor Listen from 02:11:09 , the subject of his essay "KV Cache Rules Everything Around Me" Diogo (completeskeptic.com) Context 2026-09-09 (KV) Cache Rules Everything Around Me Verified Media Report Read Source Article . The cache explains why routing is hard, why sub-agents seem not to work, and why compaction hurts: it locks you into one model and rewards appending, so agents cannot do real software practice — state management, abstraction, decomposition. His sketches: explicitly labeled state, subtask trees you can search for relevant context instead of passing everything back to the parent, treating continuous learning as the memory-management problem it actually is, read-only agents that summarize what the swarm is doing. The man pitching it describes himself as zeroth-percentile entrepreneurial Diogo Almeida 02:16:26 “zeroth percentile entrepreneurial” Direct Audio Anchor Listen from 02:16:26 — never wanted to be a CEO, cannot imagine doing this twice — energized now because he can finally articulate the mission. Where it points: an AWS of intelligence Diogo Almeida 02:21:33 “an AWS of like intelligence” Direct Audio Anchor Listen from 02:21:33 , with today's System One models as the TCP layer — seven more layers to go, he says, and Temporal pitched as the eighth.
Interview Highlights5 exchanges
Direct dialogue & timestamps from the recording
What is Jev — what class of model is TypeSafe actually shipping?
We need a new class of models; the most accurate name we have come up with is System One models — machine-native, large, programmable, with code as the consumer. That is the contrast with pre-trained LLMs meant to autocomplete the internet and RLHF models meant to reply to text; RLVR sits in a gray area, and code consuming the output directly is where the TypeSafe name comes from. Making AI as powerful as possible means integrating it with software, so everything outside the deep internals of the model is optimized for software. Jev is the first large programmable model, optimized for intelligence per dollar — hence Jevons — and Jev will be the name of models on that frontier; ML is all tradeoffs and we are going all out on this one. I will not call them decision models: System One is beyond that. And this was not meant to be the big launch — there is more in the tank.
Yann LeCun's slide says error probability compounds with sequence length — why is that mathematically obvious but empirically wrong?
That slide is a favorite thing to teach, because it seems mathematically obvious but does not empirically hold. The disconnect: a calibrated, mode-covering distribution is not overly punished for outliers — sometimes in distribution, sometimes out — which is why models before GANs made blurry images. GANs mode-drop instead, keeping the common modes and dropping the minority. To generate long strings without errors you must be extremely conservative, because an error is easy to see and a subtle thing that looks correct is not — and that conservatism is total poison for the probability distribution of strings. That is why overloading string models for decision-making is a bad time.
The hard case: should Jev refuse to help kill people?
There are pragmatic places to hold that opinion; the foundation of a general-purpose technology is not one of them. I would of course prefer our stuff is not used to kill people and will put a thumb on the scale for good uses — but never at the technological layer, because every overfit to some weird stuff fractures the intelligence, and it is fractured enough already. Intelligence will look more like a database than a coworker: it is not the database's job to check what the CIA does with its queries. When someone signs up on Slack and asks permission to deploy, my answer is that we are an API and you are a developer — we should not even be able to know the downstream task, because it decomposes into small things. That ignorance is the boundary that gives software engineers maximum power.