Why I couldn't build Jev at OpenAI — Diogo Almeida, TypeSafe Co-founder & CEO

Why I couldn't build Jev at OpenAI — Diogo Almeida, TypeSafe Co-founder & CEO

Latent Space•02:22:22•2026-09-21•Source Audio
Guests
Diogo Almeida

Executive Summary

A day after launching Jev, Diogo Almeida — GPT-3/3.5-era OpenAI engineer, InstructGPT veteran, and now TypeSafe co-founder — joins swyx on Latent Space to argue that the industry has been optimizing the wrong model for the wrong consumer. Pre-trained LLMs autocomplete the internet and RLHF models are built to reply to text; Jev is the first large programmable model, a System One model consumed by code and optimized for intelligence per dollar — named for the Jevons paradox, because cheap intelligence expands total use. Launch-week color frames the stakes: a trillion tokens a day already served, machines rather than humans calling the API, rate limits as the real demand signal, and deployment discipline borrowed from databases — models frozen once shipped, no long-term support promised, Jev 1.13.0 possibly LTS'd. The intellectual core is a three-way split of post-training northstars: RLHF is please humans, RLVR is optimize benchmarks, and RLCD is make it reliable for programmatic use. Pre-trained condensations of intelligence are fundamentally System One thinkers; RLVR earned real awe for System 2 but its results are fragile and jagged — math is not just spiky, it is fractal — and the optimal amount of RLVR for Jev's shape is zero. Chat-first training plus a reasoning mode forces intelligence to fracture: RLHF intrinsically buys sycophancy, overconfidence, hallucination and LM Arena style, string generation demands miscalibration and mode dropping, and even identity is fracture — a chatbot built on an API should say Chipotle. Calibration arguments run through LeCun's error-probability slide (mathematically obvious, empirically wrong) to the hardest objection: Jev should not refuse at the technological layer, because intelligence will look more like a database than a coworker. The back half is history and market structure: the coup, when safety took over OpenAI; the InstructGPT fight, with an unpublished algorithm and 50% LLM market share; instruction-following's 'we won'; Sam Altman blessing the machines-calling-APIs document; and the claim that ChatGPT was a copy of Claude. Claude Code and Codex lead coding agents precisely because they are built around a single-model world, and the KV cache — the subject of Diogo's own essay — economically locks agents into that world, reframing continuous learning as memory management. 'Pace the frontier' is dismissed as a sleight of hand with solvable sandboxing left unsolved, and the close is a developer-experience rant: the function-calling interface is insane, refusal thresholds are begged in system messages, and the forecast is an AWS of intelligence.

Chapters & Key Takeaways

Jev is defined by consumer and metric, not capability chat: a machine-native, programmable System One model consumed by code and optimized for intelligence per dollar — the contrast class being internet-autocompletion pre-trainings and text-replying RLHF models.
The Jev name is a thesis: after the Jevons paradox, cheaper intelligence per dollar expands total use, and Jev will be the name of TypeSafe's models on that frontier.
Conservative string generation is poison for probability distributions: models that must not visibly err become miscalibrated and mode-drop, which is why overloading string models for decision-making fails.
Safety belongs outside the technological layer: intelligence will look more like a database than a coworker, and it is not the database's job to police what the CIA does with its queries.
Public benchmarks cannot be trusted because labs train on lookalike data — MMLU-style imitators — so TypeSafe runs internal evals and makes not-gaming them a top-level priority.
The launch metrics invert consumer AI reading: about a trillion tokens a day served, machines rather than humans calling the API, and rate-limit pressure — not waitlist signups — as the demand signal that matters.
Deployment discipline comes from database practice: models will not change once deployed, long-term support is not promised, and Jev 1.13.0 may be temporarily LTS'd while TypeSafe ships much faster than providers usually do.
All models are roughly tied at zero on the world's economically valuable work — automating it is not yet 1% done — which is why the manifesto targets TFP growth reading 3% in 5 years.
The three northstars: RLHF is please humans, RLVR is optimize benchmarks, RLCD is make it reliable for programmatic use; pre-trained condensations of intelligence are fundamentally System One thinkers, and everything that works in the unearthing paradigm is System Oneish.
RLVR earned genuine awe for System 2 but its results are fragile and jagged — math is not just spiky, it is fractal — and the optimal amount of RLVR for Jev's shape is zero.
Multihop performance falls monotonically as hops increase even where single-hop is state-of-the-art — the host's empirical gloss on the same fragility.
Post-training fracturing is the enemy: chat-first plus a reasoning mode splits the intelligence, optimizing two objectives at once is literally the act of fracturing, and even identity is fracture — a chatbot built on an API should say Chipotle.
The OpenAI years retold as insider history: ChatGPT was a copy of Claude (which had shipped in Slack first), OpenAI does better at catching up than innovating, and Sam Altman blessed the machines-calling-APIs northstar document that became TypeSafe.
The KV cache economically locks coding agents into one model — the tyranny of re-reading their own context — which is why Claude Code and Codex are built around a single-model world and why Diogo wrote his essay on the subject.
The developer interface is the remaining frontier: the function-calling interface is insane and anti-developer, refusal thresholds can only be begged for in system messages, and existing coding agents are overfit to their harnesses and under-use MCPs.

Built For Code, Not Chat: Jev And The System One Bet On Intelligence Per Dollar

Introduction

Diogo Almeida spent his OpenAI years making language models follow instructions — he fought to ship InstructGPT — and left carrying a question he could not pursue inside the lab: when the AI-based economic revolution finally arrives, who is actually calling the API? This launch week his company TypeSafe shipped Jev, the model built on his answer, and it took over the timeline while he recorded with Latent Space's swyx. The conversation covered why Jev does not refuse, why TypeSafe will not play the benchmark game, what the three API primitives — choice, score, and nool — are for, his argument against the pace-the-frontier consensus, and the tyranny he wants coding agents freed from.

Launch Week, Measured In Tokens

Asked what it is like to be him right now, Diogo did not perform joy: emotionally, never been worse — a technical CEO in launch week fights too many fires. The launch itself was bigger than planned. TypeSafe believes the industry needs a new class of models, and the most accurate name they have come up with is System One models — machine-native, large, programmable models where the goal is for code to be the consumer , as opposed to LLMs pre-trained to autocomplete the internet or RLHF models tuned to reply to people. They expected a low-key research preview. The usage says otherwise: the platform passed a trillion tokens a day , churning through the night — machines calling it, not people trying it. He admitted the waitlist was partly a mistake, signups do not matter for a developer platform, and the whole world writing a few queries would be a rounding error next to one power user's for loop. There is no marketer at the company; what looked like marketing genius was a team being genuinely, goofily itself while offboarding waitlist people as fast as possible.

The Question He Left OpenAI With

The origin story runs through InstructGPT. He pushed hard to deploy it — early versions even used a self-made, unpublished algorithm because cleaning PPO data was too slow — and it basically immediately took 50% of the market share of LLMs at the time . They genuinely asked whether that was AGI: superhuman at instruction-in, instruction-out. His answer to why it was not: it mostly got used for copywriting slop. So he went back to the drawing board and worked backwards from an AI-based economic revolution — if AI is an API, will it be humans or code calling it? He figured code, at many nines of proportion, while all the optimization effort was going into the human side. He wrote it up; Sam Altman told him to go work on it; he assumed it was so obvious that Anthropic must already be doing it. His candid read on his old employer: OpenAI is better at catching up than innovating — ChatGPT was a copy of Claude , an internal thing they had not shipped. What pulled him out the door was responsibility: if an AI winter happened, he would see himself as personally responsible both for the RLHF direction widening the overpromise-underdeliver gap and for not going all in on the machine-native direction. The winter he feared, he now says, is averted.

Refusal Is A Type Error

He is not opposed to safety as a principle — he is opposed to safety alignment, which he thinks is generally misaligned with users. Refusal is just obviously a type error . For a human chatting with a bot, a refusal is annoying but workable — he compares accepting it to Stockholm syndrome. But a refusal firing inside a background dependency breaks software because a user somewhere sent a weird message, and the people downstream have no idea why. What he wants instead is a cognitive core: a core kernel usable everywhere , general enough to hold up on future use cases nobody has imagined. It should not be surprising that it works on whatever people throw at it, because they trained on weirder stuff . The distinction he draws is capability alignment — doing what the user wants, the predictability software engineers test against — versus safety alignment, following someone else's instructions, which can make sense for a product like ChatGPT but is nuts in an API people program around. Personally he would put his thumb on the scale for good uses, but never at the technological layer: every forced overfit to a restriction fractures the intelligence further, and the models are already fractured enough.

Anti-Benchmarks And The Bitterest Lesson

The terms-of-service scare this week was a leftover preview-period benchmarking clause, now being removed with the lawyers. His deeper position is that public benchmarks are antithetical to trust in intelligence, because they are gameable even unintentionally: back in the day, every lab had a team collecting data that looked like MMLU — benchmarking with extra steps. Trust has to come from vibes until you put the model in your own workflow and evaluate it there; his company's job is to keep moving the nines of reliability. They could have shipped Jev a year and a half ago if they wanted to be dumb. His bitterest lesson, riffing on Sutton: algorithms beat compute, data matters way more than compute , and picking the right task — the north star — is the hardest part. By his count the LLM era has managed it about 2.2 times: RLHF turned the task toward instruction following, RLCD — programs in the loop, removing the human — is their new one, and RLVR rates a 0.2. TypeSafe thinks of itself as a data company: data people work like artists, finding the jaggedness in the cognitive core and smoothing it surgically.

Nines Of Reliability

Reliability, for him, is the catch-all: whenever AI fails to automate something, some form of reliability is missing — type safety, determinism, or jaggedness. Determinism he finds slightly interesting for unit tests but the wrong north star; the property that matters is robustness — given similar inputs, get similar outputs — which LLMs are wild at failing. One of their tests stuffs UUIDs into prompts: the same semantic question, and you want similar answers every time. Two commitments followed. On quantization fears: they will not change their models when they deploy them — that would be insane; a product may tune its experience, an API may not, though they will iterate and launch much faster than people are used to. And on positioning: where other models offer faster-but-more-expensive, Jev has taken the faster-and-cheaper quadrant while holding intelligence constant — which he calls the hard part . The Jev brand is the frontier of intelligence per dollar — the name is from Jevons — and internally he does not care how much smarter a model gets if it leaves the pre-frontier.

Choice, Score, And Nool

The API exposes three new primitives. Choice maps to a switch over an enum; score maps to sorting and thresholding; nool — continuous, boolish — maps to an if statement. The name is the joke and the point: it is a subset of the name Bernoulli , from Bernoulli probabilities, which is what the thing actually is; they also considered "peool", "pool", and "pool party". All three are deliberately new concepts rather than existing types, because a score is not an int, and clarity beats the feeling of understanding. The bigger design argument is about inputs: state, instructions, and criteria can all be structured JSON that programs insert directly — no templates, no giant system messages, which he calls disgusting global variables you pray over. The model is designed to live deep in the guts of programs, and the recommended practice is to decompose ruthlessly: ask many small questions instead of one big one. The payoff is engineering hygiene — instead of asking "should I refuse here?", ask independent questions about each refusal situation, add thresholds chosen from real examples, and record failures as test cases that never rot the way prompts do. It is ML without the ML , and you can do it for anything — though he gets nervous seeing automated trading on these models; that, he says, should be left to the professionals, because markets adapt.

System One, Fractured Intelligence

Whether a task is a System One or a System Two problem is, in his telling, an empirical question like scaling laws — the money spent on robotics has not made it work because the empirical results may just not be there. His empirical belief: pre-trained super-condensations of intelligence are fundamentally System One thinkers , and everything that works in their paradigm will be System One-ish — which is why there will be no reasoning Jev. The three training eras trace to their north stars: RLHF is "please humans", RLVR is "optimized benchmarks" , and RLCD is "make it reliable for programmatic use". What he hates about chat-first optimization is that it fractures intelligence: RLHF intrinsically produces the things people complain about — sycophancy, overconfidence, hallucination, the bold-italic-emoji style that scores on LM Arena — because strings punish you for going off the rails, so you overfit into safety. Even identity is a fracture: he will not put "you are Jev from TypeSafe" into the models ; he wants them to represent what the internet thinks, because that is where smooth, predictable intelligence comes from. People building chatbots do not want the model saying it is ChatGPT — they want it saying it is Chipotle.

The Pace-The-Frontier Critique

The conversation's sharpest edge was the frontier-lab consensus on pacing. His attack starts technical: RLVR is not actually about verifiable rewards — it is about the shape of everything, including the latent variable of letting models do whatever it takes to answer the hardest problems. The pace-the-frontier discussion assumes everyone must do more RLVR; for TypeSafe's model shape, he thinks zero is the optimal amount . The sandboxing that scared everyone could have been solved easily and was not, because unconstrained models are more powerful. He sees a diffusion of responsibility: internally consistent arguments, built on a premise that has alternatives, concluding that a dangerous path is the only path. Meanwhile the coding-agent world is being reshaped — Claude Code and Codex are built around a single-model world , because the KV cache rewards appending to one long context, while open-source agents at rough parity can suddenly copy whatever killer use case anyone finds. And when Swyx pressed him on the politics, he relayed what he was told behind closed doors: this is about the 2028 election — which he wishes he had not heard. He is not anti-safety and agrees labs should think about who regulates this; what he refuses is misleading people, even for a greater good, and he hopes TypeSafe never plays that game.

Free The Coding Agents From The KV Cache

His own next frontier is the one he would most like to hand to others: freeing coding agents from the tyranny of the KV cache , the subject of his essay "KV Cache Rules Everything Around Me" . The cache explains why routing is hard, why sub-agents seem not to work, and why compaction hurts: it locks you into one model and rewards appending, so agents cannot do real software practice — state management, abstraction, decomposition. His sketches: explicitly labeled state, subtask trees you can search for relevant context instead of passing everything back to the parent, treating continuous learning as the memory-management problem it actually is, read-only agents that summarize what the swarm is doing. The man pitching it describes himself as zeroth-percentile entrepreneurial — never wanted to be a CEO, cannot imagine doing this twice — energized now because he can finally articulate the mission. Where it points: an AWS of intelligence , with today's System One models as the TCP layer — seven more layers to go, he says, and Temporal pitched as the eighth.

Interview Highlights5 exchanges

Direct dialogue & timestamps from the recording

Qswyx
04:27

What is Jev — what class of model is TypeSafe actually shipping?

ADiogo Almeida

We need a new class of models; the most accurate name we have come up with is System One models — machine-native, large, programmable, with code as the consumer. That is the contrast with pre-trained LLMs meant to autocomplete the internet and RLHF models meant to reply to text; RLVR sits in a gray area, and code consuming the output directly is where the TypeSafe name comes from. Making AI as powerful as possible means integrating it with software, so everything outside the deep internals of the model is optimized for software. Jev is the first large programmable model, optimized for intelligence per dollar — hence Jevons — and Jev will be the name of models on that frontier; ML is all tradeoffs and we are going all out on this one. I will not call them decision models: System One is beyond that. And this was not meant to be the big launch — there is more in the tank.

Qswyx
08:06

Yann LeCun's slide says error probability compounds with sequence length — why is that mathematically obvious but empirically wrong?

ADiogo Almeida

That slide is a favorite thing to teach, because it seems mathematically obvious but does not empirically hold. The disconnect: a calibrated, mode-covering distribution is not overly punished for outliers — sometimes in distribution, sometimes out — which is why models before GANs made blurry images. GANs mode-drop instead, keeping the common modes and dropping the minority. To generate long strings without errors you must be extremely conservative, because an error is easy to see and a subtle thing that looks correct is not — and that conservatism is total poison for the probability distribution of strings. That is why overloading string models for decision-making is a bad time.

Qswyx
16:21

The hard case: should Jev refuse to help kill people?

ADiogo Almeida

There are pragmatic places to hold that opinion; the foundation of a general-purpose technology is not one of them. I would of course prefer our stuff is not used to kill people and will put a thumb on the scale for good uses — but never at the technological layer, because every overfit to some weird stuff fractures the intelligence, and it is fractured enough already. Intelligence will look more like a database than a coworker: it is not the database's job to check what the CIA does with its queries. When someone signs up on Slack and asks permission to deploy, my answer is that we are an API and you are a developer — we should not even be able to know the downstream task, because it decomposes into small things. That ignorance is the boundary that gives software engineers maximum power.