Claude Code vs Codex vs Jev: Speed, Cost, and When to Use Each
Who this is forWorking developers who have tried an AI coding tool and want to know what Jev is without reading about model architecture.
TL;DR: Jev, the model that swept X last week, is not an AI that writes text and code like Claude Code or Codex. It is a judgment-only model that picks answers from a fixed set of options. On official benchmarks its accuracy was 5 points below Opus 5, but its cost was 1/440th and its response time was 0.4 seconds. This post explains what is different and where to use it, written for beginners. If you need sentences, use Claude Code or Codex. If you only need to choose from a set of options, use Jev.
Contents
- What is Jev? An AI that doesn’t write
- The three at a glance: Claude Code, Codex, Jev
- Which one for which job
- How different is the performance? Official benchmarks
- Where it fits: a gatekeeper inside a coding agent
- What people who tried it say
- Try it yourself: waitlist and playground
- Summary
On September 15, 2026, Diogo Almeida, who co-built ChatGPT, unveiled Jev, the first model from his company TypeSafe AI. The announcement post scored 1,890 points on Hacker News, and on X it drew more than 35 million views. Given the buzz, few posts answer the basic question of what it actually is. This post answers where someone who has used Claude Code and Codex should place Jev.
1. What is Jev? An AI that doesn’t write
Jev is a model built by TypeSafe AI. The company calls it a new category called “System One models,” a name borrowed from a psychology term for fast, intuitive judgment.
Its defining trait is that it does not write a single word of text. When you ask Claude Code or Codex something, they write out replies, code, and explanations in sentences. Jev instead returns only answers in one of three forms:
| Question type | What it asks | What comes back |
|---|---|---|
| Choice | One of several options | The chosen item + confidence |
| Score | A score on a fixed scale | Score + confidence |
| Noul | Yes or no | Probability of yes |
Confidence is a number between 0 and 1 that shows how sure the model is of its own answer. Some reviewers call it reliability.
The founder spelled out this tradeoff on X. The second post in the announcement thread says: “The benefit is not free. Jev cannot generate text.” (translated, @CompleteSkeptic)
2. The three at a glance: Claude Code, Codex, Jev
Put side by side, the names look like rival products, but they do different jobs. Let’s start with the differences in one table.
| Item | Claude Code | Codex | Jev |
|---|---|---|---|
| Made by | Anthropic | OpenAI | TypeSafe AI |
| What it does | Reads code, edits it, runs commands | Reads code, edits it, runs commands | Picks answers to fixed questions |
| What it returns | Text, code, file changes | Text, code, file changes | Choices, scores, probabilities |
| Underlying model | Sonnet 5 (Pro), Opus 5 (Max) | GPT-5.6 family (Sol, Terra, Luna) | Jev 1.13 |
| Personal pricing | From $20/month (Pro) | From $20/month (Plus), partly on a free plan | $0.042 per 1M input tokens, output free |
| Where you use it | Terminal, IDE, desktop app, web | Terminal, IDE, app, web, cloud | API calls from your own code |
| Availability | Generally available | Generally available | Early access (waitlist) |
Prices and default models are based on each company’s official documentation as of September 2026. Sources: Claude pricing, Claude Code model config, Codex pricing, and TypeSafe model docs.
3. Which one for which job
Here is the conclusion first. The sections below back up this diagram.
Jobs that suit Jev include classifying support tickets, judging whether something is risky, and scoring. They have answers within a fixed set and need to be called several times a second. Section 5 covers what jev-guard is.
4. How different is the performance? Official benchmarks
TypeSafe runs a separate evaluation page. It breaks four tasks, security alert triage, counseling record review, invoice processing, and customer service, into judgment problems, and measures accuracy, cost per case, and time taken for each model.
The averages across the four tasks are shown in the table below. Every model was run with the same workflow Jev uses, meaning each task is split into small judgment questions. Accuracy is not measured against human-written answers. It is the match rate against reference answers made by GPT-6 Astra and Claude Fable 5.1.
| Model | Accuracy | Cost per case | Time taken |
|---|---|---|---|
| GPT-5.6 Sol | 74.1% | $0.0836 | 23.3 seconds |
| Claude Opus 5 | 73.1% | $0.1761 | 37.8 seconds |
| GPT-5.6 Terra | 67.9% | $0.0304 | 10.1 seconds |
| Jev | 67.8% | $0.0004 | 0.4 seconds |
| Claude Sonnet 5 | 67.8% | $0.1174 | 78.1 seconds |
| GPT-5.6 Luna | 66.8% | $0.0033 | 12.9 seconds |
| Claude Haiku 4.5 | 53.6% | $0.0195 | 12.5 seconds |
Source: workflow values from the overview table on evals.typesafe.ai, accessed September 19, 2026. Each model ran with its provider’s default settings. The same page says the workflow approach was more accurate, cheaper, and faster than a prompt approach that asks everything at once, across all models.
It is also worth not taking the announcement’s emphasis on “0% hallucination” at face value. The announcement post itself says this figure is not an experimental result but a structural guarantee that the model cannot output answers outside the fixed format. It can still pick a wrong answer within the format.
5. Where it fits: a gatekeeper inside a coding agent
If Jev can’t write code, why did Claude Code and Codex users react to it? Because it can slot into the many small judgment points a coding agent hits while it works.
Within four days of the announcement, several open-source projects built hooks like this. From the collection awesome-jev, I picked only the ones that plug directly into Claude Code and Codex.
| Project | What it does | Works with | Measured in README |
|---|---|---|---|
| jev-guard | Judges the risk of every tool call and checks for prompt injection | Claude Code, Codex, terminal tools for Copilot and Gemini, Cursor, and others | About 0.75 seconds and about $0.00004 per call |
| jev-judgment | Sends only yes/no judgments to Jev and keeps the main model focused on code | Coding agent skills | About 250ms |
| limpet | When an agent tries to stop before the job is done, checks rules and makes it continue | Claude Code-style stop hooks | About 0.7 seconds, a few cents per day |
6. What people who tried it say
The Hacker News thread for the announcement has 495 comments. It mixes people who have used the model with people still waiting in the queue, so I note which is which when quoting.
I got approved from the waitlist and it’s really good. For reference, it’s now also available on the Vercel gateway. (translated)
(Hacker News, porr****, hands-on use)
I tried it in early access and it was pretty useful. Asking several questions in yes/no form and using it as a second check raised my confidence in other models’ output. Models like this don’t replace LLMs; they work very well when used alongside them. (translated)
(Hacker News, tyle****, hands-on use)
The speed comparison seems misleading. Can’t Jev only produce structured output? It would be very useful for classification, routing, and scoring, but it’s a completely different thing from the code-generation models we use today for code and automation. (translated)
(Hacker News, jaco****, reviewing the announcement)
It’s on OpenRouter now, and I’ve spent the whole day building a guard. I have a lot of other ideas too. It’s incredibly convenient. (translated)
(Reddit r/PiCodingAgent, fing****, hands-on use)
I don’t understand the hype around this model. Any LLM can be constrained in its output and made to answer in parallel. I built something similar and run it on my laptop. (translated)
(Reddit r/PiCodingAgent, No_I****, similar implementation)
In Korea, an introduction post and comments appeared on GeekNews (a Korean tech news aggregator).
I just checked with the actual Jev: in 62ms it returned an 84% probability of 1 with 83% confidence.
(GeekNews comment, dice question test, hands-on use)
The commenter says Jev is more consistent than that, reaching roughly GPT-5.6 Terra level at lower cost and latency. They’re testing it now and it looks promising.
(GeekNews comment, compared with a BART classifier, under testing)
There is also a counter-experiment. A developer on X tested the small open model Gemma 3 270M by batching questions and reading only the probability of the first token. They reported it was about 77 times faster than having the same model write the full JSON (translated, @nwnwnyo). This is a counterpoint: Jev is not special, and models are simply fast when they don’t write text.
7. Try it yourself: waitlist and playground
For now, you have to go through the waitlist. Here is the order from signing up to your first question.
- On typesafe.ai, click Join Waitlist and enter your email.
- When the approval email arrives, sign in at console.typesafe.ai with that email or your Google account.
- Open Playground in the console and paste in any text. The official example is a customer message saying “Stripe has not been connected for three days.”
- Add a question. The example is a Noul-type question: “Does this message show urgency?”
- When the result shows a yes probability and confidence, you’re done. Mix Choice and Score questions and send several at once.
To call it from code, install one Python package. The snippet below is a shortened version of the official quickstart example.
from typesafe_sdk import Choice, Noul, TypeSafeClient
client = TypeSafeClient() # reads the TYPESAFE_API_KEY environment variable
response = client.system_one(
state="Stripe has not been connected for three days. Payments keep failing.",
questions={
"department": Choice(
instructions="Which team should take this?",
criteria={"billing": "Payment issue", "tech": "Technical issue"},
),
"is_urgent": Noul(instructions="Is this a message that shows urgency?"),
},
)
8. Summary
Frequently asked questions
- What is Jev, and can it write text or code like Claude Code?
- Jev is a model from TypeSafe AI that does not write text or code. It returns only choices, scores, or yes/no probabilities for fixed questions.
- How does Jev's cost compare with Claude Opus 5 in TypeSafe's benchmark?
- In TypeSafe's own test, Jev cost $0.0004 per case versus $0.1761 for Claude Opus 5, about 1/440th, and took 0.4 seconds. Its accuracy was 67.8% versus 73.1%.
BuildnWrite helps teams build AI agents that keep running. About BuildnWrite ›