Guides

Claude Code vs Codex vs Jev: Speed, Cost, and When to Use Each

10 min read#claude-code#codex#jev#ai-models#typesafe

Who this is forWorking developers who have tried an AI coding tool and want to know what Jev is without reading about model architecture.

TL;DR: Jev, the model that swept X last week, is not an AI that writes text and code like Claude Code or Codex. It is a judgment-only model that picks answers from a fixed set of options. On official benchmarks its accuracy was 5 points below Opus 5, but its cost was 1/440th and its response time was 0.4 seconds. This post explains what is different and where to use it, written for beginners. If you need sentences, use Claude Code or Codex. If you only need to choose from a set of options, use Jev.

Contents

  1. What is Jev? An AI that doesn’t write
  2. The three at a glance: Claude Code, Codex, Jev
  3. Which one for which job
  4. How different is the performance? Official benchmarks
  5. Where it fits: a gatekeeper inside a coding agent
  6. What people who tried it say
  7. Try it yourself: waitlist and playground
  8. Summary

On September 15, 2026, Diogo Almeida, who co-built ChatGPT, unveiled Jev, the first model from his company TypeSafe AI. The announcement post scored 1,890 points on Hacker News, and on X it drew more than 35 million views. Given the buzz, few posts answer the basic question of what it actually is. This post answers where someone who has used Claude Code and Codex should place Jev.

1. What is Jev? An AI that doesn’t write

TypeSafe AI official homepage. Above a pink cloud background, the headline reads 'The First (Public) System One Model' with a Join Waitlist button.
The TypeSafe AI homepage. The Join Waitlist button in the upper right is the only way in.

Jev is a model built by TypeSafe AI. The company calls it a new category called “System One models,” a name borrowed from a psychology term for fast, intuitive judgment.

Its defining trait is that it does not write a single word of text. When you ask Claude Code or Codex something, they write out replies, code, and explanations in sentences. Jev instead returns only answers in one of three forms:

Question type What it asks What comes back
Choice One of several options The chosen item + confidence
Score A score on a fixed scale Score + confidence
Noul Yes or no Probability of yes

Confidence is a number between 0 and 1 that shows how sure the model is of its own answer. Some reviewers call it reliability.

Side-by-side diagram: given the same customer support message, an LLM generates a reply draft and code one word at a time, while Jev answers three fixed questions at once, returning the owning department, whether it is urgent, and a complaint score, each with confidence.
Left: how Claude Code and Codex work. Right: how Jev works.

The founder spelled out this tradeoff on X. The second post in the announcement thread says: “The benefit is not free. Jev cannot generate text.” (translated, @CompleteSkeptic)

Diogo Almeida's X announcement post. The text says that after co-building ChatGPT, he spent two years building Jev and a new training method called RLCD. It includes a claim of 20 to 200 times faster and shows 35 million views.
This is the first post in the announcement thread. The number next to the bar icon at the bottom is the view count, and the date is shown in Korea Standard Time (KST).

2. The three at a glance: Claude Code, Codex, Jev

Put side by side, the names look like rival products, but they do different jobs. Let’s start with the differences in one table.

Item Claude Code Codex Jev
Made by Anthropic OpenAI TypeSafe AI
What it does Reads code, edits it, runs commands Reads code, edits it, runs commands Picks answers to fixed questions
What it returns Text, code, file changes Text, code, file changes Choices, scores, probabilities
Underlying model Sonnet 5 (Pro), Opus 5 (Max) GPT-5.6 family (Sol, Terra, Luna) Jev 1.13
Personal pricing From $20/month (Pro) From $20/month (Plus), partly on a free plan $0.042 per 1M input tokens, output free
Where you use it Terminal, IDE, desktop app, web Terminal, IDE, app, web, cloud API calls from your own code
Availability Generally available Generally available Early access (waitlist)

Prices and default models are based on each company’s official documentation as of September 2026. Sources: Claude pricing, Claude Code model config, Codex pricing, and TypeSafe model docs.

3. Which one for which job

Here is the conclusion first. The sections below back up this diagram.

Which tool for which job. Claude Code, Codex, Jev. If the job requires writing sentences or code, such as building features, fixing bugs, or writing docs, use Claude Code or Codex. If you need an explanation of why a decision was made, use Claude Code or Codex, since Jev does not give reasons. If you want to block risky actions by a coding agent, add jev-guard and separate the writer from the judge. If none of these apply, use Jev for frequent choices within a fixed set, such as classification, judgment, and scoring.

Jobs that suit Jev include classifying support tickets, judging whether something is risky, and scoring. They have answers within a fixed set and need to be called several times a second. Section 5 covers what jev-guard is.

4. How different is the performance? Official benchmarks

TypeSafe runs a separate evaluation page. It breaks four tasks, security alert triage, counseling record review, invoice processing, and customer service, into judgment problems, and measures accuracy, cost per case, and time taken for each model.

Scatter plot from TypeSafe's evaluation page. The x-axis is cost per case and the y-axis is accuracy. Jev sits alone at the far left at $0.0004 with 67.8% accuracy, while Opus 5 and Sol sit in the upper right.
The x-axis is cost and the y-axis is accuracy. The further toward the upper left, the cheaper and more accurate.

The averages across the four tasks are shown in the table below. Every model was run with the same workflow Jev uses, meaning each task is split into small judgment questions. Accuracy is not measured against human-written answers. It is the match rate against reference answers made by GPT-6 Astra and Claude Fable 5.1.

Model Accuracy Cost per case Time taken
GPT-5.6 Sol 74.1% $0.0836 23.3 seconds
Claude Opus 5 73.1% $0.1761 37.8 seconds
GPT-5.6 Terra 67.9% $0.0304 10.1 seconds
Jev 67.8% $0.0004 0.4 seconds
Claude Sonnet 5 67.8% $0.1174 78.1 seconds
GPT-5.6 Luna 66.8% $0.0033 12.9 seconds
Claude Haiku 4.5 53.6% $0.0195 12.5 seconds

Source: workflow values from the overview table on evals.typesafe.ai, accessed September 19, 2026. Each model ran with its provider’s default settings. The same page says the workflow approach was more accurate, cheaper, and faster than a prompt approach that asks everything at once, across all models.

It is also worth not taking the announcement’s emphasis on “0% hallucination” at face value. The announcement post itself says this figure is not an experimental result but a structural guarantee that the model cannot output answers outside the fixed format. It can still pick a wrong answer within the format.

5. Where it fits: a gatekeeper inside a coding agent

If Jev can’t write code, why did Claude Code and Codex users react to it? Because it can slot into the many small judgment points a coding agent hits while it works.

Diagram of a coding agent flow running from user instruction to LLM planning, tool calls, and result application. A Jev judgment box sits above the tool-call step and splits the flow into three outcomes: allow, ask for confirmation, or block.
The LLM still writes the text and code. Jev only checks whether a tool may be run.

Within four days of the announcement, several open-source projects built hooks like this. From the collection awesome-jev, I picked only the ones that plug directly into Claude Code and Codex.

Project What it does Works with Measured in README
jev-guard Judges the risk of every tool call and checks for prompt injection Claude Code, Codex, terminal tools for Copilot and Gemini, Cursor, and others About 0.75 seconds and about $0.00004 per call
jev-judgment Sends only yes/no judgments to Jev and keeps the main model focused on code Coding agent skills About 250ms
limpet When an agent tries to stop before the job is done, checks rules and makes it continue Claude Code-style stop hooks About 0.7 seconds, a few cents per day

6. What people who tried it say

The Hacker News thread for the announcement has 495 comments. It mixes people who have used the model with people still waiting in the queue, so I note which is which when quoting.

Top of the Hacker News thread. Below the title 'Introducing System One Models and Jev', it shows 1890 points and 495 comments.
Points and comment count as of September 18, 2026.

I got approved from the waitlist and it’s really good. For reference, it’s now also available on the Vercel gateway. (translated)

(Hacker News, porr****, hands-on use)

I tried it in early access and it was pretty useful. Asking several questions in yes/no form and using it as a second check raised my confidence in other models’ output. Models like this don’t replace LLMs; they work very well when used alongside them. (translated)

(Hacker News, tyle****, hands-on use)

The speed comparison seems misleading. Can’t Jev only produce structured output? It would be very useful for classification, routing, and scoring, but it’s a completely different thing from the code-generation models we use today for code and automation. (translated)

(Hacker News, jaco****, reviewing the announcement)

It’s on OpenRouter now, and I’ve spent the whole day building a guard. I have a lot of other ideas too. It’s incredibly convenient. (translated)

(Reddit r/PiCodingAgent, fing****, hands-on use)

I don’t understand the hype around this model. Any LLM can be constrained in its output and made to answer in parallel. I built something similar and run it on my laptop. (translated)

(Reddit r/PiCodingAgent, No_I****, similar implementation)

In Korea, an introduction post and comments appeared on GeekNews (a Korean tech news aggregator).

GeekNews introduction post on Jev. The title describes an AI model that returns judgments and probabilities instead of sentences, with summary bullet points below.
The GeekNews introduction post. The five summary lines under the title capture the model's core points.

I just checked with the actual Jev: in 62ms it returned an 84% probability of 1 with 83% confidence.

(GeekNews comment, dice question test, hands-on use)

The commenter says Jev is more consistent than that, reaching roughly GPT-5.6 Terra level at lower cost and latency. They’re testing it now and it looks promising.

(GeekNews comment, compared with a BART classifier, under testing)

There is also a counter-experiment. A developer on X tested the small open model Gemma 3 270M by batching questions and reading only the probability of the first token. They reported it was about 77 times faster than having the same model write the full JSON (translated, @nwnwnyo). This is a counterpoint: Jev is not special, and models are simply fast when they don’t write text.

7. Try it yourself: waitlist and playground

For now, you have to go through the waitlist. Here is the order from signing up to your first question.

  1. On typesafe.ai, click Join Waitlist and enter your email.
  2. When the approval email arrives, sign in at console.typesafe.ai with that email or your Google account.
  3. Open Playground in the console and paste in any text. The official example is a customer message saying “Stripe has not been connected for three days.”
  4. Add a question. The example is a Noul-type question: “Does this message show urgency?”
  5. When the result shows a yes probability and confidence, you’re done. Mix Choice and Score questions and send several at once.
TypeSafe console login screen. Under the heading 'Welcome to TypeSafe' are a Continue with Google button and an email field.
The console login screen. Until your waitlist request is approved, you can't go any further from here.
TypeSafe docs Quick start page. It shows a JSON example of pasting state into the Playground and adding a noul question named urgency.
The official quickstart. The example text from steps 3 and 4 appears as-is.

To call it from code, install one Python package. The snippet below is a shortened version of the official quickstart example.

from typesafe_sdk import Choice, Noul, TypeSafeClient

client = TypeSafeClient()  # reads the TYPESAFE_API_KEY environment variable
response = client.system_one(
    state="Stripe has not been connected for three days. Payments keep failing.",
    questions={
        "department": Choice(
            instructions="Which team should take this?",
            criteria={"billing": "Payment issue", "tech": "Technical issue"},
        ),
        "is_urgent": Noul(instructions="Is this a message that shows urgency?"),
    },
)

8. Summary

Frequently asked questions

What is Jev, and can it write text or code like Claude Code?
Jev is a model from TypeSafe AI that does not write text or code. It returns only choices, scores, or yes/no probabilities for fixed questions.
How does Jev's cost compare with Claude Opus 5 in TypeSafe's benchmark?
In TypeSafe's own test, Jev cost $0.0004 per case versus $0.1761 for Claude Opus 5, about 1/440th, and took 0.4 seconds. Its accuracy was 67.8% versus 73.1%.