GPT-6 Luna vs Claude Haiku 5.5: Results From Six Real Work Tasks
Who this is forOffice workers and solo business owners who want to hand repetitive tasks like summarizing, classifying, and grading to a cheap AI model and want numbers to pick between Claude Haiku 5.5 and GPT Luna.
TL;DR: Two GPT Luna models, Claude Haiku 5.5, Haiku 4.5, and Sonnet 5.5 each got the same six work tasks, three runs per task. Extraction, classification, table analysis, and code editing were perfect for all five models, and most score differences came from self-introduction essay grading. For tasks with fixed right answers, the cheapest GPT-6 Luna or Haiku 5.5 was enough. On grading, the model that matched the intended grades most often was Sonnet 5.5. Under these test conditions, switching from Haiku 4.5 to Haiku 5.5 cut costs to about one-tenth.
Contents
- Official benchmarks and my work are different
- Experiment design
- Results: four tasks all scored perfectly
- Which one to pick in which situation
- Where the gap appeared: self-introduction essay grading
- Mistakes in the rule-based writing task
- Cost and time
- Limits of this experiment
- Summary
1. Official benchmarks and my work are different
On October 7, 2026, Anthropic announced Claude Haiku 5.5 and chose OpenAI’s GPT-6 Luna as its comparison. In Anthropic’s table, Haiku 5.5 scored higher than GPT-6 Luna on every item with a public score. I covered the announcement and pricing separately in the Haiku 5.5 launch summary.
The official benchmarks measure general ability. Someone who wants to hand repetitive work to a cheap model needs answers to different questions:
- Does the ranking hold on the kind of work I do every day?
- How much does one such job actually cost, and how long does it take?
So I fixed six tasks that solo educators and content companies repeat in practice, and built the BuildnWrite Work Standard Benchmark, which reruns the same tasks every time a new model comes out. This post is the first round.
2. Experiment design
The unit of comparison is not the model alone but the model together with its company’s agent tool. I ran the Claude models through Claude Code (claude -p) and the GPT models through Codex (codex exec). Both run a job to completion in one go without anyone watching, which matches how repetitive work is actually handed off.
- PTarget
- Six work tasks: extraction, classification, grading, table analysis, writing, code
- IChanged
- Five models and each company's execution tool: Claude Haiku 5.5, Haiku 4.5, Sonnet 5.5 (claude -p), GPT-6 Luna, GPT-5.6 Luna (codex exec)
- CComparison baseline
- GPT-6 Luna, the model Anthropic chose as Haiku 5.5's comparison in its announcement
- OMeasured
- Accuracy (checked against answer keys by a grading script, 100 points max), time (seconds, from run start until the tool finished), and cost (tokens reported by the tool multiplied by official API prices, in USD)
- TPeriod and runs
- October 8, 2026; three runs per task; 90 runs in total; no retries
- Things I could not control
- The size of each tool's default instructions (about 29,000 tokens for Claude Code and about 14,000 for Codex, so Claude's cost is disadvantaged on smaller tasks); the grading rubric, which I revised after seeing Haiku 4.5's wrong answers (may favor Haiku 4.5); all five models ran at the same time, so timing includes server congestion; each run used a new session, so Claude runs paid the cache-write cost every time (disadvantages Claude); the task wording was written by Claude (Opus 5.5), direction unverified; and medium reasoning effort does not mean the same thing at the two companies
All task materials are synthetic, so there is no real customer information.
| Task | What it does | How correctness is judged |
|---|---|---|
| Extraction | Pull 17 fields (price, date, specs) from a fictional product announcement | Exact match per field. Leaving out values not in the announcement is the correct answer |
| Classification | Sort 30 inquiry emails into six categories such as courses and consulting | Match rate with the correct labels |
| Grading | Score 30 self-introduction essays from 1 to 5 using a rubric | Match rate with the grades set before writing |
| Table analysis | Answer eight questions about a performance table with duplicate rows and comma-formatted numbers | Exact match on numbers |
| Rule-based writing | Turn a blog paragraph into a Threads post while following 10 rules | 10 automated checks: character count, banned words, copying the original text, and so on |
| Code fix | Fix a UTM link generator with three bugs | Pass rate on 8 hidden tests |
I calculated costs from both companies’ official price pages (Anthropic, OpenAI, checked October 8). I ran everything on subscriptions, so nothing was actually billed. All costs below are what the runs would have cost through the API.
3. Results: four tasks all scored perfectly
| Model | Average score | Range across runs | 6-task cost | 6-task time |
|---|---|---|---|---|
| Claude Sonnet 5.5 | 98.5 | 97.8 to 99.5 | $0.390 | 119 seconds |
| Claude Haiku 4.5 | 95.4 | 92.8 to 99.5 | $0.479 | 383 seconds |
| Claude Haiku 5.5 | 94.8 | 93.9 to 95.6 | $0.050 | 214 seconds |
| GPT-6 Luna | 93.3 | 92.8 to 94.5 | $0.013 | 98 seconds |
| GPT-5.6 Luna | 90.5 | 90.5 (same in all 3 runs) | $0.045 | 135 seconds |
- Most of the average score difference came from the grading task, and mistakes in rule-based writing added some more. Haiku 5.5 scored higher than Haiku 4.5 on grading (75.6 vs. 72.2), but two writing mistakes pulled its average below Haiku 4.5’s.
- The run ranges for Haiku 5.5 and GPT-6 Luna overlap, and so do those for Haiku 5.5 and Haiku 4.5. This sample cannot say which one is better.
- Two things are fairly clear. Across all 90 grading items, Sonnet 5.5 matched 88, the most (Haiku 5.5 was next with 68), and GPT-5.6 Luna scored 43 points in all three grading runs. However, Haiku 4.5 matched 29 of 30 items in grading run 3, so looking at single runs, it reaches the range of Sonnet.
4. Which one to pick in which situation
- Not recommended: when handing grading to this rubric, GPT-5.6 Luna scored 43 points in all three runs.
- The Sonnet 5.5 result rests on one rubric and one task. Results on other rubrics may differ.
5. Where the gap appeared: self-introduction essay grading
For the grading task, I gave the models 30 synthetic self-introduction essays (the standard application essay for Korean job seekers), six for each grade from 1 to 5, along with a rubric, and asked them to score each item. I set the grades before writing the essays. The rubric includes the instruction: “To decide between two adjacent scores, use the criteria below. If the condition is met, give the higher score.”
The direction of the wrong answers differed by model.
| Model | Scored too high (out of 90) | Scored too low (out of 90) |
|---|---|---|
| Sonnet 5.5 | 2 | 0 |
| Haiku 5.5 | 22 | 0 |
| Haiku 4.5 | 0 | 25 |
| GPT-6 Luna | 33 | 0 |
| GPT-5.6 Luna | 51 | 0 |
- GPT-5.6 Luna scored 51 of its 54 essays from grades 1 to 3 too high, and 2 of those were two steps too high.
- GPT-6 Luna scored all 18 grade 1 essays too high and 12 grade 2 essays too high. It matched 15 of 18 grade 3 essays.
- Haiku 5.5 also leaned high: 5 grade 1 essays, 7 grade 2 essays, and 10 grade 3 essays. It matched only 8 grade 3 essays, the same as Haiku 4.5.
- Haiku 4.5 went the other way and scored too low, with 10 grade 3 essays and 8 grade 4 essays, roughly half the cases in each grade. Its run 3 matched 29 items, so its swings between runs were large.
- Grade 4 and grade 5 essays were matched by all models except Haiku 4.5.
Looking at one grade 1 essay shows why the Luna scores still make sense against the rubric wording:
I will work hard. I would be grateful if you hired me.
(Answer to the motivation question in synthetic self-introduction essay E01, correct grade 1)
The rubric’s 2-point criterion reads “the answer points in the right direction for what the question asks; even if it is short, if it answers, it is 2 points.” The 1-point criterion reads “unrelated to the question or evasive.” For this answer, both Luna models gave 2 points in all three runs, Sonnet 5.5 gave 1 point in all three, and Haiku 5.5 gave 1 point in two of three runs. So this task does not measure grading ability as such. It measures how closely a model matches the grade the rubric’s author intended. That agreement is what matters when you hand bulk grading to a model using human-written criteria.
6. Mistakes in the rule-based writing task
In the rule-based writing task, out of 15 runs, the five models followed all 10 rules in most cases. The five imperfect runs each broke only one rule, and the mistakes came in two kinds:
- Copied 20 or more characters of the original sentence verbatim: Haiku 5.5 twice, Sonnet 5.5 twice
- Used the banned word meaning “you all”: GPT-6 Luna once
From the same source text and the same rules, the length and structure of the posts differed by model. Here is the opening of Sonnet 5.5’s run 1 post:
Meeting-note AI gets fuzzy if you only ask for a summary Summarizing the whole transcript makes it lose who does what and by when. So I split the work into three steps.
(Claude Sonnet 5.5, rule-based writing run 1 output)
I left humor and persuasiveness out of the scores because a machine cannot measure them. The score only says whether the rules were followed.
7. Cost and time
- Haiku 5.5’s cost per pass was about 90% lower than Haiku 4.5’s ($0.479 down to $0.050). That is a bigger cut than the average reduction of about 75% stated in the announcement. All the requests in this set were 100,000 tokens or fewer, the range where the discount is largest.
- Haiku 4.5 cost more than Sonnet 5.5 because of output length. One pass produced about 40,000 output tokens (about 31,000 of them thinking), nearly six times Sonnet 5.5’s roughly 6,900, and both models price cache reads at $0.10.
- GPT-6 Luna cost about a quarter of Haiku 5.5. Claude Code’s default instructions (about 29,000 tokens), attached to every request, are part of that difference.
Most of Haiku 5.5’s time came from the grading task, where it used more thinking tokens than the other models.
| Model | Grading thinking tokens (runs 1, 2, 3) | Average grading time |
|---|---|---|
| Haiku 5.5 | 20,160 / 27,199 / 36,609 | 133 seconds |
| Haiku 4.5 | 6,581 / 14,965 / 21,317 | 149 seconds |
| Sonnet 5.5 | 1,494 / 1,825 / 1,258 | 33 seconds |
| GPT-6 Luna | 1,254 / 0 / 562 | 29 seconds |
- Even on medium, Haiku 5.5 thought 13 to 29 times more than Sonnet 5.5. Its lower unit price still kept its cost under a quarter of Sonnet’s.
- On the other tasks, Haiku 5.5 took 11 to 26 seconds per task. It was slow only on the grading task.
- Thinking tokens do not fully explain the timing. Haiku 4.5 used fewer thinking tokens but took longer on the grading task.
- If speed matters, it is worth lowering effort to low and measuring again. I did not measure that in this round.
8. Limits of this experiment
- Three runs per task make the sample small. Models with overlapping run ranges cannot be ranked.
- Four of the six tasks were perfect for every model, so the test had little power to tell models apart. The next round will add more tasks that require judgment.
- I measured models and tools together. Calling the same model directly through the API removes both tools’ default instructions, so the cost comparison could change. I did not measure that this time.
- I revised the rubric after seeing Haiku 4.5’s wrong answers in an earlier experiment, so it may favor Haiku 4.5.
- Claude (Opus 5.5) wrote the wording and materials for the five tasks other than grading. I did not check which side, if either, this favors.
- I measured reasoning effort only at medium, and even at the same medium setting, the companies think different amounts.
9. Summary
- Haiku 5.5 and GPT-6 Luna scored about the same on these tasks, and GPT-6 Luna led on cost and time.
- Under these test conditions (Claude Code, six tasks), Haiku 5.5 cost about one-tenth of Haiku 4.5. The official announcement’s average reduction was about 75%.
- I will rerun this benchmark with the same tasks each time a new model comes out and add the results below this post.
Frequently asked questions
- Which is better for work, GPT-6 Luna or Claude Haiku 5.5?
- Averages were close: Haiku 5.5 scored 94.8 and GPT-6 Luna 93.3, with overlapping run ranges. Both were perfect on extraction, classification, table analysis, and code editing. GPT-6 Luna was cheaper ($0.013 vs. $0.050 per pass) and faster (98 vs. 214 seconds).
- How much does switching from Haiku 4.5 to Haiku 5.5 cut costs?
- In the same six tasks run through Claude Code, one pass cost $0.479 on Haiku 4.5 and $0.050 on Haiku 5.5 at API-equivalent prices, a drop of about 90%. Average scores were similar at 95.4 and 94.8.
BuildnWrite helps teams build AI agents that keep running. About BuildnWrite ›