Experiments

GPT-6 Luna vs Claude Haiku 5.5: Results From Six Real Work Tasks

11 min read#llm-benchmark#ai-model-comparison#claude-haiku-5-5#gpt-6-luna#claude-code

Who this is forOffice workers and solo business owners who want to hand repetitive tasks like summarizing, classifying, and grading to a cheap AI model and want numbers to pick between Claude Haiku 5.5 and GPT Luna.

TL;DR: Two GPT Luna models, Claude Haiku 5.5, Haiku 4.5, and Sonnet 5.5 each got the same six work tasks, three runs per task. Extraction, classification, table analysis, and code editing were perfect for all five models, and most score differences came from self-introduction essay grading. For tasks with fixed right answers, the cheapest GPT-6 Luna or Haiku 5.5 was enough. On grading, the model that matched the intended grades most often was Sonnet 5.5. Under these test conditions, switching from Haiku 4.5 to Haiku 5.5 cut costs to about one-tenth.

Contents

  1. Official benchmarks and my work are different
  2. Experiment design
  3. Results: four tasks all scored perfectly
  4. Which one to pick in which situation
  5. Where the gap appeared: self-introduction essay grading
  6. Mistakes in the rule-based writing task
  7. Cost and time
  8. Limits of this experiment
  9. Summary

1. Official benchmarks and my work are different

On October 7, 2026, Anthropic announced Claude Haiku 5.5 and chose OpenAI’s GPT-6 Luna as its comparison. In Anthropic’s table, Haiku 5.5 scored higher than GPT-6 Luna on every item with a public score. I covered the announcement and pricing separately in the Haiku 5.5 launch summary.

Benchmark table from Anthropic's announcement. In order of Haiku 5.5, Haiku 4.5, GPT-6 Luna, and Sonnet 5.5 (for reference): GDPval-AA v2.1 scores are 1620, 735, 1437, and 1840; OSWorld 2.1 offline subset scores are 72.4%, 15.7%, 48.9%, and 83.9%; Terminal-Bench 4.0 scores are 39.2%, 0.0%, 16.4%, and 70.6%.
In the official table, Haiku 5.5 (pink) scores higher than GPT-6 Luna on every public item.

The official benchmarks measure general ability. Someone who wants to hand repetitive work to a cheap model needs answers to different questions:

  • Does the ranking hold on the kind of work I do every day?
  • How much does one such job actually cost, and how long does it take?

So I fixed six tasks that solo educators and content companies repeat in practice, and built the BuildnWrite Work Standard Benchmark, which reruns the same tasks every time a new model comes out. This post is the first round.

2. Experiment design

The unit of comparison is not the model alone but the model together with its company’s agent tool. I ran the Claude models through Claude Code (claude -p) and the GPT models through Codex (codex exec). Both run a job to completion in one go without anyone watching, which matches how repetitive work is actually handed off.

PTarget
Six work tasks: extraction, classification, grading, table analysis, writing, code
IChanged
Five models and each company's execution tool: Claude Haiku 5.5, Haiku 4.5, Sonnet 5.5 (claude -p), GPT-6 Luna, GPT-5.6 Luna (codex exec)
CComparison baseline
GPT-6 Luna, the model Anthropic chose as Haiku 5.5's comparison in its announcement
OMeasured
Accuracy (checked against answer keys by a grading script, 100 points max), time (seconds, from run start until the tool finished), and cost (tokens reported by the tool multiplied by official API prices, in USD)
TPeriod and runs
October 8, 2026; three runs per task; 90 runs in total; no retries
Things I could not control
The size of each tool's default instructions (about 29,000 tokens for Claude Code and about 14,000 for Codex, so Claude's cost is disadvantaged on smaller tasks); the grading rubric, which I revised after seeing Haiku 4.5's wrong answers (may favor Haiku 4.5); all five models ran at the same time, so timing includes server congestion; each run used a new session, so Claude runs paid the cache-write cost every time (disadvantages Claude); the task wording was written by Claude (Opus 5.5), direction unverified; and medium reasoning effort does not mean the same thing at the two companies
Same task files and same instructions, medium reasoning effort (Haiku 4.5 doesn't support it), an empty environment without personal settings or external connected tools, and a new working folder for each run.

All task materials are synthetic, so there is no real customer information.

Task What it does How correctness is judged
Extraction Pull 17 fields (price, date, specs) from a fictional product announcement Exact match per field. Leaving out values not in the announcement is the correct answer
Classification Sort 30 inquiry emails into six categories such as courses and consulting Match rate with the correct labels
Grading Score 30 self-introduction essays from 1 to 5 using a rubric Match rate with the grades set before writing
Table analysis Answer eight questions about a performance table with duplicate rows and comma-formatted numbers Exact match on numbers
Rule-based writing Turn a blog paragraph into a Threads post while following 10 rules 10 automated checks: character count, banned words, copying the original text, and so on
Code fix Fix a UTM link generator with three bugs Pass rate on 8 hidden tests

I calculated costs from both companies’ official price pages (Anthropic, OpenAI, checked October 8). I ran everything on subscriptions, so nothing was actually billed. All costs below are what the runs would have cost through the API.

3. Results: four tasks all scored perfectly

Scores by task, 3-run average out of 100. A grid where only the cells that are not perfect are dark gray. Sonnet 5.5: extraction 100, classification 100, grading 98, table analysis 100, writing 93, code 100. Haiku 4.5: grading 72, the rest 100. Haiku 5.5: grading 76, writing 93, the rest 100. GPT-6 Luna: grading 63, writing 97, the rest 100. GPT-5.6 Luna: grading 43, the rest 100.
Extraction, classification, table analysis, and code were perfect for all five models in all three runs.
Model Average score Range across runs 6-task cost 6-task time
Claude Sonnet 5.5 98.5 97.8 to 99.5 $0.390 119 seconds
Claude Haiku 4.5 95.4 92.8 to 99.5 $0.479 383 seconds
Claude Haiku 5.5 94.8 93.9 to 95.6 $0.050 214 seconds
GPT-6 Luna 93.3 92.8 to 94.5 $0.013 98 seconds
GPT-5.6 Luna 90.5 90.5 (same in all 3 runs) $0.045 135 seconds
  • Most of the average score difference came from the grading task, and mistakes in rule-based writing added some more. Haiku 5.5 scored higher than Haiku 4.5 on grading (75.6 vs. 72.2), but two writing mistakes pulled its average below Haiku 4.5’s.
  • The run ranges for Haiku 5.5 and GPT-6 Luna overlap, and so do those for Haiku 5.5 and Haiku 4.5. This sample cannot say which one is better.
  • Two things are fairly clear. Across all 90 grading items, Sonnet 5.5 matched 88, the most (Haiku 5.5 was next with 68), and GPT-5.6 Luna scored 43 points in all three grading runs. However, Haiku 4.5 matched 29 of 30 items in grading run 3, so looking at single runs, it reaches the range of Sonnet.

4. Which one to pick in which situation

Which one to pick in which situation. If you run tasks on Haiku 4.5, move them to Haiku 5.5: similar scores and about one-tenth the six-task cost. If the work has fixed right answers, like extraction, classification, or table math, use GPT-6 Luna or Haiku 5.5: all five models scored perfectly in 12 runs, at about $0.002 and $0.004 per run. If the work means reading criteria and assigning scores, use Sonnet 5.5: 88 of 90 items matched the intended grades on this rubric. If none of these applies, measure with your own work: this is a small experiment with six tasks run three times each.
I split the choices only where the numbers in this experiment actually diverged.
  • Not recommended: when handing grading to this rubric, GPT-5.6 Luna scored 43 points in all three runs.
  • The Sonnet 5.5 result rests on one rubric and one task. Results on other rubrics may differ.

5. Where the gap appeared: self-introduction essay grading

For the grading task, I gave the models 30 synthetic self-introduction essays (the standard application essay for Korean job seekers), six for each grade from 1 to 5, along with a rubric, and asked them to score each item. I set the grades before writing the essays. The rubric includes the instruction: “To decide between two adjacent scores, use the criteria below. If the condition is met, give the higher score.”

Grading task hits by grade, out of 18 total across 3 runs. A grid where only the cells under half are dark gray. Sonnet 5.5: grade 1 essays 17, grade 2 17, grade 3 18, grade 4 18, grade 5 18. Haiku 4.5: 18, 13, 8, 10, 16. Haiku 5.5: 13, 11, 8, 18, 18. GPT-6 Luna: 0, 6, 15, 18, 18. GPT-5.6 Luna: 1, 1, 1, 18, 18.
Both Luna models missed most on low-grade essays, and Haiku 4.5 missed many mid-grade essays. Haiku 5.5 matched only 8 of 18 grade 3 essays.

The direction of the wrong answers differed by model.

Model Scored too high (out of 90) Scored too low (out of 90)
Sonnet 5.5 2 0
Haiku 5.5 22 0
Haiku 4.5 0 25
GPT-6 Luna 33 0
GPT-5.6 Luna 51 0
  • GPT-5.6 Luna scored 51 of its 54 essays from grades 1 to 3 too high, and 2 of those were two steps too high.
  • GPT-6 Luna scored all 18 grade 1 essays too high and 12 grade 2 essays too high. It matched 15 of 18 grade 3 essays.
  • Haiku 5.5 also leaned high: 5 grade 1 essays, 7 grade 2 essays, and 10 grade 3 essays. It matched only 8 grade 3 essays, the same as Haiku 4.5.
  • Haiku 4.5 went the other way and scored too low, with 10 grade 3 essays and 8 grade 4 essays, roughly half the cases in each grade. Its run 3 matched 29 items, so its swings between runs were large.
  • Grade 4 and grade 5 essays were matched by all models except Haiku 4.5.

Looking at one grade 1 essay shows why the Luna scores still make sense against the rubric wording:

I will work hard. I would be grateful if you hired me.

(Answer to the motivation question in synthetic self-introduction essay E01, correct grade 1)

The rubric’s 2-point criterion reads “the answer points in the right direction for what the question asks; even if it is short, if it answers, it is 2 points.” The 1-point criterion reads “unrelated to the question or evasive.” For this answer, both Luna models gave 2 points in all three runs, Sonnet 5.5 gave 1 point in all three, and Haiku 5.5 gave 1 point in two of three runs. So this task does not measure grading ability as such. It measures how closely a model matches the grade the rubric’s author intended. That agreement is what matters when you hand bulk grading to a model using human-written criteria.

6. Mistakes in the rule-based writing task

In the rule-based writing task, out of 15 runs, the five models followed all 10 rules in most cases. The five imperfect runs each broke only one rule, and the mistakes came in two kinds:

  • Copied 20 or more characters of the original sentence verbatim: Haiku 5.5 twice, Sonnet 5.5 twice
  • Used the banned word meaning “you all”: GPT-6 Luna once
Raw file of the Threads post written by Haiku 5.5. Under a first line reading 'What happens when you hand meeting notes to AI in full,' there is a three-step numbered list, a sentence saying meeting-note cleanup time dropped to within 15 minutes per meeting, and a last line asking how your team handles meeting notes.
Haiku 5.5 run 1 original. The phrase "cut to within 15 minutes per meeting" is copied verbatim from the source text.
Raw file of the Threads post written by GPT-6 Luna. It opens with the first line 'Meeting notes matter more for action than for summaries,' followed by five lines, and ends with a question asking which task you most want to cut down after a meeting.
GPT-6 Luna run 2 original. The banned word "you all" appears in the last line.

From the same source text and the same rules, the length and structure of the posts differed by model. Here is the opening of Sonnet 5.5’s run 1 post:

Meeting-note AI gets fuzzy if you only ask for a summary Summarizing the whole transcript makes it lose who does what and by when. So I split the work into three steps.

(Claude Sonnet 5.5, rule-based writing run 1 output)

I left humor and persuasiveness out of the scores because a machine cannot measure them. The score only says whether the rules were followed.

7. Cost and time

Six-task cost per full pass, API-equivalent, 3-run average, drawn as horizontal bars starting at 0. Haiku 4.5 $0.479, Sonnet 5.5 $0.390, Haiku 5.5 $0.050 (blue), GPT-5.6 Luna $0.045, GPT-6 Luna $0.013.
Haiku 4.5 cost even more than Sonnet 5.5, a higher-tier model.
  • Haiku 5.5’s cost per pass was about 90% lower than Haiku 4.5’s ($0.479 down to $0.050). That is a bigger cut than the average reduction of about 75% stated in the announcement. All the requests in this set were 100,000 tokens or fewer, the range where the discount is largest.
  • Haiku 4.5 cost more than Sonnet 5.5 because of output length. One pass produced about 40,000 output tokens (about 31,000 of them thinking), nearly six times Sonnet 5.5’s roughly 6,900, and both models price cache reads at $0.10.
  • GPT-6 Luna cost about a quarter of Haiku 5.5. Claude Code’s default instructions (about 29,000 tokens), attached to every request, are part of that difference.
Six-task time per full pass, 3-run average, drawn as horizontal bars starting at 0. Haiku 4.5 383 seconds, Haiku 5.5 214 seconds (blue), GPT-5.6 Luna 135 seconds, Sonnet 5.5 119 seconds, GPT-6 Luna 98 seconds.
On this task set and the medium setting, Haiku 5.5 was slower than Sonnet 5.5.

Most of Haiku 5.5’s time came from the grading task, where it used more thinking tokens than the other models.

Model Grading thinking tokens (runs 1, 2, 3) Average grading time
Haiku 5.5 20,160 / 27,199 / 36,609 133 seconds
Haiku 4.5 6,581 / 14,965 / 21,317 149 seconds
Sonnet 5.5 1,494 / 1,825 / 1,258 33 seconds
GPT-6 Luna 1,254 / 0 / 562 29 seconds
  • Even on medium, Haiku 5.5 thought 13 to 29 times more than Sonnet 5.5. Its lower unit price still kept its cost under a quarter of Sonnet’s.
  • On the other tasks, Haiku 5.5 took 11 to 26 seconds per task. It was slow only on the grading task.
  • Thinking tokens do not fully explain the timing. Haiku 4.5 used fewer thinking tokens but took longer on the grading task.
  • If speed matters, it is worth lowering effort to low and measuring again. I did not measure that in this round.

8. Limits of this experiment

  • Three runs per task make the sample small. Models with overlapping run ranges cannot be ranked.
  • Four of the six tasks were perfect for every model, so the test had little power to tell models apart. The next round will add more tasks that require judgment.
  • I measured models and tools together. Calling the same model directly through the API removes both tools’ default instructions, so the cost comparison could change. I did not measure that this time.
  • I revised the rubric after seeing Haiku 4.5’s wrong answers in an earlier experiment, so it may favor Haiku 4.5.
  • Claude (Opus 5.5) wrote the wording and materials for the five tasks other than grading. I did not check which side, if either, this favors.
  • I measured reasoning effort only at medium, and even at the same medium setting, the companies think different amounts.

9. Summary

  • Haiku 5.5 and GPT-6 Luna scored about the same on these tasks, and GPT-6 Luna led on cost and time.
  • Under these test conditions (Claude Code, six tasks), Haiku 5.5 cost about one-tenth of Haiku 4.5. The official announcement’s average reduction was about 75%.
  • I will rerun this benchmark with the same tasks each time a new model comes out and add the results below this post.

Frequently asked questions

Which is better for work, GPT-6 Luna or Claude Haiku 5.5?
Averages were close: Haiku 5.5 scored 94.8 and GPT-6 Luna 93.3, with overlapping run ranges. Both were perfect on extraction, classification, table analysis, and code editing. GPT-6 Luna was cheaper ($0.013 vs. $0.050 per pass) and faster (98 vs. 214 seconds).
How much does switching from Haiku 4.5 to Haiku 5.5 cut costs?
In the same six tasks run through Claude Code, one pass cost $0.479 on Haiku 4.5 and $0.050 on Haiku 5.5 at API-equivalent prices, a drop of about 90%. Average scores were similar at 95.4 and 94.8.