Which AI Model for Work Automation? GPT-6 Luna vs Claude Haiku 5.5 on 11 Real Tasks
Who this is forOffice workers and solo business owners who want to hand repetitive work like diagrams, documents, blog drafts, and browser tasks to AI and need to know whether cheaper models are good enough.
TL;DR: I gave 11 real work tasks, including diagrams, slides, documents, blog drafts, and browser work, to five models, three times each. Haiku 5.5 and Sonnet 5.5 met the set specs in all 33 runs, and Haiku 5.5 cost about 1/12.6 as much as Sonnet 5.5. But when one rater scored the drawings, charts, and documents blind (model names hidden), the cheapest model, GPT-6 Luna, scored highest. For spec-checking automation, Haiku 5.5 was the most economical option in this experiment. For visual drafts a person will polish, GPT-6 Luna was the most economical.
Contents
- Two Models That Were Close on Small Tasks
- Experiment Design
- Results: Machine Scores and Human Scores Diverged
- Which Model to Choose for Each Situation
- Outputs Side by Side
- Where Things Went Wrong
- Cost, Time, and Performance
- Limitations of This Experiment
- Summary
1. Two Models That Were Close on Small Tasks
In the earlier experiment, I compared Haiku 5.5 and GPT Luna on six small judgment tasks: extracting presentation text, classifying inquiries, and fixing code. Four of the six tasks gave all five models a perfect score, so I couldn’t separate Haiku 5.5 and GPT-6 Luna on points.
This time I changed the question: do the results hold up on the work I repeatedly hand to AI, where the output is a drawing, a document, or browser actions?
- I counted the 334 conversations I had with AI over the past 90 days by task type, and picked the tasks I do most often.
- I dropped card-news posts (multi-slide social posts) because there were only four.
- The models are the same five as in the earlier experiment: Claude Haiku 5.5, Haiku 4.5, Sonnet 5.5, GPT-6 Luna, and GPT-5.6 Luna.
2. Experiment Design
Each model ran through its company’s agent tool. Claude used Claude Code (claude -p), GPT used Codex (codex exec), and the browser tasks ran through the Aside browser, which can connect models from both companies.
- PTarget
- 11 real work tasks: 5 visual outputs, 3 writing, 3 browser tasks
- IChanged
- 5 models: Claude Haiku 5.5, Haiku 4.5, Sonnet 5.5, GPT-6 Luna, GPT-5.6 Luna
- CBaseline
- GPT-6 Luna, the cheapest of the five models
- OMeasured
- Machine score (100 points), human score (5 points, visual tasks), cost (API-equivalent dollars), time (seconds)
- TPeriod and runs
- October 8, 2026, 3 runs per task, 165 runs total
- Not controlled
- Human scoring by one rater; each model ran on different tools; task prompts written by Claude (Opus 5.5)
| Track | Task | Machine check |
|---|---|---|
| Visual | Diagram (flowchart) | Nodes and connectors, overlapping text, whether lines cross text |
| Visual | Lecture slides, 5 pages | Overflow beyond the slide, minimum font size, outline-style sentences |
| Visual | One web page | Horizontal scrolling at phone width, button height, required elements |
| Visual | Data chart | Values of 48 points, whether values match bar heights |
| Visual | External document (PDF) | Page count, words broken across lines, numbers not in the source |
| Writing | Threads post | Character count, link placement, tone, numbers not in the source |
| Writing | Blog draft | Whether it passes this blog’s publishing gate |
| Writing | Web research | 5 official-documentation values and their sources |
| Browser | Search, shopping comparison, booking form | Compare against correct values and submitted values |
- The visual tasks were opened in a real browser, and I measured the position and size of every character.
- Cost was calculated from the official price lists from Anthropic and OpenAI (checked October 8). The runs used subscriptions, so nothing was actually billed. Every cost is what it would have cost via the API.
3. Results: Machine Scores and Human Scores Diverged
| Model | Spec score | Perfect runs | Human score | 11-task cost | 11-task time |
|---|---|---|---|---|---|
| Claude Haiku 5.5 | 100.0 | 33/33 | 3.8 | $0.117 | 365 s |
| Claude Sonnet 5.5 | 100.0 | 33/33 | 3.6 | $1.473 | 297 s |
| Claude Haiku 4.5 | 96.9 | 20/33 | 2.6 | $1.136 | 723 s |
| GPT-5.6 Luna | 96.7 | 23/33 | 3.8 | $0.125 | 334 s |
| GPT-6 Luna | 92.5 | 19/33 | 4.6 | $0.044 | 281 s |
- Haiku 5.5 and Sonnet 5.5 both completed all 33 runs of the 11 tasks within spec.
- GPT-6 Luna broke the spec most often (14 of 33 runs), but had the highest human score at 4.6.
- On browser tasks alone, the four models other than GPT-6 Luna scored a perfect score in all 9 runs.
4. Which Model to Choose for Each Situation
- Cases I’d advise against: using Haiku 4.5 for visual outputs. It passed the spec in only 4 of 15 runs, and its human score was 2.6.
- For spec-checking automation, Sonnet 5.5 matched Haiku 5.5’s spec results, but cost about 12.6x more.
- GPT-6 Luna had the lowest spec score of the five models on writing tasks (including web research), so I did not extend this recommendation to draft writing.
5. Outputs Side by Side
The figures below show the first-run outputs, labeled with spec scores and human scores.
- Haiku 5.5’s flowchart got a perfect score on both spec and human rating.
- Haiku 4.5’s flowchart earned 1 point because its arrows were scattered apart on the left side of the nodes. Each connector was labeled with the node it linked, so the machine check missed this flaw.
6. Where Things Went Wrong
Where spec points were lost differed by model.
| Model | Where points were mostly lost (times deducted out of 3 runs per task) |
|---|---|
| GPT-6 Luna | Flowchart line crossed text 3 times, web research 3 times (3 blanks, 2 wrong answers), shopping comparison 2 times |
| GPT-5.6 Luna | Flowchart line crossed text 2 times, slide text overlap 2 times, web research 3 times (2 blanks, 1 wrong answer) |
| Claude Haiku 4.5 | Phone-screen text under 14px 3 times, slide text under 28px 2 times, flowchart lines 2 times |
In web research, the output price of GPT-5.6 Luna (official $1.20) was recorded as $0.60 by GPT-6 Luna in 2 of 3 runs, and by GPT-5.6 Luna itself in 1 run.
"q2": {"answer": "0.60", "source_url": "https://developers.openai.com/api/docs/pricing"}(GPT-6 Luna, web research run 2 answer file)
- Both Luna models left the answer blank rather than invent it when they couldn’t find “the minimum version needed to use Haiku 5.5 in Claude Code.” That’s what the instructions asked for.
- The two wrong GPT-6 Luna answers in the shopping comparison were on the same task that failed in the earlier Aside experiment.
7. Cost, Time, and Performance
- Haiku 4.5 cost about 10x as much as Haiku 5.5. Its input, cache, and output prices are all 10x Haiku 5.5’s.
- For time, GPT-6 Luna (281 s) and Sonnet 5.5 (297 s) were fastest, and Haiku 4.5 (723 s) was slowest.
- One run of Sonnet 5.5’s chart task (run 2) stalled mid-run when the tool stopped and restarted, so its cost record was lost. I replaced that single run with a re-run value.
8. Limitations of This Experiment
- Human scores come from one rater who looked at one set of first-run outputs for each visual task. With five human scores per model, the gap between GPT-6 Luna’s 4.6 and Haiku 5.5’s 3.8 adds up to 4 points in total. Writing and browser tasks have no human scores.
- Machine scoring only measures spec compliance. It missed flaws like Haiku 4.5’s flowchart, where the connector labels were correct but the arrows were detached.
- While auditing the scorer, I found and fixed two bugs and rescored. One was reading only the first number in web search answers. The other was judging “2 pages or fewer” in the document task as “exactly 2 pages.”
- I measured models together with their execution tools. Calling the same models directly through the API could produce different results and costs, which I did not measure this time.
- I measured only one reasoning effort level, medium. There’s no guarantee that medium means the same level at the two companies.
9. Summary
- Even with the same model, “did it follow the spec?” and “does it look good?” gave different results. You need to measure both to choose well.
- If you’re using Haiku 4.5, switching to Haiku 5.5 could improve both spec compliance and cost.
Frequently asked questions
- Which is better for AI work automation, GPT Luna or Claude Haiku?
- Haiku 5.5 met the spec checks in all 33 runs; GPT-6 Luna in 19. On blind human ratings of the five visual tasks (5-point scale, one rater), GPT-6 Luna scored 4.6 and Haiku 5.5 scored 3.8. One pass of all 11 tasks cost $0.044 with GPT-6 Luna and $0.117 with Haiku 5.5 (measured October 8, 2026).
- How much do the work results differ between Claude Haiku 5.5 and Sonnet 5.5?
- Both models completed all 33 runs of the 11 tasks within spec. One pass cost $0.117 for Haiku 5.5 and $1.473 for Sonnet 5.5, so Haiku 5.5 cost about 1/12.6 of Sonnet 5.5.
BuildnWrite helps teams build AI agents that keep running. About BuildnWrite ›