Experiments

Which AI Model for Work Automation? GPT-6 Luna vs Claude Haiku 5.5 on 11 Real Tasks

8 min read#ai-automation#model-comparison#claude-haiku-5-5#gpt-6-luna#llm-benchmark

Who this is forOffice workers and solo business owners who want to hand repetitive work like diagrams, documents, blog drafts, and browser tasks to AI and need to know whether cheaper models are good enough.

TL;DR: I gave 11 real work tasks, including diagrams, slides, documents, blog drafts, and browser work, to five models, three times each. Haiku 5.5 and Sonnet 5.5 met the set specs in all 33 runs, and Haiku 5.5 cost about 1/12.6 as much as Sonnet 5.5. But when one rater scored the drawings, charts, and documents blind (model names hidden), the cheapest model, GPT-6 Luna, scored highest. For spec-checking automation, Haiku 5.5 was the most economical option in this experiment. For visual drafts a person will polish, GPT-6 Luna was the most economical.

Contents

  1. Two Models That Were Close on Small Tasks
  2. Experiment Design
  3. Results: Machine Scores and Human Scores Diverged
  4. Which Model to Choose for Each Situation
  5. Outputs Side by Side
  6. Where Things Went Wrong
  7. Cost, Time, and Performance
  8. Limitations of This Experiment
  9. Summary

1. Two Models That Were Close on Small Tasks

In the earlier experiment, I compared Haiku 5.5 and GPT Luna on six small judgment tasks: extracting presentation text, classifying inquiries, and fixing code. Four of the six tasks gave all five models a perfect score, so I couldn’t separate Haiku 5.5 and GPT-6 Luna on points.

This time I changed the question: do the results hold up on the work I repeatedly hand to AI, where the output is a drawing, a document, or browser actions?

  • I counted the 334 conversations I had with AI over the past 90 days by task type, and picked the tasks I do most often.
  • I dropped card-news posts (multi-slide social posts) because there were only four.
  • The models are the same five as in the earlier experiment: Claude Haiku 5.5, Haiku 4.5, Sonnet 5.5, GPT-6 Luna, and GPT-5.6 Luna.

2. Experiment Design

Each model ran through its company’s agent tool. Claude used Claude Code (claude -p), GPT used Codex (codex exec), and the browser tasks ran through the Aside browser, which can connect models from both companies.

PTarget
11 real work tasks: 5 visual outputs, 3 writing, 3 browser tasks
IChanged
5 models: Claude Haiku 5.5, Haiku 4.5, Sonnet 5.5, GPT-6 Luna, GPT-5.6 Luna
CBaseline
GPT-6 Luna, the cheapest of the five models
OMeasured
Machine score (100 points), human score (5 points, visual tasks), cost (API-equivalent dollars), time (seconds)
TPeriod and runs
October 8, 2026, 3 runs per task, 165 runs total
Not controlled
Human scoring by one rater; each model ran on different tools; task prompts written by Claude (Opus 5.5)
Same task files, same instructions, reasoning effort set to medium, and a blank environment without personal settings.
Track Task Machine check
Visual Diagram (flowchart) Nodes and connectors, overlapping text, whether lines cross text
Visual Lecture slides, 5 pages Overflow beyond the slide, minimum font size, outline-style sentences
Visual One web page Horizontal scrolling at phone width, button height, required elements
Visual Data chart Values of 48 points, whether values match bar heights
Visual External document (PDF) Page count, words broken across lines, numbers not in the source
Writing Threads post Character count, link placement, tone, numbers not in the source
Writing Blog draft Whether it passes this blog’s publishing gate
Writing Web research 5 official-documentation values and their sources
Browser Search, shopping comparison, booking form Compare against correct values and submitted values
  • The visual tasks were opened in a real browser, and I measured the position and size of every character.
  • Cost was calculated from the official price lists from Anthropic and OpenAI (checked October 8). The runs used subscriptions, so nothing was actually billed. Every cost is what it would have cost via the API.
Blind rating screen. Under the diagram task, flowchart outputs A through D sit side by side. Under each one is a score button from 1 to 5, with a scale running from unusable to usable as is.
On the rating screen, only the letters A through E appear instead of model names.

3. Results: Machine Scores and Human Scores Diverged

Model Spec score Perfect runs Human score 11-task cost 11-task time
Claude Haiku 5.5 100.0 33/33 3.8 $0.117 365 s
Claude Sonnet 5.5 100.0 33/33 3.6 $1.473 297 s
Claude Haiku 4.5 96.9 20/33 2.6 $1.136 723 s
GPT-5.6 Luna 96.7 23/33 3.8 $0.125 334 s
GPT-6 Luna 92.5 19/33 4.6 $0.044 281 s
Dot plots in two panels, top and bottom, placing machine scores and human scores side by side. Machine scores (100-point scale, x-axis from 88 to 100): Haiku 5.5 100.0, Sonnet 5.5 100.0, Haiku 4.5 96.9, GPT-5.6 Luna 96.7, GPT-6 Luna 92.5. Human scores (5-point scale): GPT-6 Luna 4.6, Haiku 5.5 3.8, GPT-5.6 Luna 3.8, Sonnet 5.5 3.6, Haiku 4.5 2.6.
GPT-6 Luna, last on machine scores, had the highest human score. The machine-score axis starts at 88.
  • Haiku 5.5 and Sonnet 5.5 both completed all 33 runs of the 11 tasks within spec.
  • GPT-6 Luna broke the spec most often (14 of 33 runs), but had the highest human score at 4.6.
  • On browser tasks alone, the four models other than GPT-6 Luna scored a perfect score in all 9 runs.

4. Which Model to Choose for Each Situation

Which model to choose for each situation. If you're building automation that checks specs with a script, choose Haiku 5.5: all 33 runs passed spec, and cost is about 1/12.6 of Sonnet's. If you need drawings, charts, or document drafts that a person will polish, and cost must be lowest, choose GPT-6 Luna: human score 4.6 (one rater), $0.044, and check the specs yourself. If you're running tasks on Haiku 4.5, move to Haiku 5.5: perfect runs rise from 20 to 33, and cost is about one-tenth. If none of these apply, measure on your own work: this is a small experiment of 11 tasks, run 3 times each.
I split the picks between spec-checking automation and visual drafts a person will polish. Human scores are based on one rater.
  • Cases I’d advise against: using Haiku 4.5 for visual outputs. It passed the spec in only 4 of 15 runs, and its human score was 2.6.
  • For spec-checking automation, Sonnet 5.5 matched Haiku 5.5’s spec results, but cost about 12.6x more.
  • GPT-6 Luna had the lowest spec score of the five models on writing tasks (including web research), so I did not extend this recommendation to draft writing.

5. Outputs Side by Side

The figures below show the first-run outputs, labeled with spec scores and human scores.

Diagram task outputs from five models. Haiku 5.5: machine score 100, human score 5. Sonnet 5.5: machine score 100, human score 4. Haiku 4.5: machine score 92, human score 1; arrows are scattered on the left, away from the nodes. GPT-6 Luna: machine score 92, human score 4; a return line crosses a 'critical issue' label. GPT-5.6 Luna: machine score 92, human score 2; thick arrowheads cover the nodes, and a curve crosses the text FAIL.
The Haiku 4.5 diagram had a spec score of 92 even though its arrows float away from the nodes.
  • Haiku 5.5’s flowchart got a perfect score on both spec and human rating.
  • Haiku 4.5’s flowchart earned 1 point because its arrows were scattered apart on the left side of the nodes. Each connector was labeled with the node it linked, so the machine check missed this flaw.
Data chart task outputs from five models. Monthly visit line chart for four channels. Haiku 5.5, GPT-6 Luna, and GPT-5.6 Luna: human score 5. Sonnet 5.5 and Haiku 4.5: human score 2. The Haiku 4.5 chart has lines overflowing above the graph.
The Sonnet chart had all numbers correct, yet its human score was 2.
External document task outputs: five models' course introduction PDFs (a training program brochure). Haiku 5.5 and Sonnet 5.5 are two pages long, with human scores of 2 and 3. Haiku 4.5, GPT-6 Luna, and GPT-5.6 Luna are one page each, all with human score 5.
All three documents condensed to one page scored 5.

6. Where Things Went Wrong

Where spec points were lost differed by model.

Model Where points were mostly lost (times deducted out of 3 runs per task)
GPT-6 Luna Flowchart line crossed text 3 times, web research 3 times (3 blanks, 2 wrong answers), shopping comparison 2 times
GPT-5.6 Luna Flowchart line crossed text 2 times, slide text overlap 2 times, web research 3 times (2 blanks, 1 wrong answer)
Claude Haiku 4.5 Phone-screen text under 14px 3 times, slide text under 28px 2 times, flowchart lines 2 times

In web research, the output price of GPT-5.6 Luna (official $1.20) was recorded as $0.60 by GPT-6 Luna in 2 of 3 runs, and by GPT-5.6 Luna itself in 1 run.

"q2": {"answer": "0.60", "source_url": "https://developers.openai.com/api/docs/pricing"}

(GPT-6 Luna, web research run 2 answer file)

  • Both Luna models left the answer blank rather than invent it when they couldn’t find “the minimum version needed to use Haiku 5.5 in Claude Code.” That’s what the instructions asked for.
  • The two wrong GPT-6 Luna answers in the shopping comparison were on the same task that failed in the earlier Aside experiment.

7. Cost, Time, and Performance

Scatter plot putting cost, time, and performance together. The x-axis is the cost of one pass of the 11 tasks on a reversed log scale, so the right side is cheaper. The y-axis is the machine score. The upper-right area of cheap and accurate contains only Haiku 5.5 ($0.117, 365 s, 100 points). Sonnet 5.5 is at the top left: $1.473, 297 s, 100 points. GPT-5.6 Luna: $0.125, 334 s, 96.7 points. Haiku 4.5: $1.136, 723 s, 96.9 points. GPT-6 Luna is at the bottom right: $0.044, 281 s, 92.5 points.
The cost axis is reversed, so the upper right is the better side (the price-axis convention used by Artificial Analysis). The shaded region covers under $0.3 per pass and 98 points or higher. Only Haiku 5.5, which falls inside it, is highlighted in the accent color.
  • Haiku 4.5 cost about 10x as much as Haiku 5.5. Its input, cache, and output prices are all 10x Haiku 5.5’s.
  • For time, GPT-6 Luna (281 s) and Sonnet 5.5 (297 s) were fastest, and Haiku 4.5 (723 s) was slowest.
  • One run of Sonnet 5.5’s chart task (run 2) stalled mid-run when the tool stopped and restarted, so its cost record was lost. I replaced that single run with a re-run value.

8. Limitations of This Experiment

  • Human scores come from one rater who looked at one set of first-run outputs for each visual task. With five human scores per model, the gap between GPT-6 Luna’s 4.6 and Haiku 5.5’s 3.8 adds up to 4 points in total. Writing and browser tasks have no human scores.
  • Machine scoring only measures spec compliance. It missed flaws like Haiku 4.5’s flowchart, where the connector labels were correct but the arrows were detached.
  • While auditing the scorer, I found and fixed two bugs and rescored. One was reading only the first number in web search answers. The other was judging “2 pages or fewer” in the document task as “exactly 2 pages.”
  • I measured models together with their execution tools. Calling the same models directly through the API could produce different results and costs, which I did not measure this time.
  • I measured only one reasoning effort level, medium. There’s no guarantee that medium means the same level at the two companies.

9. Summary

  • Even with the same model, “did it follow the spec?” and “does it look good?” gave different results. You need to measure both to choose well.
  • If you’re using Haiku 4.5, switching to Haiku 5.5 could improve both spec compliance and cost.

Frequently asked questions

Which is better for AI work automation, GPT Luna or Claude Haiku?
Haiku 5.5 met the spec checks in all 33 runs; GPT-6 Luna in 19. On blind human ratings of the five visual tasks (5-point scale, one rater), GPT-6 Luna scored 4.6 and Haiku 5.5 scored 3.8. One pass of all 11 tasks cost $0.044 with GPT-6 Luna and $0.117 with Haiku 5.5 (measured October 8, 2026).
How much do the work results differ between Claude Haiku 5.5 and Sonnet 5.5?
Both models completed all 33 runs of the 11 tasks within spec. One pass cost $0.117 for Haiku 5.5 and $1.473 for Sonnet 5.5, so Haiku 5.5 cost about 1/12.6 of Sonnet 5.5.