Which AI Model for Work Automation? GPT-6 Luna vs Claude Haiku 5.5 on 11 Real Tasks
11 real work tasks, from diagrams and slides to browser work, ran through five GPT and Claude models, three runs each, with scores and costs.
The same task built with several tools, measured on cost and time with the same yardstick.
11 real work tasks, from diagrams and slides to browser work, ran through five GPT and Claude models, three runs each, with scores and costs.
I ran five AI models through six real work tasks, 90 runs total. Four tasks were perfect for all; most gaps came from self-introduction grading.
I ran three web tasks nine times each on Aside with Claude and GPT models. Three models scored perfectly; the cheapest, Luna, missed twice on shopping.
I built the same 5 slides seven ways and measured cost, time, and spec compliance. The most expensive path followed the spec the least.