Which AI Model for Work Automation? GPT-6 Luna vs Claude Haiku 5.5 on 11 Real Tasks
11 real work tasks, from diagrams and slides to browser work, ran through five GPT and Claude models, three runs each, with scores and costs.
3 posts tagged model-comparison
11 real work tasks, from diagrams and slides to browser work, ran through five GPT and Claude models, three runs each, with scores and costs.
I ran three web tasks nine times each on Aside with Claude and GPT models. Three models scored perfectly; the cheapest, Luna, missed twice on shopping.
Fable 5 costs twice as much per token as Opus 4.8, but measured sessions cost 1.0 to 1.4 times as much. See the token, cost, and intervention data behind it.