Which AI Model for Work Automation? GPT-6 Luna vs Claude Haiku 5.5 on 11 Real Tasks
11 real work tasks, from diagrams and slides to browser work, ran through five GPT and Claude models, three runs each, with scores and costs.
2 posts tagged llm-benchmark
11 real work tasks, from diagrams and slides to browser work, ran through five GPT and Claude models, three runs each, with scores and costs.
I ran five AI models through six real work tasks, 90 runs total. Four tasks were perfect for all; most gaps came from self-introduction grading.