BuildnWriteBlogProductAbout BuildnWrite ›

Categories

  • All posts 100
  • Guides 28
  • Concepts 14
  • Insights 54
  • Experiments 4

Topics

  • n8n automation 14
  • Integrations and auth 2
  • Deploy and servers 1
  • Data and APIs 1

#llm-benchmark

2 posts tagged llm-benchmark

  • Experiments

    Which AI Model for Work Automation? GPT-6 Luna vs Claude Haiku 5.5 on 11 Real Tasks

    11 real work tasks, from diagrams and slides to browser work, ran through five GPT and Claude models, three runs each, with scores and costs.

    October 8, 20268 min read

  • Experiments

    GPT-6 Luna vs Claude Haiku 5.5: Results From Six Real Work Tasks

    I ran five AI models through six real work tasks, 90 runs total. Four tasks were perfect for all; most gaps came from self-introduction grading.

    October 8, 202611 min read

← All posts

Hands-on notes on AI agents and work automation.

BuildnWrite · Seoul, South Korea · About us · RSS

© 2026 BuildnWrite. All rights reserved.