Building a Solo Board Game Server in One Weekend with Claude Code: A Retrospective
Who this is forDevelopers and board game fans who want to build a solo game client or server with an AI coding agent and need realistic numbers on cost, effort, and limits.
When young children make it impossible to schedule board game nights, a strategy game player has few options. This project shows what one person can build in a single weekend with Claude Code. The author, a strategy board game player who could no longer meet a regular group, built a web version of Age of Steam in about 36 hours. The result was a playable 12,308-line TypeScript codebase with 50 commits, deployed to a live server at a total API-equivalent cost of about $327. This retrospective covers the timeline, the code and test numbers, the cost breakdown, and the lessons on game design, quality assurance, and LLM-driven game AI. Readers who want to build a solo game with an AI coding agent will get concrete figures and the reasoning behind each decision.
Weekend timeline

Key data
All figures were measured from the git log, session transcripts, and repository files. The aggregation script was analyze_usage_v2.py, kept in the session scratchpad.
Project scale (July 4, 2026 22:30 to July 6, 2026 10:36)
| Item | Measured value |
|---|---|
| Calendar time | About 36 hours (Saturday night setup to Monday morning final sealing commit) |
| Claude Code session time | 13.5 hours (5 main sessions, based on July 5–6 logs) |
| Commits | 50 (July 4: 5, July 5: 37, July 6: 8) |
| Code | TypeScript 12,308 lines (engine 3,717 / ui 4,198 / ai 2,184 / server 707 / scripts 1,359) |
| Tests | 67 cases (engine 9 files + ai 3 files) |
| Runtime dependencies | Only 3: react, react-dom, express |
| Deployment | aos.ggplab.xyz (Mac mini launchd + Cloudflare Tunnel, invite-code gate) |
Token and cost measurements (5 sessions + 12 subagents, snapshot at July 6, 2026 11:10)
| Item | Value |
|---|---|
| API calls (assistant messages) | 1,088 (deduplicated globally by message.id) |
| Input / output | 250K / 869K tokens |
| Cache writes / cache reads | 4.32M / 257.4M tokens |
| Cache hit share | 97.9% of all processed tokens |
API list-price equivalent cost (Fable 5 at $10/$50; cache writes at 1.25× for 5-minute TTL and 2.0× for 1-hour TTL; cache reads at 0.1×):
| Model | Cost | Breakdown |
|---|---|---|
| claude-fable-5 (main) | $279.83 | Cache reads $198.16, cache writes (all 1-hour TTL) $54.77 |
| claude-opus-4-8 | $44.15 | |
| claude-sonnet-5 (subagents) | $3.20 | Introductory price $2/$10 |
| Total | $327.19 | Cache reads $226.41 (69%) + cache writes $65.40 (20%) |
Three-way cross-check: (1) No duplicate message.id values appeared across session files, so resumed sessions were not double-counted. (2) The author’s own re-aggregation gave $327.19. (3) The independent tool ccusage gave $325.80, within 0.4%, with the gap attributed to snapshot timing. The first estimate of $287.52 undercounted because it treated all cache writes as 5-minute TTL at 1.25×. In fact, all Fable cache writes used the 1-hour TTL at 2×. Development sessions were still running while the retrospective was written, so these numbers keep growing.
The author is on a Max subscription, so the actual out-of-pocket spend was KRW 0. Even so, the conversion shows a clear structure: 89% of the cost of 13.5 hours of agent operation came not from new reasoning (output) but from writing context to the cache and reading it back.
Game AI measurements
| Item | Value |
|---|---|
| Difficulty structure | easy (weighted random among heuristic top 3) / normal (heuristic top 1) / hard (LLM hybrid) |
| Self-play evaluation | 6 bot strategy profiles × 1,020 games in total (bot-strategy-eval-g17.json) |
| Weight tuning | About 40 weights in a linear evaluation function tuned over 2 generations with a (1+λ) evolution strategy |
| LLM backend | Switched from codex CLI (local) to Gemini 2.5 Flash (deployed server) |
| Final decision | HARD_DIFFICULTY_ENABLED = false: sealed because LLM hard was weaker than heuristic medium |
Insights
1. A solo game does not need a database: localStorage is enough
The architecture reflects the single-player assumption. The entire game state is stored in localStorage (34 call sites across 10 files, with 8 keys including aos-rustbelt-save-v1), and the server has no database. The server only handles the LLM bridge and writes telemetry as JSONL. As a result, deployment reduced to rsync plus launchd, and the whole category of state-synchronization bugs disappeared. Narrowing down “who will use this service” first removes half the stack.
2. Freeze the references before the code: a single source of truth for rules and decisions
On the first night, the work was not coding but contract freezing. The original rulebook PDF and a survey of existing implementations went into docs/research/. The engine type contract and the map data single source of truth (rust-belt.json) were frozen at the fourth commit. Twenty-two points where rule interpretations diverged were logged in DECISIONS.md with cross-checked sources, including auction details and the timing of ownership loss. These documents became guardrails that kept the agent from wandering over the next two days. For apps with tight domain rules, like board games, creating a single source of truth for the rules first greatly reduces trial and error.
3. Half the commits were QA: polishing over building
More than 28 of the 50 commits were QA batches. The project started on PC, but the author wanted to play while moving, so two rounds of real mobile play QA were run. They found a dock overflow, a white-screen crash in the production phase, and animation issues. The final commit was also “10 individual mobile QA fixes.” With AI coding, feature implementation is fast, but the real-use QA loop still requires a person to play the game. This time was the largest in perceived effort and the most valuable.
4. Fan-made maps: print them if physical, generate and validate if self-hosted
Fan-made maps for Age of Steam were originally playable only by printing the board. On the author’s own server, Codex was given the Rust Belt schema and generated three fan maps (Heartland Plains, Great Lakes North, Ohio Valley). Only maps that passed the coordinate duplicate and cargo growth-chain validator (validate-map.ts) and the self-play full-completion gate (selfplay-map.ts) were accepted. The pattern “the LLM generates, a deterministic validator filters” can be reused not only for maps but for content-style data in general.
5. The wall of LLM game AI: shallow thinking is fragmented, deep thinking is slow
This was the biggest frustration and the biggest takeaway. The advanced difficulty was built with an LLM, but in live play it was weaker than the heuristic medium difficulty, so it was shelved. The author diagnosed four causes, based on the HANDOFF document:
- The hybrid scope was narrow: Because of latency, only three phases (stock, auction, action selection) went to the LLM, while building and movement were answered instantly by heuristics. The strategic thread was broken.
- Reasoning was turned off: codex’s default effort took over 90 seconds in the build phase, so it was forced to
low, and Gemini ran withthinkingBudget: 0. Shallow thinking cannot read multi-layer move sequences. - The state was stateless: Each move built a new prompt, so there was no “plan made three turns ago.” Planning and execution were disconnected.
- The candidate space depended on the heuristic: Legal moves were cut down to the heuristic’s top 24 before being passed on, so the LLM’s freedom was effectively only “the freedom to deviate from rank 1.”
Latency and depth of thinking are a direct trade-off. The next attempt is planned as a nested structure: deep thinking only for phases where slowness is acceptable (auctions and plan formation), with the plan stored as state, and a fast heuristic that follows the plan during quick phases. On the other hand, the safety boundary that confines LLM output to a finite set by enumerating legal moves worked reliably. There were zero cases of the LLM making an illegal move.
6. Stop cost spikes by reverting
A note in the July 5 QA handoff document records that the project was reverted to commit 60c7436 because work done with Opus had become too expensive. When an agent goes far in the wrong direction, a git revert followed by a new instruction is cheaper than correcting it through conversation. Keeping commits small served as insurance.
Open issues
- CLAUDE.md says “local only, no public deployment,” but the project is actually live behind an invite-code gate. Mismatches between documentation and reality confuse the agent in later sessions, so this needs cleanup.
- Trademark and copyright risk was investigated and recorded in
docs/LEGAL.md(game name trademark risk: medium to high). Keeping the invite code private is a precondition.
Sources
- Repository:
~/Projects/age-of-steam-web(50 commits in git log, July 4–6, 2026, measured) - Usage measurements: 5 session files
~/.claude/projects/-Users-limjung-Projects-age-of-steam-web/*.jsonlplus 12 subagent files (0 parse errors, 8,070 lines total) - Evidence for sealing LLM hard:
src/ai/index.ts(HARD_DIFFICULTY_ENABLED = false),docs/HANDOFF-2026-07-06-mobile-qa.md - Self-play figures:
data/bot-strategy-eval-g17.json(1,020 games),data/learned-weights.json,data/strategy-notes.md - Rules and legal research:
docs/DECISIONS.md(22 items),docs/LEGAL.md,docs/knowledge/PROJECT-OVERVIEW.md - Prices: Claude API official price list (Fable 5 $10/$50 per MTok, cache writes 1.25× and reads 0.1×, Sonnet 5 introductory $2/$10)
Bottom line
A single person with a weekend and Claude Code can ship a playable, deployed board game with about 12,000 lines of code, provided the domain rules are frozen early and enough time is reserved for real-play QA. The API-equivalent cost was about $327, and most of it came from re-reading cached context rather than generating new output, so long sessions are where the money goes. The evidence also shows that an LLM is not a reliable game opponent on its own. Its outputs were safe when constrained to legal moves, but its play was weaker than a simple heuristic until planning and deeper reasoning were added.
Frequently asked questions
- How much did the Claude Code build of Age of Steam web cost in API-equivalent terms?
- About $327.19 at API list prices, of which 69% was cache reads. The author was on a subscription (Max), so the actual out-of-pocket spend was KRW 0. The total was cross-checked against the independent tool ccusage, which reported $325.80.
- Why was the LLM-powered hard difficulty shelved?
- In live play, LLM hard was weaker than heuristic medium, so the author set HARD_DIFFICULTY_ENABLED to false. Causes: a narrow hybrid scope, reasoning turned off for latency, no persistent plan across turns, and only the heuristic top 24 moves offered.
BuildnWrite helps teams build AI agents that keep running. About BuildnWrite ›