Insights

Is ChatGPT Plus Enough for Codex? A 72-Day Usage Review

8 min read#codex#chatgpt-plus#chatgpt-pro#usage-analysis#ai-coding-agent#cost-optimization

Who this is forDevelopers and solo builders who use Codex in intense bursts and need to choose between ChatGPT Plus, Pro, and extra credits.

If you use Codex in bursts, the plan decision is harder than it looks. Some weeks you barely touch it, and then a single sprint with several agents running at once can push you against the usage limit. Choosing between ChatGPT Plus at $20 per month, the Pro tiers at $100 and $200, and paying for extra credits depends on when your usage peaks, not only on how much you use in total. This article walks through one person’s 72 days of local Codex records. It shows how the usage broke down by month, by task type, by model, and by reasoning setting, and how close the peaks came to the limit. You will get a concrete way to decide which plan fits an irregular workload, along with a monthly review rule you can apply to your own data.

Short Summary

Codex was not barely used. From February onward, it was used heavily for automation, development, and lecture production across 72 days. The limit only mattered during a few concentrated sprints. Normally, Plus at $20 per month is enough. Extra credits cover the peak months. Pro only matches the cost structure in months when several projects run in parallel every day.

Summary Diagram

Codex usage retrospective and Plus decision summary

Original HTML summary

Key Data

Measurement Scope

  • Measurement date: July 19, 2026 (KST, Korea Standard Time)
  • Full recorded range: September 15, 2025 to July 19, 2026
  • Effective continuous use began: February 13, 2026
  • Data sources: the local Codex state_5.sqlite database and sessions/**/*.jsonl files
  • Authentication and secret files, along with conversation content, were excluded from aggregation and outputs.

Excluding conversation content matters here. Every number below comes from metadata and counts, not from the text of the sessions.

Overall Usage

Metric Measured value Interpretation
Session files / threads 346 / 344 Includes subagents and automatic reviews
User message events 1,589 1,516 when support threads are excluded
Active days 72 days About 46% of the days since February
Active days, last 30 days 15 days Usage is intermittent rather than low
Messages / threads, last 30 days 185 / 84 Parallel runs and benchmarks increased the thread count
Consolidated thread tokens 1.081 billion Cumulative per-thread totals from the Codex state database
Raw token events 1.274 billion Includes internal model calls and handoff events
Cache share of raw input 93.1% Agent-style pattern that reuses large repository context
Output tokens 6.84 million Generated output is small relative to the input context
Internal model calls 14,905 calls Internal reasoning steps, not user messages
Session length median / top 10% 4 minutes / 61 minutes or more Short checks and long sprints coexist

Converting the 1.081 billion consolidated tokens directly into API cost would be a mistake. Most of the input was cached, and under a ChatGPT subscription, token counts affect the included limit and credit consumption rather than the bill itself.

Monthly Activity

Month Active days Message events Threads created Raw token events
February 2026 8 188 16 162 million
March 2026 16 514 52 197 million
April 2026 11 334 83 245 million
May 2026 9 258 46 245 million
June 2026 14 114 66 81 million
July 1–19, 2026 12 180 79 344 million

March had the largest conversation volume. July had fewer conversations than March, but each task was heavier. The Age of Steam development work and the photo migration created many long contexts and parallel tasks. This pattern explains why raw token counts can rise even when message counts do not.

Work Breakdown

I first classified the threads automatically, using project names and the titles of the first requests. The shares below measure each category’s share of consolidated thread tokens. They do not measure how valuable the work was.

Category Threads Token share Representative work
Automation, agents, and infrastructure operations 74 27.9% Mac mini/SSH, Trigger.dev to n8n, Discord summarization and deduplication, Notion to PostgreSQL
Product and software development 47 19.9% Age of Steam map, AI self-play, and i18n; Brass engine and data; web UI
Lectures, books, documents, and presentations 63 14.9% Healthcare lecture outline, n8n book index, Hanbit Media book writing (Hanbit is a Korean IT publisher), Remotion and slides
Subagents, reviews, and benchmarks 55 9.9% Parallel photo migration work, game AI benchmark, automatic review
Personal IT and media management 9 8.0% iCloud/OneDrive to Google Photos, timestamp metadata, AhnLab removal (AhnLab is a Korean security software vendor)
Codex, Claude, and AI work environments 54 7.3% Skills and plugins, Obsidian session logs, OpenUsage
Research, content, and information gathering 22 7.1% n8n trends, Inflearn case studies (Inflearn is a Korean online course platform), YouTube sources, sales rankings and job postings
Work management, data analysis, and collaboration 12 2.8% Mentoring and calendar, assignment grading, meeting agendas, real estate transaction prices
Other / insufficient context 8 2.3% Follow-up requests whose earlier context cannot be restored from titles alone

The automation and development categories account for nearly half of all token usage. The subagent and review category is only about 10%, which matters later when we discuss parallelism.

Models and Credit Mix

Model Share of consolidated tokens
GPT-5.5 42.6%
GPT-5.4 25.6%
GPT-5.6 Sol 17.8%
GPT-5.3-Codex 12.8%
Other 1.2%

The reasoning effort was high for 49.3% of tokens and xhigh for 17.8%. Together they account for 67.1% of consolidated tokens. Keeping this setting for simple monitoring and file checks uses up the Plus limit faster than necessary, so the reasoning setting is one of the main levers for stretching a limit.

After April 2, 2026, when OpenAI moved most Plus and Pro Codex usage to token-based credits, applying the official rate card to the local token mix gives the following results:

  • April 2 to July 19, 2026: about 16,600 credit equivalents, with a model identification rate of 98.0%
  • Last 30 days: about 6,700 credit equivalents, with a model identification rate of 95.2%

These figures show consumption in the same units as the included limit. They do not represent additional payments.

Current Plan Cost

The local records store the plan only as pro. They do not distinguish Pro $100 from Pro $200. Because no receipts are stored, the actual total amount paid cannot be confirmed from local data.

Plan Monthly price Limit relative to Plus Savings when switching to Plus
Plus $20 1x Baseline
Pro 5x $100 5x $80 per month / $960 per year
Pro 20x $200 20x $180 per month / $2,160 per year

The first local Pro record is dated March 12, 2026. The $100 Pro tier was announced on April 9, 2026. Unless the plan was changed directly, the account was probably on the original $200 Pro tier. Promotions or account transfers are possible, so the receipt is the final reference.

Insights

1. Usage Frequency Is Not Low, but Intensity Is Uneven

Usage occurred on 15 of the last 30 days, so the feeling of barely using Codex is not accurate. Still, a few days pushed usage up sharply, such as the Age of Steam pipeline on July 7 and the photo migration on July 11. For a subscription decision, the 5-hour peak matters more than the monthly total, because the peak determines whether you hit the limit mid-task.

2. Plus for Normal Work, with Limits in Some Sprints

The local limit snapshots show that the highest Pro 5-hour window usage was 30%, followed by 20% and 19%. Converting the 50 days with recorded Pro usage to Plus on a Pro 5x basis, only 2 days reach or exceed the limit. The highest recent weekly window was 4% of Pro. Even when the Pro 20x basis is applied most conservatively, that is about 80% of the Plus weekly limit.

In practice, routine checks, documentation, and small fixes fit comfortably within Plus. Focused development days that run several agents in parallel may need extra credits or a temporary Pro plan.

3. Pro’s Value Was Concurrency and Duration, Not Features

The CLI, app, IDE, skills, plugins, MCP, web search, and SSH work used so far all remain available on Plus. The record shows no use of GPT-5.3-Codex-Spark, which the note lists as Pro-only. The real difference was not whether a task was possible. It was how long and how many sessions in parallel you can run in a day.

4. Cost Optimization Starts with Model Choice and Parallelism

Of the consolidated tokens, 67.1% used high or xhigh reasoning, and subagents, reviews, and benchmarks accounted for about 10%. Lowering the default model to Terra, moving repeated checks to Luna or mini, and limiting parallel agents to one or two at a time can substantially increase the effective limit, even on Plus.

  1. Switch to Plus at $20 starting with the next payment.
  2. Use Terra by default, Luna or mini for repeated checks, and reserve Sol with xhigh reasoning for complex design and final review.
  3. Limit normal parallel agents to one or two.
  4. When the limit is reached, switch to a lighter model first, and buy extra credits only in that month.
  5. Reevaluate after one month. Keep Plus if limit-hit months number two or fewer. Consider Pro $100 if extra credit spending approaches $80 per month. Consider Pro $200 only if you run several projects in parallel every day.

Analysis Limits

  • Token counts are not the same as productivity or actual billing.
  • Raw event tokens are about 18% larger than consolidated thread tokens because they include internal calls and handoffs. Totals use the consolidated values, while the model and cache mix use the raw event values.
  • Work classification is automatic, based on project names and request titles. The 2.3% with insufficient context was left in the Other category.
  • The local plan_type=pro value alone cannot distinguish Pro $100 from Pro $200. The Plus conversions considered both cases.

Bottom Line

For this usage pattern, the evidence supports ChatGPT Plus at $20 per month as the default plan. Usage is frequent but bursty, concentrated in a few sprints where the 5-hour window peaks. Pro’s distinct benefit is longer and more parallel sessions, which only pays off if you run that kind of workload every day. Peak months can be handled with lighter models, fewer parallel agents, and short-term extra credits. Review the decision after one month, using the count of limit-hit months and your extra credit spending as the deciding measures.

Sources

Official Documentation

Local Measurements

  • ~/.codex/state_5.sqlite: threads, models, reasoning effort, and consolidated tokens
  • ~/.codex/sessions/**/*.jsonl: user messages, raw tokens, limit snapshots, and tool calls
  • Aggregation date: July 19, 2026 (KST)

Frequently asked questions

Is ChatGPT Plus enough for heavy but irregular Codex use?
For everyday work, the data says yes. Plus covers routine checks, documentation, and small fixes. Peak sprints with many parallel agents may need extra credits or a temporary Pro plan, so review the choice after one month.
When does ChatGPT Pro make sense for Codex?
Pro fits users who run several projects in parallel every day for long sessions. The advantage was concurrency and session length, not extra features. Choose Pro $100 if extra credits approach $80 per month, and Pro $200 only for daily multi-project parallel work.