Codex vs Claude Code.
Two terminal agents at the same $20 a month and the same API prices, apart on benchmarks, permission defaults and usage limits. Every published figure below was read on 11 October 2026 from the page linked under Sources, and the session costs come from 30 days of our own logs. Starkomand, which runs both, is ours; the comparison does not depend on it.
v1.0.6 · Apple silicon · free for 3 days · no account
The short version.
- BenchmarksCodex CLI has the top Terminal-Bench 2.1 run, 87.4% with GPT-6 Astra, against 83.8% for Claude Code with Claude Fable 5. Neither vendor’s newest model has a published run yet.
- SWE-bench VerifiedNo longer separates them. On Vals’ last run, Claude Opus 5 scored 97.0% and GPT-5.6 Sol 96.2%.
- PriceThe same. Both start at $20 a month, and the API ladders match rung for rung: $10 in and $50 out per million tokens at the top, $2 and $10 for the everyday model, $0.10 and $0.50 for the small one.
- Measured costIn 30 days of our own sessions, the median session cost $2.44 in Codex and $2.38 in Claude Code at API list price. Most of it paid for re-reading cached context, not for writing code.
- DefaultsCodex starts inside an OS-enforced sandbox with network access off. Claude Code starts in auto mode, where a second model reviews actions, and its sandbox is opt-in.
- If you pay for bothThe useful question is which one gets this task today, and where the work goes when one of them hits a limit.
Side by side.
Codex CLI next to Claude Code, from both vendors’ public pages on 11 October 2026.
| Criterion | Codex CLIOpenAI | Claude CodeAnthropic |
|---|---|---|
| Current models | GPT-6 Astra, GPT-6.1 Sol, GPT-6 Sol and GPT-6 LunaGPT-6.1 Sol arrived on 29 September; GPT-5.5 leaves Codex on 14 October. | Fable 5.1, Opus 5.5, Sonnet 5.5 and Haiku 5.5Anthropic suggests Opus 5.5 for most work and Fable 5.1 for long-horizon agentic work. |
| Where it runs | The terminal, IDEs, the web and iOS, plus cloud tasks on OpenAI’s machines | The terminal, VS Code, JetBrains, a desktop app and the web |
| Context window | 272K in the CLI’s bundled settings, as reported in August for GPT-5.6 SolContested; see the note under the table. | 1M tokens by default on every current modelThe API charges no premium past 200K. |
| Terminal-Bench 2.1, best published run | 87.4% with GPT-6 Astra | 83.8% with Claude Fable 5 |
| Starting permissions | Auto preset: edits and runs commands in the working directory, asks before leaving it or reaching the networkRecommended for version-controlled folders; read-only elsewhere. | Auto mode: a second model, the classifier, reviews actions instead of youSince v2.1.283. Manual mode asks before most edits and commands. |
| OS sandbox | On by default, with network access offSeatbelt on macOS, bwrap and seccomp on Linux. | Opt-in with /sandboxOn macOS, Linux and WSL2. |
| Project instructions | AGENTS.md |
CLAUDE.md |
| Cheapest plan with the agent | ChatGPT Free and Go ($8/month) for light tasks; Plus at $20/month | Claude Pro, $20/monthClaude’s Free plan does not include Claude Code. |
| Top individual plan | ChatGPT Pro, $500/month | Claude Max 20x, $200/month |
| API, everyday model | GPT-6.1 Sol: $2 in, $10 out per million tokens | Sonnet 5.5: $2 in, $10 out per million tokens |
| Usage windows | Five-hour windows on Plus; none on Pro at present; weekly limits may apply | A five-hour session plus a weekly limit, shared with Claude chat |
Same on both. MCP servers, hooks, skills and subagents, so “the extensible one” is no longer a reason to pick either.
The context row. In openai/codex#38917, users reported in August that the CLI’s bundled settings held GPT-5.6 Sol to 272,000 tokens even after the documented 1M override. The issue was closed on 19 August, and comments in the thread report different caps for different models. OpenAI’s API pricing still charges a separate rate for prompts above 272K. Run /status in your own install before you plan a task around a number.
Every figure on this page except the measured session costs was read on 11 October 2026 from the page listed under Sources, and both products change monthly; product names belong to their owners. Spotted a change? Tell us.
Benchmarks.
Terminal-Bench measures agents completing real command-line tasks. Version 2.1 fixes 28 of the 89 tasks in 2.0: some depended on internet resources that had changed, some had budgets too tight for valid solutions, and some had instructions that did not match their tests.
| Harness | Model | Score | Run date |
|---|---|---|---|
| Codex CLI | GPT-6 Astra | 87.4% | 3 Sep 2026 |
| Claude Code | Claude Fable 5 | 83.8% | 9 Jun 2026 |
| Codex CLI | GPT-5.5 | 83.1% | 23 Apr 2026 |
| Terminus 2 | Claude Fable 5 | 80.4% | 9 Jun 2026 |
| Claude Code | Claude Opus 4.8 | 78.9% | 28 May 2026 |
| Codex CLI | GPT-5.6 Terra | 78.4% | 26 Jun 2026 |
| Terminus 2 | GPT-5.5 | 78.0% | 23 Apr 2026 |
| Claude Code | Claude Sonnet 5 | 74.6% | 30 Jun 2026 |
Selected rows from the Terminal-Bench 2.1 leaderboard.
- The newest models are missing
- Fable 5.1, Opus 5.5, Sonnet 5.5 and GPT-6.1 Sol have no runs on the board yet, so the headline compares a September OpenAI model with a June Anthropic one.
- Each CLI adds points to its own model
- Claude Fable 5 scores 83.8% in Claude Code and 80.4% in Terminus 2, the benchmark’s neutral harness; GPT-5.5 scores 83.1% in Codex CLI and 78.0% in Terminus 2. That gap of 3 to 5 points is the closest public measure of the CLIs themselves, as opposed to the models.
- Older comparisons quote older runs
- A page that quotes 77.3% against 65.4% is quoting GPT-5.3-Codex and Claude Opus 4.6 on Terminal-Bench 2.0, two generations and one benchmark revision ago.
- SWE-bench Verified has saturated
- On Vals’ leaderboard, last updated 1 September, Claude Opus 5 scores 97.0%, GPT-5.6 Sol 96.2% and Claude Fable 5 95.0% in the same mini-SWE-agent harness, and Vals no longer runs it on new models. A 0.8-point gap near the ceiling is noise.
Neither benchmark measures your repository. The better test is one real task run in both: note the wall-clock time, which allowance it drew from and whether the diff shipped. It is a sample of one, so repeat it after the next model release.
Prices.
Both sell the agent inside a chat plan, and both start at $20 a month.
| Plan | Price | What it includes |
|---|---|---|
| ChatGPT Free | $0 | Quick coding tasks; GPT-6 Luna in the desktop app, subject to rollout |
| ChatGPT Go | $8/month | Lightweight coding tasks |
| ChatGPT Plus | $20/month | Codex on the web, CLI, IDE and iOS; GPT-6.1 Sol and GPT-6 Luna |
| ChatGPT Pro | $100, $200 or $500/month | No five-hour limit at present; Ultrafast mode on the $500 plan |
| Claude Free | $0 | No Claude Code |
| Claude Pro | $20/month, or $17/month billed annually | Claude Code |
| Claude Max 5x | $100/month | Five times Pro’s per-session allowance |
| Claude Max 20x | $200/month | Twenty times Pro’s per-session allowance |
For teams, a ChatGPT Business seat and a Claude Team standard seat both cost $25 a month billed monthly, or $20 billed annually.
At the API
At the API the two price ladders match almost exactly, per million tokens of input and output.
| Rung | OpenAI | Anthropic |
|---|---|---|
| Top | GPT-6 Astra: $10 / $50 | Fable 5.1: $10 / $50 |
| Everyday | GPT-6.1 Sol: $2 / $10 | Sonnet 5.5: $2 / $10 |
| Small | GPT-6 Luna: $0.10 / $0.50 | Haiku 5.5: from $0.10 / $0.50 |
Opus 5.5, at $4 and $20, sits between the top and everyday rungs. OpenAI’s rates apply to prompts up to 272K input tokens, and Haiku 5.5’s lowest rate to prompts up to 100K.
So “which is cheaper” has no answer at list price. The answer that matters to a subscriber is different: if you already pay $20 for Claude Pro and $20 for ChatGPT Plus, the second agent costs nothing extra, and you hold two allowances on two model families. The waste is the one you are not using.
Usage limits.
Both throttle on rolling windows, and neither publishes one number that holds for every account.
| Criterion | Codex CLIOpenAI | Claude CodeAnthropic |
|---|---|---|
| Short window | A five-hour window on Plus; none on Pro at present | A five-hour session limit |
| Long window | Weekly limits may apply; no published numbers | A weekly limit; Max adds a separate weekly limit for Fable |
| Shared with | Your ChatGPT plan’s Codex allowance; ChatGPT credits extend it | Claude chat and the IDE extensions |
| Check it with | /status in a session; /usage for account totals |
/usage |
OpenAI does publish per-model ranges for Plus: 15 to 160 local GPT-6.1 Sol messages per five hours, 5 to 45 for GPT-6 Astra and 350 to 3,000 for GPT-6 Luna. The model you pick moves your runway more than the plan does.
If you pay for both, there is a better move than waiting for the reset: continue the task in the other agent with its context intact, meaning the file list, the failing test and the decisions made so far. A short written brief beats pasting a wall of scrollback, and keeps the second agent from rereading the whole repository on the allowance you just switched to.
What a session costs.
Measured, not estimated: every interactive session we ran in the Starkomand repository from 11 September to 11 October 2026, priced at the API list rates above. The two agents did different work in those sessions, so read this as one team’s month, not as a benchmark.
| Criterion | Codex CLIOpenAI | Claude CodeAnthropic |
|---|---|---|
| Sessions | 198 | 336 |
| Model responses per session | 24 at the median | 22 at the median |
| Cost per session at list price | $2.44 at the medianHalf of sessions fell between $0.95 and $6.03; one in ten cost more than $10.04. | $2.38 at the medianHalf of sessions fell between $0.83 and $6.61; one in ten cost more than $11.88. |
| 30 days at list price | $901 | $1,695 |
| Where the money went | Cached input 63%, uncached input 25%, output 12% | Cache reads 39%, cache writes 36%, output 25%Uncached input rounds to 0%. |
| Models behind the spend | GPT-6 Astra 71%, GPT-5.6 Sol 20%, GPT-6 Sol 7% | Fable 5.1 38%, Opus 5.5 35%, Opus 5 27% |
What the numbers show. In both agents, re-reading context cost more than writing: output was a quarter of Claude Code’s spend and an eighth of Codex’s, although about 97% of all input tokens came from cache, which both vendors bill at a small fraction of the input price. Per response, long sessions were no dearer than short ones ($0.10 against $0.13 in Claude Code and $0.08 against $0.09 in Codex, longest quarter of sessions against shortest), so a session cost what it did because of how many responses it ran. The costliest tenth of sessions took 42% of each agent’s spend.
How we measured. Both CLIs log every model response’s token counts on your own machine: Claude Code under ~/.claude/projects/, Codex under ~/.codex/sessions/. We counted each response once, priced it at the vendor’s published rates, including Anthropic’s cache-write and OpenAI’s cached-input rates, and kept interactive sessions in this one repository. Claude Code’s own cost records for the same sessions came within 4% of our total. Codex’s auto-review checks run on codex-auto-review, which has no published API price, so their tokens, 8% of Codex’s total, are left out.
On a plan, none of this is a bill. A Claude or ChatGPT plan meters the same work against its usage windows instead of charging per token, and Claude Code’s docs say the session cost in /usage is meant for API users. The list-price figure is what the same work would cost on an API key, which is the number to hold against a $20, $100 or $200 plan.
Your own numbers. In Claude Code, /usage shows the current session’s tokens and their cost at list price, and on Pro and Max what counts against your plan limits. In Codex, /status shows the session’s token usage and remaining limits, and /usage shows your account’s token activity by day or week.
A third tool.
The three-way searches, Cursor vs Claude Code vs Codex and Gemini CLI vs Claude Code vs Codex among them, ask whether a third tool displaces the top two. For paid daily terminal work it does not, but each one changes something.
-
Cursor
An AI editor first, which also ships a CLI,
cursor-agent. Cursor CLI with Grok 4.5 scores 79.3% on Terminal-Bench 2.1. Hobby is free with limited agent requests; Pro is $20 a month.Pick it if you want to read and edit diffs inline in an editor.
-
Gemini CLI
On 18 June 2026 Google stopped serving Gemini CLI requests for free, Google AI Pro and Google AI Ultra individual accounts, and asked users to move to Antigravity CLI. Gemini Code Assist licenses and API keys still work, but the old free-tier advice no longer applies to individuals.
-
Antigravity
Google’s agent-first IDE and the way on for Gemini CLI users. The Individual plan is $0, with unlimited Tab completions and Command requests and “basic weekly rate limits” on agent work, for which Google publishes no number.
Pick it if you want an agent IDE for nothing and can live inside an unquantified cap.
-
OpenCode
An open-source terminal agent that works with any provider, including models on your own machine. The software is free and the inference is not: you pay whichever provider you point it at, and you own the setup.
Pick it if you want to choose the provider yourself.
Cursor changes the interface, Antigravity changes the price and OpenCode changes who you pay. Codex and Claude Code stay at the top for paid terminal work.
Which one.
Pick Codex if
you want an OS-enforced sandbox with network access off from the first run, the top published Terminal-Bench run, a free or $8 way in for light tasks, or a Pro plan without a five-hour window.
Pick Claude Code if
you want a documented 1M-token context on every model for large repositories, a permission mode that lets a classifier approve routine actions, or the Claude models you already use in chat.
Pick both if
you already pay for both. $20 plus $20 buys two independent allowances, two model families that review each other’s work, and a way to keep going when either one throttles.
MoreCodex CLI in StarkomandClaude Code in StarkomandRun multiple Claude Code agents in parallel
Running both by hand.
The plain version needs no app. Open two terminals, run claude in one and codex in the other, and keep one set of rules: write them in AGENTS.md, which Codex reads, and make CLAUDE.md the single line @AGENTS.md, which Claude Code expands as an import.
# CLAUDE.md, the whole file: Claude Code imports the shared rules
@AGENTS.md
# a worktree each when both write code
git worktree add ../app-codex -b codex-task
# then, in a terminal each
claude
cd ../app-codex && codex
Give the two agents different jobs, and when both write code, give each its own git worktree so neither overwrites the other’s changes.
Running both in Starkomand.
Starkomand is ours, so this is the one section about it. What plain terminals do not give you is the movement of work between them, and that is what we built it for. It runs on macOS with Apple silicon only and needs no Starkomand account: each CLI runs in a real terminal window on one canvas, signed in with your own Claude and ChatGPT plans.

- Plan in one, build in the other
- Open Agent actions in the Claude Code window, choose Hand off… and pick Codex. The brief carries the objective and acceptance criteria, and the Codex window lands wired to the source.
- Review with the other model family
- In Settings → Workflows, set the Review workflow’s receiving agent to the other CLI. Build and review then sends the finished work to a fresh reviewer from the other vendor.
- See which allowance has room
- Settings → Usage meters shows one ring per signed-in agent with its five-hour and weekly resets. Past 90%, that agent’s model chip turns amber and offers Hand off to keep going →.
- Continue after a limit
- When Claude Code cannot take another turn, the model chip’s Continue from recorded conversation starts Codex from the conversation the app recorded.
- Read across windows
- An agent can read another window’s transcript by name, so the plan does not need retyping.
- Redirect by voice
- Hold Option and speak to an agent by name.
We have not measured whether this makes anyone faster. What it changes is that both allowances get used, and a task survives either one running out.
Questions
Is Codex better than Claude Code?
Not overall, and the published data cannot say otherwise. Codex CLI holds the top Terminal-Bench 2.1 run at 87.4% against Claude Code’s 83.8%, but the newest models from both vendors have no runs yet, and SWE-bench Verified has saturated at 96 to 97% for both families.
Which is cheaper, Codex or Claude Code?
At list price, neither. Both entry plans are $20 a month, and the API ladders match at $10/$50, $2/$10 and $0.10/$0.50 per million tokens. ChatGPT’s Free and $8 Go plans include Codex for light tasks; Claude’s Free plan does not include Claude Code.
Which is better for large repositories?
On paper, Claude Code: its docs say every current model runs with a 1M-token window by default. Codex users reported a 272K cap in the CLI’s bundled settings in August, so check /status on your own install.
Which has better usage limits?
ChatGPT Pro currently has no five-hour window; every Claude plan has a five-hour session limit and a weekly limit. Below that, both depend on your account and model, and the better strategy is to continue the task in the other agent when one runs out.
How much does a Claude Code or Codex session cost?
In our own 30 days of sessions in one repository, the median session cost $2.38 in Claude Code and $2.44 in Codex at API list price, and one session in ten cost more than $10 to $12. On a Claude or ChatGPT plan the same work is metered against usage windows instead of billed per token.
How do I check token usage in Claude Code and Codex?
In Claude Code, run /usage: it shows the current session’s tokens and their cost at list price, and on Pro and Max what counts against your plan limits. In Codex, /status shows the session’s token usage and remaining limits, and /usage shows account token activity. Both CLIs also log every response’s token counts locally, under ~/.claude/projects/ and ~/.codex/sessions/.
Is ChatGPT Codex the same thing as the Codex CLI?
Yes, in the sense that matters. Codex is one agent with several surfaces: the CLI, the IDE extension, the web and iOS apps, and cloud tasks. Signing in with ChatGPT bills all of them to your ChatGPT plan; an API key bills the CLI, SDK and IDE at API rates, with no cloud features. The name also used to cover models such as GPT-5.3-Codex, but the current lineup is general GPT-6 models.
Can I use both in one project?
Yes. Keep the shared rules in AGENTS.md with CLAUDE.md importing it, give each agent its own job, and use separate git worktrees when both write code.
Do I need separate subscriptions?
Yes. Claude Code runs on a Claude plan or an Anthropic API key; Codex runs on a ChatGPT plan or an OpenAI API key.
What does Reddit say about Codex vs Claude Code?
Read the threads as sentiment, not measurement. Their verdicts flip with each model release, and the r/ClaudeAI thread that still ranks for the question is from October 2025, older than every model on this page. Check a thread’s date before you trust its verdict.
Sources.
- AnthropicModels overviewPlans and pricingAPI pricingMax planClaude Code with Pro or MaxModel configurationPermission modesSandboxingManage costs
- OpenAICodex pricingCodex changelogApprovals and securitySlash commandsAPI pricingopenai/codex#38917
- BenchmarksTerminal-Bench 2.1 leaderboardVals SWE-bench Verified
- OthersCursor pricingGemini CLI transitionAntigravity pricing
Try it against your own setup.
Every feature, every provider, free for 3 days. Your CLIs are already installed; the canvas takes a minute.
v1.0.6 · Apple silicon · no account