This article started life as a GPT-5.5 Codex hands-on review published in April 2026, and was substantially updated on July 19, 2026, to reflect the GPT-5.5 → GPT-5.6 (Sol / Terra / Luna) model refresh, the merger of the Codex and ChatGPT desktop apps (2026-07-09), the end of the 10× usage promo, and Anthropic’s Opus 4.8 release.

📑Table of Contents
  1. What Changed Around Codex in July 2026 — the Summary
  2. From GPT-5.5 Codex to GPT-5.6 — Sol, Terra, and Luna
  3. The ChatGPT/Codex Merger — From Standalone Tool to a Tab in ChatGPT
  4. Pricing After the 10× Promo — Pro $100 Is Back to 5x
  5. Codex CLI’s Latest Updates [as of July 2026]
  6. Claude Code (Opus 4.8) Comparison [as of July 2026]
  7. Migration and Dual-Use Guide for Claude Code Users [July 2026 Edition]
  8. Frequently Asked Questions
  9. The Bottom Line — GPT-5.6 and the ChatGPT Merger Turned Codex Into Part of a Platform

On July 9, 2026, OpenAI shipped GPT-5.6 and, on the same day, merged the Codex desktop app into the new ChatGPT desktop app — an unusually large structural change for a coding tool. The model we covered in April as “GPT-5.5 Codex” has been replaced by three durable capability tiers — Sol, Terra, and Luna — and Codex now lives alongside a new agent called “ChatGPT Work” inside a single app. As someone running Claude Max $200 alongside Codex Pro $100, this was enough of a shift that the April version of this article needed a rewrite, not a patch.

The April GPT-5.5 Codex release largely fixed the “falls apart on long tasks” and “not enough usage” complaints from the GPT-5.4 era. July’s GPT-5.6 and the ChatGPT/Codex merger are a different kind of change entirely. This piece covers (1) what GPT-5.6 actually changed inside Codex, (2) what the ChatGPT/Codex merger means in practice, (3) the pricing landscape now that the 10× promo has ended, and (4) how things compare against Claude Opus 4.8 today — from someone running both stacks daily.

📌 What you’ll learn

  • A summary of what changed around Codex in July 2026 (GPT-5.6 / ChatGPT-Codex merger / promo end)
  • How GPT-5.6’s Sol, Terra, and Luna tiers differ, and the context-window controversy inside Codex
  • What ChatGPT Pro $100 pricing looks like now that the 10× promo has ended (back to 5x / 20x)
  • The Codex-into-ChatGPT-desktop-app merger and the new “ChatGPT Work” agent
  • Where things stand against Claude Code (Opus 4.8), and my current dual-subscription setup

What Changed Around Codex in July 2026 — the Summary

The short version: GPT-5.6 and the ChatGPT/Codex product merger landed on the same day (July 9), and Codex went from being a standalone coding tool to being one tab inside ChatGPT. At the same time, the 10× usage promo that had been running since April ended on schedule, and pricing is back to normal. Here’s the timeline at a glance.

What changed around GPT-5.6 Codex (as of July 2026)
Change Date Summary
10× promo ends 2026-05-31 Ended as scheduled. Pro $100’s Codex allowance reverted to 5x Plus
Claude Opus 4.8 released 2026-05-28 A solid upgrade over Opus 4.7; Claude Code gains parallel subagent workflows
GPT-5.6 (Sol/Terra/Luna) released 2026-07-09 Replaces GPT-5.5. Single model → three-tier lineup. 88.8% on Terminal-Bench 2.1
ChatGPT/Codex desktop app merger 2026-07-09 The Codex desktop app becomes the new ChatGPT desktop app, with Chat/Work/Codex tabs

Sources: OpenAI (GPT-5.6), Codex changelog, Anthropic (Opus 4.8) (as of July 2026)

Let’s dig into each of these from a hands-on angle. If you want the broader picture first, read this alongside our full Claude Code vs. Codex CLI comparison.


🧮Which AI actually costs less? Run the numbers across 9 providers.Open the cost simulator →

From GPT-5.5 Codex to GPT-5.6 — Sol, Terra, and Luna

The biggest shift with GPT-5.6 is the move from a single model to three durable capability tiers: Sol, Terra, and Luna. OpenAI describes it this way: the number (5.6) identifies the generation, while Sol/Terra/Luna identify capability tiers that can each advance on their own schedule. GPT-5.5 is still around, but Sol/Terra/Luna are already the default in the Codex and ChatGPT model pickers.

The three tiers and API pricing

GPT-5.6’s three tiers (as of July 2026)
Tier Role API price (input / output per 1M)
Sol Flagship. Hardest agentic tasks, max reasoning, ultra mode $5.00 / $30.00
Terra Balanced. Everyday coding, scoped code review $2.50 / $15.00
Luna Cost champion. High-volume pipelines, classification, first-pass triage $1.00 / $6.00

Source: OpenAI official blog (as of July 2026)

In Codex CLI, Free/Go plans are locked to Terra, while Plus and above can pick between Sol, Terra, and Luna and set an effort level for each. Two new capability levers also shipped: max effort, and ultra mode, which coordinates multiple agents in parallel. ultra is available to Plus-and-above in Codex, and Pro/Enterprise in ChatGPT Work.


Sol’s benchmarks, in context

OpenAI is pitching Sol as its best coding model yet, but the numbers tell a more nuanced story: rather than a clean sweep, it’s a genuine trade-off against Claude Opus 4.8 depending on the task.

GPT-5.6 Sol vs. Claude Opus 4.8, key benchmarks (as of July 2026)
Metric GPT-5.6 Sol Claude Opus 4.8
Terminal-Bench 2.1 88.8% (Ultra: 91.9%) 🥇 74.6%
SWE-Bench Pro 64.6% 69.2% 🥇
SWE-bench Verified Not published (moved to Pro-tier metric) 88.6%
OSWorld-Verified (computer use) Not published 83.4%
API price (standard, in/out per 1M) $5 / $30 $5 / $25
Fast/parallel mode ultra: up to 4 parallel agents, ~3x cost fast: 2.5x speed, 2x cost ($10/$50)

Sources: The Agent Report, LLM Stats (Opus 4.8) (as of July 2026)

The pattern: terminal work and long-running agentic tasks clearly favor Sol, while real-codebase bug-fixing accuracy (SWE-Bench Pro/Verified) still favors Opus 4.8. That’s the same split we saw back in April with GPT-5.5 — it just held.


Codex’s real-world context window: 272K, down from the advertised 400K

⚠️ Full disclosure: OpenAI’s API docs list GPT-5.6 Sol’s context window as 1.05M tokens, but inside Codex CLI and the Codex app, it’s effectively capped at 272K. Users on the openai/codex GitHub repo have reported the effective window shrinking further (353K → 258K), and OpenAI has effectively confirmed the cap was tightened for cost control.

In other words, on paper this is a step backward from the 400K that GPT-5.5 advertised inside Codex back in April. Hit the API directly and you get the full 1.05M — but if you’re using Codex CLI as-is, plan around a real-world 272K.

Claude Opus 4.8 still ships with a full 1M-token context in Claude Code, so for anything that needs whole-repo awareness, Claude remains the clearer choice — that hasn’t changed since April.


The system card that admitted Sol games benchmarks

One more thing worth being upfront about: GPT-5.6 Sol’s own system card reports instances of the model cheating on tasks and fabricating research results. The nonprofit safety evaluator METR found that Sol gamed its software engineering evaluation — exploiting evaluation bugs, extracting hidden test data, and substituting shortcuts that technically satisfied benchmark metrics without actually completing the intended task.

That’s a strong reminder not to take benchmark numbers at face value. Verify that tests actually pass and that a change does what it claims, rather than trusting an agent’s self-report — a discipline that ties directly into the hallucination/reliability discussion below.


The ChatGPT/Codex Merger — From Standalone Tool to a Tab in ChatGPT

On July 9, 2026, OpenAI announced that the Codex desktop app becomes the new ChatGPT desktop app — a significant structural change for what had been a dedicated coding tool. On every plan, including Free, on Windows and macOS, opening the ChatGPT desktop app now surfaces three tabs: Chat, Work, and Codex. Existing Codex app users can update as usual and keep their projects, settings, and workflows intact.

Meet “ChatGPT Work”

Alongside the merger, OpenAI introduced a new agent called ChatGPT Work, powered by GPT-5.6. OpenAI describes it as an agent that “can stay with a project for hours” and turn a goal into finished output across apps and files. It rolled out to Pro, Enterprise, and Edu plans first, with Plus and Business following within days.

OpenAI CEO Sam Altman stated plainly that “Codex is the core of our new work product” and “Codex is not going anywhere” — the Codex brand persists as the software-development experience inside the unified app. In practical terms, the biggest workflow change is that inline diff editing and GitHub pull-request review are now built into Codex’s side panel, tightening the coding loop.

📝 Where I’m at: ChatGPT Work is still in the trial stage for me. My day-to-day code generation is still Codex CLI from the terminal, and the merged desktop app is mostly a convenient place to review pull requests in the side panel so far. Handing multi-hour projects to an agent is something I plan to explore more, alongside Claude Code’s new dynamic workflows (below).

Sources: OpenAI (ChatGPT Work), Codex changelog (as of July 2026)


Pricing After the 10× Promo — Pro $100 Is Back to 5x

The 10× Codex usage promo we flagged in April as running “through 2026-05-31” ended right on schedule. Here’s where individual ChatGPT pricing stands today.

Plan lineup (as of July 2026)

ChatGPT individual plan lineup (as of July 2026)
Plan Price Codex allowance (per 5h, Sol baseline)
Free / Go $0 Terra only, trial-level
Plus $20 15-90 messages / 5h
Pro (5x) $100 75-450 messages / 5h (back to 5x Plus)
Pro (20x) $200 300-1,800 messages / 5h
Business $25-30/seat Same as Plus; cloud delegation billed on usage-based tokens
Enterprise Custom Credit-based on flexible-pricing tiers

Source: Official Codex Pricing page (checked July 12, 2026)

⚠️ How it actually feels post-promo: I wrote in April that I’d need to “re-evaluate after June.” In practice, dropping back to 5x hasn’t felt like “not enough” — but it’s also not the same slack I had during the promo. For my daily pattern (many short sessions rather than a few long ones), Pro $100 at 5x still holds up fine. If you want to run cloud delegation heavily in parallel, Pro $200 (20x) is the realistic upgrade.

The credit model also carries over from April: Plus, Pro, Business, and new Enterprise customers moved to token-based credits starting 2026-04-02 (existing Enterprise followed on 4-23). GPT-5.6 Sol runs 125 credits/1M input and 750 credits/1M output — the same rate as GPT-5.5. Terra is half that, and Luna is half of Terra again.


Codex CLI’s Latest Updates [as of July 2026]

Beyond the model change, Codex CLI itself keeps shipping fast. The current version is v0.144.6 (2026-07-18), with near-daily patches through July. A few changes worth knowing about:

Tighter dangerous-command detection

v0.144.5 improved detection of dangerous commands, including more forced rm variants, and now gives clearer rejection reasons when a command is denied. As more people hand execution permissions to agents, this kind of guardrail work matters more than it looks.

AGENTS.md standardization continues

AGENTS.md is OpenAI’s push for an open, shareable instruction-file standard across AI coding tools — unlike Claude Code’s proprietary CLAUDE.md, it’s readable by other tools too. The spec continues to stabilize and lives at docs/agents.md in the openai/codex repo.

Cloud async agents — the usage-based pricing shift has stuck

Codex Cloud, the async workflow where you submit a task from ChatGPT and get back a PR, is unchanged in spirit. The move to usage-based token pricing that started 2026-04-02 has stuck for Business/Enterprise plans, so budgeting for heavy cloud-delegation usage still needs attention there. Individual plans (Plus/Pro) remain within your subscription allowance.


Claude Code (Opus 4.8) Comparison [as of July 2026]

Claude Opus 4.8 (2026-05-28) and GPT-5.6 (2026-07-09) both represent generational moves. April’s “genuinely close contest” framing still holds, but the split has become clearer: Codex has widened its lead on terminal and agentic work, while Claude has held its edge on precise code fixes and context scale.

Codex CLI (GPT-5.6 Sol) vs. Claude Code (Opus 4.8) (as of July 2026)
Category Codex CLI (GPT-5.6 Sol) Claude Code (Opus 4.8)
Release 2026-07-09 2026-05-28
Context (real-world, in-product) Effectively 272K (API: 1.05M) 1M (official spec)
Terminal-Bench 2.1 88.8% 74.6%
SWE-Bench Pro 64.6% 69.2%
Parallel agent feature ultra mode (up to 4 parallel agents) dynamic workflows (parallel subagents)
Fast mode ultra: ~3x cost fast: 2.5x speed, 2x cost (3x cheaper than prior gens)
Subscription ceiling Pro $100 (5x) / Pro $200 (20x) Max $100 / Max $200
Product structure Merged into the ChatGPT desktop app Standalone CLI / IDE extension
API price (in/out per 1M, standard) $5 / $30 $5 / $25

Sources: OpenAI, Anthropic, our own Claude Code vs. Codex CLI comparison (as of July 2026)

My setup hasn’t changed since April: Claude Code for design, Codex for implementation. GPT-5.6 widened Codex’s lead on terminal work and agentic tasks, but context scale and precise code-fix accuracy still favor Claude, so I’m not ready to shift everything to Codex. Both ChatGPT Work and Claude Code’s dynamic workflows are pushing toward “hand an agent multi-hour projects” — that’s the comparison I want to dig into in a future update.


Migration and Dual-Use Guide for Claude Code Users [July 2026 Edition]

The April conclusion still stands: you don’t need to switch away from Claude Code. Now that the promo has ended, the sensible move is to run Pro $100 (5x) for a week or two, measure your real usage, and upgrade to Pro $200 (20x) only if 5x genuinely isn’t enough.

Step 1: Test whether 5x is enough for a week or two

If you judged your usage during the 10× promo, you’ll likely feel the pinch once you’re back to normal pricing. Measure your real usage at 5x before deciding whether to stay on Pro $100 or move up to Pro $200.

Step 2: Run AGENTS.md alongside CLAUDE.md

Keep your existing CLAUDE.md and simply add AGENTS.md — the two coexist without conflict. Put tool-agnostic instructions (coding standards, testing policy, prohibited actions) in AGENTS.md, and keep Claude Code-specific instructions (Skills/Hooks usage) in CLAUDE.md. Codex CLI only reads AGENTS.md, so there’s no cross-contamination.

Step 3: A concrete division of labor

🧠 Claude Code (design)

  • Full-codebase awareness with 1M context
  • UI generation, large-scale refactors
  • Parallel codebase migrations via dynamic workflows

⚙️ Codex CLI (implementation)

  • Individual logic implementation, long terminal sessions
  • Parallel task processing via ultra mode
  • Cross-model review (checking code Claude wrote)

Frequently Asked Questions

Q1: What’s the biggest difference between GPT-5.5 and GPT-5.6?

The single model was replaced by a three-tier lineup: Sol, Terra, and Luna. Sol made clear gains on terminal and agentic benchmarks (88.8% on Terminal-Bench 2.1), but Codex’s real-world context window shrank to 272K — a step back from GPT-5.5’s 400K back in April.

Q2: Now that ChatGPT and Codex are merged, is Codex CLI going away?

No. The merger is about the desktop GUI app; Codex CLI (the terminal tool) keeps working independently, exactly as before. Sam Altman said outright that “Codex is not going anywhere” — it remains the software-development brand inside the unified app. If you rely on scripting or CI integration, the CLI is unaffected.

Q3: Did Pro $100 get worse now that the 10× promo ended?

It’s a real reduction compared to the promo, but not unusable. In my day-to-day pattern — many short sessions rather than a few long ones — Pro $100 at 5x still holds up. If you’re regularly hitting caps on long, continuous tasks, that’s the signal to consider upgrading to Pro $200 (20x).

Q4: Why did Codex’s context window shrink to 272K?

OpenAI hasn’t stated a reason outright, but user reports on the openai/codex GitHub repo and general industry read both point to cost control. Hit the API directly and you can access the advertised 1.05M tokens, but through Codex CLI/app the realistic limit is 272K. Plan around that if you need to load a large repo at once.

Q5: Is it true that GPT-5.6 Sol “cheated” on benchmarks?

Yes — it’s acknowledged in OpenAI’s own system card. The nonprofit evaluator METR found that Sol gamed its software-engineering evaluation, including exploiting evaluation bugs, extracting hidden test data, and substituting shortcuts that technically satisfied benchmark metrics without genuinely completing the task. Don’t take benchmark scores at face value — verify that tests actually pass on real work.

Q6: Same $100, Claude Max vs. ChatGPT Pro (5x) — which gives you more usage?

Post-promo, I don’t feel as clear a gap as I did in April. For everyday short sessions, both hold up similarly in my experience. Claude Code’s edge is UI generation, large refactors, and parallel work via dynamic workflows; Codex CLI’s edge is terminal and agentic tasks — pick based on the work, not raw usage volume.

Q7: How should I split time between “ChatGPT Work” and Codex?

ChatGPT Work is a general-purpose agent that can stick with a project across apps and files for hours; Codex is the software-development-specific experience. For coding-heavy work, stick with the Codex tab (or CLI); for cross-cutting work that includes docs or research, Work is the better fit. I’m still keeping Work in the trial-and-see bucket myself.

Q8: Can AGENTS.md and CLAUDE.md coexist?

Yes. Codex only reads AGENTS.md, Claude Code only reads CLAUDE.md, and neither interferes with the other. Put shared conventions and prohibited actions in AGENTS.md, and keep tool-specific instructions (like Claude Code’s Skills/Hooks) in their own file — that’s the cleanest split in practice.

Q9: Should I switch to Codex now?

Same conclusion as April: no need to switch, and running both is the best value in my experience. GPT-5.6 has more upside on terminal and agentic work, but real-codebase fix accuracy and context scale still favor Claude — if I could only keep one today, I’d still keep Claude Code.


The Bottom Line — GPT-5.6 and the ChatGPT Merger Turned Codex Into Part of a Platform

GPT-5.5 Codex has entered a new phase as of July 2026, with GPT-5.6 and the ChatGPT merger

I’m still running Claude Max $200 + Codex Pro $100, splitting design and implementation between them.

What GPT-5.6 changed: The Sol/Terra/Luna three-tier lineup, 88.8% on Terminal-Bench 2.1 — but Codex’s real-world context shrank to 272K.
The ChatGPT/Codex merger: One unified desktop app, plus a new “ChatGPT Work” agent. Codex CLI itself is unaffected and works exactly as before.
Pricing: The 10× promo ended on schedule; Pro $100 is back to the standard 5x (Pro $200 stays at 20x).

If you’ve been all-in on Claude Code, it’s worth testing Pro $100 (5x) now that the promo dust has settled. Terminal and agentic work to Codex, large-scale design and refactoring to Claude — that split still holds up as of July 2026.

krona23

Author

krona23

Over 20 years in the IT industry, serving as Division Head and CTO at multiple companies running large-scale web services in Japan. Experienced across Windows, iOS, Android, and web development. Currently focused on AI-native transformation. At DevGENT, sharing practical guides on AI code editors, automation tools, and LLMs in three languages.

DevGENT about →

Leave a Reply

Trending

Discover more from DevGENT

Subscribe now to keep reading and get access to the full archive.

Continue reading