Anthropic shipped Claude Opus 5.5 and OpenAI shipped GPT-6 Sol within 24 hours of each other on 22 September 2026. Both sell the same story — better and cheaper — and both published the benchmark charts they win. Shortest answer: GPT-6 Sol costs half as much per token, but Opus 5.5 wins the one benchmark both vendors measured the same way. Every other number was measured under different conditions and cannot be placed side by side, which is exactly what most comparison tables published this week get wrong.
Half the price per token — but identical cache pricing
Per million tokens, checked against both vendors' official pricing pages on 23 September 2026:
| Model | Input | Output | Cache read |
|---|---|---|---|
| Claude Opus 5.5 | $4 | $20 | $0.20 |
| GPT-6 Sol | $2 | $10 | $0.20 |
| Claude Fable 5.1 | $10 | $50 | $0.25 |
| GPT-6 Luna | $0.10 | $0.50 | $0.01 |
Barely anyone has noticed: cache reads cost exactly $0.20 on both. Anthropic prices Opus 5.5 cache reads at 5% of input ($4 × 5%); OpenAI discounts cached input by 90% ($2 × 10%). They land on the same number. For an agent with a long fixed system prompt, the input-side gap nearly disappears and the difference becomes output-only.

The one benchmark that actually compares: AutomationBench
AutomationBench measures business workflows across multiple applications, and it is the only test both vendors have published where the reference scores match exactly — Anthropic and OpenAI both report Fable 5.1 at 31.4% and Opus 5 at 26.9%. When the anchors agree, the new models can be compared.
| Model | AutomationBench | Source |
|---|---|---|
| Claude Opus 5.5 | 40.0% | Anthropic |
| GPT-6 Sol (xhigh) | 33.2% | OpenAI |
| Claude Fable 5.1 | 31.4% | both agree |
| GPT-6 Astra (low) | 30.3% | OpenAI |
| Claude Opus 5 | 26.9% | both agree |
| GPT-5.6 Sol | 18.1% | OpenAI |

Score is not the whole story — cost per task differs 11×
OpenAI published a figure Anthropic did not: on AutomationBench, Sol at xhigh effort costs $0.27 per task, and Opus 5 costs 11.1× that per task. That is why "which is smarter" and "which is worth it" never resolve into the same answer. Opus 5.5 leads by 6.8 points, but Anthropic has not published cost per task for 5.5, so nobody can yet calculate what the real gap costs. Anyone reading 40.0% as proof of better value is drawing a conclusion the data does not support.
The trap: OSWorld 2.0 scores that do not compare
Several comparison tables this week put both vendors' OSWorld 2.0 numbers side by side. Check the reference model and the problem is immediate:
- •Anthropic reports OSWorld 2.0 — Opus 5.5 at 81.8%, Fable 5.1 at 80.7%, Opus 5 at 74.0%
- •OpenAI reports OSWorld 2.0 Offline — GPT-6 Sol (xhigh) at 60.5%, Opus 5 (med) at 60.3%
- •The same model, Opus 5, scores 74.0% and 60.3% — 13.7 points apart, because it is a different variant of the suite run at a different effort level
A quick test before trusting any comparison table: find a model that appears in both vendors' reports and check whether its score matches. If it does, the comparison holds. If it does not, the two were measured differently and that table cannot support a decision.

Real monthly cost
At 33.19 THB per USD (23 Sep 2026), two workloads common in Thai teams. The "90% cache" rows model an agent with a long fixed system prompt, which is the normal shape of production work:
| Workload / month | Opus 5.5 | GPT-6 Sol | Gap |
|---|---|---|---|
| Support bot, 2M/0.4M tokens | ฿531 | ฿266 | 2.0× |
| Support bot + 90% cache | ฿304 | ฿158 | 1.9× |
| Dev team coding agent, 20M/3M | ฿4,647 | ฿2,323 | 2.0× |
| Dev team + 90% cache | ฿2,376 | ฿1,248 | 1.9× |
The number to remember: the gap stays around 2× regardless of caching. For a typical Thai SME that is a few hundred baht a month — less than an hour of staff time. Choosing on price alone means deciding on the smallest variable in the equation.

Spec differences, and what OpenAI has not disclosed
Opus 5.5 documents its specs fully: a 1M token context window, 128K max output, default effort medium, a June 2026 knowledge cutoff, model ID claude-opus-5-5, and a commitment not to retire it before 22 September 2027.
For GPT-6 Sol (gpt-6-sol), no source has published the context window, which matters enormously for agents that read whole codebases. If your work depends on long context, that is a gap you have to test yourself rather than read about.
- •Opus 5.5 is Anthropic's own recommended default, not its top model — Fable 5.1 ($10/$50) still sits above it for demanding reasoning and long-horizon agents
- •GPT-6 Sol is likewise OpenAI's mid-tier — GPT-6 Astra, released 19 days earlier, is the flagship and beats Sol on OSWorld Offline (72.6% vs 60.5%)
- •So Opus 5.5 versus Sol compares each vendor's working workhorse, which is fairer than pitting a flagship against a mid-tier
- •Anthropic's Batch API is 50% off, narrowing the gap further for work that does not need an immediate answer
Which should a Thai business choose?
My read, with the caveat that both launched a day ago and nobody has long-run production data yet: Choose GPT-6 Sol for high-volume, low-complexity work — summarising documents, answering customer chats, translation, classification. A 6.8-point AutomationBench gap barely shows up in what your customers see, and you halve the bill every month. Choose Opus 5.5 for long-running agents where mistakes are expensive — code that ships, automations spanning several systems that fail silently. A few thousand baht a month costs less than cleaning up after one bad agent run. Do not migrate on launch-day news. The real cost of switching models is not the API bill; it is rebuilding and re-testing every prompt. If what you have works, wait a month for real usage data.
- •Further reading: AI pricing compared for Thai teams
- •The previous round: GPT-5.6 vs Claude Fable 5
- •On the Anthropic side: Claude Fable 5 · Claude pricing for Thai business
Want to know which model fits your workload without testing it yourself? Free 30-minute consultation or see AI consulting — we cost it out against your actual volumes first.
Arm - CherCode
Full-Stack Developer & Founder
Software developer with 5+ years of experience in Web Development, AI Integration, and Automation. Specializing in Next.js, React, n8n, and LLM Integration. Founder of CherCode, building systems for Thai businesses.
Portfolio


