Skip to main content
AISep 23, 20269 min

Claude Opus 5.5 vs GPT-6 Sol: Half the Price, But Worth It?

Anthropic and OpenAI shipped within 24 hours of each other on 22 Sep 2026. GPT-6 Sol is half the price per token, but Opus 5.5 wins the one benchmark both vendors measured the same way — and the cache-read price is identical at $0.20.

Claude Opus 5.5 vs GPT-6 Sol เทียบราคาและ benchmark - CherCode

Anthropic shipped Claude Opus 5.5 and OpenAI shipped GPT-6 Sol within 24 hours of each other on 22 September 2026. Both sell the same story — better and cheaper — and both published the benchmark charts they win. Shortest answer: GPT-6 Sol costs half as much per token, but Opus 5.5 wins the one benchmark both vendors measured the same way. Every other number was measured under different conditions and cannot be placed side by side, which is exactly what most comparison tables published this week get wrong.

Half the price per token — but identical cache pricing

Per million tokens, checked against both vendors' official pricing pages on 23 September 2026:

ModelInputOutputCache read
Claude Opus 5.5$4$20$0.20
GPT-6 Sol$2$10$0.20
Claude Fable 5.1$10$50$0.25
GPT-6 Luna$0.10$0.50$0.01

Barely anyone has noticed: cache reads cost exactly $0.20 on both. Anthropic prices Opus 5.5 cache reads at 5% of input ($4 × 5%); OpenAI discounts cached input by 90% ($2 × 10%). They land on the same number. For an agent with a long fixed system prompt, the input-side gap nearly disappears and the difference becomes output-only.

API pricing per million tokens for Opus 5.5, GPT-6 Sol, Fable 5.1 and GPT-6 Luna

The one benchmark that actually compares: AutomationBench

AutomationBench measures business workflows across multiple applications, and it is the only test both vendors have published where the reference scores match exactly — Anthropic and OpenAI both report Fable 5.1 at 31.4% and Opus 5 at 26.9%. When the anchors agree, the new models can be compared.

ModelAutomationBenchSource
Claude Opus 5.540.0%Anthropic
GPT-6 Sol (xhigh)33.2%OpenAI
Claude Fable 5.131.4%both agree
GPT-6 Astra (low)30.3%OpenAI
Claude Opus 526.9%both agree
GPT-5.6 Sol18.1%OpenAI
AutomationBench scores for Opus 5.5, GPT-6 Sol, Fable 5.1 and Opus 5

Score is not the whole story — cost per task differs 11×

OpenAI published a figure Anthropic did not: on AutomationBench, Sol at xhigh effort costs $0.27 per task, and Opus 5 costs 11.1× that per task. That is why "which is smarter" and "which is worth it" never resolve into the same answer. Opus 5.5 leads by 6.8 points, but Anthropic has not published cost per task for 5.5, so nobody can yet calculate what the real gap costs. Anyone reading 40.0% as proof of better value is drawing a conclusion the data does not support.

The trap: OSWorld 2.0 scores that do not compare

Several comparison tables this week put both vendors' OSWorld 2.0 numbers side by side. Check the reference model and the problem is immediate:

  • Anthropic reports OSWorld 2.0 — Opus 5.5 at 81.8%, Fable 5.1 at 80.7%, Opus 5 at 74.0%
  • OpenAI reports OSWorld 2.0 Offline — GPT-6 Sol (xhigh) at 60.5%, Opus 5 (med) at 60.3%
  • The same model, Opus 5, scores 74.0% and 60.3% — 13.7 points apart, because it is a different variant of the suite run at a different effort level

A quick test before trusting any comparison table: find a model that appears in both vendors' reports and check whether its score matches. If it does, the comparison holds. If it does not, the two were measured differently and that table cannot support a decision.

Opus 5 scores 74.0% and 60.3% on OSWorld 2.0 because the variants differ

Real monthly cost

At 33.19 THB per USD (23 Sep 2026), two workloads common in Thai teams. The "90% cache" rows model an agent with a long fixed system prompt, which is the normal shape of production work:

Workload / monthOpus 5.5GPT-6 SolGap
Support bot, 2M/0.4M tokens฿531฿2662.0×
Support bot + 90% cache฿304฿1581.9×
Dev team coding agent, 20M/3M฿4,647฿2,3232.0×
Dev team + 90% cache฿2,376฿1,2481.9×

The number to remember: the gap stays around 2× regardless of caching. For a typical Thai SME that is a few hundred baht a month — less than an hour of staff time. Choosing on price alone means deciding on the smallest variable in the equation.

Monthly cost in Thai baht for Opus 5.5 versus GPT-6 Sol

Spec differences, and what OpenAI has not disclosed

Opus 5.5 documents its specs fully: a 1M token context window, 128K max output, default effort medium, a June 2026 knowledge cutoff, model ID claude-opus-5-5, and a commitment not to retire it before 22 September 2027. For GPT-6 Sol (gpt-6-sol), no source has published the context window, which matters enormously for agents that read whole codebases. If your work depends on long context, that is a gap you have to test yourself rather than read about.

  • Opus 5.5 is Anthropic's own recommended default, not its top model — Fable 5.1 ($10/$50) still sits above it for demanding reasoning and long-horizon agents
  • GPT-6 Sol is likewise OpenAI's mid-tier — GPT-6 Astra, released 19 days earlier, is the flagship and beats Sol on OSWorld Offline (72.6% vs 60.5%)
  • So Opus 5.5 versus Sol compares each vendor's working workhorse, which is fairer than pitting a flagship against a mid-tier
  • Anthropic's Batch API is 50% off, narrowing the gap further for work that does not need an immediate answer

Which should a Thai business choose?

My read, with the caveat that both launched a day ago and nobody has long-run production data yet: Choose GPT-6 Sol for high-volume, low-complexity work — summarising documents, answering customer chats, translation, classification. A 6.8-point AutomationBench gap barely shows up in what your customers see, and you halve the bill every month. Choose Opus 5.5 for long-running agents where mistakes are expensive — code that ships, automations spanning several systems that fail silently. A few thousand baht a month costs less than cleaning up after one bad agent run. Do not migrate on launch-day news. The real cost of switching models is not the API bill; it is rebuilding and re-testing every prompt. If what you have works, wait a month for real usage data.

Want to know which model fits your workload without testing it yourself? Free 30-minute consultation or see AI consulting — we cost it out against your actual volumes first.

Share:
Arm - CherCode

Arm - CherCode

Full-Stack Developer & Founder

Software developer with 5+ years of experience in Web Development, AI Integration, and Automation. Specializing in Next.js, React, n8n, and LLM Integration. Founder of CherCode, building systems for Thai businesses.

Portfolio

Related Service

AI Consulting

Learn More