Claude Opus 5.5 vs GPT-6 Sol: Which Model for Agentic Work?

Claude Opus 5.5 ($4/$20) and GPT-6 Sol ($2/$10) launched on the same day and score level on capability. Pricing, caching, the 272K surcharge and a worked monthly cost comparison.

Published 23 September 2026 by Jake Hissitt, Founder of Stob.AI.

Two frontier models launched on the same day, 22 September 2026, and between them they now cover most serious production work.

Claude Opus 5.5 at $4 in / $20 out, and GPT-6 Sol at $2 in / $10 out.

They are close enough on capability that price, caching and failure behaviour decide the answer for most teams.

Here is the comparison we use when choosing for client systems.

The specification table   Claude Opus 5.5 GPT-6 Sol Input price / 1M $4.00 $2.00 Output price / 1M $20.00 $10.00 Cached input / 1M $0.20 $0.20 Cache write / 1M $5.00 (5-minute) $2.50 Context window 1,000,000 1,050,000 Max output 128,000 128,000 Long-prompt penalty None published Above 272K input: $4 in / $15 out Reasoning control Adaptive thinking, always on, medium default Reasoning effort none to max Knowledge cutoff June 2026 2026 Batch discount 50% Available Where each one wins GPT-6 Sol wins on price at the same capability tier Sol is half the price of Opus 5.5 on both input and output while scoring at the same level on our reasoning and coding measures.

It also gives you an explicit reasoning-effort dial from none to max, which matters when a workflow has a mix of trivial and hard steps and you do not want to pay for deliberation on the trivial ones.

That dial is the single most useful cost lever in the GPT-6 line.

Set effort low for classification and routing, high for the two or three steps that genuinely need it.

Claude Opus 5.5 wins on long-running agent discipline Opus 5.5 has adaptive thinking always on, with the model deciding how much deliberation a step deserves rather than waiting for you to set it.

In long agent runs that tends to produce fewer wasted turns and fewer confidently wrong tool calls.

It is also the model with the better-documented agent tooling around it, and the one we reach for when an agent runs unattended against real business systems for hours rather than seconds.