How We Rank AI Models: Our Comparison Tool Methodology
The scoring formula, weightings and availability penalties behind our 230+ model comparison tool, why cost is a tie-breaker rather than a driver, and how recalibration works when a new generation lands.
Published 23 September 2026 by Jake Hissitt, Founder of Stob.AI.
Our AI model comparison tool ranks more than 230 models.
People ask how the order is decided, so here it is in full.
The six scores Every model carries six scores from 1 to 10: Reasoning — multi-step problem solving, analysis and judgement under ambiguity.
Coding — code generation, debugging and agentic software work.
Speed — practical latency and throughput for interactive use.
Cost efficiency — capability delivered per dollar at published rates.
Context length — usable input window and maximum output.
Multimodal — image, audio and video handling.
How the rank is calculated Rank is capability-first.
Reasoning and coding carry triple weight, context length and multimodal single weight, and cost efficiency and speed act as tie-breakers rather than drivers: Score = (reasoning x 3) + (coding x 3) + context + multimodal + (cost efficiency x 0.4) + (speed x 0.3) − availability penalty The availability penalty is what stops the list filling up with models you cannot responsibly build on: Status Penalty Superseded by a newer model −6 Active previous generation −3 Limited, invitation-only or trusted access −2 Preview −1.5 A superseded model with perfect capability scores therefore sits well below a current model with slightly lower ones.
That is deliberate.