GPT-6 Luna vs Gemini 3.8 Flash: Cheapest Model for Volume
GPT-6 Luna costs $0.10/$0.50; Gemini 3.8 Flash costs $0.75/$3.75. A cost-per-million-tasks table, where Flash earns its premium, and the routing rule for high-volume production work.
Published 23 September 2026 by Jake Hissitt, Founder of Stob.AI.
Most production AI spend is not going on hard reasoning.
It goes on millions of small, repetitive calls: classify this email, extract these fields, route this record, summarise this note.
For that work, two models now dominate the shortlist.
GPT-6 Luna at $0.10 in / $0.50 out, and Gemini 3.8 Flash at $0.75 in / $3.75 out.
Side by side GPT-6 Luna Gemini 3.8 Flash Input / 1M $0.10 $0.75 Output / 1M $0.50 $3.75 Cached input / 1M $0.01 Available Context 1,050,000 1,048,576 Max output 128,000 65,536 Long-prompt rule Above 272K: $0.20 / $0.75 Published rate applies to prompts up to 200K Strengths Cheapest frontier-family tier, reasoning-effort control Multimodal, computer use, search grounding, speed Published benchmarks Shares the GPT-6 family controls and context Terminal-Bench 2.1 89.4%, HLE-Verified 54.9%, OSWorld 2.0 59.0% Cost per million tasks Assume a typical classification call: 800 input tokens, 120 output tokens, no caching.
Model Input cost Output cost Per 1M tasks GPT-6 Luna $80 $60 $140 Gemini 3.8 Flash $600 $450 $1,050 GPT-6 Sol (for reference) $1,600 $1,200 $2,800 Claude Opus 5.5 (for reference) $3,200 $2,400 $5,600 Luna is roughly 7.5 times cheaper than Gemini 3.8 Flash on this shape of work, and 20 times cheaper than Sol.
With cached input at $0.01 per million, a Luna workflow with a stable prompt prefix gets cheaper again.
So why would you pay for Flash?
Because Gemini 3.8 Flash is not really the same product.
It is a multimodal model with computer use and search grounding built in, and it performs at a level Luna is not aimed at.
Images, video and audio in.