GPT-5.6 Luna vs Gemini 3.6 Flash vs Claude Opus 5

GPT-5.6 Luna vs Gemini 3.6 Flash vs Claude Opus 5: Which Tier Fits Your Workload?

2026-09-23

Key Takeaways

  • These three models sit at deliberately different price points: Luna at $0.20 and $1.20 per million tokens, Gemini 3.6 Flash at $0.75 and $3.75 under current promotional pricing, Claude Opus 5 at $5.00 and $25.00.
  • Opus 5 costs 25 times more per input token than Luna, and its 96.0% SWE-bench Verified score explains why teams still pay it.
  • Luna beats Gemini 3.6 Flash on three of four shared coding benchmarks, including DeepSWE v1.1 at 67% against 49%.
  • Gemini 3.6 Flash leads on vision tasks, averaging 83% against Luna’s 74% in independent testing, and on long-context retrieval.
  • All three offer roughly a 1M-token context window, so context size is not a deciding factor.
  • Gemini 3.6 Flash pricing returns to $1.50 and $7.50 on 1 January 2027, which changes the arithmetic for anyone planning past this year.
  • Independent intelligence rankings disagree sharply on Luna and Flash, so run your own evaluation rather than trusting a single index.
GPT-5.6 Luna vs Gemini 3.6 Flash vs Claude Opus 5 - artistic impression. Image source: Alius Noreika / Google Gemini

GPT-5.6 Luna vs Gemini 3.6 Flash vs Claude Opus 5 – artistic impression. Image source: Alius Noreika / Google Gemini

Picking the Right Tier in One Paragraph

Choose GPT-5.6 Luna when you are running the same well-defined task thousands of times a day and cost per call dominates everything else. Choose Gemini 3.6 Flash when your pipeline handles images, documents, audio or video, or when raw output speed at around 304 tokens per second matters more than the last few points of coding accuracy. Choose Claude Opus 5 when a task has to finish correctly without a human checking it, particularly agentic coding, long refactors and multi-step workflows where a failed run costs more than the tokens saved.

The gap between these tiers is not marketing. A workload consuming 400 million input tokens and 75 million output tokens a month costs roughly $170 on Luna, about $581 on Gemini 3.6 Flash at promotional rates, and around $3,875 on Opus 5. That is a 23-fold spread on identical volume, which means the question is never simply which model is smartest, but how often a cheaper model’s mistakes cost more than the difference.

Price and Specification Comparison

Specification GPT-5.6 Luna Gemini 3.6 Flash Claude Opus 5
Developer OpenAI Google Anthropic
Release date 9 July 2026 21 July 2026 24 July 2026
Input per 1M tokens $0.20 $0.75 (promotional) $5.00
Output per 1M tokens $1.20 $3.75 (promotional) $25.00
Cached input per 1M $0.02 $0.075 Standard Opus rate
Context window 1,050,000 1,048,576 1,000,000
Maximum output 128,000 65,536 128,000
Knowledge cutoff February 2026 Not disclosed May 2026
Input modalities Text, image Text, image, audio, video, PDF Text, image
Effort control none to max, six levels Thinking, medium default low to max, five levels

One pricing detail deserves attention. Gemini 3.6 Flash launched at $1.50 and $7.50 per million tokens. When Gemini 3.7 Flash arrived on 13 August 2026, Google moved 3.6 Flash onto the same introductory rate of $0.75 and $3.75. Both return to $1.50 and $7.50 on 1 January 2027, which doubles the cost of anything you build on it today. Anyone budgeting beyond this year should model the standard rate, not the promotional one. Our guide to how token pricing translates into real monthly bills covers this kind of planning in more depth.

Benchmark Results, With the Disagreements Left In

Coding and Agentic Work

On the shared coding benchmarks, Luna holds a consistent lead over Gemini 3.6 Flash despite costing less.

Benchmark GPT-5.6 Luna Gemini 3.6 Flash Claude Opus 5
SWE-Bench Pro 62.7% 58.7% Not published
SWE-bench Verified Not published Not published 96.0%
DeepSWE v1.1 67% 49% 74% (public leaderboard)
Terminal-Bench 2.1 84.7% 78.0% Not published
GDPval-AA v2 1584 Elo 1421 Elo 1824 Elo
Frontier-Bench v0.1 Not published Not published 43.3%
ARC-AGI-3 Not published Not published 30.2%
MRCR v2 (8-needle) Loses Wins Not published

Opus 5’s separation is largest exactly where failure is most expensive. Its 96.0% SWE-bench Verified result and 1824 GDPval-AA v2 Elo put it in a different category from either budget model, and Anthropic more than doubled Opus 4.8’s Frontier-Bench score with it. For teams choosing a model to write production code, our tier list of the best LLMs for vibe coding works through where each model breaks down in practice.

Vision and Multimodal

Gemini 3.6 Flash reverses the picture on anything visual. Roboflow’s independent vision evaluations put it at 83% average against Luna’s 74%, with the widest margins on reasoning at 78% against 60% and identification at 99% against 83%. Luna edges ahead only on OCR, 90.7% against 88.2%. Gemini also accepts audio, video and PDFs natively, where Luna handles text and images only. If your pipeline processes invoices, screenshots, product photos or recorded calls, this difference outweighs the coding scores.

Where the Rankings Contradict Each Other

Composite intelligence indexes disagree enough that none should be treated as settled. Artificial Analysis has published Gemini 3.6 Flash at 34 against Luna at 32 in one high-effort comparison, and at 50 against 33 in another, while separate coverage records Gemini 3.6 Flash at 50 on the same index. LLM Stats places the two within two points, at 43.5 for Gemini and 45.4 for Luna, and notes Luna wins three of four shared benchmarks while Gemini wins long-context retrieval. Output speed figures vary the same way: one measurement gives Gemini 213 tokens per second against Luna’s 152, while Google’s own reporting cites 304 tokens per second for 3.6 Flash.

These inconsistencies come from different effort settings, different test dates and different harnesses rather than from bad measurement. The practical response is to run your own evaluation on your own prompts before committing a production workload. Our roundup of which generative AI model responds fastest shows how much latency figures shift once reasoning modes enter the picture.

Real Monthly Cost on Identical Volume

Monthly workload GPT-5.6 Luna Gemini 3.6 Flash Claude Opus 5
1M input / 100K output (single job) $0.32 $1.13 $7.50
400M input / 75M output $170 $581 $3,875
Same, with 80% cache hits on input $112 Varies by cache rate Varies by cache rate
Same, at Gemini’s 2027 standard rate $170 $1,162 $3,875

Figures are calculated from published list prices and exclude tool call fees, image inputs and batch discounts. They illustrate scale rather than predicting an invoice.

Availability Differences That Catch Teams Out

Luna is not selectable in standard ChatGPT conversations. It is a developer and workflow model reached through the OpenAI API, Codex, ChatGPT Work or GitHub Copilot. Free and Go users receive Luna as their default model and it powers the Think button, but paid users cannot pick it from the model menu, which caused visible confusion on launch day. GitHub Copilot added all three GPT-5.6 tiers on 9 July, with Luna available on Pro, Pro+, Max, Business and Enterprise plans.

Gemini 3.6 Flash is available through the Gemini API with a free tier alongside paid usage. Claude Opus 5 is the default model on Claude Max subscriptions and the strongest option on Claude Pro, and reaches developers through the Claude API, Amazon Bedrock, Google Cloud and Microsoft Foundry. It supports zero data retention, which matters for regulated workloads. Readers comparing what each subscription actually unlocks may find our guide to using generative AI without a paid plan useful.

A Routing Strategy Beats a Single Choice

Most production teams running meaningful volume end up using two or three of these rather than picking one. The pattern that works is straightforward: send classification, routing, extraction, short-form generation and moderation to Luna; send document, image and video processing to Gemini 3.6 Flash; reserve Opus 5 for the work where an error propagates, such as code that ships, analysis that informs a decision, or agent runs nobody is watching.

The economics favour this arrangement more than they used to. OpenAI cut Luna’s price by 80% on 30 July 2026, three weeks after launch, dropping it from $1.00 and $6.00 to $0.20 and $1.20. Sam Altman framed the goal as offering “the best price/intelligence tradeoff at every level.” The company attributed the cut to rewritten production GPU kernels that reduced serving costs by about 20%, a redesigned speculative decoding system that improved token generation efficiency by over 15%, and better prompt caching across multi-step agent workflows. Pressure from lower-cost open-weight models, covered in our piece on whether cheap Chinese AI models can rival OpenAI and Anthropic, helped force the move.

The Bottom Line by Use Case

Workload Best fit Why
High-volume classification and routing GPT-5.6 Luna Cheapest per call by a wide margin, adequate reasoning
Document, image or video processing Gemini 3.6 Flash Native multimodal input and stronger vision scores
Production code generation Claude Opus 5 96.0% SWE-bench Verified
Long unattended agent runs Claude Opus 5 Failure cost exceeds token savings
Long-context retrieval across large documents Gemini 3.6 Flash Wins MRCR v2 8-needle against Luna
Terminal and agentic coding on a budget GPT-5.6 Luna 84.7% Terminal-Bench 2.1 at a fraction of Opus pricing
Latency-critical interactive features Either budget model Test both; published speed figures conflict

Prices, benchmark scores and availability described here are current as of September 2026. Promotional rates expire and benchmark leaderboards move; verify figures with each vendor before committing production budget.

If you are interested in this topic, we suggest you check our articles:

Sources: OpenAI — GPT-5.6 Luna model documentation, The New Stack — OpenAI slashes API costs, OpenRouter — Gemini 3.6 Flash pricing and benchmarks, Fello AI — Gemini 3.6 Flash pricing and benchmarks, Roboflow — Gemini 3.6 Flash vs GPT-5.6 Luna vision evals, LLM Stats — Gemini 3.6 Flash vs GPT-5.6 Luna

Written by Alius Noreika

GPT-5.6 Luna vs Gemini 3.6 Flash vs Claude Opus 5: Which Tier Fits Your Workload?
We use cookies and other technologies to ensure that we give you the best experience on our website. If you continue to use this site we will assume that you are happy with it..
Privacy policy