GPT-5.6 Luna vs Gemini 3.6 Flash: Which Model Wins?

GPT-5.6 Luna vs Gemini 3.6 Flash: Which Comes Out on Top?

2026-09-28

Key Takeaways

  • GPT-5.6 Luna wins on coding and price. It leads SWE-Bench Pro, DeepSWE v1.1 and Terminal-Bench 2.1 while costing roughly a quarter as much per token.
  • Gemini 3.6 Flash wins on multimodal work and throughput, averaging 83% on independent vision evaluations against Luna’s 74%, with native audio, video and PDF input.
  • Luna costs $0.20 and $1.20 per million tokens; Gemini 3.6 Flash costs $0.75 and $3.75 under promotional pricing that expires on 1 January 2027.
  • Gemini’s price doubles to $1.50 and $7.50 next January, widening the gap to 7.5 times on input.
  • Both offer roughly a 1M-token context window, but Gemini caps output at 65,536 tokens against Luna’s 128,000.
  • Composite intelligence rankings contradict each other on these two models, so neither index should decide your choice.
  • Gemini 3.6 Flash has been superseded twice, by 3.7 Flash in August and 3.8 Flash in September 2026.
Generative AI coding systems - artistic impression. Image source: Alius Noreika / AI

Generative AI coding systems – artistic impression. Image source: Alius Noreika / AI

The Verdict in Two Paragraphs

For text-in, text-out work at scale, GPT-5.6 Luna comes out on top. It wins three of the four benchmarks both models publish, posting 62.7% on SWE-Bench Pro against 58.7%, 67% on DeepSWE v1.1 against 49%, and 84.7% on Terminal-Bench 2.1 against 78.0%. It also scores 1584 Elo on GDPval-AA v2 against Gemini’s 1421. It does this while costing $0.20 per million input tokens against Gemini’s $0.75, and that gap widens to 7.5 times when Google’s promotional pricing ends.

For anything involving images, documents, audio or video, Gemini 3.6 Flash comes out on top. Roboflow’s independent vision evaluations put it at 83% average against Luna’s 74%, with the widest margins on visual reasoning at 78% against 60% and identification at 99% against 83%. Gemini accepts audio, video and PDFs natively where Luna handles text and images only, and Google reports output speeds around 304 tokens per second with built-in computer use scoring 83% on OSWorld. Gemini also wins the single long-context benchmark both publish, MRCR v2 at eight needles.

Specification Comparison

Specification GPT-5.6 Luna Gemini 3.6 Flash
Developer OpenAI Google
Release date 9 July 2026 21 July 2026
Model ID gpt-5.6-luna gemini-3.6-flash
Input per 1M tokens $0.20 $0.75 promotional, $1.50 from January 2027
Output per 1M tokens $1.20 $3.75 promotional, $7.50 from January 2027
Cached input per 1M $0.02 $0.075 promotional, $0.15 standard
Context window 1,050,000 tokens 1,048,576 tokens
Maximum output 128,000 tokens 65,536 tokens
Knowledge cutoff 16 February 2026 Not published
Input modalities Text, image Text, image, audio, video, PDF
Reasoning control none, low, medium, high, xhigh, max Thinking supported, medium default
Fine-tuning Not supported Not supported

Benchmark Results Side by Side

Benchmark GPT-5.6 Luna Gemini 3.6 Flash Winner
SWE-Bench Pro 62.7% 58.7% Luna
DeepSWE v1.1 67% 49% Luna
Terminal-Bench 2.1 84.7% 78.0% Luna
GDPval-AA v2 1584 Elo 1421 Elo Luna
MRCR v2 (8-needle) Lower Higher Gemini
Vision average 74% 83% Gemini
Visual reasoning 60.5% 77.7% Gemini
Identification 83.3% 99.0% Gemini
OCR 90.7% 88.2% Luna
Data extraction 80.4% 95.9% Gemini
Counting 67.1% 80.2% Gemini
Object detection 61.0% 57.1% Luna

The split is clean. Luna owns text-based coding and agentic work; Gemini owns everything that starts with a picture. OCR and object detection are the two exceptions where Luna edges ahead on visual input, which suggests its weakness is visual reasoning rather than visual perception.

Why the Composite Rankings Disagree

This is the part most comparisons skip, and it matters. Artificial Analysis has published Gemini 3.6 Flash at 34 against Luna at 32 on its Intelligence Index in one high-effort comparison, and 50 against 33 in another. Separate coverage records Gemini 3.6 Flash at 50 on the same index, and elsewhere Luna appears at 51 against Gemini 3.5 Flash-Lite’s 36. LLM Stats places them within two points of each other at 43.5 and 45.4.

Output speed is equally unsettled. One measurement gives Gemini 213 tokens per second against Luna’s 152, with Luna responding far faster on time to first token at 2.00 seconds against 16.09. Google’s own reporting cites 304 tokens per second for 3.6 Flash. These are not errors so much as different effort settings, different test dates and different serving conditions producing different answers.

The conclusion for a buying decision is that composite indexes cannot separate these two models reliably. Individual task benchmarks can, and your own evaluation on your own prompts can. Our overview of which generative AI models respond fastest shows how much reasoning modes distort latency figures across the board, and our survey of how AI benchmark results are produced explains why single-source scores move so much.

Cost Over a Real Workload

Workload GPT-5.6 Luna Gemini 3.6 Flash (promotional) Gemini 3.6 Flash (from Jan 2027)
1M input / 100K output $0.32 $1.13 $2.25
50M input / 10M output $22 $75 $150
400M input / 75M output $170 $581 $1,162
4B input / 500M output $1,400 $4,875 $9,750

Luna’s advantage is 3.4 times at current Gemini promotional rates and 6.8 times once those rates expire. On a four billion token monthly pipeline, that is a difference of roughly $8,350 a month. Two caveats work against Luna, though: prompts exceeding 272,000 input tokens are billed at double input and 1.5 times output for the entire request, and Luna’s 128K output ceiling doubles Gemini’s, which can mean fewer calls for long generation tasks. Readers modelling this for their own stack can start from our guide to token pricing in practice.

Availability and Access

Luna is not selectable inside standard ChatGPT conversations. It reaches developers through the OpenAI API, Codex, ChatGPT Work and GitHub Copilot, where it arrived on 9 July across Pro, Pro+, Max, Business and Enterprise plans. Free and Go users get it as their default model with unlimited text chats and a Think button that routes harder questions through higher reasoning. There is no free API tier, so evaluation requires a funded account.

Gemini 3.6 Flash is available through the Gemini API with a free tier alongside paid usage, and through Google’s broader Gemini surfaces. Its position in Google’s own lineup has moved twice since launch: Gemini 3.7 Flash arrived on 13 August 2026 scoring 43.6% on FrontierCode 1.1 Main against 3.6 Flash’s 34.4%, and Gemini 3.8 Flash followed in September with 90.8% on Terminal-Bench 2.1 against 3.7 Flash’s 81.6%. Teams selecting 3.6 Flash today should confirm that a newer Flash model at the same price is not simply the better purchase. Our coverage of how the Gemini Flash line fits production workflows gives additional background on the family.

Practical Recommendations

Task Pick Reason
Code generation and repair GPT-5.6 Luna Wins every shared coding benchmark
Terminal and agentic workflows GPT-5.6 Luna 84.7% Terminal-Bench 2.1
Invoice, receipt or form processing Gemini 3.6 Flash 95.9% data extraction against 80.4%
Image and video understanding Gemini 3.6 Flash Native multimodal input, stronger visual reasoning
Pure OCR GPT-5.6 Luna 90.7% against 88.2%, at lower cost
Retrieval across very long documents Gemini 3.6 Flash Wins MRCR v2 at eight needles
Long-form generation GPT-5.6 Luna 128K output ceiling against 65K
Highest volume, lowest cost GPT-5.6 Luna 3.4x cheaper now, 6.8x from January 2027
Latency-critical interactive features Test both Published speed figures conflict

The Honest Summary

Luna comes out on top overall, on the strength of winning most shared benchmarks while costing substantially less. That verdict holds for the majority of production workloads, which are text in and text out. But it flips completely the moment a pipeline touches images, audio, video or PDFs, where Gemini 3.6 Flash is the better instrument and the price difference stops being the deciding factor.

The stronger play for teams with real volume is routing rather than selection: text and code to Luna, multimodal to Gemini, and neither to the harder reasoning problems that belong on a frontier model. Both companies have been cutting prices under pressure from lower-cost open-weight competitors, a dynamic we examine in our look at whether cheap Chinese AI models can rival the established labs, and the direction of travel favours anyone willing to spread a workload across more than one provider.

Prices, benchmark scores and model availability described here are current as of September 2026. Promotional rates expire and newer Flash models have since shipped; verify current figures with each vendor before committing.

If you are interested in this topic, we suggest you check our articles:

Sources: OpenAI — GPT-5.6 Luna model documentation, OpenRouter — Gemini 3.6 Flash pricing and benchmarks, Roboflow — Gemini 3.6 Flash vs GPT-5.6 Luna vision evals, MyClaw — Gemini 3.6 Flash vs GPT-5.6 Luna: coding and cost, LLM Stats — Gemini 3.6 Flash vs GPT-5.6 Luna, Fello AI — Gemini 3.6 Flash pricing and benchmarks, DataCamp — Gemini 3.8 Flash features and benchmarks

Written by Alius Noreika

GPT-5.6 Luna vs Gemini 3.6 Flash: Which Comes Out on Top?
We use cookies and other technologies to ensure that we give you the best experience on our website. If you continue to use this site we will assume that you are happy with it..
Privacy policy