Key Takeaways
- GPT-5.6 Luna wins on coding and price. It leads SWE-Bench Pro, DeepSWE v1.1 and Terminal-Bench 2.1 while costing roughly a quarter as much per token.
- Gemini 3.6 Flash wins on multimodal work and throughput, averaging 83% on independent vision evaluations against Luna’s 74%, with native audio, video and PDF input.
- Luna costs $0.20 and $1.20 per million tokens; Gemini 3.6 Flash costs $0.75 and $3.75 under promotional pricing that expires on 1 January 2027.
- Gemini’s price doubles to $1.50 and $7.50 next January, widening the gap to 7.5 times on input.
- Both offer roughly a 1M-token context window, but Gemini caps output at 65,536 tokens against Luna’s 128,000.
- Composite intelligence rankings contradict each other on these two models, so neither index should decide your choice.
- Gemini 3.6 Flash has been superseded twice, by 3.7 Flash in August and 3.8 Flash in September 2026.
The Verdict in Two Paragraphs
For text-in, text-out work at scale, GPT-5.6 Luna comes out on top. It wins three of the four benchmarks both models publish, posting 62.7% on SWE-Bench Pro against 58.7%, 67% on DeepSWE v1.1 against 49%, and 84.7% on Terminal-Bench 2.1 against 78.0%. It also scores 1584 Elo on GDPval-AA v2 against Gemini’s 1421. It does this while costing $0.20 per million input tokens against Gemini’s $0.75, and that gap widens to 7.5 times when Google’s promotional pricing ends.
For anything involving images, documents, audio or video, Gemini 3.6 Flash comes out on top. Roboflow’s independent vision evaluations put it at 83% average against Luna’s 74%, with the widest margins on visual reasoning at 78% against 60% and identification at 99% against 83%. Gemini accepts audio, video and PDFs natively where Luna handles text and images only, and Google reports output speeds around 304 tokens per second with built-in computer use scoring 83% on OSWorld. Gemini also wins the single long-context benchmark both publish, MRCR v2 at eight needles.
Specification Comparison
| Specification | GPT-5.6 Luna | Gemini 3.6 Flash |
|---|---|---|
| Developer | OpenAI | |
| Release date | 9 July 2026 | 21 July 2026 |
| Model ID | gpt-5.6-luna |
gemini-3.6-flash |
| Input per 1M tokens | $0.20 | $0.75 promotional, $1.50 from January 2027 |
| Output per 1M tokens | $1.20 | $3.75 promotional, $7.50 from January 2027 |
| Cached input per 1M | $0.02 | $0.075 promotional, $0.15 standard |
| Context window | 1,050,000 tokens | 1,048,576 tokens |
| Maximum output | 128,000 tokens | 65,536 tokens |
| Knowledge cutoff | 16 February 2026 | Not published |
| Input modalities | Text, image | Text, image, audio, video, PDF |
| Reasoning control | none, low, medium, high, xhigh, max | Thinking supported, medium default |
| Fine-tuning | Not supported | Not supported |
Benchmark Results Side by Side
| Benchmark | GPT-5.6 Luna | Gemini 3.6 Flash | Winner |
|---|---|---|---|
| SWE-Bench Pro | 62.7% | 58.7% | Luna |
| DeepSWE v1.1 | 67% | 49% | Luna |
| Terminal-Bench 2.1 | 84.7% | 78.0% | Luna |
| GDPval-AA v2 | 1584 Elo | 1421 Elo | Luna |
| MRCR v2 (8-needle) | Lower | Higher | Gemini |
| Vision average | 74% | 83% | Gemini |
| Visual reasoning | 60.5% | 77.7% | Gemini |
| Identification | 83.3% | 99.0% | Gemini |
| OCR | 90.7% | 88.2% | Luna |
| Data extraction | 80.4% | 95.9% | Gemini |
| Counting | 67.1% | 80.2% | Gemini |
| Object detection | 61.0% | 57.1% | Luna |
The split is clean. Luna owns text-based coding and agentic work; Gemini owns everything that starts with a picture. OCR and object detection are the two exceptions where Luna edges ahead on visual input, which suggests its weakness is visual reasoning rather than visual perception.
Why the Composite Rankings Disagree
This is the part most comparisons skip, and it matters. Artificial Analysis has published Gemini 3.6 Flash at 34 against Luna at 32 on its Intelligence Index in one high-effort comparison, and 50 against 33 in another. Separate coverage records Gemini 3.6 Flash at 50 on the same index, and elsewhere Luna appears at 51 against Gemini 3.5 Flash-Lite’s 36. LLM Stats places them within two points of each other at 43.5 and 45.4.
Output speed is equally unsettled. One measurement gives Gemini 213 tokens per second against Luna’s 152, with Luna responding far faster on time to first token at 2.00 seconds against 16.09. Google’s own reporting cites 304 tokens per second for 3.6 Flash. These are not errors so much as different effort settings, different test dates and different serving conditions producing different answers.
The conclusion for a buying decision is that composite indexes cannot separate these two models reliably. Individual task benchmarks can, and your own evaluation on your own prompts can. Our overview of which generative AI models respond fastest shows how much reasoning modes distort latency figures across the board, and our survey of how AI benchmark results are produced explains why single-source scores move so much.
Cost Over a Real Workload
| Workload | GPT-5.6 Luna | Gemini 3.6 Flash (promotional) | Gemini 3.6 Flash (from Jan 2027) |
|---|---|---|---|
| 1M input / 100K output | $0.32 | $1.13 | $2.25 |
| 50M input / 10M output | $22 | $75 | $150 |
| 400M input / 75M output | $170 | $581 | $1,162 |
| 4B input / 500M output | $1,400 | $4,875 | $9,750 |
Luna’s advantage is 3.4 times at current Gemini promotional rates and 6.8 times once those rates expire. On a four billion token monthly pipeline, that is a difference of roughly $8,350 a month. Two caveats work against Luna, though: prompts exceeding 272,000 input tokens are billed at double input and 1.5 times output for the entire request, and Luna’s 128K output ceiling doubles Gemini’s, which can mean fewer calls for long generation tasks. Readers modelling this for their own stack can start from our guide to token pricing in practice.
Availability and Access
Luna is not selectable inside standard ChatGPT conversations. It reaches developers through the OpenAI API, Codex, ChatGPT Work and GitHub Copilot, where it arrived on 9 July across Pro, Pro+, Max, Business and Enterprise plans. Free and Go users get it as their default model with unlimited text chats and a Think button that routes harder questions through higher reasoning. There is no free API tier, so evaluation requires a funded account.
Gemini 3.6 Flash is available through the Gemini API with a free tier alongside paid usage, and through Google’s broader Gemini surfaces. Its position in Google’s own lineup has moved twice since launch: Gemini 3.7 Flash arrived on 13 August 2026 scoring 43.6% on FrontierCode 1.1 Main against 3.6 Flash’s 34.4%, and Gemini 3.8 Flash followed in September with 90.8% on Terminal-Bench 2.1 against 3.7 Flash’s 81.6%. Teams selecting 3.6 Flash today should confirm that a newer Flash model at the same price is not simply the better purchase. Our coverage of how the Gemini Flash line fits production workflows gives additional background on the family.
Practical Recommendations
| Task | Pick | Reason |
|---|---|---|
| Code generation and repair | GPT-5.6 Luna | Wins every shared coding benchmark |
| Terminal and agentic workflows | GPT-5.6 Luna | 84.7% Terminal-Bench 2.1 |
| Invoice, receipt or form processing | Gemini 3.6 Flash | 95.9% data extraction against 80.4% |
| Image and video understanding | Gemini 3.6 Flash | Native multimodal input, stronger visual reasoning |
| Pure OCR | GPT-5.6 Luna | 90.7% against 88.2%, at lower cost |
| Retrieval across very long documents | Gemini 3.6 Flash | Wins MRCR v2 at eight needles |
| Long-form generation | GPT-5.6 Luna | 128K output ceiling against 65K |
| Highest volume, lowest cost | GPT-5.6 Luna | 3.4x cheaper now, 6.8x from January 2027 |
| Latency-critical interactive features | Test both | Published speed figures conflict |
The Honest Summary
Luna comes out on top overall, on the strength of winning most shared benchmarks while costing substantially less. That verdict holds for the majority of production workloads, which are text in and text out. But it flips completely the moment a pipeline touches images, audio, video or PDFs, where Gemini 3.6 Flash is the better instrument and the price difference stops being the deciding factor.
The stronger play for teams with real volume is routing rather than selection: text and code to Luna, multimodal to Gemini, and neither to the harder reasoning problems that belong on a frontier model. Both companies have been cutting prices under pressure from lower-cost open-weight competitors, a dynamic we examine in our look at whether cheap Chinese AI models can rival the established labs, and the direction of travel favours anyone willing to spread a workload across more than one provider.
Prices, benchmark scores and model availability described here are current as of September 2026. Promotional rates expire and newer Flash models have since shipped; verify current figures with each vendor before committing.
If you are interested in this topic, we suggest you check our articles:
- Which GenAI Model Is the Fastest to Respond?
- AI Benchmarks: Performance Metrics Show Record Gains
- Tokens Explained: The Currency of Generative AI
- Gemini 3.5 Flash and the SEO Agency Playbook for 2026
- Can Cheap Chinese AI Models Rival OpenAI and Anthropic?
Sources: OpenAI — GPT-5.6 Luna model documentation, OpenRouter — Gemini 3.6 Flash pricing and benchmarks, Roboflow — Gemini 3.6 Flash vs GPT-5.6 Luna vision evals, MyClaw — Gemini 3.6 Flash vs GPT-5.6 Luna: coding and cost, LLM Stats — Gemini 3.6 Flash vs GPT-5.6 Luna, Fello AI — Gemini 3.6 Flash pricing and benchmarks, DataCamp — Gemini 3.8 Flash features and benchmarks
Written by Alius Noreika

