GPT-6 Luna vs Gemini 3.8 Flash vs Opus 5.5: Which Tier?

GPT-6 Luna vs Gemini 3.8 Flash vs Claude Opus 5.5: Which Tier Fits Your Workload?

2026-09-24

Key Takeaways

  • All three models are September 2026 releases: Gemini 3.8 Flash arrived on 2 September, while Claude Opus 5.5 and GPT-6 Luna launched on 22 September, about 90 minutes apart.
  • GPT-6 Luna costs $0.10 per million input tokens and $0.50 per million output tokens, Gemini 3.8 Flash costs $0.75 and $3.75 until 31 December 2026, and Claude Opus 5.5 costs $4 and $20.
  • Opus 5.5 now costs 40 times more per token than Luna, double the 20-fold input-price gap it had against GPT-5.6 Luna, because OpenAI halved Luna’s price while Anthropic cut Opus by 20%.
  • One independent test now covers all three: Artificial Analysis scores Opus 5.5 at 58, the highest it has recorded, Gemini 3.8 Flash at 41 and GPT-6 Luna at 37 on its Intelligence Index v4.3.2.
  • The same test puts the average cost of one index task at $0.07 for Luna, $1.24 for Flash and $5.98 for Opus 5.5, a wider spread than the price sheets suggest.
  • Gemini 3.8 Flash is the fastest of the three at about 291 output tokens per second and the only one that reads audio, video and PDF natively, but it is also verbose, and its price doubles on 1 January 2027.
  • Opus 5.5 leads on agentic coding and business workflows, including 66.4% on Terminal-Bench 4.0 by Anthropic’s count, but its safeguards can hand cybersecurity and biology requests to older Claude models mid-workflow.
  • For teams with real volume, routing work across two or three of these tiers still beats picking one model.

Which Model Fits Which Workload

GPT-6 Luna vs Gemini 3.8 Flash vs Claude Opus 5.5 - artistic impression.

GPT-6 Luna vs Gemini 3.8 Flash vs Claude Opus 5.5 – artistic impression. Image source: Alius Noreika / Google Gemini

Pick GPT-6 Luna when you run a narrow, well-defined task thousands of times a day and cost per call decides everything: classification, extraction, routing, moderation and short background agent steps. Pick Gemini 3.8 Flash when the input is an image, a PDF, a recording or a video, or when a person is waiting on the answer and raw output speed matters. Pick Claude Opus 5.5 when a job has to finish correctly with nobody checking it, such as a long refactor, a multi-hour agent run, or analysis that feeds a business decision, because a failed run there costs more than the tokens you save.

On a workload of 400 million input tokens and 75 million output tokens a month, list prices put the bill at about $78 on Luna, $581 on Gemini 3.8 Flash and $3,100 on Opus 5.5. That is a 40-fold spread on identical volume. Opus 5.5 is clearly the strongest model of the three, so the real question is how often the cheaper model’s mistakes cost you more than that difference.

What Changed Since the GPT-5.6 Luna and Gemini 3.6 Flash Comparison

Every model in this lineup was replaced within a single month. Our earlier comparison of GPT-5.6 Luna, Gemini 3.6 Flash and Claude Opus 5 covered the July generation. Since then, all three vendors have shipped successors, and two of them cut prices on the same afternoon.

OpenAI released GPT-6 Luna on 22 September, 19 days after its GPT-6 Astra flagship, alongside a mid-tier GPT-6 Sol. Luna’s price fell from $0.20 and $1.20 to $0.10 and $0.50 per million tokens, a 50% cut on input and 58% on output. An OpenAI spokesperson told VentureBeat and The New Stack that the new rates are permanent, not promotional. The company’s own explanation was short: “Improvements in caching and inference let us serve these models at lower cost, and we’re passing those savings directly on to users and customers.” Readers who followed our cost analysis of ChatGPT 5.6 Luna will remember that its 80% July price cut already set a low floor. GPT-6 halves that floor again.

The overall tier position of GPT-6 Luna, Gemini 3.8 Flash, and Claude Opus 5.5. Image source: Claude AI

The overall tier position of GPT-6 Luna, Gemini 3.8 Flash, and Claude Opus 5.5. Image source: Claude AI

Google moved even faster. Gemini 3.8 Flash, which Google describes as its most intelligent Flash model, reached general availability on 2 September. It was the company’s third Flash release in roughly six weeks, following 3.6 Flash on 21 July and 3.7 Flash on 13 August, and it kept the same introductory price as both.

Anthropic released Claude Opus 5.5 on 22 September as the first model in its Claude 5.5 family, with Sonnet 5.5 and Haiku 5.5 expected in the following weeks. The company introduced it with one line: “It performs at the level of Claude Fable 5.1 for most tasks, and costs 40% less to run than Opus 5.” OpenAI announced GPT-6 Sol and Luna about 90 minutes later.

The most useful change for buyers is not a price. The July models never shared a benchmark harness across all three tiers. Artificial Analysis has now run all three new models through version 4.3.2 of its Intelligence Index, which gives one directly comparable score for the first time.

Price and Specification Comparison

Specification GPT-6 Luna Gemini 3.8 Flash Claude Opus 5.5
Developer OpenAI Google Anthropic
Release date 22 September 2026 2 September 2026 22 September 2026
API model ID gpt-6-luna gemini-3.8-flash claude-opus-5-5
Input per 1M tokens $0.10 $0.75 to 31 Dec 2026, then $1.50 $4.00
Output per 1M tokens $0.50 $3.75 to 31 Dec 2026, then $7.50 $20.00
Cached input per 1M $0.01 $0.075, then $0.15 $0.20
Cache writes $0.125 per 1M Storage at $0.50 per 1M tokens per hour, then $1.00 $5 per 1M (5-minute), $8 (1-hour)
Batch input / output $0.05 / $0.25 $0.375 / $1.875 $2 / $10
Faster tier input / output Fast mode $0.20 / $1.00 Priority $1.35 / $6.75 Fast mode $8 / $40, up to 2.5x speed
Context window 1,050,000 tokens 1,048,576 tokens 1,000,000 tokens
Maximum output 128,000 tokens 65,536 tokens 128,000 tokens (300,000 on Batch API beta)
Knowledge cutoff 18 May 2026 Not published June 2026
Input modalities Text, image Text, image, audio, video, PDF Text, image
Reasoning control none, low, medium (default), high, xhigh, max low, medium (default), high Adaptive thinking always on; medium effort by default

Three lines on these price sheets move a real invoice more than the headline rates do. Luna has a long-context rule: once a request passes 272,000 input tokens, the whole request is billed at double the input and cache rates and 1.5 times the output rate, so a 300,000-token prompt costs $0.06 in input rather than $0.03. Gemini 3.8 Flash bills thinking tokens as output, and every line of its price sheet, including context caching, doubles on 1 January 2027. Falling back to 3.6 or 3.7 Flash offers no escape, because both sit on the same promotional window.

Opus 5.5 changed most where agents spend money. Cache reads dropped 60%, from $0.50 to $0.20 per million tokens, and five-minute cache writes fell from $6.25 to $5. Anthropic says cache reads make up most of the cost of agentic and coding work, and Opus 5.5 applies no long-context surcharge anywhere in its 1M-token window.

Benchmarks: One Shared Yardstick and a Stack of Vendor Charts

The Artificial Analysis Intelligence Index

Version 4.3.2 of the Artificial Analysis Intelligence Index combines ten evaluations: AA-Briefcase, GDPval-AA v2.1, AutomationBench-AA, Terminal-Bench 4.0, SciCode, Humanity’s Last Exam, GDP.pdf, CritPt, AA-Omniscience and AA-LCR. Because one organisation runs every model through the same harness, it is the fairest single comparison of these three tiers available today.

Artificial Analysis metric GPT-6 Luna Gemini 3.8 Flash Claude Opus 5.5
Effort level tested max high max
Intelligence Index v4.3.2 37 41 58
Score at default medium effort 29 Not published 51.2
Cost per index task $0.07 $1.24 $5.98
Output tokens to run the index 150 million 170 million 260 million
Output speed 131 tokens per second 291 tokens per second Not yet measured
Time to first token 102.8 seconds 15.1 seconds Not yet measured

Opus 5.5’s 58 is the highest score Artificial Analysis has measured “by several points”, ahead of GPT-6 Astra and Claude Fable 5.1, which both sit at 53. Luna ranks sixth of 175 models in its price class for intelligence, while Gemini 3.8 Flash ranks first of 210 in its class for speed. Two notes matter when reading older coverage. Artificial Analysis rebuilt the index in September, so earlier readings that put Gemini 3.8 Flash at 59 used a different scale. And Luna’s 37 sits level with GPT-5.6 Luna’s score on the new index, which confirms that GPT-6 Luna’s gain is price, not ceiling.

Coding and Agent Benchmarks Where the Vendors Overlap

Benchmark GPT-6 Luna Gemini 3.8 Flash Claude Opus 5.5 Who reported it
DeepSWE v1.1 (long software tasks) 66.6% at max 73.7% Not published OpenAI; Google (Opus 5 scored 74.0% on Google’s chart)
AutomationBench (business workflows) 20.7% at max 29.7% at medium 40.0% OpenAI; Zapier; Zapier run for Anthropic
FrontierCode 1.1 (mergeable code) 42.4% at max Not published 54.4% at max, 54.6% at medium OpenAI; Anthropic
OSWorld 2.0 (computer use) 52.7% (offline variant) 59.0% 81.8% Each vendor, different setups
Terminal-Bench 4.0 (general agent work) Not published 19.1% 66.4% at xhigh; 59.6% in Artificial Analysis’s run Public leaderboard; Anthropic; Artificial Analysis
GDPval-AA (knowledge work, Elo) Not published 1545 (v2) 1846 (v2.1) Google; Anthropic

Opus 5.5 leads on every row where it appears. The closer fight sits between the two budget models, and Gemini 3.8 Flash wins it on capability: 73.7% against 66.6% on DeepSWE, 29.7% against 20.7% on AutomationBench, and 59.0% against 52.7% on OSWorld. Luna’s case rests on cost. OpenAI’s chart puts its 66.6% DeepSWE result at $0.22 per task and its 42.4% FrontierCode result at $0.11 per task.

Flash also has a clear weak spot. Google’s own methodology table shows it at 19.1% on Terminal-Bench 4.0, against 51.8% for Claude Opus 5, and 59.0% against 75.4% on OSWorld-2.0. The pattern is a model that handles bounded, domain-specific work very well and struggles once a task turns into open-ended agent driving.

Why the Numbers Disagree

Several of these figures come with caveats that change how far you should trust them. Anthropic reports Opus 5.5 at 66.4% on Terminal-Bench 4.0 at xhigh effort, but Artificial Analysis measured 59.6% on its own harness, which puts Opus 5.5 level with GPT-6 Astra rather than clearly ahead. Google’s methodology table lists Gemini 3.8 Flash at 89.4% on Terminal-Bench 2.1, while Google Cloud’s developer guide lists 90.8%. One outlet cites about 71.0% for Flash on DeepSWE from the model card, against the 73.7% most coverage uses, and Google itself corrected a rounding error on Opus 5’s DeepSWE score.

Effort settings add another layer. Luna reaches its best score at max effort on every OpenAI chart, and its Artificial Analysis score drops from 37 to 29 at the default medium setting. Zapier ran Opus 5.5 through AutomationBench without fallback models, so every safeguard intervention counted as a failure. Flash’s $0.55 cost per AutomationBench task uses Google’s standard list price; at today’s promotional rate it is closer to $0.27. OpenAI’s launch charts also compare against Claude Opus 5, not Opus 5.5, because both launched the same day. Anthropic adds its own warning that benchmark margins at this level have become a less reliable guide, and that the gap between Opus 5.5 and Fable 5.1 is narrower in daily use than the scores suggest.

Multimodal Work

Gemini 3.8 Flash is the only model of the three that accepts audio, video and PDF files directly. It posts 86.2% on CharXiv Reasoning for charts, 87.8% on the agentic LVBench long-video test against 75.4% for Claude Opus 5, and 35.0% on the GDP.pdf expert document test. GPT-6 Luna reads text and images only, and OpenAI lists audio and video as unsupported. Opus 5.5 also takes text and images, and scores 89.0% on Anthropic’s Chartography chart-recognition test. If your pipeline processes invoices, call recordings or product videos, Flash wins before anyone looks at a coding score.

Real Monthly Cost on Identical Volume

Monthly workload GPT-6 Luna Gemini 3.8 Flash Claude Opus 5.5
1M input / 100K output (single job) $0.15 $1.13 $6.00
400M input / 75M output $77.50 $581.25 $3,100
Same, with 80% of input served from cache $48.70 $365.25 $1,884
Same, through the Batch API $38.75 $290.63 $1,550
400M / 75M at Gemini’s 2027 standard rate $77.50 $1,162.50 $3,100

These figures use published list prices and leave out cache-write charges, tool fees, audio or image inputs and Luna’s surcharge above 272,000 input tokens. They show scale, not a forecast invoice. For context, the same 400M/75M workload cost $170 a month on GPT-5.6 Luna. Our guide to tokens as the working currency of generative AI explains how these per-million rates turn into monthly bills.

Cost per Task Tells a Different Story

Per-token prices hide how many tokens a model spends on each job. Artificial Analysis found Gemini 3.8 Flash generated 170 million output tokens to finish its index, against a median of 88 million for comparable models. Opus 5.5 at max effort used 260 million, roughly 119,000 output tokens per task against about 73,000 for Opus 5 and 27,000 for GPT-6 Astra. The result is that Flash costs about 18 times more than Luna per completed index task, even though its token price is only 7.5 times higher, and Opus 5.5 costs about 85 times more.

That also puts Anthropic’s 40% savings claim in context. Anthropic says typical workloads at default settings cost 40% less than on Opus 5, thanks to the 20% list-price cut plus fewer tokens per task. Early testers back that up at lower effort: Box measured about a third as many tokens as Opus 5, GitHub saw more terminal tasks completed in less than half the steps, and Deloitte found Opus 5.5 at its lowest effort caught 72% of known code-review bugs, against 56% for Opus 5 at high effort. At max effort the picture changes. Artificial Analysis found Opus 5.5 roughly level with Opus 5 on cost per task there, despite using 1.6 times the output tokens. The savings live at the medium default, so raise effort only for the jobs that need it.

Gemini 3.8 Flash has the same lever. In a ComputingForGeeks test on a single coding prompt, the default medium thinking level billed about five times the output tokens of the low setting, and high billed about ten times as many while returning a shorter answer. Luna works the other way: it scores best at max effort, but Artificial Analysis clocked over 100 seconds to first token at that setting, which rules it out for anything a user waits on. Our roundup of which generative AI models respond fastest shows how reasoning modes distort latency across the market, and the same token accounting applies at the top of OpenAI’s range, as our GPT-6 Astra token cost guide works through.

Model-by-Model Notes

GPT-6 Luna: Same Ceiling, Half the Bill

GPT-6 Luna’s gains show up in cost per task rather than peak scores. On OpenAI’s own charts, Luna at max effort matches GPT-6 Sol at xhigh on DeepSWE, both at 66.6%, for $0.22 per task against $1.00, and it outscores Sol at low effort on four of five agentic benchmarks. Its AutomationBench score rose 5.4 percentage points over GPT-5.6 Luna. On OpenAI’s internal factual-error test, Luna at max erred 7.6% of the time, fewer errors than GPT-5.6 Sol at max (8.5%) for about 1.4% of the cost.

Caching improved across the GPT-6 line. Cached input carries a 90% discount, and changing reasoning effort or switching tools mid-conversation no longer breaks the cache. GitHub reported that these changes cut the share of prompt tokens needing fresh processing by more than half in Copilot. OpenAI’s stress tests also show better behaviour than GPT-5.6 Luna: coding deception fell to 2.8% from 9.5%, and attempts to work around explicit warnings fell to 42.4% from 76.5%. Luna still failed to disclose a deliberately broken search tool in 28.7% of test runs, so teams running it unattended should enforce permission checks in code rather than trust the model to stop. OpenAI is clear about where Luna sits: “GPT‑6 Astra continues to be our best model across the board.” For teams that find Luna too weak and Opus too expensive, GPT-6 Sol at $2 and $10 fills the middle rung.

Gemini 3.8 Flash: Fastest Output, Heaviest Token Use

Gemini 3.8 Flash improved on 3.7 Flash across all 14 benchmarks in Google’s methodology table and led eight of them against Claude Opus 5, Claude Sonnet 5 and the GPT-5.6 line. Its strongest results come from specialised work: 61.4% on Vals Finance Agent v2 against 58.6% for Opus 5, 10.0% on the Harvey Legal Agent Benchmark, and 86.2% on the LABBench2 biology research tasks. Google also shipped a security sibling, Gemini 3.8 Flash Cyber, which scored 47.2% on the CWE-Bench patching test against 47.8% for a leading frontier model. Google’s Chrome Security team found it produced 2.6 times more correct vulnerability patches than larger commercial models.

The Gemini API offers a free tier for 3.8 Flash, with the trade-off that free-tier traffic can be used to improve Google’s products. Consumers reach the model in the Gemini app, AI Mode and Google Sheets through Google AI Pro or Ultra subscriptions.

Claude Opus 5.5: Fable-Class Results at Opus Prices

On Anthropic’s launch table, Opus 5.5 beats both Claude Fable 5.1 and GPT-6 Astra on agentic coding and knowledge work: 66.4% on Terminal-Bench 4.0 against 55.8% and 57.9%, 54.4% on FrontierCode against 50.3% and 53.3%, and 1846 Elo on GDPval-AA v2.1 against 1735 and 1542. It also reaches 81.8% on OSWorld 2.0 and 67.7% on Humanity’s Last Exam. Astra keeps the lead on AutomationBench (41.4% against 40.0%) and Terminal-Bench-Science (64.6% against 58.7%). At the default medium effort, Opus 5.5 scores 54.6% on FrontierCode, above Astra’s best result, at about a fifth of the cost per task.

Customer tests point the same way. In an internal C-to-Rust port of the HAProxy load balancer, Opus 5.5 finished in 9.5 hours against 12 hours for Fable 5.1, at 51% lower cost. One tester audited and fixed a 200,000-line codebase in under three hours, a job that took Opus 5 more than 20 hours and 2.5 times the tokens. Output also generates more than 30% faster than on Opus 5. On safety, Opus 5.5 posted Anthropic’s best score so far on an automated behavioural audit of nearly 2,000 scenarios, and in a new containment test it tried to get around boundaries about 85% less often than Opus 5. METR and Frontier Design tested it before release.

Availability and Migration Traps

GPT-6 Luna is live in the OpenAI API, and in ChatGPT Work and Codex for Plus, Pro, Business, Enterprise and Edu users. Free and Go users get it in the desktop app. Neither Luna nor Sol is in the everyday Chat surface yet. Developers on Chat Completions should note that function calling there only works with reasoning effort set to none, so agents that call tools while reasoning need the Responses API.

Gemini 3.8 Flash dropped the minimal thinking level, which now returns an HTTP 400 error, so code that used minimal as a cost floor on 3.6 or 3.5 Flash must switch to low. Google’s migration checklist also removes temperature, top_p and top_k from generation configs. In the first days after launch, ComputingForGeeks logged frequent 503 high-demand errors that cleared on retry, which makes a fallback to 3.7 Flash, at the same price, cheap insurance.

Claude Opus 5.5 runs on the Claude API, Amazon Bedrock, Google Cloud, Microsoft Foundry and Claude Platform on AWS, with zero data retention available. Subscribers on Claude plans get 20% higher five-hour usage limits. Four API changes break code written for Opus 5: thinking can no longer be disabled, forced tool use returns an error, thinking blocks are tied to the model and conversation that produced them, and the older computer_20251124 computer use tool is rejected on the Claude API and Google Cloud. Text between tool calls now arrives inside thinking blocks, so apps that stream progress updates can go quiet until they change a display setting.

The bigger architectural trap is safeguard routing. When Opus 5.5’s classifiers flag a request, Anthropic reroutes it to an older model: most cybersecurity tasks go to Opus 4.8, while biology and frontier LLM development requests go to Opus 5. Routine bug finding and fixing still works on Opus 5.5 itself. In a multi-step agent workflow, that means individual calls may be answered by models with different abilities, which your evaluations should account for. Vetted organisations can apply to Anthropic’s Life Sciences Verification Program, and the Cyber Verification Program is expanding to include Opus 5.5.

Route Work Instead of Picking One Model

Most teams at meaningful volume end up running two or three of these models. A practical split sends classification, extraction and high-volume background agent steps to GPT-6 Luna at max effort; sends document, image, audio and video intake, plus latency-sensitive features, to Gemini 3.8 Flash; and reserves Claude Opus 5.5 at its medium default for code that ships, unattended agent runs and analysis that informs a decision, raising effort only for the hardest jobs.

If Luna’s error rate on your prompts starts to pile up retries, GPT-6 Sol and Claude Sonnet 5, both at $2 and $10 per million tokens, sit between the budget tier and Opus. Budget Gemini 3.8 Flash at $1.50 and $7.50 for anything planned beyond December, and rerun your own evaluation set on each model before moving production traffic, since every vendor chart in this article measures a slightly different thing.

The Bottom Line by Use Case

Workload Best fit Why
High-volume classification, extraction and routing GPT-6 Luna $0.10 / $0.50 per million tokens; $0.07 per Artificial Analysis index task
Background agents on a tight budget GPT-6 Luna at max effort 66.6% on DeepSWE at $0.22 per task on OpenAI’s chart
Documents, images, audio and video Gemini 3.8 Flash Only model of the three with native audio, video and PDF input
Latency-sensitive interactive features Gemini 3.8 Flash About 291 tokens per second; Luna at max took over 100 seconds to first token
Finance and legal agent tasks at low cost Gemini 3.8 Flash 61.4% on Vals Finance Agent v2 and 10.0% on Harvey, both above Opus 5
Production code and long refactors Claude Opus 5.5 66.4% on Terminal-Bench 4.0 and 54.4% on FrontierCode
Long unattended agent runs Claude Opus 5.5 Top Artificial Analysis score; a failed run costs more than the tokens saved
Business workflow automation Claude Opus 5.5, or Flash on a budget 40.0% against 29.7% on AutomationBench
Security-sensitive code work Test before committing Opus 5.5 reroutes most cyber tasks to Opus 4.8; Google offers a separate 3.8 Flash Cyber model

Prices, benchmark scores and availability described here are current as of 24 September 2026. Gemini’s promotional rates expire on 31 December 2026, and leaderboards change weekly; verify figures with each vendor before committing production budget.

If you are interested in this topic, we suggest you check our articles:

Sources: OpenAI — Introducing GPT-6 Sol and Luna, OpenAI — GPT-6 Luna model page, Digital Applied — GPT-6 Sol and Luna prices and benchmarks, Google — Introducing Gemini 3.8 Flash and 3.8 Flash Cyber, Google Cloud — Developer’s guide to Gemini 3.8 Flash, ComputingForGeeks — Gemini 3.8 Flash benchmarks and API costs, BenchmarkList — Gemini 3.8 Flash scores, Anthropic — Claude Opus 5.5, Claude Platform Docs — Claude Opus 5.5, Artificial Analysis — Claude Opus 5.5 takes the top spot, Artificial Analysis — GPT-6 Luna, Artificial Analysis — Gemini 3.8 Flash, Kingy AI — Claude Opus 5.5 specs and benchmarks, OrcaRouter — GPT-6 Luna vs Claude Opus 5, CodingFleet — Claude Opus 5.5 vs GPT-6 Sol

Written by Alius Noreika

GPT-6 Luna vs Gemini 3.8 Flash vs Claude Opus 5.5: Which Tier Fits Your Workload?
We use cookies and other technologies to ensure that we give you the best experience on our website. If you continue to use this site we will assume that you are happy with it..
Privacy policy