Gemini 3.8 Flash is Google's newest Flash model, released on 2 September 2026. The API name is gemini-3.8-flash and it shipped as New Stable, not preview.
The thing to know before anything else: the price you see today is promotional — $0.75 input / $3.75 output per 1M tokens, reverting to $1.50 / $7.50 on 1 January 2027. Build next year's budget on today's number and the bill doubles within four months.
This was written the same day the model shipped. Everything here comes from Google's own pricing page and changelog, and I will also be explicit about what Google has not disclosed.
What Google Said, and What It Left Out
The official docs describe Gemini 3.8 Flash as "our most intelligent Flash model, engineered for long-horizon software engineering, autonomous agents, and complex enterprise workflows", and immediately reclassify Gemini 3.7 Flash as the previous generation. But as of launch day, three things are missing from the documentation:
- •No benchmark numbers — not a single score against 3.7 Flash or any competitor
- •No context window — the token limit is not stated
- •No knowledge cutoff — the training data date is not given
Be sceptical of any article quoting 3.8 Flash benchmarks today. Google has not published them. If you see numbers, ask where they came from.
Every Tier, per 1M Tokens
Google prices four tiers by urgency. The left column is the promotional rate, valid through 31 December 2026:
| Tier | Promo, to 31 Dec 2026 | From 1 Jan 2027 | Best for |
|---|---|---|---|
| Batch / Flex | $0.375 / $1.875 | $0.75 / $3.75 | Work that can wait, processed in bulk |
| Standard | $0.75 / $3.75 | $1.50 / $7.50 | General use, live chatbots |
| Priority | $1.35 / $6.75 | $2.70 / $13.50 | Latency-critical work |

The Odd Part — Three Versions, One Price
Open the pricing page and something stands out: Gemini 3.6 Flash, 3.7 Flash and 3.8 Flash cost exactly the same across all four tiers.
So if you are still calling gemini-3.6-flash or gemini-3.7-flash, you are paying the same money for an older model. Moving to 3.8 costs nothing extra — all that is left is testing whether the output changed.
More counterintuitive still: the older Gemini 3.5 Flash is more expensive, at $1.50 / $9.00 — double the promotional input price of 3.8 Flash, and nearly triple on output.
If your project still points at Gemini 3.5 Flash, swap the model ID for gemini-3.8-flash and run your existing test set — it gets cheaper immediately with no other code changes.

The 1 January Cliff
This is the part that matters most if you are budgeting. Promotional pricing runs to 31 December 2026. After that the list price is double what you pay now — and it lands in the middle of many companies' fiscal year. If you are running a proof of concept this month and taking a number to your management, quote the post-promo price. Otherwise you will spend January explaining why the bill doubled on identical volume.

What It Actually Costs
Take a LINE OA chatbot answering 8,000 customer messages a month, assuming roughly 700 input tokens and 250 output tokens per message including prompt and product context, at about ฿36.5 to the dollar:
| Model | API cost / month | In baht |
|---|---|---|
| Gemini 3.8 Flash (promo) | $11.70 | ~฿427 |
| Gemini 3.8 Flash (from 1 Jan 2027) | $23.40 | ~฿854 |
| Gemini 3.5 Flash | $26.40 | ~฿964 |
| GPT-5.6 Luna | $3.52 | ~฿128 |
This is a worked hypothetical, not an invoice. The per-message token counts are estimates — count with your own tokenizer before budgeting. Note also that Thai text costs more tokens than English of the same meaning, because Thai has no spaces between words. The GPT-5.6 Luna price was verified 23 August 2026; these prices move fast, so check the official page before deciding.
So Which One Should You Use
Answering from today's numbers rather than the hype. Use 3.8 Flash if you are already inside Google's ecosystem — Workspace, Vertex, BigQuery — or your workload is long-running agent work, which is precisely what Google says this release was built for. Don't rush the migration if your workload is high-volume short chatbot replies and cost is the deciding factor, because GPT-5.6 Luna remains roughly 3x cheaper even against the promotional rate. Don't decide on the phrase "most intelligent" alone — Google has published no benchmarks. The more reliable move is to take 20–30 of your real prompts, run them side by side, and judge on the outcome your business actually needs.
Summary: Gemini 3.8 Flash shipped 2 September 2026 at a promotional $0.75/$3.75 per 1M tokens — cheaper than the older 3.5 Flash and identical to 3.6 and 3.7, so moving up from those is free. The catch is that the price doubles on 1 January 2027, so budget on the post-promo figure. Google has published no benchmarks or context window, so test against your own workload before migrating. Compare providers at AI cost comparison for Thai businesses and every Gemini plan and price.

Arm - CherCode
Full-Stack Developer & Founder
Software developer with 5+ years of experience in Web Development, AI Integration, and Automation. Specializing in Next.js, React, n8n, and LLM Integration. Founder of CherCode, building systems for Thai businesses.
Portfolio


