AI

Gemini 3.8 Flash Released: Benchmarks, Pricing and API Costs

Three Flash releases went GA in about six weeks. Gemini 3.6 Flash on 21 July, 3.7 Flash on 13 August, and Gemini 3.8 Flash on 2 September, the last of which Google calls its “most intelligent Flash model”. That cadence is fast enough that most teams are still calling a model two versions behind whatever the docs now recommend.

Original content from computingforgeeks.com - post 171394

This covers what the new Flash model actually changes, all 14 benchmarks Google published against Claude Opus 5, Claude Sonnet 5 and the GPT-5.6 line, the API pricing and the increase scheduled for January, and the thinking-level setting that turned out to matter more for cost than anything on the spec sheet. Google measures it against Opus 5 rather than the newest OpenAI flagship, so the Opus 5 breakdown is the useful companion piece here. If you are tracking the wider release cadence, the GPT-6 Astra notes cover a model that appears nowhere in Google’s comparison set.

The published spec sheet is identical to the previous Flash model, line for line. The differences are in benchmark quality, in what the model does with your token budget, and in one capability that was quietly removed. Every call below was run on gemini-3.8-flash in September 2026, so the token counts and errors are measured rather than copied from the model card.

What changed between 3.7 Flash and 3.8 Flash

Nothing on the spec sheet. Both models publish the same context window, the same output ceiling, the same input modalities, and the same capability list. Google’s own model reference pages put them side by side with no divergence except the release month. The thinking-level column below comes from the thinking guide’s support table rather than the model pages, which only say “Thinking: Supported” for the two older models.

ItemGemini 3.8 FlashGemini 3.7 FlashGemini 3.6 FlashGemini 3.5 Flash
Model IDgemini-3.8-flashgemini-3.7-flashgemini-3.6-flashgemini-3.5-flash
GA date2 September 202613 August 202621 July 2026Earlier line
Input token limit1,048,5761,048,5761,048,5761,048,576
Output token limit65,53665,53665,53665,536
Input modalitiesText, Image, Video, Audio, PDFText, Image, Video, Audio, PDFText, Image, Video, Audio, PDFText, Image, Video, Audio, PDF
Thinking levelslow, medium, highlow, medium, highminimal, low, medium, highminimal, low, medium, high
Input price per million$0.75$0.75$0.75$1.50
Output price per million$3.75$3.75$3.75$9.00

One row in that table is not a tie. The minimal thinking level works on 3.6 Flash and 3.5 Flash and is rejected outright by 3.7 and 3.8, which makes the newer models the first in the line that cannot be told to stop reasoning. That is a removal, not an addition, and it has a direct cost consequence covered further down.

Everything else matches. Neither current Flash model supports audio generation, image generation, or the Live API. Both support caching, batch, function calling, structured outputs, search grounding, code execution, URL context, file search, grounding with Google Maps, and computer use in preview. Both offer flex and priority inference tiers.

The practical read: this is a quality bump inside a fixed envelope. If you were waiting for a bigger context window or native image output on the Flash line, it did not arrive. What changed is how well the model uses the window it already had, which shows up in the agentic benchmarks rather than the capability matrix.

Benchmark scores against Opus 5 and the GPT-5.6 line

The model page shows four charts. Google’s evaluation methodology document, a four page PDF published alongside it, carries the full scored table: 14 benchmarks against the same six models, with the API prices printed in the same grid. That table is the honest basis for a verdict, and it is less flattering and more useful than the four charts.

Benchmark3.8 Flash3.7 FlashOpus 5Sonnet 5GPT-5.6 SolGPT-5.6 Terra
Input price per million$0.75$0.75$5.00$2.00$4.00$2.00
Output price per million$3.75$3.75$25.00$10.00$20.00$12.00
DeepSWE v1.1 (long-horizon SWE)73.7%65.3%74.0%53.8%72.7%69.6%
GDPVal-AA v2 (knowledge work, Elo)154514821824158417101528
Vals Finance Agent v261.4%59.0%58.6%53.9%53.8%54.4%
Harvey Legal Agent Benchmark10.0%8.8%6.7%5.0%2.5%0.8%
Terminal-bench 2.1 (agentic terminal coding)89.4%85.8%89.1%80.4%88.8%87.4%
Terminal-bench 4.0 (general agent capabilities)19.1%11.2%51.8%12.4%37.3%23.6%
GDP.PDF (expert PDF comprehension)35.0%34.0%37.0%28.0%40.0%29.0%
CharXiv Reasoning (no tools)86.2%84.5%83.7%70.1%85.8%85.9%
LVBench (long video, agentic and static)87.8% and 87.1%85.4%75.4%68.5%82.1%78.9%
HLE-Verified (expert reasoning)54.9%53.6%54.4%31.0%54.5%51.1%
OSWorld-2.0 (agentic computer use)59.0%50.6%75.4%42.6%62.6%50.2%
BioMysteryBench (human-solvable)88.8%87.1%90.1%87.5%79.5%83.8%
BioMysteryBench (human-difficult)56.5%43.5%49.4%34.1%44.7%49.4%
LABBench2 (biology research tasks)86.2%82.1%84.2%80.1%82.1%81.2%

LVBench is the one row with two figures for the new model, an agentic and a static score, and it leads on both. Counting leaders across those 14 rows: the new Flash model takes 8, Claude Opus 5 takes 5, and GPT-5.6 Sol takes 1. Opus 5 costs $25.00 per million output tokens against $3.75, so Google is claiming eight wins at roughly a seventh of the price. That is the actual pitch, and the table supports it.

Where it loses, it loses badly, and the pattern is consistent. Terminal-bench 4.0 measures general agent capability and the score is 19.1% against 51.8% for Opus 5, a gap of nearly 33 points. OSWorld-2.0 for agentic computer use is 59.0% against 75.4%. Knowledge work on GDPVal-AA v2 is 1545 Elo against 1824. So the shape is a model that is strong on bounded, domain-specific work and weak the moment the task becomes open-ended agent driving. Pick it for financial analysis, terminal coding, chart reasoning, long video and biology research. Do not pick it to run a computer unattended.

Version over version the picture is better than the promoted charts suggest. The new model improved on all 14 benchmark rows against 3.7 Flash. Three of the four charts on the model page are head-to-head bar charts, and they moved 2.4 points on Vals Finance, 1.3 on HLE-Verified and 1.2 on Harvey. Rank those against the 13 rows measured in percentage points and two of the three sit at the very bottom, second and third smallest, beaten only by the uncharted GDP.PDF at 1.0. The third, Vals Finance, lands exactly on the median. Meanwhile three of the four biggest gains land on benchmarks with no chart on the model page: 13.0 points on the human-difficult BioMysteryBench split, 8.4 on OSWorld-2.0, and 7.9 on Terminal-bench 4.0. The remaining one of the four is DeepSWE, tied at 8.4, which does get a chart, though as a cost-efficiency scatter rather than a head-to-head bar. Six weeks produced a real step, mostly on benchmarks with no bar chart attached.

Three caveats belong on this table before anyone quotes it. Gemini scores are pass@1 except where the methodology notes otherwise, and they are not all Google’s own measurements: GDPVal-AA v2 comes from the Artificial Analysis leaderboard, Terminal-bench 4.0 from the official public leaderboard, and DeepSWE from the Datacurve leaderboard with only the new Flash model’s own score self-computed. Non-Gemini figures are “sourced from providers’ self reported numbers” except where stated, and for GPT-5.6 Terra and Sonnet 5 Google defaults to “maximum thinking/reasoning settings available”. Vals Finance Agent v2 and the Harvey benchmark also come from Vals.AI rather than Google. And DeepSWE carries a correction: Google notes it “originally incorrectly reported Opus 5’s score as 74% due to rounding on the Datacurve public leaderboard” while the table still prints 74.0%, so treat the 73.7% against 74.0% ordering on that row as unresolved rather than a loss by three tenths of a point.

The Sonnet 5 column needs one more warning. Its 31.0% on HLE-Verified looks like a capability collapse next to everything else in that row, but Google’s methodology states that “a significant proportion of questions were blocked by content policy filters for Sonnet 5 and a small number for Opus 5”. That number is partly a refusal artifact. It is not evidence about reasoning ability, and the vendor publishing it says so.

Gemini 3.8 Flash pricing and the January increase

Gemini 3.8 Flash bills at $0.75 per million input tokens and $3.75 per million output tokens, and both figures are introductory. Google’s pricing page states the rate plainly: “$0.75 through December 31, 2026. $1.50 starting January 1, 2027.” Output follows the same pattern, moving from $3.75 to $7.50.

That is a straight doubling on both sides of the meter, with a date attached. Context caching doubles alongside it on both counts, from $0.075 to $0.15 per million tokens and from $0.50 to $1.00 per million tokens per hour of storage. Anyone sizing a budget off today’s rate is sizing it off a number with a four month shelf life. The 3.6 and 3.7 Flash models carry the same promotional pricing and the same expiry, so downgrading a version does not buy an escape from the increase.

Older models are priced differently. Gemini 3.5 Flash sits at $1.50 and $9.00 with no promotional rate, which makes the current 3.8 Flash rate cheaper than the model three versions back, for now. The Flash-Lite tier undercuts all of them, and the floor is not where most people assume: 3.5 Flash-Lite is $0.30 input, but 3.1 Flash-Lite is $0.25 and 2.5 Flash-Lite is $0.10, both still on the current pricing page. Those two carry a modality premium the headline figure hides, charging $0.50 and $0.30 respectively for audio input.

One line on the pricing page settles a question people keep asking about reasoning models. The output row is labelled “Output price (including thinking tokens)”. Thoughts bill as output, at the output rate, which is what makes the next section a cost discussion rather than a latency one.

Call the model from the API

Access is a plain REST call against the Generative Language endpoint with the key in a header. Export the key once so it stays out of your shell history and out of the command itself:

export GEMINI_API_KEY="your-api-key-here"

A minimal generation request needs the model ID in the path and nothing else configured:

curl -s "https://generativelanguage.googleapis.com/v1beta/models/gemini-3.8-flash:generateContent" \
  -H "x-goog-api-key: ${GEMINI_API_KEY}" \
  -H 'Content-Type: application/json' \
  -d '{"contents":[{"parts":[{"text":"Reply with exactly: ok"}]}]}'

The response carries the answer plus a usage block, and the usage block is the part worth reading:

"usageMetadata": {
  "promptTokenCount": 5,
  "candidatesTokenCount": 1,
  "thoughtsTokenCount": 72,
  "totalTokenCount": 78
}

A five token prompt asking for a single word came back with one token of answer and 72 tokens of thinking. Thinking is on by default at the medium level, so that request billed 73 tokens of output to produce a two letter reply, plus five tokens of input at the cheaper rate. Repeating the identical prompt returned 65 and 79 thought tokens, so treat any single measurement as approximate.

Google’s current quickstart leads with a newer Interactions API rather than generateContent. Both are supported, and the September changelog names them together, but the field names differ: the interactions surface reports usage.total_thought_tokens where the call above reports usageMetadata.thoughtsTokenCount. If you are reading a token count that does not match this article, check which surface you are on.

To confirm which Flash models a key can reach, list them rather than guessing at IDs:

curl -s "https://generativelanguage.googleapis.com/v1beta/models?pageSize=200" \
  -H "x-goog-api-key: ${GEMINI_API_KEY}" | grep -o '"name": "models/gemini-3\.[0-9]*-flash"'

Our key returned 54 models in total, four of which matched that pattern:

"name": "models/gemini-3.5-flash"
"name": "models/gemini-3.6-flash"
"name": "models/gemini-3.7-flash"
"name": "models/gemini-3.8-flash"

All four are callable, which is why the thinking-level difference between them matters in practice rather than in theory. If an ID you know is spelled correctly does not appear in that list, the 404 from a generation call is a visibility problem on the key rather than a bad path.

What each thinking level actually costs

This is the setting that decides your bill. We sent one identical 28 token coding prompt at each supported thinking level and recorded the usage block returned by the API. Figures are one representative run per level, and the cost columns are output charges only, excluding the 28 input tokens. The token counts were measured on a free-tier key, and the dollar figures apply the paid rate for each column, introductory on the left and standard from January on the right.

Thinking levelThought tokensAnswer tokensBilled outputOutput cost nowFrom 1 JanLatency
minimalRejected with HTTP 400 on 3.8 and 3.7 Flash
low0192192$0.00072$0.001443.2s
medium (default)792171963$0.00361$0.007225.1s
high1,8101531,963$0.00736$0.014728.4s

The high setting billed 1,963 output tokens against 192 for low, a factor of ten on the same prompt, and it returned a shorter answer. Thought tokens accounted for 92% of the billed output at that level. The default medium setting costs five times low.

Thought counts move between runs, so read those ratios as orders of magnitude rather than constants. A repeat of the low run returned 167 answer tokens against the 192 recorded above, and a second medium run came in at 779 thought tokens and 961 billed output against the 792 and 963 in the table. The shape of the result held every time: low spends nothing on thinking, medium spends several hundred tokens, high spends thousands.

Extrapolate to 10,000 requests of the same shape and the gap stops being academic. At low the output charge is $7.20, at medium $36.11, at high $73.61. After the January rate change those become $14.40, $72.23, and $147.23. That is arithmetic on a single measured request rather than a measured fleet, but the ranking does not change: the thinking level is a bigger lever on spend than the choice between Flash versions.

None of that makes low the automatic answer. It produced a working answer on a self-contained coding prompt, which is exactly the shape of task where reasoning tokens buy little. On the long-horizon agentic work the DeepSWE chart advertises, the thinking budget is the product. Set it per call site rather than globally, and pass the level explicitly instead of inheriting the default:

curl -s "https://generativelanguage.googleapis.com/v1beta/models/gemini-3.8-flash:generateContent" \
  -H "x-goog-api-key: ${GEMINI_API_KEY}" \
  -H 'Content-Type: application/json' \
  -d '{"contents":[{"parts":[{"text":"Summarise this log line"}]}],
       "generationConfig":{"thinkingConfig":{"thinkingLevel":"low"}}}'

Token accounting across a fleet of agents is its own discipline once you are paying for thoughts you never read. The context engineering notes go into where those tokens accumulate.

Error: “Thinking level MINIMAL is not supported for this model”

Passing minimal returns HTTP 400 with status INVALID_ARGUMENT and this message:

Thinking level MINIMAL is not supported for this model. Please retry with other thinking level.

This is a regression across the Flash line, and it is the one change in this release that can break working code. We sent the same minimal request to all four callable Flash models:

ModelResult with thinkingLevel: minimal
gemini-3.8-flashHTTP 400, rejected
gemini-3.7-flashHTTP 400, rejected, identical message
gemini-3.6-flashHTTP 200, 6 total tokens, no thinking
gemini-3.5-flashHTTP 200, 6 total tokens, no thinking

The level was dropped starting with 3.7 Flash. Anyone upgrading from 3.6 or 3.5 with minimal set as a cost floor gets a hard 400 rather than a graceful downgrade, and the cheapest setting still available on the new models is low. Map minimal to low in your client before the request goes out.

The four results side by side, summarised from the response bodies the loop returns:

Gemini 3.8 and 3.7 Flash rejecting thinkingLevel minimal while 3.6 and 3.5 Flash accept it

Note the two accepting models return a 6 token total with no thoughtsTokenCount field at all, which is what a genuinely non-thinking response looks like on this API. The level itself remains valid in the API surface, and 3.5 Flash-Lite still accepts it and even defaults to it, so this is a per-model restriction rather than a removed feature. Do not assume that generalises across Flash-Lite: 2.5 Flash-Lite supports only low, medium and high.

Error 503: “This model is currently experiencing high demand”

Five days after GA, 5 of roughly 14 calls came back with a 503 before succeeding on retry:

{
  "error": {
    "code": 503,
    "message": "This model is currently experiencing high demand. Spikes in demand are usually temporary. Please try again later.",
    "status": "UNAVAILABLE"
  }
}

Every one of ours cleared on a later attempt within a few seconds, and the older Flash model answered normally during the same window. A status of UNAVAILABLE is the tell that this is capacity on a new model rather than a malformed request or the quota problem covered next. The official SDKs already retry 429 and 5xx with exponential backoff, and the Python SDK retries transient errors up to four times, so this mainly bites raw HTTP and curl clients where you add the backoff yourself. Wiring the previous Flash version as a fallback is cheap insurance because the pricing is identical, though it rejects minimal too. It can help with the quota errors below as well, since limits are applied per project but “vary depending on the specific model being used”, so a per-model ceiling on the newest model does not necessarily bind on the older one.

Error 429: “You exceeded your current quota”

Enough testing on a free key and the failure mode changes shape entirely. The error block below is trimmed to the three fields that matter here. Google’s API error reference does not document the full payload for this case, but a standard google.rpc details array carrying QuotaFailure and RetryInfo, which is where the specific metric and a retry delay would appear if present:

{
  "error": {
    "code": 429,
    "message": "You exceeded your current quota, please check your plan and billing details.",
    "status": "RESOURCE_EXHAUSTED"
  }
}

The standard Flash model is free-tier available, input, output and context caching all at no charge, so a free key is a realistic way to evaluate it. The trade is that free-tier traffic is used to improve Google’s products, which the paid tier states it does not do.

Backoff does not help here the way it helps a 503, and RESOURCE_EXHAUSTED on its own does not tell you which limit you hit. Google splits 429 into three causes with different remedies, published as string code values on the newer Interactions surface rather than in the numeric-code block above: rate_limit_exceeded for per-minute or per-second pressure, where waiting and retrying with exponential backoff is correct; quota_exceeded for a daily cap, where retrying is pointless until the quota resets or you request an increase; and too_many_requests, which is also a backoff case. Limits run on requests per minute, input tokens per minute and requests per day, they apply per project rather than per key, and the daily counter resets at midnight Pacific. Google documents the three cause codes but does not map the human-readable message text to them, so the message above pointing at plan and billing details only suggests the daily-quota variant. If it is that one, the documented remedies are waiting for the midnight reset or requesting an increase, and enabling billing is the way to stop meeting the ceiling at all.

Google no longer publishes the numeric free-tier limits, stating they depend on usage tier and should be read from AI Studio. Moving off the free tier means setting up billing in AI Studio, and the jump to the first paid tier takes effect more or less immediately.

Moving from 3.7 Flash to 3.8 Flash

Swapping the model ID is the first of five items on Google’s own migration checklist, not the whole job. That checklist sits under the “what’s new” page rather than on one of its own, which is why it is easy to miss. The spec sheet being identical is what misleads here: the limits and capabilities carry over, but the request shape has moved and several of the changes fail loudly.

Strip temperature, top_p and top_k from your generation configs, all three deprecated in July. Replace thinking_budget with the thinking_level string enum, remembering that minimal is not among the accepted values. Remove candidate_count, unsupported since Gemini 3. Standardize multi-turn conversations on the server-side previous_interaction_id and remove prefilled model turns. Preserve thought signatures, which is a baseline Gemini 3 requirement rather than anything new here.

One checklist item is scoped specifically to the API this article uses. Google’s wording is “Only if using generateContent API: Ensure all FunctionResponse objects include call_id and name.” If you are doing tool calling through generateContent and your responses omit either field, that is a migration break waiting for you. The same section tells you to place multimodal assets inside the response payload and to format inline instructions with a double newline.

Beyond the checklist, re-measure your thought token consumption after the swap rather than assuming it carried over, since that is the axis where the models demonstrably differ and it is the axis you are billed on. Confirm your retry policy handles both 503 and 429, and that it treats them differently.

Coming from further back is a bigger job than the version numbers suggest. Off 3.6 Flash you lose minimal, which is a code change wherever it is set. Off 3.5 Flash you also gain a price cut, from $1.50 and $9.00 down to the current promotional $0.75 and $3.75, though that advantage expires in January along with everything else on the promotional rate.

For anyone driving Gemini from a terminal rather than an SDK, the Gemini CLI command reference and the CLI setup guide both track the current model line. If your integration runs through Vertex AI instead of an API key, note that it is a separate surface with its own auth and its own model availability, so the calls in this article do not port across unchanged. Our Vertex AI streaming and tool-use walkthrough covers that path, and the model IDs there need updating to the current line before the examples will run.

Keep reading

Claude Code Cheat Sheet – Commands, Shortcuts, Tips AI Claude Code Cheat Sheet – Commands, Shortcuts, Tips Ollama Models Cheat Sheet 2026 (gpt-oss, Qwen3-Coder, DeepSeek) AI Ollama Models Cheat Sheet 2026 (gpt-oss, Qwen3-Coder, DeepSeek) GPT-6 Astra: Benchmarks, Pricing and API Access, Tested AI GPT-6 Astra: Benchmarks, Pricing and API Access, Tested Get Started with LanceDB in Python AI Get Started with LanceDB in Python Install and Self-Host Karakeep with Docker AI Install and Self-Host Karakeep with Docker Install pgvector on PostgreSQL 17 (Rocky Linux 10 / Ubuntu 24.04) AI Install pgvector on PostgreSQL 17 (Rocky Linux 10 / Ubuntu 24.04)

Leave a Comment

Press ESC to close