Google announced Gemini 3.8 Flash on September 2 as an update to its 3.7 Flash model. The release also introduced a separate cybersecurity variant. This issue focuses on the general model because its documented capabilities and limitations affect a wider range of software and research workflows.
Built for Product Teams moving at AI Speed.
Your teams are moving fast, burning tokens, and shipping more than ever.
But more output doesn’t mean more impact.
Jira Product Discovery brings your ideas, customer insights, and business context together so product teams can weigh the evidence, make the tradeoffs, and decide what’s actually worth building. Then connect those decisions directly to delivery in Jira, so everyone knows what you’re building and why.
Jira Product Discovery. For better product decisions in the AI era.
The model card says Gemini 3.8 Flash accepts text, images, audio, and video. It provides a context window of up to one million tokens and a maximum output of 64,000 tokens. Its stated knowledge cutoff is March 2026. These limits describe the model interface, not the accuracy of every answer.
Gemini 3.8 Flash supports customizable effort levels. Google says the setting changes the balance among quality, cost, and latency. The launch article adds that higher effort can involve extra reasoning steps and repeated tool calls. A longer or more expensive computation can change the economics of a task even when the listed token rate stays unchanged.
Google’s model card lists introductory prices of seventy five cents per million input tokens and three dollars and seventy five cents per million output tokens. It says regular prices become one dollar and fifty cents and seven dollars and fifty cents from January 1, 2027. A token price is only one part of a workflow cost.
The published evaluation table shows an uneven pattern. Gemini 3.8 Flash scores 73.7 percent on DeepSWE v1.1, compared with 65.3 percent for Gemini 3.7 Flash. It records 89.4 percent on Terminal Bench 2.1, but 19.1 percent on Terminal Bench 4.0. Different tests measure different tasks and should not be collapsed into one ranking.
200 Ways To Make Money With AI
Ready to transform artificial intelligence from a buzzword into your personal revenue generator?
HubSpot’s groundbreaking guide "200+ AI-Powered Income Ideas" is your gateway to financial innovation in the digital age.
Inside you'll discover:
A curated collection of 200+ profitable opportunities spanning content creation, e-commerce, gaming, and emerging digital markets, each vetted for real-world potential
Step-by-step implementation guides designed for beginners, making AI accessible regardless of your technical background
Cutting-edge strategies aligned with current market trends, ensuring your ventures stay ahead of the curve
Download your guide today and unlock a future where artificial intelligence powers your success. Your next income stream is waiting.
The same table reports 35.0 percent on GDP.PDF document comprehension and 59.0 percent on OSWorld computer use. It reports 86.2 percent on CharXiv chart reasoning. These tests use different scoring rules and input conditions. A result can identify a narrow capability without predicting a complete workflow. It also reports 54.9 percent on HLE Verified and 86.2 percent on LABBench2.
Independent reporting adds a further qualification. Ars Technica described the release as Google’s third Flash model in six weeks. It reported that Google’s benchmark claims showed improvement, while computer use still trailed Claude Opus on the cited test. That does not settle the model’s usefulness, but it shows why vendor tables need outside comparison.
Google’s model card also documents ordinary foundation-model limitations. It mentions hallucinations, occasional slowness or timeouts, and higher token use at greater effort. It reports a slight regression in non-English safety evaluation relative to Gemini 3.7 Flash. The card says its automated results are not directly comparable with earlier cards because the evaluations changed.
The model is distributed through the Gemini app, Gemini Enterprise Agent Platform, Google AI Studio, Gemini API, Google AI Mode, and Google Antigravity. Availability across several surfaces does not make their tools, limits, or terms identical. A deployment record should specify the channel and model version.
The verified conclusion is limited. Gemini 3.8 Flash expands the tradeoff between effort, latency, token use, and task performance. Google’s evidence shows gains in some tests and weaker results in others. A serious evaluation therefore needs the actual documents, tools, languages, and failure checks used by the intended workflow. Independent testing remains necessary before choosing a model.
285 Free Prompts to Save Time
Your inbox isn't the problem. Neither are your meetings. It's the hours you spend on work AI could do in minutes.
285 copy-paste prompts: built for email, meetings, admin, and delegation. Each one has a clear role, a specific task, and a defined output.
No setup. No prompt engineering. Just copy, paste, and get your time back. Free when you subscribe to Mindstream.





