google

Gemini 3.8 Flash

Gemini 3.8 Flash is Google's multimodal AI offering 1M context, 64K output, and agentic coding capabilities at $0.75 per million input tokens.

MultimodalHigh-SpeedGoogle DeepMindLong ContextAgentic Coding
google logogoogleGemini 32026-09-02
Context
1Mtokens
Max Output
66Ktokens
Input Price
$0.75/ 1M
Output Price
$3.75/ 1M
Modality:TextImageAudioVideo
Capabilities:VisionToolsStreamingReasoning
Benchmarks
GPQA
59%
GPQA: Graduate-Level Science Q&A. A rigorous benchmark with 448 multiple-choice questions in biology, physics, and chemistry created by domain experts. PhD experts only achieve 65-74% accuracy, while non-experts score just 34% even with unlimited web access (hence 'Google-proof'). Gemini 3.8 Flash scored 59% on this benchmark.
HLE
45.4%
HLE: High-Level Expertise Reasoning. Tests a model's ability to demonstrate expert-level reasoning across specialized domains. Evaluates deep understanding of complex topics that require professional-level knowledge. Gemini 3.8 Flash scored 45.4% on this benchmark.
MMLU
88.2%
MMLU: Massive Multitask Language Understanding. A comprehensive benchmark with 16,000 multiple-choice questions across 57 academic subjects including math, philosophy, law, and medicine. Tests broad knowledge and reasoning capabilities. Gemini 3.8 Flash scored 88.2% on this benchmark.
MMLU Pro
76.5%
MMLU Pro: MMLU Professional Edition. An enhanced version of MMLU with 12,032 questions using a harder 10-option multiple choice format. Covers Math, Physics, Chemistry, Law, Engineering, Economics, Health, Psychology, Business, Biology, Philosophy, and Computer Science. Gemini 3.8 Flash scored 76.5% on this benchmark.
SimpleQA
42%
SimpleQA: Factual Accuracy Benchmark. Tests a model's ability to provide accurate, factual responses to straightforward questions. Measures reliability and reduces hallucinations in knowledge retrieval tasks. Gemini 3.8 Flash scored 42% on this benchmark.
IFEval
87.5%
IFEval: Instruction Following Evaluation. Measures how well a model follows specific instructions and constraints. Tests the ability to adhere to formatting rules, length limits, and other explicit requirements. Gemini 3.8 Flash scored 87.5% on this benchmark.
AIME 2025
82%
AIME 2025: American Invitational Math Exam. Competition-level mathematics problems from the prestigious AIME exam designed for talented high school students. Tests advanced mathematical problem-solving requiring abstract reasoning, not just pattern matching. Gemini 3.8 Flash scored 82% on this benchmark.
MATH
78%
MATH: Mathematical Problem Solving. A comprehensive math benchmark testing problem-solving across algebra, geometry, calculus, and other mathematical domains. Requires multi-step reasoning and formal mathematical knowledge. Gemini 3.8 Flash scored 78% on this benchmark.
GSM8k
94%
GSM8k: Grade School Math 8K. 8,500 grade school-level math word problems requiring multi-step reasoning. Tests basic arithmetic and logical thinking through real-world scenarios like shopping or time calculations. Gemini 3.8 Flash scored 94% on this benchmark.
MGSM
91%
MGSM: Multilingual Grade School Math. The GSM8k benchmark translated into 10 languages including Spanish, French, German, Russian, Chinese, and Japanese. Tests mathematical reasoning across different languages. Gemini 3.8 Flash scored 91% on this benchmark.
MathVista
68%
MathVista: Mathematical Visual Reasoning. Tests the ability to solve math problems that involve visual elements like charts, graphs, geometry diagrams, and scientific figures. Combines visual understanding with mathematical reasoning. Gemini 3.8 Flash scored 68% on this benchmark.
SWE-Bench
49%
SWE-Bench: Software Engineering Benchmark. AI models attempt to resolve real GitHub issues in open-source Python projects with human verification. Tests practical software engineering skills on production codebases. Top models went from 4.4% in 2023 to over 70% in 2024. Gemini 3.8 Flash scored 49% on this benchmark.
HumanEval
85%
HumanEval: Python Programming Problems. 164 hand-written programming problems where models must generate correct Python function implementations. Each solution is verified against unit tests. Top models now achieve 90%+ accuracy. Gemini 3.8 Flash scored 85% on this benchmark.
LiveCodeBench
65%
LiveCodeBench: Live Coding Benchmark. Tests coding abilities on continuously updated, real-world programming challenges. Unlike static benchmarks, uses fresh problems to prevent data contamination and measure true coding skills. Gemini 3.8 Flash scored 65% on this benchmark.
MMMU
69.5%
MMMU: Multimodal Understanding. Massive Multi-discipline Multimodal Understanding benchmark testing vision-language models on college-level problems across 30 subjects requiring both image understanding and expert knowledge. Gemini 3.8 Flash scored 69.5% on this benchmark.
MMMU Pro
52%
MMMU Pro: MMMU Professional Edition. Enhanced version of MMMU with more challenging questions and stricter evaluation. Tests advanced multimodal reasoning at professional and expert levels. Gemini 3.8 Flash scored 52% on this benchmark.
ChartQA
86%
ChartQA: Chart Question Answering. Tests the ability to understand and reason about information presented in charts and graphs. Requires extracting data, comparing values, and performing calculations from visual data representations. Gemini 3.8 Flash scored 86% on this benchmark.
DocVQA
93.5%
DocVQA: Document Visual Q&A. Document Visual Question Answering benchmark testing the ability to extract and reason about information from document images including forms, reports, and scanned text. Gemini 3.8 Flash scored 93.5% on this benchmark.
Terminal-Bench
89.4%
Terminal-Bench: Terminal/CLI Tasks. Tests the ability to perform command-line operations, write shell scripts, and navigate terminal environments. Measures practical system administration and development workflow skills. Gemini 3.8 Flash scored 89.4% on this benchmark.
ARC-AGI
4%
ARC-AGI: Abstraction & Reasoning. Abstraction and Reasoning Corpus for AGI - tests fluid intelligence through novel pattern recognition puzzles. Each task requires discovering the underlying rule from examples, measuring general reasoning ability rather than memorization. Gemini 3.8 Flash scored 4% on this benchmark.

About Gemini 3.8 Flash

Learn about Gemini 3.8 Flash's capabilities, features, and how it can help you achieve better results.

Model Overview

Gemini 3.8 Flash is Google DeepMind's high-efficiency multimodal model released in September 2026. Built on Google's Tensor Processing Unit hardware, the system natively ingests text, audio, high-resolution imagery, and video streams up to two hours long within a single 1,048,576-token context window. Output capacity reaches 65,536 tokens per request. The model introduces configurable reasoning budgets across low, medium, and high parameters, allowing engineers to balance execution latency against token consumption.

Agentic Execution and Architecture

The model targets long-horizon coding tasks and terminal control. Rather than finishing after one output pass, Gemini 3.8 Flash supports recursive execution loops and autonomous tool interactions. Training emphasized container manipulation, command-line interfaces, and cybersecurity defense workloads. This training produces lower error rates on test suites and self-correction during live builds in environments such as Google Antigravity.

Production Workloads and Grounding

Developers deploy Gemini 3.8 Flash for latency-sensitive applications requiring deterministic tool use. Native integrations connect the model directly to Google Search grounding and Google Maps data without separate retrieval pipelines. While frontier flagship models retain advantages on open-ended creative tasks, 3.8 Flash delivers comparable coding evaluation scores at a lower inference price.

Gemini 3.8 Flash

Use Cases

Discover the different ways you can use Gemini 3.8 Flash to achieve great results.

Autonomous Terminal Maintenance

Executes multi-file refactors, runs test suites inside containerized environments, and patches runtime errors autonomously.

High-Volume Video Ingestion

Analyzes raw two-hour video recordings and audio streams directly within its 1M token window without external preprocessing.

Financial Report Extraction

Ingests lengthy financial filings and complex visual PDFs to calculate balance sheet metrics into validated JSON.

Rapid UI Prototyping

Produces complete web applications, interactive SVG dashboards, and WebGL simulations in under fifteen seconds.

Defensive Security Validation

Scans source repositories to identify logic vulnerabilities, trace tainted variables, and generate candidate regression patches.

Low-Latency Agent Orchestration

Serves as the fast action planner in hierarchical multi-agent setups, calling external tools grounded by Google Search.

Strengths

Limitations

Agentic Coding Economics: Achieved 73.7% on DeepSWE v1.1 and 89.4% on Terminal-Bench 2.1 at an introductory cost of $0.75 per million input tokens.
High Token Consumption in Deep Thinking: High reasoning configurations execute extended internal loops that increase token output by up to 30% over 3.7 Flash.
Native Long-Context Multimodality: Processes up to 1,048,576 tokens of combined audio, video, and text without external frame extraction pipelines.
Lower Accuracy on Complex Terminal Benchmarks: Scores 19.1% on Terminal-Bench 4.0, trailing larger frontier models like Claude Opus 5 on complex container tasks.
Configurable Reasoning Budgets: Provides low, medium, and high thinking levels so developers can regulate token latency and depth on demand.
Marginal Exam Reasoning Gains: General scientific reasoning showed negligible movement, scoring 45.4% on Humanity's Last Exam compared to 45.7% on 3.7 Flash.
Grounded Tool Integration: Connects natively to Google Search and Google Maps APIs to provide verifiable factual answers with low hallucination rates.
Scheduled Rate Increase: Introductory pricing expires on December 31, 2026, doubling costs to $1.50 input and $7.50 output per million tokens.

API Quick Start

google/gemini-3.8-flash

View Documentation
google SDK
import { GoogleGenAI } from "@google/genai";

const ai = new GoogleGenAI({ apiKey: process.env.GEMINI_API_KEY });

async function main() {
  const response = await ai.models.generateContent({
    model: "gemini-3.8-flash",
    contents: "Analyze this code repository structure for memory leaks.",
    config: {
      maxOutputTokens: 8192,
      thinkingConfig: { thinkingBudget: 2048 }
    }
  });
  console.log(response.text);
}

main();

Install the SDK and start making API calls in minutes.

Community Feedback

See what the community thinks about Gemini 3.8 Flash

Gemini 3.8 Flash is another jump in agentic capabilities... our 3rd updated Flash model in only 6 weeks.
Logan Kilpatrick
twitter
The speed combined with the fact that this thing is really good at geography and geospatial skills makes it stunning.
MapDev
hackernews
Gemini 3.8 Flash is top of the Redactle LLM benchmark... it also does the evals cheaper and faster than almost all other models.
PuzzleSolver
reddit
I main antigravity for work. Over the past year we've gone from it taking minutes to outputting near perfect work in 10 seconds.
SaaSBuilder
hackernews
Gemini Flash 3.8 beats Opus 5 at 15x lower price on four self-contained three.js physics tasks.
GraphicsCoder
twitter
The 1M context window handling video directly without converting to images first saves tons of engineering work.
VideoAIPro
reddit

Related Videos

Watch tutorials, reviews, and discussions about Gemini 3.8 Flash

Deep SUI V1.1 Gemini 3.8 Flash coming in at 73.7%... effectively even with Claude Opus 5 which was just released

Then we have terminal bench 2.1. It got the number one score at 89.4. This is agentic terminal coding

good. So, Claude Opus 5 absolutely dominating the competition, coming in at 1824, second place 1710 for GPT 5.6 Soul, and then kind of a much less good score of 1545 for Gemini 3.8 Flash. So

According to artificial intelligence index, it's essentially at the Pareto frontier of cost versus performance

Compared to the previous Gemini flash model, this ate a lot more tokens for a given task... up to 30% increase

For token generation, you can expect up to 300 tokens per second according to the Artificial Intelligence Index.

It's really impressive that at this price you are getting Opus 5 level of intelligence

The big problem of Gemini is solved which is it is not adding too much of waste content on the website

And guys, make no mistake, this is an impressive model because every couple of weeks we are seeing new models from Gemini like 3.6, 3.7 Flash and now 3.8 Flash.

More than just prompts

Supercharge your workflow with AI Automation

Automatio combines the power of AI agents, web automation, and smart integrations to help you accomplish more in less time.

AI Agents
Web Automation
Smart Workflows

What to get right first

The decisions that are painful to change later in Gemini 3.8 Flash.

Select Thinking Levels Explicitly

Configure thinking effort to low for structured data transforms or high for terminal debugging to control token usage.

Enable Context Caching

Activate Gemini API context caching on prompts over 32,000 tokens to reduce input costs by up to 75 percent.

Ground Queries with Native Tools

Declare Google Search and Maps tools directly in the request to pull verified real-world facts and recent data.

Supply Direct Compiler Feedback

Pipe compiler and test runner output directly back to the model so it can fix syntax and logic errors in loops.

Testimonials

What Our Users Say

Join thousands of satisfied users who have transformed their workflow

Jonathan Kogan

Jonathan Kogan

Co-Founder/CEO, rpatools.io

Automatio is one of the most used for RPA Tools both internally and externally. It saves us countless hours of work and we realized this could do the same for other startups and so we choose Automatio for most of our automation needs.

Mohammed Ibrahim

Mohammed Ibrahim

CEO, qannas.pro

I have used many tools over the past 5 years, Automatio is the Jack of All trades.. !! it could be your scraping bot in the morning and then it becomes your VA by the noon and in the evening it does your automations.. its amazing!

Ben Bressington

Ben Bressington

CTO, AiChatSolutions

Automatio is fantastic and simple to use to extract data from any website. This allowed me to replace a developer and do tasks myself as they only take a few minutes to setup and forget about it. Automatio is a game changer!

Sarah Chen

Sarah Chen

Head of Growth, ScaleUp Labs

We've tried dozens of automation tools, but Automatio stands out for its flexibility and ease of use. Our team productivity increased by 40% within the first month of adoption.

David Park

David Park

Founder, DataDriven.io

The AI-powered features in Automatio are incredible. It understands context and adapts to changes in websites automatically. No more broken scrapers!

Emily Rodriguez

Emily Rodriguez

Marketing Director, GrowthMetrics

Automatio transformed our lead generation process. What used to take our team days now happens automatically in minutes. The ROI is incredible.

Jonathan Kogan

Jonathan Kogan

Co-Founder/CEO, rpatools.io

Automatio is one of the most used for RPA Tools both internally and externally. It saves us countless hours of work and we realized this could do the same for other startups and so we choose Automatio for most of our automation needs.

Mohammed Ibrahim

Mohammed Ibrahim

CEO, qannas.pro

I have used many tools over the past 5 years, Automatio is the Jack of All trades.. !! it could be your scraping bot in the morning and then it becomes your VA by the noon and in the evening it does your automations.. its amazing!

Ben Bressington

Ben Bressington

CTO, AiChatSolutions

Automatio is fantastic and simple to use to extract data from any website. This allowed me to replace a developer and do tasks myself as they only take a few minutes to setup and forget about it. Automatio is a game changer!

Sarah Chen

Sarah Chen

Head of Growth, ScaleUp Labs

We've tried dozens of automation tools, but Automatio stands out for its flexibility and ease of use. Our team productivity increased by 40% within the first month of adoption.

David Park

David Park

Founder, DataDriven.io

The AI-powered features in Automatio are incredible. It understands context and adapts to changes in websites automatically. No more broken scrapers!

Emily Rodriguez

Emily Rodriguez

Marketing Director, GrowthMetrics

Automatio transformed our lead generation process. What used to take our team days now happens automatically in minutes. The ROI is incredible.

Related AI Models

deepseek

DeepSeek-V3.2-Speciale

DeepSeek

DeepSeek-V3.2-Speciale is a reasoning-first LLM featuring gold-medal math performance, DeepSeek Sparse Attention, and a 131K context window. Rivaling GPT-5...

131K context
$0.28/$0.42/1M
other

MiMo V2.5 Pro

Other

MiMo V2.5 Pro is Xiaomi's open-source 1.02T parameter MoE model featuring a 1M context window, native multimodality, and elite agentic coding performance.

1M context
$1.00/$3.00/1M
deepseek

DeepSeek-V4-Flash

DeepSeek

DeepSeek-V4-Flash is an open-weight 1M context AI model scoring 54.4% on SWE-bench at $0.14 per 1M tokens, optimized for agentic coding and reasoning.

1M context
$0.14/$0.28/1M
moonshot

Kimi K2.7 Code

Moonshot

Kimi K2.7 Code is a 1T parameter MoE model from Moonshot AI. It features a 262k context window and 30% more efficient reasoning for software engineering.

262K context
$0.95/$4.00/1M
anthropic

Claude 3.7 Sonnet

Anthropic

Claude 3.7 Sonnet is Anthropic's first hybrid reasoning model, delivering state-of-the-art coding capabilities, a 200k context window, and visible thinking.

200K context
$3.00/$15.00/1M
minimax

MiniMax M2.5

minimax

MiniMax M2.5 is a SOTA MoE model featuring a 1M context window and elite agentic coding capabilities at disruptive pricing for autonomous agents.

1M context
$0.15/$1.20/1M
google

Gemini 3.6 Flash

Google

Gemini 3.6 Flash is Google's high-speed model featuring a 17% reduction in token consumption, $1.50/M input pricing, and advanced 3D visualization.

1M context
$1.50/$7.50/1M
google

Gemini 3.5 Flash

Google

Gemini 3.5 Flash is Google's high-speed multimodal model with a 1M context window, optimized for sub-second agentic loops and complex coding tasks.

1M context
$1.50/$9.00/1M

Frequently Asked Questions

Find answers to common questions about Gemini 3.8 Flash