google

Gemini 3.8 Flash

Gemini 3.8 Flash 是 Google 推出的多模态 AI,提供 1M context、64K 输出以及 agentic coding 能力,每百万输入 token 仅需 0.75 美元。

MultimodalHigh-SpeedGoogle DeepMindLong ContextAgentic Coding
google logogoogleGemini 32026-09-02
上下文
1Mtokens
最大输出
66Ktokens
输入价格
$0.75/ 1M
输出价格
$3.75/ 1M
模态:TextImageAudioVideo
能力:视觉工具流式传输推理
基准测试
GPQA
59%
GPQA: 研究生级科学问答. 由领域专家创建的448道多选题的严格基准测试,涵盖生物学、物理学和化学。博士专家仅达到65-74%的准确率。 Gemini 3.8 Flash 在此基准测试中得分 59%。
HLE
45.4%
HLE: 高级专业推理. 测试模型在专业领域展示专家级推理能力的能力。 Gemini 3.8 Flash 在此基准测试中得分 45.4%。
MMLU
88.2%
MMLU: 大规模多任务语言理解. 涵盖57个学科的16,000道多选题的综合基准测试。 Gemini 3.8 Flash 在此基准测试中得分 88.2%。
MMLU Pro
76.5%
MMLU Pro: MMLU专业版. MMLU的增强版本,包含12,032道使用更难的10选项多选格式的问题。 Gemini 3.8 Flash 在此基准测试中得分 76.5%。
SimpleQA
42%
SimpleQA: 事实准确性基准. 测试模型对直接问题提供准确、事实性回答的能力。 Gemini 3.8 Flash 在此基准测试中得分 42%。
IFEval
87.5%
IFEval: 指令遵循评估. 衡量模型遵循特定指令和约束的能力。 Gemini 3.8 Flash 在此基准测试中得分 87.5%。
AIME 2025
82%
AIME 2025: 美国数学邀请赛. 来自著名AIME考试的竞赛级数学问题。 Gemini 3.8 Flash 在此基准测试中得分 82%。
MATH
78%
MATH: 数学问题解决. 涵盖代数、几何、微积分等领域的综合数学基准测试。 Gemini 3.8 Flash 在此基准测试中得分 78%。
GSM8k
94%
GSM8k: 小学数学8K. 8,500道需要多步推理的小学水平数学应用题。 Gemini 3.8 Flash 在此基准测试中得分 94%。
MGSM
91%
MGSM: 多语言小学数学. GSM8k基准测试翻译成10种语言版本。 Gemini 3.8 Flash 在此基准测试中得分 91%。
MathVista
68%
MathVista: 数学视觉推理. 测试解决涉及图表、图形等视觉元素的数学问题的能力。 Gemini 3.8 Flash 在此基准测试中得分 68%。
SWE-Bench
49%
SWE-Bench: 软件工程基准. AI模型尝试解决开源Python项目中的真实GitHub问题。 Gemini 3.8 Flash 在此基准测试中得分 49%。
HumanEval
85%
HumanEval: Python编程问题. 164道手写编程问题,模型必须生成正确的Python函数实现。 Gemini 3.8 Flash 在此基准测试中得分 85%。
LiveCodeBench
65%
LiveCodeBench: 实时编程基准. 在持续更新的真实世界编程挑战中测试编程能力。 Gemini 3.8 Flash 在此基准测试中得分 65%。
MMMU
69.5%
MMMU: 多模态理解. 大规模多学科多模态理解基准测试,测试视觉语言模型在大学水平问题上的表现。 Gemini 3.8 Flash 在此基准测试中得分 69.5%。
MMMU Pro
52%
MMMU Pro: MMMU专业版. MMMU的增强版本,问题更具挑战性,评估更严格。 Gemini 3.8 Flash 在此基准测试中得分 52%。
ChartQA
86%
ChartQA: 图表问答. 测试理解和推理图表信息的能力。 Gemini 3.8 Flash 在此基准测试中得分 86%。
DocVQA
93.5%
DocVQA: 文档视觉问答. 测试从文档图像中提取信息的能力。 Gemini 3.8 Flash 在此基准测试中得分 93.5%。
Terminal-Bench
89.4%
Terminal-Bench: 终端/CLI任务. 测试执行命令行操作和编写shell脚本的能力。 Gemini 3.8 Flash 在此基准测试中得分 89.4%。
ARC-AGI
4%
ARC-AGI: 抽象与推理. AGI抽象和推理语料库 - 通过新颖的模式识别谜题测试流体智力。 Gemini 3.8 Flash 在此基准测试中得分 4%。

关于 Gemini 3.8 Flash

了解 Gemini 3.8 Flash 的功能、特性以及它如何帮助您获得更好的效果。

模型概述

Gemini 3.8 Flash 是 Google DeepMind 于 2026 年 9 月发布的高效多模态模型。该系统基于 Google 的 TPU 硬件构建,可在单一 1,048,576 token 的 context window 内原生摄取文本、音频、高分辨率图像以及长达两小时的视频流。单次请求的输出容量可达 65,536 token。该模型引入了跨低、中、高参数的可配置 reasoning budget,使工程师能够在执行延迟和 token 消耗之间取得平衡。

Agentic 执行与架构

该模型专为长周期(long-horizon)编码任务和终端控制而设计。Gemini 3.8 Flash 不仅限于单次输出传递,还支持递归执行循环和自主 tool interaction。训练重点强调了容器操作、命令行接口和网络安全防御工作负载。这种训练使得在测试套件上的错误率更低,并且能够在 Google Antigravity 等环境中进行实时构建时的自我纠错。

生产工作负载与 Grounding

开发者将 Gemini 3.8 Flash 部署到需要确定性工具使用的延迟敏感型应用中。原生集成将模型直接连接到 Google Search grounding 和 Google Maps 数据,而无需单独的检索管道。虽然 frontier flagship model 在开放式创意任务上仍保留优势,但 3.8 Flash 以更低的 inference 价格提供了可媲美的 coding eval 得分。

Gemini 3.8 Flash

Gemini 3.8 Flash 的使用案例

发现使用 Gemini 3.8 Flash 获得出色效果的不同方式。

自主终端维护

执行多文件重构,在容器化环境中运行测试套件,并自主修补运行时错误。

高吞吐量视频摄取

直接在其 1M token 窗口内分析长达两小时的原始视频录像和音频流,无需外部预处理管道。

财务报告提取

摄取冗长的财务文件和复杂的视觉 PDF,将资产负债表指标计算为验证过的 JSON。

快速 UI 原型设计

在 15 秒内生成完整的 Web 应用、交互式 SVG 仪表盘和 WebGL 模拟。

防御性安全验证

扫描源码仓库以识别逻辑漏洞、追踪被污染的变量,并生成候选回归补丁。

低延迟 Agent 编排

在分层多 Agent 架构中担任快速动作规划器,调用由 Google Search 支持的外部工具。

优势

局限性

极具经济效益的 Agentic Coding: 在 DeepSWE v1.1 上达到 73.7%,在 Terminal-Bench 2.1 上达到 89.4%,每百万输入 token 的尝鲜价仅为 0.75 美元。
深度思考时 token 消耗大: 高推理配置会执行扩展的内部循环,与 3.7 Flash 相比,token 输出量最多增加 30%。
原生超长上下文多模态: 可处理多达 1,048,576 token 的组合音频、视频和文本,无需外部帧提取管道。
复杂终端 benchmark 准确率较低: 在 Terminal-Bench 4.0 上得分为 19.1%,在复杂的容器任务上落后于 Claude Opus 5 等更大的 frontier model。
可配置的推理预算 (Reasoning Budgets): 提供低、中、高 thinking level,使开发者能够按需调控 token 延迟和深度。
考试推理能力提升有限: 通用科学推理能力提升微乎其微,在 Humanity's Last Exam 上得分为 45.4%,而 3.7 Flash 为 45.7%。
扎实的工具集成: 原生连接 Google Search 和 Google Maps API,以低幻觉率提供可验证的事实答案。
计划中的价格上调: 尝鲜价将于 2026 年 12 月 31 日到期,届时每百万 token 的输入和输出成本将翻倍至 1.50 美元和 7.50 美元。

API快速入门

google/gemini-3.8-flash

查看文档
google SDK
import { GoogleGenAI } from "@google/genai";

const ai = new GoogleGenAI({ apiKey: process.env.GEMINI_API_KEY });

async function main() {
  const response = await ai.models.generateContent({
    model: "gemini-3.8-flash",
    contents: "Analyze this code repository structure for memory leaks.",
    config: {
      maxOutputTokens: 8192,
      thinkingConfig: { thinkingBudget: 2048 }
    }
  });
  console.log(response.text);
}

main();

安装SDK并在几分钟内开始进行API调用。

人们对 Gemini 3.8 Flash 的评价

看看社区对 Gemini 3.8 Flash 的看法

Gemini 3.8 Flash 是 agentic 能力的又一次飞跃... 这是我们短短 6 周内推出的第 3 个更新版 Flash 模型。
Logan Kilpatrick
twitter
极快的速度,再加上它在地理和地理空间技能方面表现出色,实在令人惊叹。
MapDev
hackernews
Gemini 3.8 Flash 稳居 Redactle LLM benchmark 榜首... 而且它进行 eval 的成本和速度比几乎所有其他模型都要好。
PuzzleSolver
reddit
我工作主要用 antigravity。过去一年里,我们见证了它从需要几分钟输出到在 10 秒内产出近乎完美的工作成果。
SaaSBuilder
hackernews
在四个独立的 three.js 物理任务中,Gemini Flash 3.8 以低 15 倍的价格击败了 Opus 5。
GraphicsCoder
twitter
1M 的 context window 直接处理视频,无需先转换为图像,省去了大量的工程开发工作。
VideoAIPro
reddit

关于 Gemini 3.8 Flash 的视频

观看关于 Gemini 3.8 Flash 的教程、评测和讨论

Deep SUI V1.1 Gemini 3.8 Flash 达到了 73.7%... 实际上与刚发布的 Claude Opus 5 不相上下

然后是 terminal bench 2.1。它以 89.4% 获得了第一名。这就是 agentic 终端编码

很好。所以,Claude Opus 5 绝对统治了竞争,得分为 1824,第二名是 GPT 5.6 Soul 的 1710,而 Gemini 3.8 Flash 的得分则相对较低,为 1545。所以

根据人工智能指数,它基本上处于成本与性能的 Pareto 前沿

与上一代 Gemini flash 模型相比,它在特定任务中消耗了更多 token... 增幅高达 30%

根据人工智能指数,在 token 生成方面,你最高可以期待每秒 300 个 token 的速度。

令人印象深刻的是,以这个价格你竟然能获得 Opus 5 级别的智能水平

Gemini 的大问题解决了,那就是它没有在网站上添加太多无用的内容

伙计们,毫无疑问,这是一个令人印象深刻的模型,因为每隔几周我们就能看到来自 Gemini 的新模型,比如 3.6、3.7 Flash,现在又是 3.8 Flash。

不仅仅是提示词

用以下方式提升您的工作流程 AI自动化

Automatio结合AI代理、网页自动化和智能集成的力量,帮助您在更短的时间内完成更多工作。

AI代理
网页自动化
智能工作流

Gemini 3.8 Flash专业提示

专家提示助您充分利用Gemini 3.8 Flash。

显式选择 Thinking Level

将 thinking effort 配置为 low 以处理结构化数据转换,或配置为 high 以进行终端调试,从而有效控制 token 消耗。

启用上下文缓存 (Context Caching)

在超过 32,000 token 的 prompt 上激活 Gemini API 上下文缓存,可将输入成本降低高达 75%。

使用原生工具进行 Grounding

直接在请求中声明 Google Search 和 Maps 工具,以获取经过验证的真实世界事实和最新数据。

提供直接的编译器反馈

将编译器和测试运行器的输出直接反馈给模型,使其能够在循环中修复语法和逻辑错误。

用户评价

用户怎么说

加入数千名已改变工作流程的满意用户

Jonathan Kogan

Jonathan Kogan

Co-Founder/CEO, rpatools.io

Automatio is one of the most used for RPA Tools both internally and externally. It saves us countless hours of work and we realized this could do the same for other startups and so we choose Automatio for most of our automation needs.

Mohammed Ibrahim

Mohammed Ibrahim

CEO, qannas.pro

I have used many tools over the past 5 years, Automatio is the Jack of All trades.. !! it could be your scraping bot in the morning and then it becomes your VA by the noon and in the evening it does your automations.. its amazing!

Ben Bressington

Ben Bressington

CTO, AiChatSolutions

Automatio is fantastic and simple to use to extract data from any website. This allowed me to replace a developer and do tasks myself as they only take a few minutes to setup and forget about it. Automatio is a game changer!

Sarah Chen

Sarah Chen

Head of Growth, ScaleUp Labs

We've tried dozens of automation tools, but Automatio stands out for its flexibility and ease of use. Our team productivity increased by 40% within the first month of adoption.

David Park

David Park

Founder, DataDriven.io

The AI-powered features in Automatio are incredible. It understands context and adapts to changes in websites automatically. No more broken scrapers!

Emily Rodriguez

Emily Rodriguez

Marketing Director, GrowthMetrics

Automatio transformed our lead generation process. What used to take our team days now happens automatically in minutes. The ROI is incredible.

Jonathan Kogan

Jonathan Kogan

Co-Founder/CEO, rpatools.io

Automatio is one of the most used for RPA Tools both internally and externally. It saves us countless hours of work and we realized this could do the same for other startups and so we choose Automatio for most of our automation needs.

Mohammed Ibrahim

Mohammed Ibrahim

CEO, qannas.pro

I have used many tools over the past 5 years, Automatio is the Jack of All trades.. !! it could be your scraping bot in the morning and then it becomes your VA by the noon and in the evening it does your automations.. its amazing!

Ben Bressington

Ben Bressington

CTO, AiChatSolutions

Automatio is fantastic and simple to use to extract data from any website. This allowed me to replace a developer and do tasks myself as they only take a few minutes to setup and forget about it. Automatio is a game changer!

Sarah Chen

Sarah Chen

Head of Growth, ScaleUp Labs

We've tried dozens of automation tools, but Automatio stands out for its flexibility and ease of use. Our team productivity increased by 40% within the first month of adoption.

David Park

David Park

Founder, DataDriven.io

The AI-powered features in Automatio are incredible. It understands context and adapts to changes in websites automatically. No more broken scrapers!

Emily Rodriguez

Emily Rodriguez

Marketing Director, GrowthMetrics

Automatio transformed our lead generation process. What used to take our team days now happens automatically in minutes. The ROI is incredible.

相关 AI Models

deepseek

DeepSeek-V3.2-Speciale

DeepSeek

DeepSeek-V3.2-Speciale is a reasoning-first LLM featuring gold-medal math performance, DeepSeek Sparse Attention, and a 131K context window. Rivaling GPT-5...

131K context
$0.28/$0.42/1M
other

MiMo V2.5 Pro

Other

MiMo V2.5 Pro is Xiaomi's open-source 1.02T parameter MoE model featuring a 1M context window, native multimodality, and elite agentic coding performance.

1M context
$1.00/$3.00/1M
deepseek

DeepSeek-V4-Flash

DeepSeek

DeepSeek-V4-Flash is an open-weight 1M context AI model scoring 54.4% on SWE-bench at $0.14 per 1M tokens, optimized for agentic coding and reasoning.

1M context
$0.14/$0.28/1M
moonshot

Kimi K2.7 Code

Moonshot

Kimi K2.7 Code is a 1T parameter MoE model from Moonshot AI. It features a 262k context window and 30% more efficient reasoning for software engineering.

262K context
$0.95/$4.00/1M
anthropic

Claude 3.7 Sonnet

Anthropic

Claude 3.7 Sonnet is Anthropic's first hybrid reasoning model, delivering state-of-the-art coding capabilities, a 200k context window, and visible thinking.

200K context
$3.00/$15.00/1M
minimax

MiniMax M2.5

minimax

MiniMax M2.5 is a SOTA MoE model featuring a 1M context window and elite agentic coding capabilities at disruptive pricing for autonomous agents.

1M context
$0.15/$1.20/1M
google

Gemini 3.6 Flash

Google

Gemini 3.6 Flash is Google's high-speed model featuring a 17% reduction in token consumption, $1.50/M input pricing, and advanced 3D visualization.

1M context
$1.50/$7.50/1M
google

Gemini 3.5 Flash

Google

Gemini 3.5 Flash is Google's high-speed multimodal model with a 1M context window, optimized for sub-second agentic loops and complex coding tasks.

1M context
$1.50/$9.00/1M

关于Gemini 3.8 Flash的常见问题

查找关于Gemini 3.8 Flash的常见问题答案