deepseek

DeepSeek V4.1 Flash

DeepSeek V4.1 Flash 基于非对称 MoE 架构,提供 1M context、原生视觉以及每秒 400 token 的推理速度,每百万输入 token 仅需 $0.15。

Open WeightsMoEMultimodal VisionCode GenerationReasoning
deepseek logodeepseekDeepSeek V42026-09-10
上下文
1Mtokens
最大输出
384Ktokens
输入价格
$0.15/ 1M
输出价格
$0.60/ 1M
模态:TextImage
能力:视觉工具流式传输推理
基准测试
GPQA
90.9%
GPQA: 研究生级科学问答. 由领域专家创建的448道多选题的严格基准测试,涵盖生物学、物理学和化学。博士专家仅达到65-74%的准确率。 DeepSeek V4.1 Flash 在此基准测试中得分 90.9%。
HLE
36.8%
HLE: 高级专业推理. 测试模型在专业领域展示专家级推理能力的能力。 DeepSeek V4.1 Flash 在此基准测试中得分 36.8%。
MMLU
91%
MMLU: 大规模多任务语言理解. 涵盖57个学科的16,000道多选题的综合基准测试。 DeepSeek V4.1 Flash 在此基准测试中得分 91%。
MMLU Pro
81.2%
MMLU Pro: MMLU专业版. MMLU的增强版本,包含12,032道使用更难的10选项多选格式的问题。 DeepSeek V4.1 Flash 在此基准测试中得分 81.2%。
SimpleQA
49%
SimpleQA: 事实准确性基准. 测试模型对直接问题提供准确、事实性回答的能力。 DeepSeek V4.1 Flash 在此基准测试中得分 49%。
IFEval
89.5%
IFEval: 指令遵循评估. 衡量模型遵循特定指令和约束的能力。 DeepSeek V4.1 Flash 在此基准测试中得分 89.5%。
AIME 2025
87.5%
AIME 2025: 美国数学邀请赛. 来自著名AIME考试的竞赛级数学问题。 DeepSeek V4.1 Flash 在此基准测试中得分 87.5%。
MATH
88.5%
MATH: 数学问题解决. 涵盖代数、几何、微积分等领域的综合数学基准测试。 DeepSeek V4.1 Flash 在此基准测试中得分 88.5%。
GSM8k
96.8%
GSM8k: 小学数学8K. 8,500道需要多步推理的小学水平数学应用题。 DeepSeek V4.1 Flash 在此基准测试中得分 96.8%。
MGSM
92%
MGSM: 多语言小学数学. GSM8k基准测试翻译成10种语言版本。 DeepSeek V4.1 Flash 在此基准测试中得分 92%。
MathVista
72.5%
MathVista: 数学视觉推理. 测试解决涉及图表、图形等视觉元素的数学问题的能力。 DeepSeek V4.1 Flash 在此基准测试中得分 72.5%。
SWE-Bench
74.2%
SWE-Bench: 软件工程基准. AI模型尝试解决开源Python项目中的真实GitHub问题。 DeepSeek V4.1 Flash 在此基准测试中得分 74.2%。
HumanEval
92.5%
HumanEval: Python编程问题. 164道手写编程问题,模型必须生成正确的Python函数实现。 DeepSeek V4.1 Flash 在此基准测试中得分 92.5%。
LiveCodeBench
73.3%
LiveCodeBench: 实时编程基准. 在持续更新的真实世界编程挑战中测试编程能力。 DeepSeek V4.1 Flash 在此基准测试中得分 73.3%。
MMMU
78.9%
MMMU: 多模态理解. 大规模多学科多模态理解基准测试,测试视觉语言模型在大学水平问题上的表现。 DeepSeek V4.1 Flash 在此基准测试中得分 78.9%。
MMMU Pro
58.4%
MMMU Pro: MMMU专业版. MMMU的增强版本,问题更具挑战性,评估更严格。 DeepSeek V4.1 Flash 在此基准测试中得分 58.4%。
ChartQA
88%
ChartQA: 图表问答. 测试理解和推理图表信息的能力。 DeepSeek V4.1 Flash 在此基准测试中得分 88%。
DocVQA
93%
DocVQA: 文档视觉问答. 测试从文档图像中提取信息的能力。 DeepSeek V4.1 Flash 在此基准测试中得分 93%。
Terminal-Bench
90.6%
Terminal-Bench: 终端/CLI任务. 测试执行命令行操作和编写shell脚本的能力。 DeepSeek V4.1 Flash 在此基准测试中得分 90.6%。
ARC-AGI
11.4%
ARC-AGI: 抽象与推理. AGI抽象和推理语料库 - 通过新颖的模式识别谜题测试流体智力。 DeepSeek V4.1 Flash 在此基准测试中得分 11.4%。

关于 DeepSeek V4.1 Flash

了解 DeepSeek V4.1 Flash 的功能、特性以及它如何帮助您获得更好的效果。

DeepSeek V4.1 Flash 是一种 open-source 的 mixture-of-experts 模型,总共包含 5520 亿个参数。该模型引入了非对称因果编码器-解码器结构,旨在最小化高吞吐量工作负载期间的推理成本。在处理传入的 prompt 时,网络仅激活 80 亿个参数,在 token 生成期间增加到 160 亿个。底层骨干网结合了跨层共享的压缩 key-value 缓存、两阶段稀疏索引器以及 1960 亿参数的 Engram 查找内存,所有这些都在 45 万亿 token 的语料库上进行了训练。

与 V4 系列的早期迭代不同,原生视觉理解作为标准配置提供,无需单独的视觉检查点。该模型可以直接在标准文本 prompt 中接受图像,并评估诸如架构图、UI 布局和技术图表等视觉构件。与此同时,thinking 模式原生运行,在生成最终答案之前生成显式的 reasoning 轨迹。该模型在现代加速器集群上的生成速度在每秒 300 到 427 个 token 之间,匹配或超过了更小的 dense 模型的延迟水平,同时保留了 PhD 级别的 reasoning 能力。

DeepSeek 将 V4.1 Flash 定位为更大的 V4 Pro flagship 的直接替代品。在涵盖前端生成、终端操作和软件调试的第三方评估中,V4.1 Flash 在更低延迟和更少计算开销下,达到或超过了 V4 Pro 的准确率。该模型服务于高容量 agentic 环境、自动化终端工具以及连续的代码合成工作流,在这些场景中,flagship API 定价通常令人望而却步。

DeepSeek V4.1 Flash

DeepSeek V4.1 Flash 的使用案例

发现使用 DeepSeek V4.1 Flash 获得出色效果的不同方式。

自主终端与 Shell 操作

在容器化系统中执行系统诊断、运行构建工具并解决环境错误,在 Terminal-Bench 2.1 上获得了 90.6 的高分。

全栈 UI 与前端原型设计

通过单次文本或图像 prompt 生成交互式单页应用、WebGL 着色器、Three.js 3D 环境以及响应式仪表盘布局。

复杂多文件代码调试

在其 100 万 token 的 context window 内扫描整个多仓库软件项目,追踪跨文件导入,并修复倒置逻辑或竞态条件。

自动化视觉文档提取

检查复杂的架构蓝图、数据流图和用户界面原型,以输出结构化的 JSON 模式和可执行的 API 契约。

高吞吐量 Agentic 工具调用

运行连续的后台推理循环,轮询实时 REST 端点、查询 SQL 数据库,并在多个执行轮次中验证不变状态的变化。

多语言翻译与方言分析

翻译低资源方言中的习语、地区俚语和技术文档,同时标记不确定的翻译,而不是产生幻觉术语。

优势

局限性

非对称 MoE 效率: 在总共 5520 亿参数中,输入时仅激活 80 亿,输出时激活 160 亿,将非高峰期输入定价保持在每百万 token 0.15 美元。
逻辑过度思考: 如果用户的 thinking effort 未加约束,模型会在简单的编码 prompt 上消耗过多的推理 token。
state-of-the-art 的 agent 性能: 在 Terminal-Bench 2.1 上达到 90.6%,在 DeepSWE v1.1 上达到 74.2%,在工程任务上匹配甚至超越了 flagship 闭源选项。
运动物理学不准确: 在 3js 测试中生成复杂的 3D 运动学动画时,偶尔会出现运动路径异常和可疑的物理效果。
高吞吐量生成: 提供每秒 300 到 427 个 token 的经验证解码速度,大大缩短了重推理 agent 的周转时间。
美观 UI 局限性: 在复杂的排版和间距方面落后于闭源 flagship 模型,产生功能完善但视觉上千篇一律的仪表盘层级。
统一的多模态 context: 在 1,000,000 token 的 context window 内原生吞吐视觉和文本 token,无需二次视觉适配器的开销。
苛刻的本地内存需求: 需要专用的多 GPU 集群设置,才能容纳用于全精度本地托管的庞大 5520 亿参数内存占用。

API快速入门

deepseek/deepseek-v4.1-flash

查看文档
deepseek SDK
import OpenAI from "openai";

const client = new OpenAI({
  baseURL: "https://api.deepseek.com",
  apiKey: process.env.DEEPSEEK_API_KEY,
});

async function main() {
  const completion = await client.chat.completions.create({
    model: "deepseek-flash",
    messages: [
      { role: "system", content: "You are an expert systems engineer." },
      { role: "user", content: "Write a high-performance WebGL compute shader." },
    ],
  });

  console.log(completion.choices[0].message.content);
}

main();

安装SDK并在几分钟内开始进行API调用。

人们对 DeepSeek V4.1 Flash 的评价

看看社区对 DeepSeek V4.1 Flash 的看法

根据用户请求,它在日常设计任务中以 1.4% 的成本达到了 GPT-6 Astra 分数的 98%。除了 Astra 之外,其他所有模型的分数都更低且成本更高。
OpenDesign (@OpenDesignHQ)
twitter
Deepseek V4.1 Flash 总计 5520 亿参数,80 亿/160 亿激活,采用在 45T token 上训练的新架构……这可能是近期我见过最具创新性的架构了,相当疯狂。
Elie Bakouch (@eliebakouch)
twitter
说实话,速度比大多数人想象的要重要得多,对于大多数产品用例,我宁愿选择一个稍微差一点但速度快 2 倍的模型。
MoneyLovesSpeed
hackernews
有那么一会儿,它飙升到了每秒 427 个 token 的惊人速度。但最疯狂的是,据报道整个运行成本仅为 30 美分。
WorldofAI
youtube
V4 Pro 查询被自动重定向到 V4.1 Flash 这一事实,向你说明了该架构实际有多么优秀。
KevDev99
reddit
对于这个价格区间的模型来说,Terminal-Bench 达到 90.6 真是太疯狂了。Agentic 工具的成本刚刚大幅下降。
SysAdminHero
reddit

关于 DeepSeek V4.1 Flash 的视频

观看关于 DeepSeek V4.1 Flash 的教程、评测和讨论

这个新的 DeepSeek 4.1 flash 模型快得离谱。你的速度大约是每秒 400 个 token,说实话,这个速度简直疯狂。

对于一个具备这种能力水平的推理模型来说,这种速度确实令人印象深刻,考虑到这只是一个临时的测试版本。

有那么一会儿,它飙升到了每秒 427 个 token 的惊人速度。但是,最疯狂的部分据说是整个运行成本只需 30 美分。

现实世界的测试测得每秒 300 到 400 多个 token,达到了 GPT6 Astra 设计 benchmark 分数的 98%。

它不只是在生成代码,它实际上在验证自己的数学。它甚至自己捕捉到了一个微妙的轨道控制丢弃 bug。

权重一旦在最终版本中发布,我们就会去查看它。再次提醒,如果你想支持这个频道,请成为会员。谢谢

使用这个模型运行整个 Artificial Analysis 测试套件的成本是 72 美元。这比拥有相同 intelligence 分数的模型便宜 10 倍。

DeepSeek V4 Flash 实际上是所有模型中最便宜的,而成本相同的 GPT 5.6 Luna 在 Intelligence Index 上实际上落后了两分。

我确实很喜欢这种中国实验室进入市场、在定价上压低美国实验室同时又与其 intelligence 相匹配的趋势。

不仅仅是提示词

用以下方式提升您的工作流程 AI自动化

Automatio结合AI代理、网页自动化和智能集成的力量,帮助您在更短的时间内完成更多工作。

AI代理
网页自动化
智能工作流

DeepSeek V4.1 Flash专业提示

专家提示助您充分利用DeepSeek V4.1 Flash。

管理推理努力程度

对于简单的 CRUD 生成,将 reasoning effort 设置为 low;在解决多步数学计算或复杂的代码库 Bug 时设置为 high 或 max,以优化 token 使用量。

最大化 prompt 缓存

将连续的系统 prompt 和静态文件引用放在 context window 的靠前位置,以便在非高峰期以每百万 0.003 美元的费率最大化 prompt 缓存命中率。

直接多模态输入

直接提供原始图像和视觉原型以及 CSS 需求,而不是手动转录布局规范,以获得更好的空间准确性。

使用官方模型字符串

在 API 请求中使用官方的 deepseek-flash 模型字符串,以确保自动路由到最新的活动检查点和最高效的定价。

用户评价

用户怎么说

加入数千名已改变工作流程的满意用户

Jonathan Kogan

Jonathan Kogan

Co-Founder/CEO, rpatools.io

Automatio is one of the most used for RPA Tools both internally and externally. It saves us countless hours of work and we realized this could do the same for other startups and so we choose Automatio for most of our automation needs.

Mohammed Ibrahim

Mohammed Ibrahim

CEO, qannas.pro

I have used many tools over the past 5 years, Automatio is the Jack of All trades.. !! it could be your scraping bot in the morning and then it becomes your VA by the noon and in the evening it does your automations.. its amazing!

Ben Bressington

Ben Bressington

CTO, AiChatSolutions

Automatio is fantastic and simple to use to extract data from any website. This allowed me to replace a developer and do tasks myself as they only take a few minutes to setup and forget about it. Automatio is a game changer!

Sarah Chen

Sarah Chen

Head of Growth, ScaleUp Labs

We've tried dozens of automation tools, but Automatio stands out for its flexibility and ease of use. Our team productivity increased by 40% within the first month of adoption.

David Park

David Park

Founder, DataDriven.io

The AI-powered features in Automatio are incredible. It understands context and adapts to changes in websites automatically. No more broken scrapers!

Emily Rodriguez

Emily Rodriguez

Marketing Director, GrowthMetrics

Automatio transformed our lead generation process. What used to take our team days now happens automatically in minutes. The ROI is incredible.

Jonathan Kogan

Jonathan Kogan

Co-Founder/CEO, rpatools.io

Automatio is one of the most used for RPA Tools both internally and externally. It saves us countless hours of work and we realized this could do the same for other startups and so we choose Automatio for most of our automation needs.

Mohammed Ibrahim

Mohammed Ibrahim

CEO, qannas.pro

I have used many tools over the past 5 years, Automatio is the Jack of All trades.. !! it could be your scraping bot in the morning and then it becomes your VA by the noon and in the evening it does your automations.. its amazing!

Ben Bressington

Ben Bressington

CTO, AiChatSolutions

Automatio is fantastic and simple to use to extract data from any website. This allowed me to replace a developer and do tasks myself as they only take a few minutes to setup and forget about it. Automatio is a game changer!

Sarah Chen

Sarah Chen

Head of Growth, ScaleUp Labs

We've tried dozens of automation tools, but Automatio stands out for its flexibility and ease of use. Our team productivity increased by 40% within the first month of adoption.

David Park

David Park

Founder, DataDriven.io

The AI-powered features in Automatio are incredible. It understands context and adapts to changes in websites automatically. No more broken scrapers!

Emily Rodriguez

Emily Rodriguez

Marketing Director, GrowthMetrics

Automatio transformed our lead generation process. What used to take our team days now happens automatically in minutes. The ROI is incredible.

相关 AI Models

moonshot

Kimi k2.6

Moonshot

Kimi k2.6 is Moonshot AI's 1T-parameter MoE model featuring a 256K context window, native video input, and elite performance in autonomous agentic coding.

256K context
$0.95/$4.00/1M
anthropic

Claude Opus 4.6

Anthropic

Claude Opus 4.6 is Anthropic's flagship model featuring a 1M token context window, Adaptive Thinking, and world-class coding and reasoning performance.

1M context
$5.00/$25.00/1M
google

Gemini 3 Flash

Google

Gemini 3 Flash is Google's high-speed multimodal model featuring a 1M token context window, elite 90.4% GPQA reasoning, and autonomous browser automation tools.

1M context
$0.50/$3.00/1M
deepseek

DeepSeek v4

DeepSeek

DeepSeek v4 is a 1.6T parameter MoE model featuring a 1M token context window and native multimodal support for text, vision, and video at disruptive prices.

1M context
$1.74/$3.48/1M
anthropic

Claude Sonnet 4.6

Anthropic

Claude Sonnet 4.6 offers frontier performance for coding and computer use with a massive 1M token context window for only $3/1M tokens.

1M context
$3.00/$15.00/1M
google

Gemini 3 Pro

Google

Google's Gemini 3 Pro is a multimodal powerhouse featuring a 1M token context window, native video processing, and industry-leading reasoning performance.

1M context
$2.00/$12.00/1M
alibaba

Qwen 3.7 Max

alibaba

Qwen 3.7 Max is Alibaba’s flagship AI model for deep reasoning and autonomous agent tasks, featuring a 256k context window and top-tier coding performance.

256K context
$1.20/$6.00/1M
openai

GPT-5.2 Pro

OpenAI

GPT-5.2 Pro is OpenAI's 2025 flagship reasoning model featuring Extended Thinking for SOTA performance in mathematics, coding, and expert knowledge work.

400K context
$21.00/$168.00/1M

关于DeepSeek V4.1 Flash的常见问题

查找关于DeepSeek V4.1 Flash的常见问题答案