deepseek

DeepSeek-V4-Flash

DeepSeek-V4-Flash is an open-weight 1M context AI model scoring 54.4% on SWE-bench at $0.14 per 1M tokens, optimized for agentic coding and reasoning.

Open WeightsAgentic CodingMoE Architecture1M Context WindowLow Cost
deepseek logodeepseekDeepSeek V42026-07-31
Context
1.0Mtokens
Max Output
8Ktokens
Input Price
$0.14/ 1M
Output Price
$0.28/ 1M
Modality:TextImage
Capabilities:VisionToolsStreamingReasoning
Benchmarks
GPQA
53.6%
GPQA: Graduate-Level Science Q&A. A rigorous benchmark with 448 multiple-choice questions in biology, physics, and chemistry created by domain experts. PhD experts only achieve 65-74% accuracy, while non-experts score just 34% even with unlimited web access (hence 'Google-proof'). DeepSeek-V4-Flash scored 53.6% on this benchmark.
HLE
25.2%
HLE: High-Level Expertise Reasoning. Tests a model's ability to demonstrate expert-level reasoning across specialized domains. Evaluates deep understanding of complex topics that require professional-level knowledge. DeepSeek-V4-Flash scored 25.2% on this benchmark.
MMLU
88.7%
MMLU: Massive Multitask Language Understanding. A comprehensive benchmark with 16,000 multiple-choice questions across 57 academic subjects including math, philosophy, law, and medicine. Tests broad knowledge and reasoning capabilities. DeepSeek-V4-Flash scored 88.7% on this benchmark.
MMLU Pro
72.6%
MMLU Pro: MMLU Professional Edition. An enhanced version of MMLU with 12,032 questions using a harder 10-option multiple choice format. Covers Math, Physics, Chemistry, Law, Engineering, Economics, Health, Psychology, Business, Biology, Philosophy, and Computer Science. DeepSeek-V4-Flash scored 72.6% on this benchmark.
SimpleQA
38.2%
SimpleQA: Factual Accuracy Benchmark. Tests a model's ability to provide accurate, factual responses to straightforward questions. Measures reliability and reduces hallucinations in knowledge retrieval tasks. DeepSeek-V4-Flash scored 38.2% on this benchmark.
IFEval
88%
IFEval: Instruction Following Evaluation. Measures how well a model follows specific instructions and constraints. Tests the ability to adhere to formatting rules, length limits, and other explicit requirements. DeepSeek-V4-Flash scored 88% on this benchmark.
AIME 2025
93.1%
AIME 2025: American Invitational Math Exam. Competition-level mathematics problems from the prestigious AIME exam designed for talented high school students. Tests advanced mathematical problem-solving requiring abstract reasoning, not just pattern matching. DeepSeek-V4-Flash scored 93.1% on this benchmark.
MATH
76.6%
MATH: Mathematical Problem Solving. A comprehensive math benchmark testing problem-solving across algebra, geometry, calculus, and other mathematical domains. Requires multi-step reasoning and formal mathematical knowledge. DeepSeek-V4-Flash scored 76.6% on this benchmark.
GSM8k
96%
GSM8k: Grade School Math 8K. 8,500 grade school-level math word problems requiring multi-step reasoning. Tests basic arithmetic and logical thinking through real-world scenarios like shopping or time calculations. DeepSeek-V4-Flash scored 96% on this benchmark.
MGSM
90.5%
MGSM: Multilingual Grade School Math. The GSM8k benchmark translated into 10 languages including Spanish, French, German, Russian, Chinese, and Japanese. Tests mathematical reasoning across different languages. DeepSeek-V4-Flash scored 90.5% on this benchmark.
MathVista
63.8%
MathVista: Mathematical Visual Reasoning. Tests the ability to solve math problems that involve visual elements like charts, graphs, geometry diagrams, and scientific figures. Combines visual understanding with mathematical reasoning. DeepSeek-V4-Flash scored 63.8% on this benchmark.
SWE-Bench
54.4%
SWE-Bench: Software Engineering Benchmark. AI models attempt to resolve real GitHub issues in open-source Python projects with human verification. Tests practical software engineering skills on production codebases. Top models went from 4.4% in 2023 to over 70% in 2024. DeepSeek-V4-Flash scored 54.4% on this benchmark.
HumanEval
90.2%
HumanEval: Python Programming Problems. 164 hand-written programming problems where models must generate correct Python function implementations. Each solution is verified against unit tests. Top models now achieve 90%+ accuracy. DeepSeek-V4-Flash scored 90.2% on this benchmark.
LiveCodeBench
33%
LiveCodeBench: Live Coding Benchmark. Tests coding abilities on continuously updated, real-world programming challenges. Unlike static benchmarks, uses fresh problems to prevent data contamination and measure true coding skills. DeepSeek-V4-Flash scored 33% on this benchmark.
MMMU
69.1%
MMMU: Multimodal Understanding. Massive Multi-discipline Multimodal Understanding benchmark testing vision-language models on college-level problems across 30 subjects requiring both image understanding and expert knowledge. DeepSeek-V4-Flash scored 69.1% on this benchmark.
MMMU Pro
54%
MMMU Pro: MMMU Professional Edition. Enhanced version of MMMU with more challenging questions and stricter evaluation. Tests advanced multimodal reasoning at professional and expert levels. DeepSeek-V4-Flash scored 54% on this benchmark.
ChartQA
85.7%
ChartQA: Chart Question Answering. Tests the ability to understand and reason about information presented in charts and graphs. Requires extracting data, comparing values, and performing calculations from visual data representations. DeepSeek-V4-Flash scored 85.7% on this benchmark.
DocVQA
92.8%
DocVQA: Document Visual Q&A. Document Visual Question Answering benchmark testing the ability to extract and reason about information from document images including forms, reports, and scanned text. DeepSeek-V4-Flash scored 92.8% on this benchmark.
Terminal-Bench
82.7%
Terminal-Bench: Terminal/CLI Tasks. Tests the ability to perform command-line operations, write shell scripts, and navigate terminal environments. Measures practical system administration and development workflow skills. DeepSeek-V4-Flash scored 82.7% on this benchmark.
ARC-AGI
7%
ARC-AGI: Abstraction & Reasoning. Abstraction and Reasoning Corpus for AGI - tests fluid intelligence through novel pattern recognition puzzles. Each task requires discovering the underlying rule from examples, measuring general reasoning ability rather than memorization. DeepSeek-V4-Flash scored 7% on this benchmark.

About DeepSeek-V4-Flash

Learn about DeepSeek-V4-Flash's capabilities, features, and how it can help you achieve better results.

DeepSeek-V4-Flash is an open-weight Mixture of Experts (MoE) language model designed for software engineering and multi-step reasoning tasks. It contains 284 billion total parameters with 13 billion active parameters per token during inference. The model utilizes a native 1,048,576 token context window, allowing developers to process entire code repositories or long technical documents in a single request.

The July 31, 2026 update maintains the architecture of the initial preview model while applying post-training focused on agentic behavior. This re-post-training improved its score on Terminal-Bench 2.1 from 61.8 to 82.7 and increased its SWE-bench result to 54.4%. The model supports variable reasoning effort levels and native responses API integration for streaming intermediate commentary alongside final outputs.

Developers can access the model through DeepSeek's OpenAI-compatible API or run open weights locally on hardware configurations with 128GB to 192GB of unified memory. The standard API pricing is $0.14 per 1 million input tokens and $0.28 per 1 million output tokens, with cached input hits dropping to $0.03 per million tokens.

DeepSeek-V4-Flash

Use Cases

Discover the different ways you can use DeepSeek-V4-Flash to achieve great results.

Autonomous Agentic Coding

Running multi-step terminal and file modifications using Codex CLI or Cline harnesses to fix GitHub issues and automate test creation.

Local AI Deployment

Hosting the full model on workstations equipped with 128GB unified memory or four GPUs using 4-bit quantizations.

Large-Scale Repository Refactoring

Ingesting up to 1 million tokens of codebase context to map dependencies and execute architectural updates across multiple modules.

Interactive 3D and Front-End Web Generation

Building single-file web applications, Three.js 3D environments, and complex SVG diagrams directly from prompt specifications.

High-Volume Data Processing

Processing extensive text logs and structured documents using context caching to lower API costs to $0.03 per million tokens.

Multi-Step Terminal Automation

Executing shell commands and script pipelines where the model uses internal reasoning to handle unexpected command errors.

Strengths

Limitations

High SWE-Bench Performance: Scores 54.4% on SWE-bench Verified, outperforming many larger closed models in resolving real GitHub issues.
Variable Pricing Surcharges: API token costs double during peak usage hours, which increases expenses for real-time applications.
API Cost Efficiency: Priced at $0.14 per 1M input tokens and $0.28 per 1M output tokens, offering a lower cost per task than competing models.
High Memory Overhead for Local Execution: Local hosting requires at least 128GB of VRAM or system memory for usable quantization levels.
1M Token Context Window: Ingests up to 1,048,576 tokens in a single request, enabling codebase-wide dependency analysis and refactoring.
Overthinking on Simple Tasks: Deep reasoning effort can cause long output delays on short queries if reasoning parameters are not capped.
Local Hardware Hostability: Uses a 13B active parameter MoE structure that allows execution on 128GB unified memory setups with 4-bit quants.
Follow-Up Iteration Inconsistency: Quality can vary across long chat threads compared to initial single-shot prompt generations.

API Quick Start

deepseek/deepseek-v4-flash-0731

View Documentation
deepseek SDK
import OpenAI from 'openai';

const openai = new OpenAI({
  baseURL: 'https://api.deepseek.com',
  apiKey: process.env.DEEPSEEK_API_KEY,
});

async function main() {
  const completion = await openai.chat.completions.create({
    model: 'deepseek-v4-flash',
    messages: [{ role: 'user', content: 'Write a TypeScript function to balance a binary search tree.' }],
    temperature: 1.0,
    top_p: 0.95,
  });

  console.log(completion.choices[0].message.content);
}

main();

Install the SDK and start making API calls in minutes.

Community Feedback

See what the community thinks about DeepSeek-V4-Flash

DeepSeek V4 Flash 0731 is the undisputed price-performance leader: ~158,000 requests/month within the $60 limit, intelligence 49.9 and agentic score 45.7.
u/WegoW
reddit
We are making the updated DeepSeek V4-Flash 0731 free in Cline. This is the first flash model we've found performs at SOTA levels for autonomous coding.
@cline
twitter
The 'reasoning effort' setting is brilliant. I use low for tests and max for the actual logic. Saves so much time in daily pipelines.
BinaryBuilder
hackernews
DeepSeek V4 Flash 0731 should probably be the number one contender right now if you are into local AI.
Digital Spaceport
youtube
Models you can run locally now have the intelligence score of top frontier models from 5 months ago. This is absolutely nuts for open weights.
u/joorklee
reddit
DeepSeek V4 Flash is ~150x cheaper than proprietary options while delivering highly competitive UX and design task outputs.
@CommandCodeAI
twitter

Related Videos

Watch tutorials, reviews, and discussions about DeepSeek-V4-Flash

On the SWE-bench, a respected measure of software engineering capability, the model's score jumped from 7.3 to 54.4.

DeepSeek V4 Flash is a Mixture of Experts model with 284 billion total parameters, but only 13 billion active parameters.

This active parameter size makes the model realistically feasible to run on local hobbyist hardware with 128GB of unified memory.

The post-training updates focus primarily on agentic workflows and automated command execution.

For the price point of 14 cents per million input tokens, the performance ratio is unmatched right now.

Having at least 384,000 tokens of context size set is recommended, with a maximum context window of 1 million for this model.

If you are doing agentic work, you would want to adjust your top P to 0.95 instead of 1.0.

It is very thinky and does a lot of double-checking, triple-checking, and quadruple-checking during reasoning.

DeepSeek V4 Flash 0731 should probably be the number one contender right now if you are into local AI.

Running the quantized version locally requires a minimum of 128GB to 138GB of VRAM or system memory.

The gains aren't coming from scaling up the model size, they're coming from post-training focused on improving agentic behavior.

On Terminal Bench 2.1, it scores an 82.7, which is an enormous jump from its previous 61.8 preview score.

It offers near Luna-level intelligence at roughly 60% lower cost per task, making it one of the best performance per dollar models.

The native responses API integration allows developers to process tool output streams without custom formatting.

It handles multi-file repository edits with remarkable accuracy given its lightweight active footprint.

More than just prompts

Supercharge your workflow with AI Automation

Automatio combines the power of AI agents, web automation, and smart integrations to help you accomplish more in less time.

AI Agents
Web Automation
Smart Workflows

Pro Tips

Expert tips to help you get the most out of DeepSeek-V4-Flash and achieve better results.

Calibrate Top-P for Agent Workflows

Set top_p to 0.95 and temperature to 1.0 when deploying the model in coding harnesses to improve tool execution stability.

Set Minimum Context Size for Max Reasoning

Configure a context buffer of at least 384,000 tokens when using max reasoning effort to ensure room for deep chain-of-thought processing.

Utilize Context Caching

Structure repeated system prompts and codebase context to hit API cache layers, reducing input costs from $0.14 down to $0.03 per million tokens.

Use Off-Peak Processing Windows

Schedule bulk API batch jobs during non-peak hours, as API pricing doubles during high-traffic periods.

Testimonials

What Our Users Say

Join thousands of satisfied users who have transformed their workflow

Jonathan Kogan

Jonathan Kogan

Co-Founder/CEO, rpatools.io

Automatio is one of the most used for RPA Tools both internally and externally. It saves us countless hours of work and we realized this could do the same for other startups and so we choose Automatio for most of our automation needs.

Mohammed Ibrahim

Mohammed Ibrahim

CEO, qannas.pro

I have used many tools over the past 5 years, Automatio is the Jack of All trades.. !! it could be your scraping bot in the morning and then it becomes your VA by the noon and in the evening it does your automations.. its amazing!

Ben Bressington

Ben Bressington

CTO, AiChatSolutions

Automatio is fantastic and simple to use to extract data from any website. This allowed me to replace a developer and do tasks myself as they only take a few minutes to setup and forget about it. Automatio is a game changer!

Sarah Chen

Sarah Chen

Head of Growth, ScaleUp Labs

We've tried dozens of automation tools, but Automatio stands out for its flexibility and ease of use. Our team productivity increased by 40% within the first month of adoption.

David Park

David Park

Founder, DataDriven.io

The AI-powered features in Automatio are incredible. It understands context and adapts to changes in websites automatically. No more broken scrapers!

Emily Rodriguez

Emily Rodriguez

Marketing Director, GrowthMetrics

Automatio transformed our lead generation process. What used to take our team days now happens automatically in minutes. The ROI is incredible.

Jonathan Kogan

Jonathan Kogan

Co-Founder/CEO, rpatools.io

Automatio is one of the most used for RPA Tools both internally and externally. It saves us countless hours of work and we realized this could do the same for other startups and so we choose Automatio for most of our automation needs.

Mohammed Ibrahim

Mohammed Ibrahim

CEO, qannas.pro

I have used many tools over the past 5 years, Automatio is the Jack of All trades.. !! it could be your scraping bot in the morning and then it becomes your VA by the noon and in the evening it does your automations.. its amazing!

Ben Bressington

Ben Bressington

CTO, AiChatSolutions

Automatio is fantastic and simple to use to extract data from any website. This allowed me to replace a developer and do tasks myself as they only take a few minutes to setup and forget about it. Automatio is a game changer!

Sarah Chen

Sarah Chen

Head of Growth, ScaleUp Labs

We've tried dozens of automation tools, but Automatio stands out for its flexibility and ease of use. Our team productivity increased by 40% within the first month of adoption.

David Park

David Park

Founder, DataDriven.io

The AI-powered features in Automatio are incredible. It understands context and adapts to changes in websites automatically. No more broken scrapers!

Emily Rodriguez

Emily Rodriguez

Marketing Director, GrowthMetrics

Automatio transformed our lead generation process. What used to take our team days now happens automatically in minutes. The ROI is incredible.

Related AI Models

other

MiMo V2.5 Pro

Other

MiMo V2.5 Pro is Xiaomi's open-source 1.02T parameter MoE model featuring a 1M context window, native multimodality, and elite agentic coding performance.

1M context
$1.00/$3.00/1M
deepseek

DeepSeek-V3.2-Speciale

DeepSeek

DeepSeek-V3.2-Speciale is a reasoning-first LLM featuring gold-medal math performance, DeepSeek Sparse Attention, and a 131K context window. Rivaling GPT-5...

131K context
$0.28/$0.42/1M
minimax

MiniMax M2.5

minimax

MiniMax M2.5 is a SOTA MoE model featuring a 1M context window and elite agentic coding capabilities at disruptive pricing for autonomous agents.

1M context
$0.15/$1.20/1M
google

Gemini 3.6 Flash

Google

Gemini 3.6 Flash is Google's high-speed model featuring a 17% reduction in token consumption, $1.50/M input pricing, and advanced 3D visualization.

1M context
$1.50/$7.50/1M
zhipu

GLM-4.7

Zhipu (GLM)

GLM-4.7 by Zhipu AI is a flagship 358B MoE model featuring a 200K context window, elite 73.8% SWE-bench performance, and native Deep Thinking for agentic...

200K context
$0.60/$2.20/1M
moonshot

Kimi K2.7 Code

Moonshot

Kimi K2.7 Code is a 1T parameter MoE model from Moonshot AI. It features a 262k context window and 30% more efficient reasoning for software engineering.

262K context
$0.95/$4.00/1M
alibaba

Qwen3-Coder-Next

alibaba

Qwen3-Coder-Next is Alibaba Cloud's elite Apache 2.0 coding model, featuring an 80B MoE architecture and 256k context window for advanced local development.

262K context
$0.12/$0.75/1M
openai

GPT-4o mini

OpenAI

OpenAI's most cost-efficient small model, GPT-4o mini offers multimodal intelligence and high-speed performance at a significantly lower price point.

128K context
$0.15/$0.60/1M

Frequently Asked Questions

Find answers to common questions about DeepSeek-V4-Flash