deepseek

DeepSeek-V4-Flash

DeepSeek-V4-Flash는 100만 token context를 지원하는 open-weight AI 모델로, SWE-bench에서 54.4%를 기록했으며 100만 token당 $0.14의 비용으로 agentic coding과 reasoning에 최적화되어 있습니다.

Open WeightsAgentic CodingMoE Architecture1M Context WindowLow Cost
deepseek logodeepseekDeepSeek V42026-07-31
컨텍스트
1M토큰
최대 출력
8K토큰
입력 가격
$0.14/ 1M
출력 가격
$0.28/ 1M
모달리티:TextImage
기능:비전도구스트리밍추론
벤치마크
GPQA
53.6%
GPQA: 대학원 수준 과학 Q&A. 생물학, 물리학, 화학 분야의 448개 객관식 문제로 구성된 엄격한 벤치마크. 박사 전문가도 65-74%의 정확도만 달성합니다. DeepSeek-V4-Flash이 이 벤치마크에서 53.6%점을 기록했습니다.
HLE
25.2%
HLE: 고급 전문 추론. 전문 분야에서 전문가 수준의 추론을 보여주는 모델의 능력을 테스트합니다. DeepSeek-V4-Flash이 이 벤치마크에서 25.2%점을 기록했습니다.
MMLU
88.7%
MMLU: 대규모 다중 작업 언어 이해. 57개 학술 과목에 걸쳐 16,000개의 객관식 문제로 구성된 종합 벤치마크. DeepSeek-V4-Flash이 이 벤치마크에서 88.7%점을 기록했습니다.
MMLU Pro
72.6%
MMLU Pro: MMLU 프로페셔널 에디션. 더 어려운 10지선다형 형식의 12,032개 문제를 포함하는 MMLU의 향상된 버전. DeepSeek-V4-Flash이 이 벤치마크에서 72.6%점을 기록했습니다.
SimpleQA
38.2%
SimpleQA: 사실 정확성 벤치마크. 직접적인 질문에 정확하고 사실적인 응답을 제공하는 모델의 능력을 테스트합니다. DeepSeek-V4-Flash이 이 벤치마크에서 38.2%점을 기록했습니다.
IFEval
88%
IFEval: 지시 따르기 평가. 모델이 특정 지시와 제약 조건을 얼마나 잘 따르는지 측정합니다. DeepSeek-V4-Flash이 이 벤치마크에서 88%점을 기록했습니다.
AIME 2025
93.1%
AIME 2025: 미국 초청 수학 시험. 명문 AIME 시험의 경쟁 수준 수학 문제. DeepSeek-V4-Flash이 이 벤치마크에서 93.1%점을 기록했습니다.
MATH
76.6%
MATH: 수학 문제 해결. 대수, 기하, 미적분 등의 분야를 테스트하는 종합 수학 벤치마크. DeepSeek-V4-Flash이 이 벤치마크에서 76.6%점을 기록했습니다.
GSM8k
96%
GSM8k: 초등학교 수학 8K. 다단계 추론이 필요한 8,500개의 초등학교 수준 수학 문장제. DeepSeek-V4-Flash이 이 벤치마크에서 96%점을 기록했습니다.
MGSM
90.5%
MGSM: 다국어 초등학교 수학. GSM8k 벤치마크를 10개 언어로 번역한 것. DeepSeek-V4-Flash이 이 벤치마크에서 90.5%점을 기록했습니다.
MathVista
63.8%
MathVista: 수학적 시각 추론. 차트, 그래프 등 시각적 요소가 포함된 수학 문제를 푸는 능력을 테스트합니다. DeepSeek-V4-Flash이 이 벤치마크에서 63.8%점을 기록했습니다.
SWE-Bench
54.4%
SWE-Bench: 소프트웨어 엔지니어링 벤치마크. AI 모델이 오픈소스 Python 프로젝트의 실제 GitHub 이슈를 해결하려고 시도합니다. DeepSeek-V4-Flash이 이 벤치마크에서 54.4%점을 기록했습니다.
HumanEval
90.2%
HumanEval: Python 프로그래밍 문제. 모델이 올바른 Python 함수 구현을 생성해야 하는 164개의 수작업 프로그래밍 문제. DeepSeek-V4-Flash이 이 벤치마크에서 90.2%점을 기록했습니다.
LiveCodeBench
33%
LiveCodeBench: 라이브 코딩 벤치마크. 지속적으로 업데이트되는 실제 프로그래밍 챌린지에서 코딩 능력을 테스트합니다. DeepSeek-V4-Flash이 이 벤치마크에서 33%점을 기록했습니다.
MMMU
69.1%
MMMU: 멀티모달 이해. 대학 수준 문제에서 비전-언어 모델을 테스트하는 대규모 다분야 멀티모달 이해 벤치마크. DeepSeek-V4-Flash이 이 벤치마크에서 69.1%점을 기록했습니다.
MMMU Pro
54%
MMMU Pro: MMMU 프로페셔널 에디션. 더 도전적인 문제와 더 엄격한 평가를 갖춘 MMMU의 향상된 버전. DeepSeek-V4-Flash이 이 벤치마크에서 54%점을 기록했습니다.
ChartQA
85.7%
ChartQA: 차트 질문 응답. 차트와 그래프에 제시된 정보를 이해하고 추론하는 능력을 테스트합니다. DeepSeek-V4-Flash이 이 벤치마크에서 85.7%점을 기록했습니다.
DocVQA
92.8%
DocVQA: 문서 시각 Q&A. 문서 이미지에서 정보를 추출하는 능력을 테스트하는 문서 시각 질문 응답 벤치마크. DeepSeek-V4-Flash이 이 벤치마크에서 92.8%점을 기록했습니다.
Terminal-Bench
82.7%
Terminal-Bench: 터미널/CLI 작업. 명령줄 작업을 수행하고 셸 스크립트를 작성하는 능력을 테스트합니다. DeepSeek-V4-Flash이 이 벤치마크에서 82.7%점을 기록했습니다.
ARC-AGI
7%
ARC-AGI: 추상화 및 추론. AGI를 위한 추상화 및 추론 코퍼스 - 새로운 패턴 인식 퍼즐로 유동 지능을 테스트합니다. DeepSeek-V4-Flash이 이 벤치마크에서 7%점을 기록했습니다.

DeepSeek-V4-Flash 소개

DeepSeek-V4-Flash의 기능, 특징 및 더 나은 결과를 얻는 방법에 대해 알아보세요.

DeepSeek-V4-Flash는 소프트웨어 엔지니어링 및 다단계 reasoning 작업을 위해 설계된 open-weight Mixture of Experts (MoE) 언어 모델입니다. 추론 시 token당 총 2,840억 개의 parameter 중 130억 개의 active parameter를 사용합니다. 이 모델은 1,048,576 token의 네이티브 context window를 활용하여 개발자가 단일 요청으로 전체 코드베이스나 긴 기술 문서를 처리할 수 있도록 지원합니다.

2026년 7월 31일 업데이트는 초기 프리뷰 모델의 아키텍처를 유지하면서 에이전트 동작에 초점을 맞춘 사후 학습을 적용했습니다. 이러한 추가 사후 학습을 통해 Terminal-Bench 2.1 점수가 61.8점에서 82.7점으로 향상되었고, SWE-bench 결과는 54.4%로 높아졌습니다. 이 모델은 가변적인 reasoning effort 수준을 지원하며, 최종 출력과 함께 중간 코멘터리를 스트리밍하기 위한 네이티브 응답 API 통합을 지원합니다.

개발자는 DeepSeek의 OpenAI 호환 API를 통해 모델에 액세스하거나, 128GB~192GB의 통합 메모리를 갖춘 하드웨어 구성에서 open weights를 로컬로 실행할 수 있습니다. 표준 API 가격은 입력 token 100만 개당 $0.14, 출력 token 100만 개당 $0.28이며, 캐시된 입력의 경우 100만 token당 $0.03으로 낮아집니다.

DeepSeek-V4-Flash

DeepSeek-V4-Flash 사용 사례

DeepSeek-V4-Flash을 사용하여 훌륭한 결과를 얻는 다양한 방법을 발견하세요.

자율형 Agentic Coding

Codex CLI 또는 Cline 하니스를 사용하여 다단계 터미널 및 파일 수정을 실행하고, GitHub 이슈를 해결하며 테스트 생성을 자동화합니다.

로컬 AI 배포

128GB 통합 메모리 또는 4개의 GPU가 탑재된 워크스테이션에서 4-bit quantization을 사용하여 전체 모델을 호스팅합니다.

대규모 저장소 리팩토링

최대 100만 token의 코드베이스 context를 입력받아 여러 모듈에 걸쳐 의존성을 매핑하고 아키텍처 업데이트를 실행합니다.

인터랙티브 3D 및 프론트엔드 웹 생성

prompt 사양을 바탕으로 단일 파일 웹 애플리케이션, Three.js 3D 환경, 복잡한 SVG 다이어그램을 직접 구축합니다.

대용량 데이터 처리

context caching을 사용하여 광범위한 텍스트 로그 및 구조화된 문서를 처리하고, API 비용을 100만 token당 $0.03으로 낮춥니다.

다단계 터미널 자동화

모델이 내부 reasoning을 사용하여 예기치 않은 커맨드 오류를 처리하는 셸 명령어 및 스크립트 파이프라인을 실행합니다.

강점

제한

높은 SWE-Bench 성능: SWE-bench Verified에서 54.4%를 기록하여, 실제 GitHub 이슈 해결에 있어 많은 대형 closed 모델들을 능가합니다.
변동성 있는 가격 할증: 피크 사용 시간대에는 API token 비용이 두 배로 증가하여 실시간 애플리케이션의 비용이 늘어납니다.
API 비용 효율성: 입력 token 100만 개당 $0.14, 출력 token 100만 개당 $0.28로 책정되어 경쟁 모델 대비 작업당 비용이 저렴합니다.
로컬 실행 시 높은 메모리 오버헤드: 로컬 호스팅 시 사용할 수 있는 수준의 quantization을 위해 최소 128GB의 VRAM 또는 시스템 메모리가 필요합니다.
1M Token Context Window: 단일 요청으로 최대 1,048,576 token을 수용하여 코드베이스 전체의 의존성 분석과 리팩토링이 가능합니다.
단순 작업에서의 과도한 사고(Overthinking): reasoning 파라미터를 제한하지 않으면 짧은 쿼리에 대해 깊은 reasoning 과정으로 인해 긴 출력 지연이 발생할 수 있습니다.
로컬 하드웨어 호스팅 가능: 13B active parameter MoE 구조를 사용하여 4-bit quantization을 적용한 128GB 통합 메모리 설정에서 실행할 수 있습니다.
후속 대화 반복의 일관성 부족: 초기 단일 샷 prompt 생성에 비해 긴 대화 스레드 전반에서 품질이 달라질 수 있습니다.

API 빠른 시작

deepseek/deepseek-v4-flash-0731

문서 보기
deepseek SDK
import OpenAI from 'openai';

const openai = new OpenAI({
  baseURL: 'https://api.deepseek.com',
  apiKey: process.env.DEEPSEEK_API_KEY,
});

async function main() {
  const completion = await openai.chat.completions.create({
    model: 'deepseek-v4-flash',
    messages: [{ role: 'user', content: 'Write a TypeScript function to balance a binary search tree.' }],
    temperature: 1.0,
    top_p: 0.95,
  });

  console.log(completion.choices[0].message.content);
}

main();

SDK를 설치하고 몇 분 안에 API 호출을 시작하세요.

DeepSeek-V4-Flash에 대한 사람들의 의견

커뮤니티가 DeepSeek-V4-Flash에 대해 어떻게 생각하는지 확인하세요

DeepSeek V4 Flash 0731은 명실상부한 가성비 1위입니다: $60 한도 내에서 월 약 158,000건의 요청, 지능 49.9점 및 에이전트 점수 45.7점을 기록했습니다.
u/WegoW
reddit
업데이트된 DeepSeek V4-Flash 0731을 Cline에서 무료로 제공합니다. 이는 자율 코딩 분야에서 SOTA 수준으로 성능을 발휘하는 최초의 flash 모델입니다.
@cline
twitter
'reasoning effort' 설정은 정말 환상적입니다. 테스트에는 low를 사용하고 실제 로직에는 max를 사용합니다. 일상적인 파이프라인에서 시간을 엄청나게 아껴줍니다.
BinaryBuilder
hackernews
로컬 AI에 관심이 있다면 DeepSeek V4 Flash 0731이 현재 최고의 경쟁자가 되어야 합니다.
Digital Spaceport
youtube
이제 로컬에서 실행할 수 있는 모델들이 5개월 전 최고 프론티어 모델들의 지능 점수를 보여주고 있습니다. open weights 진영에게는 정말 미친 발전입니다.
u/joorklee
reddit
DeepSeek V4 Flash는 독점 옵션보다 약 150배 저렴하면서도 매우 경쟁력 있는 UX 및 디자인 작업 결과물을 제공합니다.
@CommandCodeAI
twitter

DeepSeek-V4-Flash에 대한 동영상

DeepSeek-V4-Flash에 대한 튜토리얼, 리뷰 및 토론 시청

DeepSeek V4 Flash는 실제로 모든 모델 중에서 가장 저렴하며, 동일한 가격의 GPT 5.6 Luna는 오히려 2점 뒤처집니다.

모델 이름은 DeepSeek V4 Flash 0731이며, 요약하자면 3,000억 개의 parameters를 가지고 있고 그중 130억 개가 active 상태이며, 1백만 token의 context window를 지원하고 가중치는 이미 Hugging Face에 공개되어 있습니다.

실제로 실행해 보았을 때 DeepSeek는 여기서 꽤 부진했습니다. 23분이 걸렸죠. 3분 42초 만에 그 인상적인 결과를 만들어낸 5.6 Sol 같은 모델과 비교해 보면, [08:47] 부분은 좀 더 개선의 여지가 있다고 생각합니다.

Terminal Bench 2.1에서 새로운 Flash는 82.7점을 기록한 반면, Flash 미리보기 버전은 61.8점이었습니다... 새로운 flash는 이제 대형 모델의 미리보기 버전보다 거의 세 배 더 뛰어납니다.

이 소형 모델은 제 벤치마크에서 가장 까다로운 단일 질문인 3D 손목시계 문제에서 Fable 5, Opus 5 및 기타 모든 frontier model을 이겼습니다.

V4 패밀리의 모델로, 총 2,840억 개의 parameters와 약 130억 개의 active parameters, 그리고 1백만 token의 context window를 갖춘 mixture of experts 모델입니다. 빠르고 저렴하게 사용할 수 있도록 설계되었으며, V4 Pro는

로컬 AI에 관심이 많고 3090 같은 장비를 가지고 계시다면, DeepSeek V4 Flash 0731은 아마도 현재 가장 강력한 1순위 후보가 되어야 할 것입니다.

최소 384,000 token의 context size를 설정한 상태입니다. 최대

따라서 관심이 있으시다면 Q3 KXL을 선택하거나, 가능하면 IQ3 XXS로 낮추는 것을 추천합니다.

단순한 프롬프트 이상

워크플로를 강화하세요 AI 자동화

Automatio는 AI 에이전트, 웹 자동화 및 스마트 통합의 힘을 결합하여 더 짧은 시간에 더 많은 것을 달성할 수 있도록 도와줍니다.

AI 에이전트
웹 자동화
스마트 워크플로

DeepSeek-V4-Flash 프로 팁

DeepSeek-V4-Flash을 최대한 활용하기 위한 전문가 팁.

에이전트 워크플로우를 위한 Top-P 보정

툴 실행 안정성을 높이기 위해 코딩 하니스에 모델을 배포할 때 top_p를 0.95로, temperature를 1.0으로 설정하세요.

최대 reasoning을 위한 최소 context 크기 설정

deep chain-of-thought 처리를 위한 공간을 확보하려면 max reasoning effort를 사용할 때 최소 384,000 token의 context 버퍼를 구성하세요.

Context Caching 활용

반복되는 system prompt와 코드베이스 context를 구조화하여 API 캐시 레이어를 활용하면, 입력 비용을 100만 token당 $0.14에서 $0.03으로 줄일 수 있습니다.

피크 타임 외 처리 윈도우 사용

고트래픽 시간대에는 API 가격이 두 배로 오르므로, 대규모 API 배치 작업은 비피크 시간에 예약하세요.

후기

사용자 후기

워크플로를 혁신한 수천 명의 만족한 사용자와 함께하세요

Jonathan Kogan

Jonathan Kogan

Co-Founder/CEO, rpatools.io

Automatio is one of the most used for RPA Tools both internally and externally. It saves us countless hours of work and we realized this could do the same for other startups and so we choose Automatio for most of our automation needs.

Mohammed Ibrahim

Mohammed Ibrahim

CEO, qannas.pro

I have used many tools over the past 5 years, Automatio is the Jack of All trades.. !! it could be your scraping bot in the morning and then it becomes your VA by the noon and in the evening it does your automations.. its amazing!

Ben Bressington

Ben Bressington

CTO, AiChatSolutions

Automatio is fantastic and simple to use to extract data from any website. This allowed me to replace a developer and do tasks myself as they only take a few minutes to setup and forget about it. Automatio is a game changer!

Sarah Chen

Sarah Chen

Head of Growth, ScaleUp Labs

We've tried dozens of automation tools, but Automatio stands out for its flexibility and ease of use. Our team productivity increased by 40% within the first month of adoption.

David Park

David Park

Founder, DataDriven.io

The AI-powered features in Automatio are incredible. It understands context and adapts to changes in websites automatically. No more broken scrapers!

Emily Rodriguez

Emily Rodriguez

Marketing Director, GrowthMetrics

Automatio transformed our lead generation process. What used to take our team days now happens automatically in minutes. The ROI is incredible.

Jonathan Kogan

Jonathan Kogan

Co-Founder/CEO, rpatools.io

Automatio is one of the most used for RPA Tools both internally and externally. It saves us countless hours of work and we realized this could do the same for other startups and so we choose Automatio for most of our automation needs.

Mohammed Ibrahim

Mohammed Ibrahim

CEO, qannas.pro

I have used many tools over the past 5 years, Automatio is the Jack of All trades.. !! it could be your scraping bot in the morning and then it becomes your VA by the noon and in the evening it does your automations.. its amazing!

Ben Bressington

Ben Bressington

CTO, AiChatSolutions

Automatio is fantastic and simple to use to extract data from any website. This allowed me to replace a developer and do tasks myself as they only take a few minutes to setup and forget about it. Automatio is a game changer!

Sarah Chen

Sarah Chen

Head of Growth, ScaleUp Labs

We've tried dozens of automation tools, but Automatio stands out for its flexibility and ease of use. Our team productivity increased by 40% within the first month of adoption.

David Park

David Park

Founder, DataDriven.io

The AI-powered features in Automatio are incredible. It understands context and adapts to changes in websites automatically. No more broken scrapers!

Emily Rodriguez

Emily Rodriguez

Marketing Director, GrowthMetrics

Automatio transformed our lead generation process. What used to take our team days now happens automatically in minutes. The ROI is incredible.

관련 AI Models

other

MiMo V2.5 Pro

Other

MiMo V2.5 Pro is Xiaomi's open-source 1.02T parameter MoE model featuring a 1M context window, native multimodality, and elite agentic coding performance.

1M context
$1.00/$3.00/1M
deepseek

DeepSeek-V3.2-Speciale

DeepSeek

DeepSeek-V3.2-Speciale is a reasoning-first LLM featuring gold-medal math performance, DeepSeek Sparse Attention, and a 131K context window. Rivaling GPT-5...

131K context
$0.28/$0.42/1M
minimax

MiniMax M2.5

minimax

MiniMax M2.5 is a SOTA MoE model featuring a 1M context window and elite agentic coding capabilities at disruptive pricing for autonomous agents.

1M context
$0.15/$1.20/1M
google

Gemini 3.6 Flash

Google

Gemini 3.6 Flash is Google's high-speed model featuring a 17% reduction in token consumption, $1.50/M input pricing, and advanced 3D visualization.

1M context
$1.50/$7.50/1M
zhipu

GLM-4.7

Zhipu (GLM)

GLM-4.7 by Zhipu AI is a flagship 358B MoE model featuring a 200K context window, elite 73.8% SWE-bench performance, and native Deep Thinking for agentic...

200K context
$0.60/$2.20/1M
moonshot

Kimi K2.7 Code

Moonshot

Kimi K2.7 Code is a 1T parameter MoE model from Moonshot AI. It features a 262k context window and 30% more efficient reasoning for software engineering.

262K context
$0.95/$4.00/1M
alibaba

Qwen3-Coder-Next

alibaba

Qwen3-Coder-Next is Alibaba Cloud's elite Apache 2.0 coding model, featuring an 80B MoE architecture and 256k context window for advanced local development.

262K context
$0.12/$0.75/1M
openai

GPT-4o mini

OpenAI

OpenAI's most cost-efficient small model, GPT-4o mini offers multimodal intelligence and high-speed performance at a significantly lower price point.

128K context
$0.15/$0.60/1M

DeepSeek-V4-Flash에 대한 자주 묻는 질문

DeepSeek-V4-Flash에 대한 일반적인 질문에 대한 답변 찾기