DeepSeek V4 Flash 0731 was officially released yesterday, and the weights were already uploaded to Hugging Face. They've been gradually becoming available on provider platforms since today.
I had originally set up a Preview version locally, but today I downloaded the official version from Hugging Face and set it up. Because I had configured the Preview somewhat strangely, the performance wasn't great, and since I was using it alone, I just used it casually. But this time, to celebrate the official release, I completely overhauled everything and... found a well-functioning vllm recipe and asked Codex to set it up for me based on that. So I just sat back and relaxed.
[1/2] Running Full Serving Throughput & Concurrency Benchmark...
=========================================================
FULL SERVING BENCHMARK: DeepSeek-V4-Flash-0731
Target URL: #########################
Model ID : deepseek-v4-flash-0731
=========================================================
Warming up inference engine...
--- [DeepSeek-V4-Flash-0731] DECODE THROUGHPUT BY CONTENT ---
Prompt Type Tokens Sec Decode Speed
-----------------------------------------------
count300 600 7.20 83.4 tok/s
mult12 900 11.52 78.1 tok/s
json60 5 0.44 11.4 tok/s
bst 600 9.27 64.7 tok/s
story 253 7.92 31.9 tok/s
-----------------------------------------------
PEAK DECODE : 83.4 tok/s (count300)
MEAN DECODE : 53.9 tok/s
--- [DeepSeek-V4-Flash-0731] CONCURRENCY SCALING (c1..c12) ---
Streams OK Aggregate tok/s Per-Stream tok/s Wall Sec
--------------------------------------------------------------
1 1 70.9 70.9 5.64
2 2 115.4 57.7 6.93
4 4 158.6 41.4 10.09
6 6 211.3 36.0 11.36
8 8 169.4 31.4 18.88
10 10 184.6 28.1 21.67
12 12 201.5 26.7 23.82
--- [DeepSeek-V4-Flash-0731] PREFILL THROUGHPUT & TTFT AT DEPTH ---
Target Depth Prompt Tokens TTFT Sec Prefill Speed
----------------------------------------------------------
8000 5892 3.71 1588 tok/s
32000 23540 10.69 2201 tok/s
100000 73540 30.31 2426 tok/s
==================== SUMMARY ====================
Decode Peak : 83.4 tok/s | Mean: 53.9 tok/s
Concurrency : c1=71t/s | c2=115t/s | c4=159t/s | c6=211t/s | c8=169t/s | c10=185t/s | c12=201t/s
Prefill : 8K=1588t/s | 32K=2201t/s | 100K=2426t/s
=================================================
[2/2] Running 4-Subagent Parallel ReAct Loop Benchmark...
=========================================================
REAL MULTI-TURN AGENTIC LOOP BENCHMARK
Target LLM : ##########################
Model ID : deepseek-v4-flash-0731
Concurrency: 4 Sub-agents running parallel ReAct loops
=========================================================
===========================================================================
Agent ID | Role | Turns | Gen Tok | Time (s) | Speed
---------------------------------------------------------------------------
Subagent-1-Frontend | Frontend Dev | 6 | 2867 | 86.46 | 33.2 tok/s [INC]
Subagent-2-Backend | Backend Dev | 6 | 3821 | 143.69 | 26.6 tok/s [INC]
Subagent-3-Database | DB Specialist | 6 | 3993 | 110.44 | 36.2 tok/s [INC]
Subagent-4-Tester | QA Engineer | 6 | 4800 | 129.34 | 37.1 tok/s [INC]
---------------------------------------------------------------------------
Total Wall Clock Time : 143.69 seconds
Total Code Generated : 15,481 tokens
Aggregate Cluster Speed: 107.7 tok/s
=========================================================
✅ Benchmarking complete!
Prefill speeds are similar, but the decoding speed seems to have increased significantly.
At half the price of 5.6 luna, it seems like this model will be able to handle anything I would have asked Luna to do. ㅋㅋ