Family DeepSeek · 284B parameters

DeepSeek V4 Flash Coder 284B-A13B (MoEspresso V2)

Community requantization (MoEspresso V2) of DeepSeek V4 Flash Coder: MoE 284B / 13B active, 1M ctx, ~165 GB VRAM Q4. Code-oriented, MIT.

🇨🇳 steadfastgaze·License MIT·Context 1024k tokens·Output 2026-08-11← Catalog

01What it can do

Strengths
  • Massive native context up to ~1M tokens
  • Efficient MoE: ~13B active out of 284B, reasonable throughput for its size
  • Specialized in code and coding agents
  • Permissive MIT license
Limitations to know
  • —~165 GB of VRAM in Q4: requires multiple GPUs or a lot of RAM
  • —Community repack, not an official DeepSeek build
  • —No Ollama tag — install via HuggingFace
Architecture
Sparse MoE · 284B total parameters, ~13B active per token · 1,048,576-token (1M) context window · quantized “MoEspresso V2” repack (~56.8 GB)
Training
Community requant/repack (steadfastgaze) of the code-oriented DeepSeek V4 Flash 0731. Official base DeepSeek; MoEspresso V2 format optimized for memory footprint. Training details not published.
Ideal for
CodeAgents / code1M long context

04Install

Install Ollama for your OS. Check the model and its quantization before downloading. Start with 4096 tokens of context, then check placement with ollama ps. A command below is not proof that a test was run on your machine.

$# HuggingFace : steadfastgaze/DeepSeek-V4-Flash-0731-Coder-56.8GB-MoEspressoV2
⚠
First download: between 2 and 40 GB depending on the selected quantization. Plan for sufficient disk space; a stable connection is recommended. Subsequent launches are instant.

02Required memory

Approximate GPU VRAM required to run this model, including 4k tokens of context overhead. For a longer context, add ~1 GB per 8k-token increment.

Q4_K_M
The lightest, ~5% loss
165 GB
Q5_K_M
Good quality/size compromise
202 GB
Q8_0
Nearly indistinguishable from FP16
304 GB
FP16
Full precision — server use
568 GB
Fallback CPU · If you don't have a GPU, allow 369 GB of RAM minimum to run this model at reduced speed.

What hardware do you need for DeepSeek V4 Flash Coder 284B-A13B (MoEspresso V2)?

To run DeepSeek V4 Flash Coder 284B-A13B (MoEspresso V2) locally with Q4 quantization, you need about 165 GB of VRAM. An option to compare: BOSGAME M5 128GB / 2TB (Ryzen AI Max+ 395) — this model exceeds this mini-PC's GPU capacity: choose a smaller model or suitable infrastructure.

Current offer: BOSGAME M5 128GB / 2TB (Ryzen AI Max+ 395)
AmazonSee price →

Why this choice? Our complete guide on BOSGAME M5 128GB / 2TB (Ryzen AI Max+ 395) →

Affiliate links — commission possible at no extra cost to you; independent recommendation. As an Amazon Associate, BestLLMfor earns from qualifying purchases.

03Expected speed

Tokens generated per second in Q4_K_M, 4k context. Beyond 20 t/s, reading is comfortable. Below 10 t/s, that's just for testing.

Entry-level
~18t/s
GTX 1650, RX 6600, MBA M2 8GB
Mid-range
~28t/s
RTX 4060, 4070, MBP M3 Pro
High-end
~45t/s
RTX 4090, M4 Max, Radeon 7900