Switch language한국어
Back to the list

DeepSeek AI Releases DeepSeek-V4: Compressed Sparse Attention and Heavily Compressed Attention Enable One-Million-Token Contexts

TL;DR AI

Key summary

2 min read
  1. DeepSeek-AI previewed DeepSeek-V4-Pro and DeepSeek-V4-Flash, both built for one-million-token contexts.

  2. The models use new efficiency techniques, including compressed/sparse attention, manifold-constrained hyper-connections, Muon optimization, and FP4 quantization-aware training.

  3. DeepSeek says the approach reduces compute and memory enough to make million-token inference more practical.

  4. Public checkpoints are available, signaling a more reproducible path toward extremely long-context LLMs.

Read the original