DeepSeek AI Releases DeepSeek-V4: Compressed Sparse Attention and Heavily Compressed Attention Enable One-Million-Token Contexts

TL;DR AI
2 min readKey summary
DeepSeek-AI previewed DeepSeek-V4-Pro and DeepSeek-V4-Flash, both built for one-million-token contexts.
The models use new efficiency techniques, including compressed/sparse attention, manifold-constrained hyper-connections, Muon optimization, and FP4 quantization-aware training.
DeepSeek says the approach reduces compute and memory enough to make million-token inference more practical.
Public checkpoints are available, signaling a more reproducible path toward extremely long-context LLMs.
