FEATUREDComputing Life · Share (鸭哥 research reports)· rssZH00:00 · 08·25
→High Fidelity Nearby, Lossy at a Distance: The Shared Intuition Behind Three Long-Context Approaches
YaRN, DeepSeek V4, and DeepSeek-OCR tackle long-context bottlenecks at the coordinate, information-pathway, and input-representation layers respectively, all converging on the same intuition: keep nearby tokens high-fidelity, compress distant ones. YaRN applies frequency-partitioned interpolation to RoPE, letting Llama 2 7B reach 128K context with ~384 A100 GPU hours. DeepSeek V4 uses full-attention within a 128-token window and heavy compression plus sparse selection beyond it, cutting V4-Pro's per-token FLOPs to 27% of V3.2 at 1M context. DeepSeek-OCR compresses full pages into dense visual tokens, hitting 97% text accuracy at 10x compression. The three were developed independently by different teams. The post also flags a new challenge—maintaining positional awareness after compression—and outlines three solutions: dual-track position scales, document-wise coordinate resets, and Kimi K3's removal of positional encoding entirely.
#Nous Research#EleutherAI#DeepSeek
why featured
Featured · importance 78 · hook + knowledge + resonance
editor take
YaRN, DeepSeek V4, and DeepSeek-OCR independently landed on the same trade-off: high fidelity nearby, lossy compression at a distance.
sharp
This piece is worth opening because it connects three long-context solutions from different layers of the stack, all driven by the same intuition. YaRN partitions RoPE frequencies—high frequencies untouched, low frequencies compressed—letting Llama 2 7B hit 128K context with ~384 A100 GPU hours. DeepSeek V4 uses full-attention within a 128-token window and heavy compression plus sparse selection beyond it, cutting V4-Pro's per-token FLOPs to 27% of V3.2 at 1M context. DeepSeek-OCR compresses full pages into dense visual tokens, hitting 97% text accuracy at 10x compression. Different teams, same trade-off: spend compute on what's close, compress what's far. The post also flags a new problem—maintaining positional awareness after compression—and outlines three approaches: dual-track position scales, document-wise coordinate resets, and Kimi K3's removal of positional encoding entirely. That last part is the real headache. Compression is straightforward; making sure the model still knows what came before what after compression is the hard part.
HKR breakdown
hook ✓knowledge ✓resonance ✓