FEATUREDAI HOT (Curated Pool)· aihot-apiZH11:35 · 08·31
→DeepSeek open-sources V4-Flash-Vision-Exp, its first vision model, with multimodal agent performance near Opus-4.8
DeepSeek released V4-Flash-Vision-Exp on Hugging Face under MIT License—the first V4 model that accepts image inputs. The repo includes a minimal PyTorch inference implementation covering the vision encoder, MoE, DFlash Attention, and other core modules. It handles JPEG, PNG, GIF, and WebP for tasks like image captioning, screenshot OCR, and chart reading. Text-only performance matches the stable V4-Flash; multimodal agent benchmarks show a big jump, nearing Opus-4.8. This is an experimental version—it hit the API on Aug 21 and now has open weights.
#Agent#DeepSeek#Hugging Face
why featured
Featured · importance 82 · hook + knowledge + resonance
editor take
DeepSeek open-sourced V4-Flash-Vision under MIT, with multimodal agent scores nearing Opus-4.8.
sharp
The reason to click: DeepSeek finally shipped a V4 model that can see. It's called V4-Flash-Vision-Exp, MIT licensed, with a full PyTorch inference implementation covering the vision encoder, MoE, and DFlash Attention. Text-only performance stays level with the stable V4-Flash, but multimodal agent benchmarks got a big bump—official word says it's nearing Opus-4.8.
I'd discount this a bit. It's marked experimental, hit the API on Aug 21, and only now got open weights. No specific benchmark numbers or comparison details were shared—just "nearing Opus-4.8," with no test set named. But the MIT license is the real hook here: you can use it commercially or fine-tune it without licensing headaches. For teams building multimodal agents, that's a solid option.
HKR breakdown
hook ✓knowledge ✓resonance ✓