
DeepSeek
AI · Computer Vision · Developer Tools
DeepSeek launches experimental vision model rivaling Anthropic's Opus 4.8
August 21, 2026
The token-capped pricing model makes screenshot-heavy AI agents dramatically cheaper to run than rivals' per-image fees.
- DeepSeek released DeepSeek-V4-Flash-Vision-Exp on August 21, 2026, an experimental multimodal model live on its API platform that adds image understanding to the text-only V4-Flash.
- The Hangzhou-based startup built the model for visual AI agents, letting it read screenshots, analyze charts and act on what it sees while preserving V4-Flash's reasoning and world-knowledge scores.
- Developers can send images as base64 data, external URLs up to 32 MiB, or through a new free Files API supporting files up to 64 MiB, with as many as 600 images per request.
- On DeepSeek's own 11-benchmark table the model beats Anthropic's Opus 4.8 on three tests (DeepSWE, Agents' Last Exam, ZeroBench) but trails by as much as 12 points on NL2Repo.
- Every image is capped at 384 tokens regardless of resolution and billed at V4-Flash's existing rate of about ¥1 (roughly $0.14) per million input tokens, keeping heavy screenshot-based agent tasks cheap.
- The launch comes as DeepSeek prepares a mainland China IPO and negotiates a new funding round targeting a $71 billion valuation, up from roughly $50 billion after its first outside round with Tencent and CATL.
- Benchmarking against a currently-supported, widely-used Opus 4.8 rather than an older model signals DeepSeek is racing to close the gap on agentic vision tasks just as it courts public investors.