Two tabs open: one for the vision-tagged endpoint your product team wants for tomorrow’s image-ingest feature, one for DeepSeek V4 Flash’s plain text pricing. The launch plan assumes you already know the multimodal surcharge. You don’t, and there’s a reason for that.
The row that isn’t there
As of this writing, deepseek-v4-flash-vision has no entry among the 3514 LLM price rows the site tracks daily. DeepSeek V4 Flash lists real input and output rates for the 284B-total, 13B-active text model. Nothing separate exists for a vision-tagged SKU. That’s not a data gap on our end, it’s the state of the market: no provider has shipped a distinct published rate for image input on this model yet.
Treat the absence as the finding, not an inconvenience. If you’re pricing a vision surcharge you saw quoted somewhere outside a live table, that number isn’t backed by anything you can audit. Check /models before you build a budget line around it. The catalog updates daily, and if a vision row lands, it’ll show up there with real input and output rates, not as a rumor in a launch doc.
Why “vision surcharge” is the wrong question anyway
DeepSeek V4 Flash is a mixture-of-experts model: 284B total parameters, 13B active per token. Active params pay the compute bill regardless of what produced the token. A patch from an image and a subword from text both route through the same expert layers at the same per-token cost. There’s no architectural reason a vision-capable variant should charge a different rate per token than the text model, because the router doesn’t know or care where the token came from.
What changes is volume. A single moderate-resolution image can tokenize into a few thousand tokens depending on the vision encoder’s patch size, where the equivalent text description might run a few hundred. The bill goes up because you’re feeding more tokens through the same pipe, not because each token got more expensive. The per-token rate is set by the provider and doesn’t move with modality; what moves is how many tokens one request generates, and images generate a lot more of them than the text they’d replace.
You can see the same pattern in provider behavior even without a V4 Flash vision row to check. Vision-capable open models that do have live pricing, like Llama 3.2’s vision variant, show provider spreads over 20x between cheapest and priciest listing, in line with the spreads you’ll find on plain text models like Qwen3 235B A22B or Llama 3.1 405B. Modality doesn’t insulate you from a bad provider pick. If a vision-tagged V4 Flash endpoint does show up, expect the same spread behavior, and shop it the same way you’d shop the text tier.
The VRAM math barely moves
If you’re self-hosting instead of waiting on an API row, the numbers are close to what you’d plan for the text-only model. At FP8, 284B params times roughly 1.1 GB per billion lands near 310 GB, plus 25% KV headroom puts you around 390 GB. That clears a single card and needs multi-GPU: three H200 cards at 141 GB each cover it with room, or five H100 SXM cards at 80 GB each. At INT4, the footprint drops to roughly 155 GB before KV, around 195 GB after, which fits on two H200s.
A vision encoder bolted onto that backbone adds maybe a few hundred million to a couple billion parameters, depending on architecture. That’s noise against a 284B MoE. The weight-storage math you already ran for the text model still holds.
Your KV cache assumption is another matter. That 25% headroom figure is calibrated against normal text session lengths. Feed the same server concurrent requests carrying two or three images each, and a single request’s context length can jump by several thousand tokens before the model writes a word of output. Eight concurrent sessions, each with three images tokenizing at around 1,500 tokens apiece, adds close to 36,000 tokens of KV cache on top of whatever text context you budgeted for. That’s the failure mode: the box that ran fine in text-only load testing OOMs the first time someone in a real demo attaches a folder of screenshots, and it happens at inference time, not in the capacity plan.
What this can’t tell you is DeepSeek’s actual vision encoder size or its real image-to-token ratio. Neither is published, and estimating them from analogous open vision models is a guess dressed up as math. Test your actual image inputs before you size hardware around them.
The decision rule
If your workload needs image input on V4 Flash today, check /models first: if a vision row exists, price it against the text tier’s rate and route by actual token volume, not a modality tax that doesn’t exist at the architecture level. If no row exists, don’t build a launch plan around a rumored surcharge. Either wait, or self-host the text backbone on the H200 math above and run your own image-to-token measurement on a captioning preprocessor before you commit a KV budget you haven’t stress-tested with real concurrent image load.