Skip to main content

Try DeepSeek V4.1 Flash in the Workbench

Run this model interactively, tune parameters, and compare outputs.
Model ID: deepseek-v4-1-flash DeepSeek V4.1 Flash is a natively multimodal Mixture-of-Experts model built for coding, reasoning, and long-horizon agent tasks. It is the cost-efficient tier of the V4.1 family, and DeepSeek reports it ahead of V4 Pro on quality, speed, cost, and total task completion time, with V4 Flash and V4 Pro both retired in its favor. The model uses DeepSeek’s Causal Encoder-Decoder architecture, which activates 8B parameters while reading input and 16B while generating output, and compresses the KV cache to roughly a quarter of V4 Flash’s footprint. It reads images and text through a DeepSeek-ViT vision encoder and supports a 1M-token context window, tool use, structured outputs, and a continuously adjustable reasoning-effort setting (1-100) that trades inference cost for accuracy. Pretraining covers 45T multimodal tokens, followed by SFT, RL, and on-policy distillation post-training. Weights are published under the MIT license.

Example request

Use the Workbench as a request builder: configure parameters for this model in the UI, then open the API tab to copy the exact cURL or Python call.

Fetch model details

The models endpoint returns the full model object, including its json_request_schema.

Request parameters

This model follows the standard OpenAI chat completions request body. See the chat completions reference for the full parameter list.