DeepSeek V4.1 Flash

Overview

DeepSeek V4.1 Flash is a multimodal Mixture-of-Experts (MoE) model with 552B backbone parameters and support for contexts of up to one million tokens. The model introduces an asymmetric Causal Encoder-Decoder (CED) architecture designed to minimize computational overhead during long-context operations.

The model natively processes images and text, and generates text autoregressively. It is designed for greater capability, faster inference, higher throughput, and scaling to larger models.


Parameters

For a more detailed explanation on possible parameters, refer to OpenAI's Create chat completion documentation.

FieldTypeDefinitionDefault Value
modelstringRequired. The model key to be used. Use deepseek-ai/deepseek-v4.1-flash.deepseek-ai/deepseek-v4.1-flash
messagesarrayRequired. Conversation history. Each object has a role (user or assistant) and content (string).-
streamboolRoutes the request to the streaming path vs non-streaming path.false
max_completion_tokensintOutput token cap.-
temperaturefloatRandomness. Greedy sampling at 0 ; identical results are not guaranteed.1
top_pfloatNucleus sampling cutoff.1
top_kintLimits sampling to the top K tokens.-
reasoning_effortstringConstrains effort on reasoning for reasoning models. Supported values are none, low, high, max.

Reducing reasoning effort can result in faster responses and fewer tokens used on reasoning in a response.

Setting this value as none disables thinking.
high
nintHow many chat completion choices to generate for each input message. Currently, we only support 1.1
presence_penaltyfloatNumber between -2.0 and 2.0. Positive values penalize new tokens based on whether they appear in the text so far, increasing the model’s likelihood to talk about new topics.0
frequency_penaltyfloatNumber between -2.0 and 2.0. Positive values penalize new tokens based on their existing frequency in the text so far, decreasing the model’s likelihood to repeat the same line verbatim.0
logprobsboolWhether to return log probabilities of the output tokens or not. If true, returns the log probabilities of each output token returned in the content of messagefalse

Example Requests

from openai import OpenAI

client = OpenAI(
    base_url="https://api-cdn.thehive.ai/api/v3/",
    api_key="<YOUR_SECRET_KEY>",
)
stream = client.chat.completions.create(
    model="deepseek-ai/deepseek-v4.1-flash",
    messages=[
        {"role": "user", 
         "content": "Explain TCP and UDP in two sentences."
        }
    ],
    stream=True,
    reasoning_effort="high",
    extra_headers={"Accept": "text/event-stream"},
)
for chunk in stream:
    if chunk.choices and chunk.choices[0].delta.content:
        print(chunk.choices[0].delta.content, end="", flush=True)
    if chunk.usage:
        print(f"\n\n[usage] {chunk.usage}")
curl https://api-cdn.thehive.ai/api/v3/chat/completions \
  -H "Authorization: Bearer <YOUR_SECRET_KEY>" \
  -H "Content-Type: application/json" \
  -H "Accept: text/event-stream" \
  -d '{
    "model": "deepseek-ai/deepseek-v4.1-flash",
    "messages": [{"role": "user", "content": "Hello!"}],
    "stream": true
  }'

Note: Virginia (VA1) customers should use the URL: https://api-va1.thehive.ai/api/v3/chat/completions

Multi-Turn Conversations

There is no server-side conversation state — each call is stateless. To continue a conversation, resend the full message history on every request:

messages = [
    {"role": "user", "content": "What's a good name for a pet turtle?"},
    {"role": "assistant", "content": "How about Shelldon?"},
    {"role": "user", "content": "Give me three more."},
]

Response Format (SSE)

The response is a stream of Server-Sent Events, each an OpenAI-style chat.completion.chunk. Content arrives incrementally in choices[0].delta.content:

data: {"id":"...","object":"chat.completion.chunk","model":"<MODEL_KEY>","choices":[{"index":0,"delta":{"content":"Hello"},"finish_reason":null}]}

data: {"id":"...","object":"chat.completion.chunk","choices":[{"index":0,"delta":{"content":" there"},"finish_reason":null}]}

The stream ends with a final usage chunk (empty choices, populated usage) followed by data: [DONE] :

data: {"id":"...","object":"chat.completion.chunk","choices":[],"usage":{"prompt_tokens":18,"completion_tokens":42,"total_tokens":60}}

data: [DONE]

Usage is always included - you don't need to request stream_options.include_usage yourself.


Billing

For pricing details, refer to Hive's pricing page under Large Language Models > DeepSeek-V4.1-Flash.