DeepSeek V4.1 Flash
Overview
DeepSeek V4.1 Flash is a multimodal Mixture-of-Experts (MoE) model with 552B backbone parameters and support for contexts of up to one million tokens. The model introduces an asymmetric Causal Encoder-Decoder (CED) architecture designed to minimize computational overhead during long-context operations.
The model natively processes images and text, and generates text autoregressively. It is designed for greater capability, faster inference, higher throughput, and scaling to larger models.
Parameters
For a more detailed explanation on possible parameters, refer to OpenAI's Create chat completion documentation.
| Field | Type | Definition | Default Value |
|---|---|---|---|
model | string | Required. The model key to be used. Use deepseek-ai/deepseek-v4.1-flash. | deepseek-ai/deepseek-v4.1-flash |
messages | array | Required. Conversation history. Each object has a role (user or assistant) and content (string). | - |
stream | bool | Routes the request to the streaming path vs non-streaming path. | false |
max_completion_tokens | int | Output token cap. | - |
temperature | float | Randomness. Greedy sampling at 0 ; identical results are not guaranteed. | 1 |
top_p | float | Nucleus sampling cutoff. | 1 |
top_k | int | Limits sampling to the top K tokens. | - |
reasoning_effort | string | Constrains effort on reasoning for reasoning models. Supported values are none, low, high, max. Reducing reasoning effort can result in faster responses and fewer tokens used on reasoning in a response. Setting this value as none disables thinking. | high |
n | int | How many chat completion choices to generate for each input message. Currently, we only support 1. | 1 |
presence_penalty | float | Number between -2.0 and 2.0. Positive values penalize new tokens based on whether they appear in the text so far, increasing the model’s likelihood to talk about new topics. | 0 |
frequency_penalty | float | Number between -2.0 and 2.0. Positive values penalize new tokens based on their existing frequency in the text so far, decreasing the model’s likelihood to repeat the same line verbatim. | 0 |
logprobs | bool | Whether to return log probabilities of the output tokens or not. If true, returns the log probabilities of each output token returned in the content of message | false |
Example Requests
from openai import OpenAI
client = OpenAI(
base_url="https://api-cdn.thehive.ai/api/v3/",
api_key="<YOUR_SECRET_KEY>",
)
stream = client.chat.completions.create(
model="deepseek-ai/deepseek-v4.1-flash",
messages=[
{"role": "user",
"content": "Explain TCP and UDP in two sentences."
}
],
stream=True,
reasoning_effort="high",
extra_headers={"Accept": "text/event-stream"},
)
for chunk in stream:
if chunk.choices and chunk.choices[0].delta.content:
print(chunk.choices[0].delta.content, end="", flush=True)
if chunk.usage:
print(f"\n\n[usage] {chunk.usage}")curl https://api-cdn.thehive.ai/api/v3/chat/completions \
-H "Authorization: Bearer <YOUR_SECRET_KEY>" \
-H "Content-Type: application/json" \
-H "Accept: text/event-stream" \
-d '{
"model": "deepseek-ai/deepseek-v4.1-flash",
"messages": [{"role": "user", "content": "Hello!"}],
"stream": true
}'Note: Virginia (VA1) customers should use the URL: https://api-va1.thehive.ai/api/v3/chat/completions
Multi-Turn Conversations
There is no server-side conversation state — each call is stateless. To continue a conversation, resend the full message history on every request:
messages = [
{"role": "user", "content": "What's a good name for a pet turtle?"},
{"role": "assistant", "content": "How about Shelldon?"},
{"role": "user", "content": "Give me three more."},
]Response Format (SSE)
The response is a stream of Server-Sent Events, each an OpenAI-style chat.completion.chunk. Content arrives incrementally in choices[0].delta.content:
data: {"id":"...","object":"chat.completion.chunk","model":"<MODEL_KEY>","choices":[{"index":0,"delta":{"content":"Hello"},"finish_reason":null}]}
data: {"id":"...","object":"chat.completion.chunk","choices":[{"index":0,"delta":{"content":" there"},"finish_reason":null}]}The stream ends with a final usage chunk (empty choices, populated usage) followed by data: [DONE] :
data: {"id":"...","object":"chat.completion.chunk","choices":[],"usage":{"prompt_tokens":18,"completion_tokens":42,"total_tokens":60}}
data: [DONE]Usage is always included - you don't need to request stream_options.include_usage yourself.
Billing
For pricing details, refer to Hive's pricing page under Large Language Models > DeepSeek-V4.1-Flash.
Updated about 5 hours ago
