GLM 5.3 Flash
Overview
GLM 5.3 Flash is the first natively multimodal model in Z.ai's GLM-5 series, featuring 320B total parameters. It introduces a hybrid architecture combining sparse and linear attention, sharply reducing long-context serving costs while preserving precise long-context capabilities. The model also adopts Manifold-Constrained Hyper-Connections (mHC) to further improve scaling efficiency.
Together with Z.ai's 30T-token multimodal pre-training corpus, these changes enable GLM-5.3-Flash to deliver strong intelligence at a low cost.
Parameters
For a more detailed explanation on possible parameters, refer to OpenAI's Create chat completion documentation.
| Field | Type | Definition | Default Value |
|---|---|---|---|
model | string | Required. The model key to be used. Use zai-org/glm-5.3-flash. | zai-org/glm-5.3-flash |
messages | array | Required. Conversation history. Each object has a role (user or assistant) and content (string). | - |
stream | bool | Routes the request to the streaming path vs non-streaming path. | false |
max_completion_tokens | int | Output token cap. | - |
temperature | float | Randomness. Greedy sampling at 0 ; identical results are not guaranteed. | 1 |
top_p | float | Nucleus sampling cutoff. | 1 |
top_k | int | Limits sampling to the top K tokens. | - |
reasoning_effort | string | Constrains effort on reasoning for reasoning models. Supported values are low, high, max. Reducing reasoning effort can result in faster responses and fewer tokens used on reasoning in a response. | max |
n | int | How many chat completion choices to generate for each input message. Currently, we only support 1. | 1 |
presence_penalty | float | Number between -2.0 and 2.0. Positive values penalize new tokens based on whether they appear in the text so far, increasing the model’s likelihood to talk about new topics. | 0 |
frequency_penalty | float | Number between -2.0 and 2.0. Positive values penalize new tokens based on their existing frequency in the text so far, decreasing the model’s likelihood to repeat the same line verbatim. | 0 |
logprobs | bool | Whether to return log probabilities of the output tokens or not. If true, returns the log probabilities of each output token returned in the content of message | false |
Example Requests
from openai import OpenAI
client = OpenAI(
base_url="https://api-cdn.thehive.ai/api/v3/",
api_key="<YOUR_SECRET_KEY>",
)
stream = client.chat.completions.create(
model="zai-org/glm-5.3-flash",
messages=[
{"role": "user",
"content": "Explain TCP and UDP in two sentences."
}
],
stream=True,
reasoning_effort="high",
extra_headers={"Accept": "text/event-stream"},
)
for chunk in stream:
if chunk.choices and chunk.choices[0].delta.content:
print(chunk.choices[0].delta.content, end="", flush=True)
if chunk.usage:
print(f"\n\n[usage] {chunk.usage}")curl https://api-cdn.thehive.ai/api/v3/chat/completions \
-H "Authorization: Bearer <YOUR_SECRET_KEY>" \
-H "Content-Type: application/json" \
-H "Accept: text/event-stream" \
-d '{
"model": "zai-org/glm-5.3-flash",
"messages": [{"role": "user", "content": "Hello!"}],
"stream": true
}'Note: Virginia (VA1) customers should use the URL: https://api-va1.thehive.ai/api/v3/chat/completions
Multi-Turn Conversations
There is no server-side conversation state — each call is stateless. To continue a conversation, resend the full message history on every request:
messages = [
{"role": "user", "content": "What's a good name for a pet turtle?"},
{"role": "assistant", "content": "How about Shelldon?"},
{"role": "user", "content": "Give me three more."},
]Response Format (SSE)
The response is a stream of Server-Sent Events, each an OpenAI-style chat.completion.chunk. Content arrives incrementally in choices[0].delta.content:
data: {"id":"...","object":"chat.completion.chunk","model":"<MODEL_KEY>","choices":[{"index":0,"delta":{"content":"Hello"},"finish_reason":null}]}
data: {"id":"...","object":"chat.completion.chunk","choices":[{"index":0,"delta":{"content":" there"},"finish_reason":null}]}The stream ends with a final usage chunk (empty choices, populated usage) followed by data: [DONE] :
data: {"id":"...","object":"chat.completion.chunk","choices":[],"usage":{"prompt_tokens":18,"completion_tokens":42,"total_tokens":60}}
data: [DONE]Usage is always included - you don't need to request stream_options.include_usage yourself.
Billing
For pricing details, refer to Hive's pricing page under Large Language Models > GLM 5.3 Flash.
Updated about 5 hours ago
