GLM 5.3 Flash

Overview

GLM 5.3 Flash is the first natively multimodal model in Z.ai's GLM-5 series, featuring 320B total parameters. It introduces a hybrid architecture combining sparse and linear attention, sharply reducing long-context serving costs while preserving precise long-context capabilities. The model also adopts Manifold-Constrained Hyper-Connections (mHC) to further improve scaling efficiency.

Together with Z.ai's 30T-token multimodal pre-training corpus, these changes enable GLM-5.3-Flash to deliver strong intelligence at a low cost.


Parameters

For a more detailed explanation on possible parameters, refer to OpenAI's Create chat completion documentation.

FieldTypeDefinitionDefault Value
modelstringRequired. The model key to be used. Use zai-org/glm-5.3-flash.zai-org/glm-5.3-flash
messagesarrayRequired. Conversation history. Each object has a role (user or assistant) and content (string).-
streamboolRoutes the request to the streaming path vs non-streaming path.false
max_completion_tokensintOutput token cap.-
temperaturefloatRandomness. Greedy sampling at 0 ; identical results are not guaranteed.1
top_pfloatNucleus sampling cutoff.1
top_kintLimits sampling to the top K tokens.-
reasoning_effortstringConstrains effort on reasoning for reasoning models. Supported values are low, high, max.

Reducing reasoning effort can result in faster responses and fewer tokens used on reasoning in a response.
max
nintHow many chat completion choices to generate for each input message. Currently, we only support 1.1
presence_penaltyfloatNumber between -2.0 and 2.0. Positive values penalize new tokens based on whether they appear in the text so far, increasing the model’s likelihood to talk about new topics.0
frequency_penaltyfloatNumber between -2.0 and 2.0. Positive values penalize new tokens based on their existing frequency in the text so far, decreasing the model’s likelihood to repeat the same line verbatim.0
logprobsboolWhether to return log probabilities of the output tokens or not. If true, returns the log probabilities of each output token returned in the content of messagefalse

Example Requests

from openai import OpenAI

client = OpenAI(
    base_url="https://api-cdn.thehive.ai/api/v3/",
    api_key="<YOUR_SECRET_KEY>",
)
stream = client.chat.completions.create(
    model="zai-org/glm-5.3-flash",
    messages=[
        {"role": "user", 
         "content": "Explain TCP and UDP in two sentences."
        }
    ],
    stream=True,
    reasoning_effort="high",
    extra_headers={"Accept": "text/event-stream"},
)
for chunk in stream:
    if chunk.choices and chunk.choices[0].delta.content:
        print(chunk.choices[0].delta.content, end="", flush=True)
    if chunk.usage:
        print(f"\n\n[usage] {chunk.usage}")
curl https://api-cdn.thehive.ai/api/v3/chat/completions \
  -H "Authorization: Bearer <YOUR_SECRET_KEY>" \
  -H "Content-Type: application/json" \
  -H "Accept: text/event-stream" \
  -d '{
    "model": "zai-org/glm-5.3-flash",
    "messages": [{"role": "user", "content": "Hello!"}],
    "stream": true
  }'

Note: Virginia (VA1) customers should use the URL: https://api-va1.thehive.ai/api/v3/chat/completions

Multi-Turn Conversations

There is no server-side conversation state — each call is stateless. To continue a conversation, resend the full message history on every request:

messages = [
    {"role": "user", "content": "What's a good name for a pet turtle?"},
    {"role": "assistant", "content": "How about Shelldon?"},
    {"role": "user", "content": "Give me three more."},
]

Response Format (SSE)

The response is a stream of Server-Sent Events, each an OpenAI-style chat.completion.chunk. Content arrives incrementally in choices[0].delta.content:

data: {"id":"...","object":"chat.completion.chunk","model":"<MODEL_KEY>","choices":[{"index":0,"delta":{"content":"Hello"},"finish_reason":null}]}

data: {"id":"...","object":"chat.completion.chunk","choices":[{"index":0,"delta":{"content":" there"},"finish_reason":null}]}

The stream ends with a final usage chunk (empty choices, populated usage) followed by data: [DONE] :

data: {"id":"...","object":"chat.completion.chunk","choices":[],"usage":{"prompt_tokens":18,"completion_tokens":42,"total_tokens":60}}

data: [DONE]

Usage is always included - you don't need to request stream_options.include_usage yourself.


Billing

For pricing details, refer to Hive's pricing page under Large Language Models > GLM 5.3 Flash.