Z.aiFastFastActive

Z.ai: GLM 5.3 FlashX

GLM-5.3-FlashX is the high-speed variant of Z.ai's GLM-5.3-Flash, a native multimodal model delivering inference speeds of up to 200 tokens/s. Built on the same hybrid sparse and linear attention architecture...

Specification
Model ID
z-ai/glm-5.3-flashx
Modality
Text+image+video >text
Context
1.0M tokens
Input
$0.370/1M
Output
$1.25/1M
Updated
Sep 20, 2026
Scores
Capabilities
  • Supports tool calling
  • Supports vision inputs
  • Supports long-context workflows
  • Supports deeper reasoning tasks