
GLM 5.3: Z.ai's open-weights coding and agentic MoE model
GLM 5.3 is Z.ai's flagship Mixture-of-Experts model, built to code and ready for cyber defense. Sharing the GLM 5.2 base with substantially scaled post-training, it sustains a 1M-token context window with up to 128K output tokens and posts open-weights state-of-the-art results on Terminal-Bench 3.0 and agentic coding benchmarks.
- Developed byZ.ai
- Model familyzai-org/GLM-5.3
- ModalityLLM,
- Context Window1M
- Total Params743B
- PrecisionBF16 / FP8
- Deployment optionsShared, Dedicated, Self-hosted
Why choose GLM 5.3 on Modular?
Run leading open models with strong default performance and the ability to optimize down to the kernel — extracting more from every GPU.
Deploy efficiently across NVIDIA and AMD hardware to reduce GPU count, increase throughput, and avoid expensive closed-model licensing.
Integrate through an OpenAI-compatible endpoint, swap models freely, and scale across clouds or hardware without redesigning your application stack.
🔥 Trending models

MiniMax M3 is an open-weight, natively multimodal frontier model with ~428B total parameters and ~23B activated parameters. It combines frontier-level coding and agentic performance, an ultra-long context window of up to 1M tokens, and mixed-modality training across text, image, and video.

Inkling is a general-purpose multimodal model that accepts text, image and audio inputs and generates text outputs. The model is designed to be used by developers building AI-powered applications, including agentic and tool-use systems, coding assistants, chatbots, and retrieval-augmented generation systems, and is suitable for general-purpose conversational use, instruction-following, and other natural language and multimodal tasks.
Similar models

Gemma 4 26B A4B is a Mixture-of-Experts (MoE)model with 26B total parameters but only 4B activated per forward pass, meaning you get the quality of a much larger model at a fraction of the compute cost. It also supports a 256K context window and is designed to fit the memory footprint of high-end servers.
Get started with Modular
Schedule a demo of Modular and explore a custom end-to-end deployment built around your models, hardware, and performance goals.
Distributed, large-scale online inference endpoints
Highest-performance to maximize ROI and latency
Deploy in Modular cloud or your cloud
View all features with a custom demo

Book a demo
Talk with our sales lead Adam!
30min demo. Evaluate with your workloads. Ask us anything.
Book a demo for a personalized walkthrough of Modular in your environment. Learn how teams use it to simplify systems and tune performance at scale.
Custom 30 min walkthrough of our platform
Cover specific model or deployment needs
Flexible pricing to fit your specific needs

Book a demo
Talk with our sales lead Jay!
Run any open source model in 5 minutes, then benchmark it. Scale it to millions yourself (for free!).
Install Mojo and get up and running in minutes. A simple install, familiar tooling, and clear docs make it easy to start writing code immediately.



