
Editions that work for everyone. Scale as you grow.

Modular Cloud
Frontier model endpoints with forward-deployed Modular engineers optimizing your workloads. Start by testing shared endpoints for free, then move to dedicated endpoints.
Always-on compute with SOTA inference performance
Shared & Dedicated Endpoints
Usage metrics and observability
Lowest cost endpoints to maximize ROI for the most demanding workloads
Secure in our environment
Forward-deployed engineers tuning your deployment
For teams that want speed, reliability, and infrastructure they can trust - without managing it yourself.

Bring Your Own Cloud
Built on our production-hardened BYOC infrastructure. Your VPC, your cloud credits, your compliance policies - with Modular engineers and our control plane inside.
Everything in Dedicated Endpoint, plus:
Deployment in your cloud or on-premise
Data never leaves your VPC
Performance optimization of your specific pipelines and workloads
Custom APIs
Secure in your environment
Forward-deployed engineers tuning your deployment
For enterprises needing compliance, control, GPU flexibility, and hands-on engineering.

Enterprise
If you want full data control and have your own compute on AWS, GCP, Azure, or Oracle, or want a hybrid approach with our cloud and yours. We work with Enterprises on advanced deployment solutions. Let’s talk.
SOTA inference performance on any GPU vendor
Run AI models and pipelines on any hardware we support.
Deploy MAX and Mojo yourself - container under 1GB
Custom kernels in Mojo for novel architectures
Community support through Discord and Github
For enterprises who need full control, want to run on diverse GPUs, and create a long term parternship.
Hosted Model API Endpoints
Input ($/1M) | Output ($/1M) | Cache Hit ($/1M) | |
|---|---|---|---|
DeepSeek V4 | 1.74 | 3.48 | 0.145 |
DeepSeek V4 Flash - Standard | 0.14 | 0.28 | 0.028 |
Gemma 4 31B | 0.25 | 0.65 | 0.08 |
Gemma 4 26B A4B | 0.15 | 0.6 | 0.07 |
GLM 5 | 0.95 | 3.15 | 0.2 |
GLM 5.1 | 1.3 | 4.3 | 0.26 |
GLM 5.2 | 1.4 | 4.4 | 0.26 |
GPT OSS 120B | 0.1 | 0.5 | N/A |
Kimi K2.5 | 0.6 | 3 | 0.12 |
Kimi K2.6 | 0.85 | 3.5 | 0.16 |
Llama Guard 4 12B | 0.2 | N/A | N/A |
MiniMax M2.5 | 0.3 | 1.2 | 0.06 |
MiniMax M3 | 0.3 | 1.2 | 0.06 |
NVIDIA Nemotron 3 Super | 0.3 | 0.75 | 0.06 |
NVIDIA Nemotron 3 Ultra | 0.6 | 3.6 | 0.2 |
Qwen 3 235B A22B FP8 | 0.2 | 0.6 | N/A |
Qwen 3.5 9B | 0.17 | 0.25 | N/A |
Qwen 3.6 Plus | 0.5 | 3 | 0.1 |
Qwen 3.7-Max | 1.25 | 3.75 | 0.13 |
Input ($/1K @ 1MP) | Output ($/1K @ 1 MP) | |
|---|---|---|
FLUX.2-dev | 10 | 10 |
FLUX.2-klein-9B | 6 | 6 |
FLUX.2-klein-4B | 1 | 1 |
Want custom pricing, or preferred token rates?
Compare deployment options
Self-Hosted | Our Cloud | Your Cloud | |
|---|---|---|---|
Support | Active community and fast responses in Discord, Discourse, Github | Dedicated support, engineering team, standard and custom SLAs/SLOs | Dedicated support, engineering team, standard and custom SLAs/SLOs |
Models | Hundreds of models in our model repo, view top performers | Top performers available for dedicated endpoint, custom model deployment | Top performers available for dedicated endpoint, custom model deployment |
AI Skills | Use our open AI skills to easily write models, or optimize code | Our engineers can help train your team & migrate your workloads | Our engineers can help train your team & migrate your workloads |
Platform access | Deploy MAX and Mojo yourself anywhere you want. Build with open source | Access Modular Platform with a console for deploying, scaling and managing your AI endpoints. | Access Modular Platform with a console for deploying, scaling and managing your AI endpoints. |
Scalability | Scale on your own with the MAX container | Auto-scaling, scale to zero, burst capacity | Auto-scaling, proven at Fortune 500 scale. |
Deployment location | Self-deployed, anywhere | Our cloud | Your cloud or hybrid |
Compute hardware | See officially accepted compatible hardware for each MAX version in our docs. Scaling restrictions apply. | Diverse GPUs providers in our cloud, optimized for speed and/or cost. | NVIDIA, AMD, Trainium, TPU, Qualcomm GPUs, Intel, AMD & ARM CPUs, ASICs - deployed in your cloud. |
Custom kernels | Your engineers write custom kernels for your workloads. | Modular engineers tune kernels for your workloads | Modular engineers write custom kernels for your workloads |
Forward Deployed Engineers | Available with support plan | Included | Included; working in your environment |
Security & Compliance | SOC 2 Type 2 certified | SOC 2 Type 2 certified | SOC 2 Type 2 certified |
Billing & Pricing | Free | Per token (shared) Per minute (dedicated) | Per minute deployed. Use your AWS/GCP/Azure credits and commits |
Enterprise Contract |
FAQ
Which models can I run on Modular?
With Modular, you can run the latest open-source models, fine-tuned variants or your own custom ones. You can also self-host any model and run it on NVIDIA, AMD, Qualcomm, TPU, Trainium, ARM, Intel, or Apple Silicon. For managed deployments, Our Cloud offers shared and dedicated endpoints with forward-deployed engineers optimizing your workloads.
What are Shared & Dedicated Endpoints?
Modular runs all the latest open models in a shared endpoint (shared GPUs, billed on a $ / token basis), or a dedicated (dedicated GPUs, billed on $ / per minute basis) on our cloud infrastructure. We can also deploy in your compute environment on a dedicated basis (Your Cloud). Feel free to reach out to us if you have questions.
Which GPUs are available on Modular?
Modular supports AI accelerators across NVIDIA, AMD, Google, AWS, and Qualcomm – including NVIDIA B200, AMD MI355X, Google TPUs, AWS Trainium, and Qualcomm AI100 and AI200. One software stack runs across them all, delivering state-of-the-art performance without locking developers or infrastructure operators to a single hardware vendor.
Can I get started with Modular Self-Hosted easily?
Yes. The Self-Hosted Community Edition is completely free and open source. Install via Docker, PIP, UV, PIXI, or Conda - the container is under 700MB and runs on any GPU we support. You can be serving models in minutes.
Is Modular hosted infrastructure secure?
Yes. Modular is SOC 2 Type 2 certified and independently audited. Our Cloud and Your Cloud editions both include enterprise-grade security. With Your Cloud (BYOC), data never leaves your VPC.
How does Modular integrate with our existing infrastructure?
Modular integrates seamlessly into your stack. Our Cloud endpoints are fully compatible with the OpenAI API standard - swap in with a single line change. For custom kernels, Mojo interoperates directly with C++, CUDA, and ROCm. Every paid tier includes forward-deployed engineers to help with migration and optimization.
What level of customer support do you offer?
Support varies by edition. The free Self-Hosted option is backed by an active Discord and GitHub community. Our Cloud and Your Cloud editions include dedicated support via email, Slack, and video calls - plus forward-deployed engineers who tune your deployments directly.
Do I pay for idle time on Modular's hosted endpoints?
Pricing depends on your edition. Modular Cloud has different products and tiers available to you. The free and starter tier, with shared endpoints only are pay-as-you-go at a per-token-consumption rate. Dedicated deployments are charged per GPU hour, so you could theoretically pay for idle time. Your Cloud (BYOC) is also billed per minute of reserved GPU capacity for guaranteed low-latency availability. The Self-Hosted option is free forever with no usage fees.
Do you offer volume discounts on compute?
Yes. We offer committed-use and volume pricing for Modular Cloud and Your Cloud editions. Every paid tier also includes forward-deployed engineers who actively optimize your workloads — not just infrastructure, but hands-on engineering support.
Can I host Modular in my own cloud or on-premises?
Yes. You can use Mojo and MAX to self-host on your own infrastructure, within the terms of their respective licenses, across any hardware Modular supports. This gives you the flexibility to deploy in your own cloud or on-premises environment while keeping control of your infrastructure, data, and operations. For production deployments on your own hardware, Modular also offers BYOC, with our control plane and engineering support running directly in your environment.

Sign up today
Signup to our Cloud Platform today to get started easily.
Sign Up
Browse open models
Browse our model catalog, or deploy your own custom model
Browse models