That's a wrap on ModCon 2026! Read the highlights ➔

Editions that work for everyone.
Scale as you grow.

  • PER TOKEN OR GPU HOUR

    Modular Cloud

    Frontier model endpoints with forward-deployed Modular engineers optimizing your workloads. Start by testing shared endpoints for free, then move to dedicated endpoints.

    • Always-on compute with SOTA inference performance

    • Shared & Dedicated Endpoints

    • Usage metrics and observability

    • Lowest cost endpoints to maximize ROI for the most demanding workloads

    • Secure in our environment

    • Forward-deployed engineers tuning your deployment

    For teams that want speed, reliability, and infrastructure they can trust - without managing it yourself.

  • PAY PER MINUTE

    Bring Your Own Cloud

    Built on our production-hardened BYOC infrastructure. Your VPC, your cloud credits, your compliance policies - with Modular engineers and our control plane inside.

    • Everything in Dedicated Endpoint, plus:

    • Deployment in your cloud or on-premise

    • Data never leaves your VPC

    • Performance optimization of your specific pipelines and workloads

    • Custom APIs

    • Secure in your environment

    • Forward-deployed engineers tuning your deployment

    For enterprises needing compliance, control, GPU flexibility, and hands-on engineering.

  • CUSTOM ENGAGEMENT

    Enterprise

    If you want full data control and have your own compute on AWS, GCP, Azure, or Oracle, or want a hybrid approach with our cloud and yours. We work with Enterprises on advanced deployment solutions. Let’s talk.

    • SOTA inference performance on any GPU vendor

    • Run AI models and pipelines on any hardware we support.

    • Deploy MAX and Mojo yourself - container under 1GB

    • Custom kernels in Mojo for novel architectures

    • Community support through Discord and Github

    For enterprises who need full control, want to run on diverse GPUs, and create a long term parternship.

Hosted Model API Endpoints

Model

Input ($/1M)

Output ($/1M)

Cache Hit ($/1M)

DeepSeek V4

1.74

3.48

0.145

DeepSeek V4 Flash - Standard

0.14

0.28

0.028

Gemma 4 31B

0.25

0.65

0.08

Gemma 4 26B A4B

0.15

0.6

0.07

GLM 5

0.95

3.15

0.2

GLM 5.1

1.3

4.3

0.26

GLM 5.2

1.4

4.4

0.26

GPT OSS 120B

0.1

0.5

N/A

Kimi K2.5

0.6

3

0.12

Kimi K2.6

0.85

3.5

0.16

Llama Guard 4 12B

0.2

N/A

N/A

MiniMax M2.5

0.3

1.2

0.06

MiniMax M3

0.3

1.2

0.06

NVIDIA Nemotron 3 Super

0.3

0.75

0.06

NVIDIA Nemotron 3 Ultra

0.6

3.6

0.2

Qwen 3 235B A22B FP8

0.2

0.6

N/A

Qwen 3.5 9B

0.17

0.25

N/A

Qwen 3.6 Plus

0.5

3

0.1

Qwen 3.7-Max

1.25

3.75

0.13

Model

Input ($/1K @ 1MP)

Output ($/1K @ 1 MP)

FLUX.2-dev

10

10

FLUX.2-klein-9B

6

6

FLUX.2-klein-4B

1

1

Want custom pricing, or preferred token rates?

Talk to Sales

Compare deployment options

Self-Hosted

Our Cloud

Your Cloud

Support

Active community and fast responses in Discord, Discourse, Github

Dedicated support, engineering team, standard and custom SLAs/SLOs

Dedicated support, engineering team, standard and custom SLAs/SLOs

Models

Hundreds of models in our model repo, view top performers

Top performers available for dedicated endpoint, custom model deployment

Top performers available for dedicated endpoint, custom model deployment

AI Skills

Use our open AI skills to easily write models, or optimize code

Our engineers can help train your team & migrate your workloads

Our engineers can help train your team & migrate your workloads

Platform access

Deploy MAX and Mojo yourself anywhere you want. Build with open source

Access Modular Platform with a console for deploying, scaling and managing your AI endpoints.

Access Modular Platform with a console for deploying, scaling and managing your AI endpoints.

Scalability

Scale on your own with the MAX container

Auto-scaling, scale to zero, burst capacity

Auto-scaling, proven at Fortune 500 scale.

Deployment location

Self-deployed, anywhere

Our cloud

Your cloud or hybrid

Compute hardware

See officially accepted compatible hardware for each MAX version in our docs. Scaling restrictions apply.

Diverse GPUs providers in our cloud, optimized for speed and/or cost.

NVIDIA, AMD, Trainium, TPU, Qualcomm GPUs, Intel, AMD & ARM CPUs, ASICs - deployed in your cloud.

Custom kernels

Your engineers write custom kernels for your workloads.

Modular engineers tune kernels for your workloads

Modular engineers write custom kernels for your workloads

Forward Deployed Engineers

Available with support plan

Included

Included; working in your environment

Security & Compliance

SOC 2 Type 2 certified

SOC 2 Type 2 certified

SOC 2 Type 2 certified

Billing & Pricing

Free

Per token (shared) Per minute (dedicated)

Per minute deployed. Use your AWS/GCP/Azure credits and commits

License

Enterprise Contract

FAQ

  • Which models can I run on Modular?

    With Modular, you can run the latest open-source models, fine-tuned variants or your own custom ones. You can also self-host any model and run it on NVIDIA, AMD, Qualcomm, TPU, Trainium, ARM, Intel, or Apple Silicon. For managed deployments, Our Cloud offers shared and dedicated endpoints with forward-deployed engineers optimizing your workloads.

  • What are Shared & Dedicated Endpoints?

    Modular runs all the latest open models in a shared endpoint (shared GPUs, billed on a $ / token basis), or a dedicated (dedicated GPUs, billed on $ / per minute basis) on our cloud infrastructure. We can also deploy in your compute environment on a dedicated basis (Your Cloud). Feel free to reach out to us if you have questions.

  • Which GPUs are available on Modular?

    Modular supports AI accelerators across NVIDIA, AMD, Google, AWS, and Qualcomm – including NVIDIA B200, AMD MI355X, Google TPUs, AWS Trainium, and Qualcomm AI100 and AI200. One software stack runs across them all, delivering state-of-the-art performance without locking developers or infrastructure operators to a single hardware vendor.

  • Can I get started with Modular Self-Hosted easily?

    Yes. The Self-Hosted Community Edition is completely free and open source. Install via Docker, PIP, UV, PIXI, or Conda - the container is under 700MB and runs on any GPU we support. You can be serving models in minutes.

  • Is Modular hosted infrastructure secure?

    Yes. Modular is SOC 2 Type 2 certified and independently audited. Our Cloud and Your Cloud editions both include enterprise-grade security. With Your Cloud (BYOC), data never leaves your VPC.

  • How does Modular integrate with our existing infrastructure?

    Modular integrates seamlessly into your stack. Our Cloud endpoints are fully compatible with the OpenAI API standard - swap in with a single line change. For custom kernels, Mojo interoperates directly with C++, CUDA, and ROCm. Every paid tier includes forward-deployed engineers to help with migration and optimization.

  • What level of customer support do you offer?

    Support varies by edition. The free Self-Hosted option is backed by an active Discord and GitHub community. Our Cloud and Your Cloud editions include dedicated support via email, Slack, and video calls - plus forward-deployed engineers who tune your deployments directly.

    Do I pay for idle time on Modular's hosted endpoints?

    Pricing depends on your edition. Modular Cloud has different products and tiers available to you. The free and starter tier, with shared endpoints only are pay-as-you-go at a per-token-consumption rate. Dedicated deployments are charged per GPU hour, so you could theoretically pay for idle time. Your Cloud (BYOC) is also billed per minute of reserved GPU capacity for guaranteed low-latency availability. The Self-Hosted option is free forever with no usage fees.

    Do you offer volume discounts on compute?

    Yes. We offer committed-use and volume pricing for Modular Cloud and Your Cloud editions. Every paid tier also includes forward-deployed engineers who actively optimize your workloads — not just infrastructure, but hands-on engineering support.

    Can I host Modular in my own cloud or on-premises?

    Yes. You can use Mojo and MAX to self-host on your own infrastructure, within the terms of their respective licenses, across any hardware Modular supports. This gives you the flexibility to deploy in your own cloud or on-premises environment while keeping control of your infrastructure, data, and operations. For production deployments on your own hardware, Modular also offers BYOC, with our control plane and engineering support running directly in your environment.

Build the future of AI with Modular

  • Person with blonde hair using a laptop with an Apple logo.

    Sign up today

    Signup to our Cloud Platform today to get started easily.

    Sign Up
  • Magnifying glass emoji with black handle and round clear lens.

    Browse open models

    Browse our model catalog, or deploy your own custom model

    Browse models