Free LLM API Models

This page lists models with both prompt and completion price reported as zero, useful for experiments, prototypes, and early evaluation.

46Models listed
1M input + 500K outputCost example tokens
USD / 1MNormalized prices

Quick shortlist

Start with Ox Alpha.

This guide starts with zero-price models for experiments and prototypes; verify limits before production use.

Lead model New🔥Ox Alpha
Providerstealth
Sample cost$0
Context1.05M

The ranking is a discovery aid, not a final recommendation. Always compare the model against your workload and verify provider pricing before production use.

How to read this ranking

Models are included when both prompt and completion prices are reported as zero. Free model availability and limits can change, so verify before production use.

Model Ranking

Browse all models
ModelProviderPromptOutputSample costYour CostContextPopularityRelease
New🔥Ox Alphastealth$0$0$0$01.05M#4
🔥Nemotron 3 Ultra (free)NVIDIA$0$0$0$01M#7
🔥Laguna S 2.1 (free)Poolside$0$0$0$0262.14K#11
New🔥Nemotron 3.5 Lightning (free)NVIDIA$0$0$0$01M#17
NewDots3-Note Preview (free)Dots Studio$0$0$0$0512K
NewLFM2.5-2.6B (free)LiquidAI$0$0$0$065.54K
Ling 3.0 Tiny (free)inclusionAI$0$0$0$0262.14K
Inkling Small (free)Thinking Machines$0$0$0$0262.14K
Ling-3.0-flash (free)inclusionAI$0$0$0$0262.14K
Inkling (free)Thinking Machines$0$0$0$0262.14K
Hy3 (free)Tencent$0$0$0$0262.14K
Laguna XS 2.1 (free)Poolside$0$0$0$0262.14K
North Mini Code (free)Cohere$0$0$0$0256K
GLM 5.2 (free)Z.ai$0$0$0$0256K
Kimi K2.7 Code (free)MoonshotAI$0$0$0$0262.14K
Nex-N2-Pro (free)Nex AGI$0$0$0$0262.14K
Nemotron 3.5 Content Safety (free)NVIDIA$0$0$0$0128K
CoBuddy (free)Baidu Qianfan$0$0$0$0131.07K
Owl AlphaOpenRouter$0$0$0$01.05M
Nemotron 3 Nano Omni (free)NVIDIA$0$0$0$0256K
Laguna XS.2 (free)Poolside$0$0$0$0262.14K
Laguna M.1 (free)Poolside$0$0$0$0262.14K
DeepSeek V4 Flash (free)DeepSeek$0$0$0$01.05M
Kimi K2.6 (free)MoonshotAI$0$0$0$0262.14K
Gemma 4 26B A4B (free)Google$0$0$0$0262.14K
Gemma 4 31B (free)Google$0$0$0$0262.14K
Trinity Large Thinking (free)Arcee AI$0$0$0$0262.14K
Lyria 3 Pro PreviewGoogle$0$0$0$01.05M
Lyria 3 Clip PreviewGoogle$0$0$0$01.05M
Nemotron 3 Super (free)NVIDIA$0$0$0$0262.14K
MiniMax M2.5 (free)MiniMax$0$0$0$0204.8K
Free Models RouterOpenRouter$0$0$0$0200K
LFM2.5-1.2B-Thinking (free)LiquidAI$0$0$0$032.77K
LFM2.5-1.2B-Instruct (free)LiquidAI$0$0$0$032.77K
Nemotron 3 Nano 30B A3B (free)NVIDIA$0$0$0$0256K
Nemotron Nano 12B 2 VL (free)NVIDIA$0$0$0$0128K
Qwen3 Next 80B A3B Instruct (free)Qwen$0$0$0$0262.14K
Nemotron Nano 9B V2 (free)NVIDIA$0$0$0$0128K
gpt-oss-120b (free)OpenAI$0$0$0$0131.07K
gpt-oss-20b (free)OpenAI$0$0$0$0131.07K
GLM 4.5 Air (free)Z.ai$0$0$0$0131.07K
Qwen3 Coder 480B A35B (free)Qwen$0$0$0$01.05M
Uncensored (free)Venice$0$0$0$032.77K
Llama 3.3 70B Instruct (free)Meta$0$0$0$0131.07K
Llama 3.2 3B Instruct (free)Meta$0$0$0$0131.07K
Hermes 3 405B Instruct (free)Nous$0$0$0$0131.07K

Pricing FAQ

How is the sample workload cost calculated?

The sample workload uses 1,000,000 input tokens plus 500,000 output tokens, then applies each model's normalized USD price per 1 million tokens.

Why do input and output token prices matter separately?

Many applications are output-token heavy, while retrieval and classification workloads may be input-token heavy. Comparing both prices helps avoid picking a model that is cheap for the wrong workload shape.

Should I verify prices before production use?

Yes. AI Model Matrix normalizes public pricing metadata for comparison, but provider availability, limits, and prices can change. Always verify the final contract or provider dashboard before production use.

Related Guides

Cheapest LLM APIs

Sort models by estimated workload cost and normalized token prices.

Open guide

Largest Context Windows

Find models for long documents, retrieval, and codebase context.

Open guide

Coding Models

Compare code-oriented models by cost, context, and practical popularity signals.

Open guide

Free Models

Browse zero-price models for prototypes and evaluation.

Open guide

RAG Models

Start from large context windows and practical input-cost constraints.

Open guide

Chatbot Costs

Find budget-sensitive models for output-heavy assistant traffic.

Open guide

Cost Calculator

Enter your own input and output token volume before narrowing the shortlist.

Estimate cost