Llama API pricing, every Meta model per million tokens
On 25 September 2026, Meta models on this list cost from $0.027 per million input tokens (Llama 3.2 1B Instruct) to $1.25 (Muse Spark 1.3), and from $0.08 per million output tokens (Llama 3.1 8B Instruct) to $4.25 (Muse Spark 1.3).
| Model | Input | Output | Cached input | Batch in / out | Context | Your month | Listed |
|---|---|---|---|---|---|---|---|
| Muse Spark 1.3 Contributor | $0.10 | $0.20 | $0.002 | 1.04858M | $9.00 | 2026-09-02 | |
| Muse Spark 1.3 | $1.25 | $4.25 | $0.15 | 1.04858M | $138.75 | 2026-09-02 | |
| Muse Spark 1.2 Contributor | $0.10 | $0.20 | $0.002 | 1.04858M | $9.00 | 2026-08-21 | |
| Muse Glimmer 30B | $0.30 | $1.20 | $0.04 | 131K | $36.00 | 2026-08-09 | |
| Muse Spark 1.2 | $1.25 | $4.25 | $0.15 | 1.04858M | $138.75 | 2026-08-05 | |
| Muse Spark 1.1 | $1.25 | $4.25 | $0.15 | 1.04858M | $138.75 | 2026-07-16 | |
| Llama Guard 4 12B | $0.18 | $0.18 | 164K | $13.50 | 2025-04-30 | ||
| Llama 4 Maverick | $0.1875 | $0.6525 | 1.04858M | $21.04 | 2025-04-05 | ||
| Llama 4 Scout | $0.10 | $0.30 | 1.31072M | $10.50 | 2025-04-05 | ||
| Llama 3.3 70B Instruct | $0.10 | $0.32 | 131K | $10.80 | 2024-12-06 | ||
| Llama 3.2 1B Instruct | $0.027 | $0.201 | 60K | $4.63 | 2024-09-25 | ||
| Llama 3.2 3B Instruct | $0.05 | $0.33 | 131K | $7.95 | 2024-09-25 | ||
| Llama 3.1 70B Instruct | $0.40 | $0.40 | 131K | $30.00 | 2024-07-23 | ||
| Llama 3.1 8B Instruct | $0.05 | $0.08 | $0.025 | 131K | $4.20 | 2024-07-23 |
Compare Meta with every other provider
Common questions
How much does the Llama API cost?
On 25 September 2026, Meta models on this list cost from $0.027 per million input tokens (Llama 3.2 1B Instruct) to $1.25 (Muse Spark 1.3), and from $0.08 per million output tokens (Llama 3.1 8B Instruct) to $4.25 (Muse Spark 1.3).
What is the cheapest Llama model?
For a typical mix of three input tokens to every output token, the cheapest is Llama 3.1 8B Instruct at $0.05 per million input tokens and $0.08 per million output tokens.
What is the newest Llama model on the list?
Muse Spark 1.3 Contributor, listed on 2 September 2026, at $0.10 per million input tokens and $0.20 per million output tokens.
Which Llama model has the largest context window?
Llama 4 Scout, with 1.31072M tokens of context.
Choosing a model for a production agent
Price is one input. Quality on your own cases, latency and the cost of retries decide the rest. Tell me what you are building.