Alexey Shurov.
Free tool

Llama API pricing, every Meta model per million tokens

On 25 September 2026, Meta models on this list cost from $0.027 per million input tokens (Llama 3.2 1B Instruct) to $1.25 (Muse Spark 1.3), and from $0.08 per million output tokens (Llama 3.1 8B Instruct) to $4.25 (Muse Spark 1.3).

Prices checked 25 September 2026 at 06.00 UTC . Source OpenRouter public model list . US dollars per million tokens
ModelInputOutputCached inputBatch in / outContextYour monthListed
Muse Spark 1.3 Contributor$0.10$0.20$0.0021.04858M$9.002026-09-02
Muse Spark 1.3$1.25$4.25$0.151.04858M$138.752026-09-02
Muse Spark 1.2 Contributor$0.10$0.20$0.0021.04858M$9.002026-08-21
Muse Glimmer 30B$0.30$1.20$0.04131K$36.002026-08-09
Muse Spark 1.2$1.25$4.25$0.151.04858M$138.752026-08-05
Muse Spark 1.1$1.25$4.25$0.151.04858M$138.752026-07-16
Llama Guard 4 12B$0.18$0.18164K$13.502025-04-30
Llama 4 Maverick$0.1875$0.65251.04858M$21.042025-04-05
Llama 4 Scout$0.10$0.301.31072M$10.502025-04-05
Llama 3.3 70B Instruct$0.10$0.32131K$10.802024-12-06
Llama 3.2 1B Instruct$0.027$0.20160K$4.632024-09-25
Llama 3.2 3B Instruct$0.05$0.33131K$7.952024-09-25
Llama 3.1 70B Instruct$0.40$0.40131K$30.002024-07-23
Llama 3.1 8B Instruct$0.05$0.08$0.025131K$4.202024-07-23

Compare Meta with every other provider

Common questions

How much does the Llama API cost?

On 25 September 2026, Meta models on this list cost from $0.027 per million input tokens (Llama 3.2 1B Instruct) to $1.25 (Muse Spark 1.3), and from $0.08 per million output tokens (Llama 3.1 8B Instruct) to $4.25 (Muse Spark 1.3).

What is the cheapest Llama model?

For a typical mix of three input tokens to every output token, the cheapest is Llama 3.1 8B Instruct at $0.05 per million input tokens and $0.08 per million output tokens.

What is the newest Llama model on the list?

Muse Spark 1.3 Contributor, listed on 2 September 2026, at $0.10 per million input tokens and $0.20 per million output tokens.

Which Llama model has the largest context window?

Llama 4 Scout, with 1.31072M tokens of context.

Choosing a model for a production agent

Price is one input. Quality on your own cases, latency and the cost of retries decide the rest. Tell me what you are building.

shurco.aiGuidesFree toolsSolutionsInsightsRSSllms.txt