← Back to the pricing map

Cheapest first

Cheapest open-model inference API

The same open-weight models, hosted by different providers — sometimes tied, sometimes 4x apart.

Open-weight models like GPT-OSS, DeepSeek, and Llama aren't tied to a single provider — Groq, Together AI, Fireworks, and the model's own originating lab can all serve the same underlying weights, and each sets its own price. That means "how much does DeepSeek V4 Pro cost" doesn't have one answer; it depends which host you route through, and the gap isn't always small.

This page prices the same model, hosted by different providers, side by side. GPT-OSS 120B is identically priced on Groq and Fireworks — a genuine tie. DeepSeek V4 Pro is not: DeepSeek's own direct API prices it at roughly a quarter of what Together AI and Fireworks charge to serve the same weights. Both outcomes are shown here rather than only the tidier one.

When two hosts tie on price, the deciding factor becomes something this page doesn't measure directly: inference speed. Groq in particular is known for its LPU hardware and unusually fast token-generation speed. But when hosts don't tie — as with DeepSeek V4 Pro — the price gap itself is usually the bigger factor, and worth checking before assuming a third-party host is simply a convenience layer over the same cost.

Live pricing snapshot

Workload: 20M input tokens, 4M output tokens per month · Last verified Jul 16, 2026
RankQuality signalUnit priceBlended $/1Mvs cheapest
01GroqGPT-OSS 120B CheapestsourcebudgetNot verifiedIn $0.15 / Out $0.60 per 1M tokens$0.23/1M$5.4
02Fireworks AIGPT-OSS 120BsourcebudgetNot verifiedIn $0.15 / Out $0.60 per 1M tokens$0.23/1M$5.4tie (1.0×)
03DeepSeekDeepSeek V4 Prosourcebalanced80.6% SWE-bench VerifiedIn $0.435 / Out $0.87 per 1M tokens$0.51/1M$12.182.3× (+$6.78)
04GroqLlama 3.3 70B VersatilesourcebalancedNot verifiedIn $0.59 / Out $0.79 per 1M tokens$0.62/1M$14.962.8× (+$9.56)
05Together AIDeepSeek V4 ProsourcebalancedNot verifiedIn $1.74 / Out $3.48 per 1M tokens$2.03/1M$48.729.0× (+$43.32)
06Fireworks AIDeepSeek V4 ProsourcebalancedNot verifiedIn $1.74 / Out $3.48 per 1M tokens$2.03/1M$48.729.0× (+$43.32)

Prices are read live from the AICostCompass catalogue at page load, not hardcoded on this page, and reflect the workload above. Adjust the workload in the Compare tool to price your own usage instead.

Quality signal links to its source and is whatever third-party benchmark is actually published for that model (SWE-bench Verified, GPQA Diamond, an Elo rating, etc.) — different metrics aren't directly comparable across rows, and "Not verified" means we found no credible published score rather than that the model scores zero.

Ties and gaps both happen — check which one you're looking at

GPT-OSS 120B ties exactly between Groq and Fireworks, so speed and reliability decide there. DeepSeek V4 Pro does not tie — its originating-lab direct API prices roughly 4x below the third-party hosts serving the same weights in this table, likely reflecting a hosting/reseller margin rather than any difference in the model itself. Don't assume "same model" means "same price": check the live table for which case you're actually in.

Rate limits, SLA & data residency

Cheapest-per-token is moot if you're throttled. Cost table above, capacity/policy context below.

DeepSeek

Rate limits: No public requests-per-minute tier table — DeepSeek instead publishes a fixed concurrent-request cap per model per account, with capacity expansion available on request at no extra cost.

SLA: Not verified: no published uptime SLA was found in official DeepSeek documentation or policy pages as of the last check.

Data residency: DeepSeek's privacy policy states personal data is collected, processed, and stored in the People's Republic of China. No published enterprise zero-data-retention program or region-choice option was found.

Source · checked 2026-07-18

Methodology

Each candidate's current per-1M-token input/output rate is applied to a 20M input / 4M output token monthly workload, read live from the catalogue, spanning the same open-weight models available from multiple hosts (including, where applicable, the model's own originating-lab API) so identical models can be compared host to host.

Frequently asked questions

Why does the same model appear more than once on this page?

GPT-OSS 120B and DeepSeek V4 Pro are open-weight models available from multiple hosts at independently set prices — showing every host as its own row lets you compare the same model host to host directly, instead of assuming one representative price.

Why is DeepSeek V4 Pro so much cheaper directly from DeepSeek?

Third-party hosts (Together AI, Fireworks) add their own infrastructure and margin on top of serving someone else's open-weight model. DeepSeek's own API has no reseller layer, which is the most likely explanation for the roughly 4x gap shown in the live table — the underlying model weights are the same either way.

If the price ties, does the host not matter?

When price does tie (as with GPT-OSS 120B here), inference speed, rate limits, and reliability vary by host even though price doesn't. Groq's dedicated inference hardware is widely cited for unusually fast token generation, which this page does not measure directly.

Are these the only open-model inference hosts?

No — this page covers the hosts currently tracked in the AICostCompass catalogue. Other providers also serve open-weight models; check the main Pricing page for the full tracked list.

Related comparisons

Related guides