Cheapest first
Cheapest LLM API for high volume
Budget-tier models ranked at 250M tokens a month, with committed-use and volume discounts applied where offered.
At 250M tokens a month, the economics shift away from what applies to a typical small integration. Every provider here offers some form of committed-use or volume discount once a workload crosses a threshold, and those discounts frequently reorder the ranking compared to undiscounted list price — a model that looks mid-pack at list price can move to cheapest once volume pricing kicks in, or vice versa if a competitor's volume tier is more aggressive.
This ranking applies committed-use, volume, and batch discounts wherever the provider currently offers them at this scale, since a genuinely high-volume buyer would negotiate or qualify for them rather than pay list price. The specific discount rates and minimum commitment terms are set independently by each provider and are not guaranteed to match what a given account can actually negotiate — treat this as a directional ranking, not a quote.
Budget and fast tiers dominate this list by design: a 250M-token-a-month workload built on frontier-tier pricing is rarely economical for any provider, so high-volume buyers overwhelmingly route through the cheapest capable tier rather than the most capable one.
Live pricing snapshot
Workload: 200M input tokens, 50M output tokens per month · Last verified Jul 14, 2026| Rank | Quality signal | Unit price | Blended $/1M | vs cheapest | ||||
|---|---|---|---|---|---|---|---|---|
| 01 | Mistral | Ministral 3B Cheapestsource | budget | Not verified | In $0.04 / Out $0.04 per 1M tokens | $0.04/1M | $10 | — |
| 02 | DeepSeek | DeepSeek V4 Flashsource | budget | 79.0% SWE-bench Verifiedunverified | In $0.14 / Out $0.28 per 1M tokens | $0.17/1M | $42 | 4.2× (+$32) |
| 03 | Anthropic | Claude Haiku 4.5source | fast | 73.3% SWE-bench Verified | In $1.00 / Out $5.00 per 1M tokens | $0.90/1M | $225 | 22.5× (+$215) |
| 04 | Gemini 3 Flashsource | budget | 90.4% GPQA Diamond | In $0.50 / Out $3.00 per 1M tokens | $1.00/1M | $250 | 25.0× (+$240) | |
| 05 | OpenAI | GPT-5.6 Lunasource | budget | 51 Artificial Analysis Intelligence Index | In $1.00 / Out $6.00 per 1M tokens | $2.00/1M | $500 | 50.0× (+$490) |
Prices are read live from the AICostCompass catalogue at page load, not hardcoded on this page, and reflect the workload above. Adjust the workload in the Compare tool to price your own usage instead.
Quality signal links to its source and is whatever third-party benchmark is actually published for that model (SWE-bench Verified, GPQA Diamond, an Elo rating, etc.) — different metrics aren't directly comparable across rows, and "Not verified" means we found no credible published score rather than that the model scores zero.
Why the ranking can shift at negotiation
The gap between the cheapest and most expensive option here is typically large enough that provider choice matters more than model choice within a tier at this scale — but confirm your actual negotiated or qualified discount rate with each provider directly before committing volume, since list-price volume tiers and negotiated enterprise rates can diverge substantially.
Rate limits, SLA & data residency
Cheapest-per-token is moot if you're throttled. Cost table above, capacity/policy context below.
OpenAI
Rate limits: Five spend-based usage tiers (Tier 1 at $5 spent up to Tier 5 at $1,000+ cumulative spend and 30+ days), enforced on requests and tokens per minute and per day. A new account starts at Tier 1 with meaningfully lower limits than an established Tier 5 or enterprise account.
SLA: No published uptime SLA on standard pay-as-you-go API access. Enterprise / high-commitment customers can negotiate priority processing and an SLA via sales.
Data residency: Official data residency program for at-rest processing/storage in specific regions (EU, UK, Japan, Canada, South Korea, Singapore, Australia, India, UAE), plus a Zero Data Retention option — both sales-gated for eligible enterprise customers.
Source · checked 2026-07-18Anthropic
Rate limits: Organization-level usage tiers (Start, Build, Scale, plus custom enterprise), tied to spend, with separate request- and token-per-minute ceilings. Self-service increase requests unlock once you hit roughly 50% of your current limit.
SLA: No uptime guarantee on the default Standard tier. A Priority Tier product targeted 99.5% uptime for committed-capacity contracts, though Anthropic's docs currently note it's no longer available for new purchases (existing contracts continue through term).
Data residency: Zero Data Retention available for eligible orgs. The direct API only supports "us" or "global" inference routing today — dedicated EU-only residency requires routing through AWS Bedrock or Google Vertex AI's EU endpoints instead.
Source · checked 2026-07-18Google (Gemini)
Rate limits: Project-based usage tiers (Free, Tier 1 once billing is linked, Tier 2 above $250 spend, Tier 3 above $1,000 spend, each after a 30-day wait), enforced simultaneously on requests/tokens per minute and requests per day.
SLA: A formal SLA exists for the Gemini API on Vertex AI, but a guarantee on processed-request availability requires Provisioned Throughput (a paid reserved-capacity product) — standard pay-as-you-go has a narrower model-availability SLA.
Data residency: Jurisdictional/regional endpoints keep processing within a chosen region (e.g. US or EU); eligible enterprise customers can get zero-data-retention-equivalent contract terms via a DPA amendment that disables caching and abuse-monitoring logging.
Source · checked 2026-07-18Mistral
Rate limits: A Free/Experiment tier (no billing card, capped monthly tokens) and a Scale (pay-as-you-go) tier whose limits increase automatically with cumulative billed spend, plus a negotiated Enterprise tier. Exact RPM/TPM figures are account-specific and shown only in Mistral's Admin Console.
SLA: Enterprise plans are described as including custom SLAs as part of a negotiated agreement; no specific public uptime percentage was found for the standard pay-as-you-go tier.
Data residency: EU-hosted by default (Mistral is an EU company), with an optional US API endpoint that hosts data in the US instead. Enterprise customers can restrict data transfers outside the EU, and self-hosted/dedicated-VPC deployment is available for full control.
Source · checked 2026-07-18DeepSeek
Rate limits: No public requests-per-minute tier table — DeepSeek instead publishes a fixed concurrent-request cap per model per account, with capacity expansion available on request at no extra cost.
SLA: Not verified: no published uptime SLA was found in official DeepSeek documentation or policy pages as of the last check.
Data residency: DeepSeek's privacy policy states personal data is collected, processed, and stored in the People's Republic of China. No published enterprise zero-data-retention program or region-choice option was found.
Source · checked 2026-07-18Methodology
Each candidate's current per-1M-token input/output rate is applied to a 200M input / 50M output token monthly workload, with committed-use, volume, and batch discounts applied where the provider currently publishes them, read live from the catalogue and ranked by resulting monthly cost.
Frequently asked questions
Do I automatically qualify for volume discounts at this scale?
Not automatically — most providers require an application, minimum spend commitment, or sales conversation to unlock committed-use or volume pricing. This ranking shows currently published discount terms; confirm your actual qualifying rate directly with each provider.
Why do only budget and fast-tier models appear here?
At 250M tokens a month, frontier-tier pricing rarely stays economical for any provider, so this ranking focuses on the tiers high-volume buyers actually use in practice. See the tier-matched comparison pages elsewhere on this site if you specifically need frontier-tier capability at volume.
Can batch and committed-use discounts stack?
On the providers shown here, current published terms allow it, and this ranking applies both where available. Stacking rules are set by each provider and can change — verify current terms before budgeting against them.