Best value first
Best value LLM for chat
Balanced-tier models at a realistic heavy-chat workload, ranked on cost with quality tier shown alongside.
"Best value" and "cheapest" are different questions. The cheapest model for chat is almost always a budget or fast tier with reduced reasoning quality; best value, as used on this page, means the balanced tier from each major provider — the tier providers themselves position as their general-purpose default — ranked by cost at a workload sized for frequent personal and work chat rather than production API traffic.
The usage figure here (2.4M input / 1.2M output tokens a month) approximates heavy daily use across writing, research, and decision support — noticeably above light occasional use, well below a production API integration. It is a proxy for a real workload, not a measurement of your own usage; the live Compare tool lets you substitute your own token estimate if you have one, and the estimate-monthly-token-usage guide linked below walks through building that estimate.
Because this page compares balanced tiers specifically, it deliberately excludes each provider's frontier tier (priced for maximum capability) and budget tier (priced for maximum cost efficiency) — see the cheapest-API and cost-vs-quality pages elsewhere on this site for those comparisons.
Live pricing snapshot
Workload: 2.4M input tokens, 1.2M output tokens per month · Last verified Jul 14, 2026| Rank | Quality signal | Unit price | Blended $/1M | vs cheapest | ||||
|---|---|---|---|---|---|---|---|---|
| 01 | DeepSeek | DeepSeek V4 Pro Cheapestsource | balanced | 80.6% SWE-bench Verified | In $0.435 / Out $0.87 per 1M tokens | $0.58/1M | $2.09 | — |
| 02 | Gemini 3 Flashsource | budget | 90.4% GPQA Diamond | In $0.50 / Out $3.00 per 1M tokens | $1.33/1M | $4.8 | 2.3× (+$2.71) | |
| 03 | xAI | Grok 4.3source | balanced | 96% Artificial Analysis coding sub-eval | In $1.25 / Out $2.50 per 1M tokens | $1.67/1M | $6 | 2.9× (+$3.91) |
| 04 | Anthropic | Claude Sonnet 5source | balanced | Not verified | In $2.00 / Out $10.00 per 1M tokens | $4.67/1M | $16.8 | 8.0× (+$14.71) |
| 05 | OpenAI | GPT-5.6 Terrasource | balanced | 55 Artificial Analysis Intelligence Index | In $2.50 / Out $15.00 per 1M tokens | $6.67/1M | $24 | 11.5× (+$21.91) |
Prices are read live from the AICostCompass catalogue at page load, not hardcoded on this page, and reflect the workload above. Adjust the workload in the Compare tool to price your own usage instead.
Quality signal links to its source and is whatever third-party benchmark is actually published for that model (SWE-bench Verified, GPQA Diamond, an Elo rating, etc.) — different metrics aren't directly comparable across rows, and "Not verified" means we found no credible published score rather than that the model scores zero.
What "balanced tier" buys you here
At this usage level the total monthly cost across every candidate here is modest in absolute terms — the ranking below is more useful as a relative comparison between providers than as a number to budget precisely around, since actual usage varies week to week for chat-style workloads.
Rate limits, SLA & data residency
Cheapest-per-token is moot if you're throttled. Cost table above, capacity/policy context below.
OpenAI
Rate limits: Five spend-based usage tiers (Tier 1 at $5 spent up to Tier 5 at $1,000+ cumulative spend and 30+ days), enforced on requests and tokens per minute and per day. A new account starts at Tier 1 with meaningfully lower limits than an established Tier 5 or enterprise account.
SLA: No published uptime SLA on standard pay-as-you-go API access. Enterprise / high-commitment customers can negotiate priority processing and an SLA via sales.
Data residency: Official data residency program for at-rest processing/storage in specific regions (EU, UK, Japan, Canada, South Korea, Singapore, Australia, India, UAE), plus a Zero Data Retention option — both sales-gated for eligible enterprise customers.
Source · checked 2026-07-18Anthropic
Rate limits: Organization-level usage tiers (Start, Build, Scale, plus custom enterprise), tied to spend, with separate request- and token-per-minute ceilings. Self-service increase requests unlock once you hit roughly 50% of your current limit.
SLA: No uptime guarantee on the default Standard tier. A Priority Tier product targeted 99.5% uptime for committed-capacity contracts, though Anthropic's docs currently note it's no longer available for new purchases (existing contracts continue through term).
Data residency: Zero Data Retention available for eligible orgs. The direct API only supports "us" or "global" inference routing today — dedicated EU-only residency requires routing through AWS Bedrock or Google Vertex AI's EU endpoints instead.
Source · checked 2026-07-18Google (Gemini)
Rate limits: Project-based usage tiers (Free, Tier 1 once billing is linked, Tier 2 above $250 spend, Tier 3 above $1,000 spend, each after a 30-day wait), enforced simultaneously on requests/tokens per minute and requests per day.
SLA: A formal SLA exists for the Gemini API on Vertex AI, but a guarantee on processed-request availability requires Provisioned Throughput (a paid reserved-capacity product) — standard pay-as-you-go has a narrower model-availability SLA.
Data residency: Jurisdictional/regional endpoints keep processing within a chosen region (e.g. US or EU); eligible enterprise customers can get zero-data-retention-equivalent contract terms via a DPA amendment that disables caching and abuse-monitoring logging.
Source · checked 2026-07-18DeepSeek
Rate limits: No public requests-per-minute tier table — DeepSeek instead publishes a fixed concurrent-request cap per model per account, with capacity expansion available on request at no extra cost.
SLA: Not verified: no published uptime SLA was found in official DeepSeek documentation or policy pages as of the last check.
Data residency: DeepSeek's privacy policy states personal data is collected, processed, and stored in the People's Republic of China. No published enterprise zero-data-retention program or region-choice option was found.
Source · checked 2026-07-18xAI
Rate limits: Team-level tiers set by cumulative API spend (tracked since Jan 1, 2026), with per-model request- and token-per-minute limits that unlock automatically as spend grows and don't downgrade.
SLA: Not verified: xAI references an enterprise uptime commitment, but a specific percentage could not be confirmed on an xAI-owned page as of the last check — confirm directly with xAI before relying on a number.
Data residency: Standard API data is retained roughly 30 days for abuse auditing then deleted, with no training on API inputs/outputs by default. An Enterprise Vault option adds dedicated infrastructure, customer-managed keys, and tenant isolation.
Source · checked 2026-07-18Methodology
Each candidate's current per-1M-token input/output rate is applied to a 2.4M input / 1.2M output token monthly workload, read live from the catalogue, restricted to each provider's balanced quality tier and ranked by resulting monthly cost.
Frequently asked questions
Why not just show the cheapest model overall?
Because the cheapest model overall is typically a budget tier with lower reasoning quality, which most people asking about "chat" usage do not actually want. This page ranks balanced-tier models specifically — see the cheapest-LLM pages elsewhere on this site if lowest absolute price is your only criterion.
How was the usage estimate chosen?
It approximates frequent daily personal and work chat use — more than occasional light use, less than a production API integration. Use the live Compare tool to substitute your own estimated token volume if you have one.
Does this include subscription chat apps like ChatGPT Plus?
No — this ranks metered API pricing at the balanced tier. For flat-rate consumer subscriptions, see the ChatGPT Plus vs Claude Pro comparison.