Rates checked October 1, 2026; USD per million tokens. This guide compares token billing and explains the arithmetic. Consumer subscriptions, cloud reseller prices, taxes, and additional tool or infrastructure charges are separate.
Available API rates versus announced Argon pricing
| Model and scope | Input | Cached input read | Cache write | Output |
|---|---|---|---|---|
| Sonnet 5.5, standard Claude API | $2 | $0.20 | $2.50 for 5 minutes; $4 for 1 hour | $10 |
| GPT-6.1 Sol, Standard; up to 272K input tokens | $2 | $0.10 | $2.50 | $10 |
| GPT-6.1 Sol, Standard; more than 272K input tokens | $4 | $0.20 | $5 | $15 |
| Argon, announced introductory launch terms | $2 | 95% off the input rate | Not specified in the announcement | $10 |
| Argon, announced after the introductory period | $4 | Check launch documentation | Check launch documentation | $20 |
Sources: Anthropic’s API price list, OpenAI’s API price list and GPT-6.1 Sol pricing conditions, and Google’s Argon announcement. An introductory-period end date is not specified.
The GPT-6.1 Sol long-context rates apply to the full request once input exceeds 272K tokens. Model capacity and billing thresholds are different. See our model access and context comparison for the specification limits.
A cost calculation you can reproduce
For a request without caching, multiply input tokens by the input rate and billed output tokens by the output rate, dividing each by one million. Add any separately billed services.
Example: 10,000 ordinary input tokens and 2,000 total billed output tokens at $2/$10 cost $0.04: (10,000 ÷ 1,000,000 × $2) + (2,000 ÷ 1,000,000 × $10). This arithmetic applies to the stated standard rates for Sonnet 5.5 and short-context GPT-6.1 Sol. It is an illustration, not a measured cost for a particular prompt.
Output length, repeated calls, and unsuccessful attempts can change the bill even when two models have the same rate. Compare cost per task that meets your quality requirements, rather than assuming equal cost from the final answer’s visible length.
Cache writes and cache hits have different costs
A repeated prefix must actually qualify for caching and produce a hit. Changing it, losing eligibility, or sending additional uncached material changes the calculation. A cache read discount does not make the initial write free.
| Request state | Sonnet 5.5, 5-minute cache | GPT-6.1 Sol, short-context Standard |
|---|---|---|
| Initial cache write plus output | $0.25 + $0.01 = $0.26 | $0.25 + $0.01 = $0.26 |
| Eligible cache hit plus output | $0.02 + $0.01 = $0.03 | $0.01 + $0.01 = $0.02 |
These are conditional examples based on the listed token rates. Check your response usage and current OpenAI caching rules before treating a repeated prompt as a billed cache hit.
Thinking tokens belong in the output budget
OpenAI bills reasoning tokens as output, and Anthropic also bills Claude’s thinking tokens as output. Neither provider’s bill can be inferred solely from visible answer text. This is documented token billing, rather than a separate surcharge unique to OpenAI. Sources: OpenAI’s reasoning guide and Anthropic’s thinking and cost guide.
Record actual input, cache, output, and thinking or reasoning usage. Set suitable output limits, then check whether requests complete successfully. A smaller budget that frequently ends an answer early can increase the cost of a usable result through retries.
Batch rates and costs beyond tokens
Anthropic lists a 50% input/output discount for Batch API requests. OpenAI lists Batch and Flex at 50% below Standard; other processing modes and regional processing can change prices. Check the applicable provider documentation rather than combining discounts by assumption. We have not established Argon Batch rates from its announcement.
Budget for tools, storage, application hosting, and review time where applicable. Run a representative sample, count successful results, and estimate monthly spending from the measured workload. The same approach applies to coding assistants and agents.
Frequently asked questions
Which model is cheapest?
There is no single answer from headline prices. Standard uncached rates tie for the two available APIs compared here; cache hits, long-context charges, generated token counts, and retries can change the result. Argon’s announced pricing should not be treated as a current production option without access.
Does a ChatGPT or Claude subscription include API credit?
Do not assume it does. This page compares API token charges. Verify the billing terms of your product subscription and developer account separately.
Methodology and correction
This AI-assisted pricing guide uses primary vendor documentation checked October 1, 2026. Calculations use stated token quantities; GadgetsFocus did not run a comparative API cost benchmark. Prices and launch terms can change, so check the linked pages before committing a budget.
Correction, October 1: Earlier prices, context figures, latency estimates, and monthly cost simulations were inaccurate or unsupported. They have been replaced with sourced rates and explicit arithmetic. Submit a sourced correction through the GadgetsFocus contact page.
About GadgetsFocus · Editorial standards · Suggest a correction


