How Much Do AI APIs Cost? A 2026 Pricing Guide

June 29, 2026
by Sudipto Paul
Sudipto Paul
SP

Sudipto Paul

Sudipto Paul leads the SEO content team at G2 in India. He focuses on shaping SEO content strategies that drive high-intent traffic and ensure your brand is front-and-center as LLMs change the way buyers discover software. He also tests and evaluates software products, translating first-hand experience into content that guides buyers toward the right solutions. He runs Content Strategy Insider, a newsletter where he regularly breaks down his insights on content and search. Want to connect? Say hi to him on LinkedIn.

AI API pricing ranges from $0.01 to $30 per million input tokens and $0.02 to $180 per million output tokens, depending on the model. As of June 2026, the cheapest production-grade large language model API is Liquid LFM2-8B at $0.01 input / $0.02 output per million tokens. The most expensive flagship is GPT-5.5 Pro at $30 input / $180 output per million tokens. Most production workloads fall in the $0.05-$5 input range.

That's the headline number, but it's also the most misleading number you can plan against. Output tokens cost two to eight times more than input tokens. Prompt caching cuts repeat-context cost by 90 percent. Batch APIs cut everything in half. Long-context requests trigger separate, higher tiers above 200K tokens. The difference between sticker price and effective price is routinely an order of magnitude.

This API development software guide covers what you actually pay across 25+ providers, every modality, and three real workload patterns.

AI API pricing by provider and model: June 2026 comparison

The table below covers the most-used commercial LLM APIs at standard rates. All prices are USD per million tokens. Context windows shown are the largest currently available at standard pricing.

Flagship and frontier tier

Provider Model Input ($/1M) Output ($/1M) Context
OpenAI  GPT-5.5 $5.00 $30.00 1.05 M
OpeAI GPT-5.4 $2.50 $15.00 1.05 M
Anthropic Claude Opus 4.8 $5.00 $25.00 1M
Anthropic Claude Opus 4.7 $5.00 $25.00 1M
Anthropic 2.5 Claude Sonnet 4.6 $3.00 $15.00 1M
Google Gemini 3.1 Pro $2.00 $12.00 2M
Google Gemini 3.5 Flash $1.50 $9.00 1M
Google Gemini 2.5 Pro $1.25 $10.00 1M
xAI Grok 4 $3.00 $15.00 256K
xAI Grok 4.20 $2.00 $6.00 256K
Cohere Command R+ $2.50 $10.00 128K
Mistral Mistral Large 3 $0.50 varies 128K
Amazon Nova Pro $0.80 $3.20 300K
Z.ai GLM-5.2 $0.60 $2.20 200K

Mid-tier (general production workloads)

Provider Model Input ($/1M) Output ($/1M) Context
OpenAI GPT-5 mini $0.125 $1.00 400K
OpenAI GPT-5.4-nano $0.20 $1.25 -
OpenAI GPT-4o-mini $0.15 $0.60 128K
Anthropic Claude Haiku 4.5 $1.00 $5.00 200K
Google Gemini 2.5 Flash $0.30 $2.50 1M
xAI Grok 4 Fast $0.20 $0.50 2M
xAI Grok Code Fast 1 $0.20 $1.50 256K
DeepSeek DeepSeek V3.2 $0.23 $0.34 164K
DeepSeek DeepSeek Chat V3.1 $0.21 $0.79 33K
Alibaba Qwen3 Coder $0.22 $0.90 262K
Alibaba Qwen3-235B-A22B $0.09 $0.10 262K
Meta Llama 4 Maverick $0.15 $0.60 1M
Meta Llama 4 Scout $0.10 $0.30 328K
Meta Llama 3.3 70B Instruct $0.10 $0.32 131K
Z.ai GLM-4.5 Air $0.13 $0.85 131K
Minimax MiniMax M2.7 $0.24 $0.96 205K
Minimax MiniMax-01 $0.20 $1.10 1M
Reka Reka Flash 3 $0.10 $0.20 66K
AI21 Jamba 1.5 Mini $0.20 $0.40 -

Budget tier (high-volume, classification, routing)

Provider Model Input ($/1M) Output ($/1M) Context
Liquid LFM2-8B0A1B $0.010 $0.020 33K
Inclusionai Ling 2.6 Flash $0.010 $0.030 262K
Deepseek DeepSeek-Chat $0.014 $0.028 164K
IBM Granite 4.0-H-Micro $0.017 $0.112 131K
Meta Llama 3.2 1B Instruct $0.020 $0.020 60K
Mistral Mistral Nemo $0.020 $0.030 131K
OpenAI GPT-OSS-20B $0.029 $0.130 131K
OpenAI GPT-OSS-120B $0.039 $0.100 131K
Amazon Nova Micro $0.035 $0.140 128K
Cohere Command R7B (12-2024) $0.037   128K
Mistral Ministral 3B $0.040 $0.040 128K
Google Gemma 3 4B $0.040 $0.080 131K
OpenAI GPT-5 nano $0.050 $0.400 400K
OpenAI GPT-4.1 nano $0.050 $0.200 1M
Microsoft Phi-4 $0.070 $0.140 16K
Google Gemma 3 27b $0.080 $0.160 128K
Amazon Nova Lite $0.060 $0.240 300K
Z.ai GLM-4.7 Flash $0.060 $0.400 203K
NVIDIA Nemotron Nano 9B v2  $0.060 $0.200 131K
Google Gemini 2.5 Flash-Lite $0.10 $0.40 1M
Groq Llama 3.1 8B Instant $0.050 $0.080 128K

Most of these models can be found and compared on G2's Large Language Models (LLMs), AI Gateways, and Generative AI Infrastructure category pages.

How did we evaluate AI API pricing across models?

Pricing in this guide reflects publicly listed standard API rates as of June 25, 2026. Rates were cross-checked against official provider pricing pages and independent pricing aggregators, including Price Per Token and AI Pricing Guru.

 

Because AI API pricing changes frequently, buyers should verify current rates directly with each provider before making budgeting, contract, or vendor-selection decisions. Providers can also be compared across G2 categories such as Large Language Models, AI Gateways, AI SDK, Generative AI Infrastructure, and LLMOps.

What are the five pricing dimensions every AI APIs buyer needs to understand?

The cost of AI API depends on five core dimensions: model tier, context window, cached input, batch processing, and tool or reasoning token usage.

  1. Model tier: Within a single provider, the price spread between the cheapest and most expensive model is typically 25x to 100x. OpenAI's GPT-4.1 nano is $0.05 input per million tokens; GPT-5.5 Pro is $30. That's a 600x spread on input alone. Choosing the wrong tier for routine tasks is the single biggest source of overspending.
  2. Context window and long-context premiums: Most providers charge a flat per-token rate up to a threshold (commonly 200K tokens), then switch to higher rates above it. Google Gemini 3.1 Pro charges $2 input under 200K but jumps to $4 above it. Anthropic's Claude Opus 4.7, Opus 4.8, and Sonnet 4.6 are the notable exceptions. They price 1M-token context windows at flat rates with no surcharge.
  3. Cached input: When you reuse the same context (system prompt, document, conversation history) across requests, providers offer 75 to 90 percent discounts on the cached portion. This is the single largest cost lever for RAG pipelines and agents.
  4. Batch processing: Submitting requests asynchronously with a 24-hour return window cuts the bill by exactly 50 percent across OpenAI, Anthropic, Google, and most other providers. No quality difference, no model restrictions.
  5. Tool calls and reasoning tokens: Modern models charge separately for web search ($10 per 1,000 calls on OpenAI), file search, code execution containers, and image processing. Reasoning models also charge for internal "thinking" tokens that the model generates but doesn't expose to you. A complex Gemini 2.5 Pro request with extended reasoning can quietly multiply the visible-output bill by 3x to 5x.

AI API Pricing formula

The base API pricing formula stays simple: (input tokens × input rate) + (output tokens × output rate). The real complexity comes from everything layered on top, including cached tokens, reasoning tokens, context windows, tool calls, and other usage-based charges.

How does AI APIs pricing vary by provider?

AI API pricing differs widely by provider, model tier, context length, caching discounts, batch-processing options, and compliance needs. This section compares major providers by input and output token costs, highlights where each one is most cost-effective, and explains the trade-offs buyers should consider before choosing an AI API.

1. OpenAI

OpenAI offers one of the broadest API portfolios in the market, with prices spanning from roughly $0.03 to $180 per million tokens. Its current flagship model, GPT-5.5, is priced at $5 input / $30 output per million tokens, while GPT-5.5 Pro reaches $30 input / $180 output for the hardest reasoning workloads. For most production use cases, GPT-5.4 is the practical workhorse at $2.50 input / $15 output.

Lower-cost options such as GPT-5-mini, GPT-4o-mini, GPT-OSS-20B, and GPT-4.1-nano cover mid-tier and high-volume routing workloads, competing more directly with models such as Gemini Flash-Lite.

OpenAI also offers substantial cost reductions through caching and batch processing. On GPT-5.5 and GPT-5.4, cached input tokens are priced 90% below standard input rates, while batch processing can reduce eligible costs by a further 50%. Long-context requests above 270K input tokens fall under a separate, higher-priced schedule. Beyond model breadth and pricing, OpenAI benefits from one of the most widely adopted developer interfaces in the AI ecosystem, with the OpenAI SDK listed on G2 under AI SDK.

Want to go deeper on OpenAI API? Our complete 2026 OpenAI API pricing breakdown covers every model rate, Batch API discounts, image generation, Sora video, and tool costs in one place.

2. Anthropic

Anthropic prices its current-generation Claude models with a simple and predictable 5x output-to-input pricing ratio. Claude Haiku 4.5 is priced at $1 input / $5 output per million tokens, Claude Sonnet 4.6 at $3 / $15, and Claude Opus 4.7 at $5 / $25. Claude Opus 4.8, released May 28, 2026, retains the same $5 / $25 pricing.

Anthropic’s strongest pricing differentiator is long context: Opus 4.7, Opus 4.8, and Sonnet 4.6 include 1M-token context windows at standard rates, with no separate long-context surcharge. That makes Anthropic one of the most predictable providers for long-document analysis and agentic workloads.

Anthropic also offers meaningful savings through caching and batch processing. Cache reads are priced at 10% of base input cost across Claude models, while batch processing reduces eligible costs by 50%.

Together, these discounts can make high-cache workloads highly competitive: for example, a Sonnet 4.6 request with an 80% cache hit rate and batch processing runs at roughly $0.30 per million effective input tokens, putting it in range of mid-tier Gemini pricing and below many OpenAI tiers. Claude also shows strong market traction, with 350+ reviews on its G2 LLM category listing, making it the most-reviewed Anthropic product on G2.

Want to go deeper on Claude specifically? Read more on Claude API pricing breakdown for 2026 covering every model from Haiku 4.5 to Fable 5, Batch API discounts, and what Managed Agents and web search add to your bill.

3. Google Gemini

Gemini API pricing splits cleanly into three tiers. Gemini 3.1 Pro is the flagship model at $2 input / $12 output per million tokens. Gemini 3.5 Flash, launched at Google I/O 2026, is the recommended default for most production workloads at $1.50 / $9 and is especially strong on coding and agentic benchmarks. Gemini 2.5 Flash-Lite, at $0.10 / $0.40, is the cheapest viable production model in the major-provider lineup. Google’s open Gemma models go even lower, with Gemma 3 4B priced at $0.04 / $0.08 and Gemma 3 27B at $0.08 / $0.16.

The main pricing caveat is that Google charges for thinking tokens, or internal reasoning tokens, as part of output on the 2.5 and 3.x model families. That can quietly raise the bill when a request produces relatively little visible output but uses substantial hidden reasoning.

Long-context requests above 200K tokens on 2.5 Pro and 3.1 Pro also move to a higher pricing schedule, at roughly 2x input and 1.5x output rates. Even with those caveats, Google’s free tier remains unusually generous for prototyping, with 5 to 15 RPM, 1,000 daily requests, and access to six models.

Recommended reading: Wondering how ready your organization is for AI? Read AI maturity model: How to assess and scale to learn how to evaluate your current AI capabilities.

4. Meta Llama (via hosted providers)

Meta does not sell Llama through a direct first-party API. Instead, Llama models are available through third-party inference providers such as Groq, Together AI, Fireworks, and Replicate.

Pricing varies by host, but hosted Llama APIs typically cost between $0.05 and $0.90 per million input tokens. For example, Llama 4 Maverick on Together is priced at $0.15 input / $0.60 output per million tokens, Llama 4 Scout at $0.10 / $0.30, and Llama 3.3 70B Instruct at $0.10 / $0.32.

Groq is a common choice for low-latency Llama inference because it runs models on custom LPU silicon and can deliver more than 500 tokens per second, roughly 10x faster than many GPU-based hosts. Llama 3.1 8B Instant on Groq costs $0.05 per million input tokens. At a very large scale, usually 10 billion or more tokens per month, self-hosting Llama on owned GPUs can reduce effective costs below $0.10 per million tokens. However, that changes the cost model from variable API usage to fixed infrastructure spend.

5. Mistral AI

Mistral AI is one of the lowest-cost first-party model providers, especially at the flagship tier. Its top model, Mistral Large 3, costs roughly $0.50 per million input tokens, making it the cheapest flagship-class first-party API in this comparison. Mistral Small costs $0.10 input / $0.30 output per million tokens, while Mistral Small 3.2 24B costs $0.075 / $0.20.

The Ministral edge family is even cheaper, starting at $0.04 / $0.04 for the 3B model. Mistral Nemo, at $0.02 / $0.03, is one of the cheapest production-ready models available from a first-party provider.

Mistral’s main advantage is not just price. It also offers European GDPR-compliant hosting and strong multilingual performance, which makes it especially relevant for buyers with EU data residency or compliance requirements.

6. Cohere

Cohere is built primarily for enterprise RAG and retrieval workflows. It offers a full toolkit from one vendor, including Command R+ as its flagship LLM at $2.50 input / $10 output per million tokens, Command R as a mid-tier option at $0.15 / $0.60, Embed v3 for embeddings at $0.10 input with no output charge, and Rerank v3 for retrieval optimization.

Cohere also offers lower-cost options for high-volume workloads. Command R7B, released in December 2024, costs $0.0375 input / $0.15 output per million tokens, making it one of the cheapest first-party APIs available. Cohere is also listed on G2’s AI SDK category.

7. DeepSeek

DeepSeek is one of the most aggressively priced AI API providers in the market. DeepSeek-Chat costs $0.014 input / $0.028 output per million tokens, while DeepSeek V3.2 costs $0.23 / $0.34. DeepSeek Chat V3.1 is priced at $0.21 / $0.79, and DeepSeek V4 Flash costs $0.09 / $0.18.

The main caveat is compliance. DeepSeek is based in China, which can create geopolitical, regulatory, and operational concerns for some enterprise buyers. For teams that can use it, the pricing is especially attractive when combined with cached-input discounts. Cached-input pricing on V4 Flash dropped 10x in mid-2026 to $0.0028 per million tokens. DeepSeek is also listed in G2’s LLM category.

8. xAI Grok

xAI’s Grok lineup is priced across flagship, fast, and coding-focused tiers. Grok 4 is the flagship model at $3 input / $15 output per million tokens. Grok 4.20 costs $2 / $6, making its output pricing 60% lower than GPT-5.4.

Grok 4 Fast costs $0.20 / $0.50 and includes a 2M-token context window, making it one of the cheapest frontier-adjacent models available. Grok Code Fast 1, priced at $0.20 / $1.50, is positioned specifically for coding workloads.

The details above cover xAI Grok pricing at a glance. For the full picture, our xAI Grok API pricing breakdown goes deeper on every model tier, Priority Processing, Batch discounts, and multimodal costs.

Other notable AI API providers

Beyond the seven providers above, several second-tier players matter for specific buyer profiles.

  • Amazon Nova(AWS Bedrock first-party): Amazon Nova is AWS’s first-party LLM family, available through Amazon Bedrock alongside third-party models. Nova Micro is priced at $0.035 input / $0.14 output per million tokens, Nova Lite at $0.06 / $0.24, and Nova Pro at $0.80 / $3.20. Its main advantage is tight integration with the broader AWS ecosystem, including IAM, VPC, KMS, CloudWatch, and centralized enterprise billing. AWS Bedrock is also listed in G2’s LLMOps category with 71 reviews.
  • Alibaba Qwen: Qwen offers one of the broadest model catalogs of any AI provider. Its lineup includes Qwen3-235B-A22B at $0.09 input / $0.10 output, Qwen3 Coder at $0.22 / $0.90, Qwen3.5 Flash at $0.065 / $0.26, and many specialized models for vision, coding, and long-horizon agent workflows. For buyers comfortable using Chinese open-weight models, Qwen is one of the strongest price-to-quality options. Qwen3-235B-A22B, with 262K context and low token pricing, is among the cheapest near-frontier models available.
  • Z.ai GLM: Z.ai’s GLM family offers competitive reasoning and coding performance at relatively low prices. GLM-5.2, released in June 2026, is the company’s current frontier model at roughly $0.60 input / $2.20 output per million tokens. GLM-4.5 Air, priced at $0.13 / $0.85, and GLM-4.7 Flash, priced at $0.06 / $0.40, cover mid-tier and budget use cases. The main appeal is strong benchmark performance at a fraction of typical OpenAI flagship pricing.
  • IBM Granite: IBM Granite is an open-weight model family designed for enterprise governance, compliance, and on-premises deployment. Granite 4.0-H-Micro, priced at $0.017 input / $0.112 output per million tokens, is one of the cheapest enterprise-focused options. Granite 4.1-8B is priced at $0.05 / $0.10. Granite also powers IBM watsonx.ai, which is listed on G2 under
    Generative AI Infrastructure.
  • NVIDIA Nemotron:  NVIDIA’s Nemotron models are open-weight LLMs optimized for NVIDIA’s own inference stack. Nemotron Nano 9B v2 is priced at $0.06 input / $0.20 output per million tokens, while Nemotron 3 Super 120B costs $0.09 / $0.45. These models are especially relevant for organizations self-hosting AI workloads on NVIDIA H100 or B200 GPU clusters.
  • Microsoft Phi: Microsoft Phi is a small but capable model family aimed at cost-efficient AI workloads. Phi-4 is priced at $0.07 input / $0.14 output per million tokens, while Phi-4-mini reasoning extends the family into reasoning-focused tasks. Phi is most useful for buyers already committed to Azure who want a Microsoft-native model option below the GPT-5 pricing tier.
  • AI21 Jamba: AI21’s Jamba models use a hybrid Mamba-Transformer architecture designed for efficient long-context processing. Jamba 1.5 Mini is priced at $0.20 input / $0.40 output per million tokens, while Jamba 1.5 Large is positioned as the higher-end option. AI21 is especially relevant for regulated industries such as finance and healthcare, where data residency and enterprise controls are important buying factors.
  • MiniMax: MiniMax is a low-cost long-context provider. MiniMax-01 costs $0.20 input / $1.10 output per million tokens and supports a 1M-token context window, making it one of the cheapest options for long-context workloads. MiniMax M2.7, priced at $0.24 / $0.96, is the company’s current flagship. Because MiniMax is based in China, buyers should apply the same compliance review used for DeepSeek and Qwen.
  • Reka: Reka Flash 3 is a strong mid-tier multimodal model priced at $0.10 input / $0.20 output per million tokens. It is useful for teams that need capable multimodal performance with EU or UK data-residency considerations, without paying the higher prices often associated with larger frontier providers.
  • Liquid AI: Liquid AI’s Liquid Foundation Models use a non-Transformer architecture and are priced for high-volume, latency-sensitive workloads. The LFM2 family starts very low, with LFM2-8B priced at $0.01 input / $0.02 output per million tokens. These models are best suited for constrained, repetitive, or structured tasks where speed and cost matter more than frontier-level general reasoning.
  • Specialty Chinese providers: Providers such as ByteDance Seed, Tencent Hunyuan, Stepfun Step, Moonshot Kimi, and Xiaomi MiMo primarily serve the Chinese domestic market, though some are increasingly accessible through international gateways. Their pricing usually falls between $0.05 and $0.30 per million input tokens. However, buyers outside China should conduct a careful compliance, data residency, and security review before using these models in production.

Recommended reading: Estimating AI costs beyond API usage? Explore our guide to chatbot pricing to compare leading chatbot platforms.

How does pricing differ by AI APIs modality?

Token pricing mainly applies to text generation. Other AI modalities, such as embeddings, speech, image, and video, use different billing units. Some are priced per token, while others are priced per minute, character, image, or second of generated video.

  • Embeddings: Embeddings are usually priced per million input tokens and are much cheaper than LLM generation. OpenAI’s text-embedding-3-small costs $0.02 per million tokens, while text-embedding-3-large costs $0.13. Cohere Embed v3 costs $0.10 for input, with no output charge. Google Gemini Embeddings is free within AI Studio rate limits, with the paid tier at roughly $0.025 per million tokens. In a typical RAG pipeline, embeddings account for under 5% of the budget, while the generation step accounts for about 95%.
  • Speech-to-text: Speech-to-text APIs are usually priced per minute of audio. OpenAI Whisper API costs $0.006 per minute, making it the cheapest mainstream option. OpenAI GPT-4o Transcribe costs $6 per 1,000 minutes, while AWS Transcribe costs $0.024 per minute. Deepgram Nova-3 is also priced per minute with volume discounts, and Deepgram Flux is the current voice-agent latency leader. AssemblyAI Universal-2 leads on transcript intelligence, including sentiment, entity, and topic extraction. ElevenLabs Scribe v2 leads in multilingual real-time transcription at 150ms latency across 30 languages.
  • Text-to-speech: Text-to-speech pricing is usually based on characters, and it has one of the widest price ranges across AI modalities. ElevenLabs Multilingual v3 costs $206 per million characters and is positioned for premium realism. ElevenLabs Flash v2.5 costs $103 per million characters. OpenAI TTS costs $15 per million characters for standard voices and $30 for HD. Deepgram Aura-2 costs $30 per million characters with sub-90ms latency. Hume Octave 2 costs $7.60 per million characters, making it the cheapest emotionally adaptive option. Inworld TTS-1.5 Max costs $10 per million characters and is the current quality-per-dollar leader.
  • Image generation: Image generation can be priced either per output token or per image, depending on the provider. Gemini 3.1 Flash Image costs $60 per million output tokens, which works out to roughly $0.067 per 1024×1024 image. Gemini 3 Pro Image costs $0.134 per 1K/2K image. Stability AI, OpenAI’s GPT Image, and Replicate-hosted models usually price per image instead of per token, typically ranging from $0.02 to $0.10 per image.
  • Video generation: Video generation is usually priced per second of generated video. Google Veo 3.1 uses per-second pricing, with Veo 3.1 Lite costing $0.05 to $0.08 per second. Sora API access via OpenAI follows a similar per-second pricing model.

For routing across modalities, gateway products on G2's AI Gateways category, Kong Konnect, TrueFoundry, KrakenD, Helicone, liteLLM, let you unify billing and observability across providers.

Recommended reading: Comparing API costs? Read our guide to software pricing transparency to learn why pricing is often difficult to compare and what hidden costs to watch for. 

What are the hidden costs and discounts of running AI APIs in production?

AI costs are not fixed; they can be reduced significantly with the right optimization strategy. The biggest savings usually come from caching repeated inputs, batching async work, managing long-context usage, and monitoring hidden tool or reasoning costs.

  • Prompt caching discounts: OpenAI, Anthropic, Google, and DeepSeek all offer 75 to 99 percent discounts on prompt portions that hit the cache. For agent loops with stable system prompts, caching is the single largest cost lever. For example, a 2,000-token system prompt repeated across 10,000 daily Sonnet 4.6 requests drops from $1,800 per month uncached to $180 per month cached.
  • Batch API discounts: When requests do not need an immediate response, you can submit them asynchronously, receive results within 24 hours, and pay exactly half the standard price. Batch processing is available on OpenAI, Anthropic, Google, and most major providers. It also stacks with caching, meaning a cached batch request can cost as little as 5 percent of a synchronous, non-cached request.
  • Long-context surcharges: When usage exceeds 200K tokens on OpenAI's latest GPT-5.4/5.5, Google's Gemini 2.5 Pro and 3.1 Pro, and legacy Claude models, prices move to a higher schedule. Claude Opus 4.7, Claude Opus 4.8, and Claude Sonnet 4.6 are the exceptions because they offer flat 1M-token pricing with no premium.
  • Tool call overhead: Tool use can add separate charges beyond standard token costs. OpenAI charges $10 per 1,000 web search calls, plus the tokens consumed by retrieved content. File search costs $0.10 per GB per day for storage, plus $2.50 per 1,000 tool calls. Code interpreter containers will be billed by 20-minute session starting March 31, 2026. Google's Grounding with Search costs $14 per 1,000 prompts on Gemini 3.x, or $35 on Gemini 2.x, after free quotas.
  • Reasoning token charges: Models that use extended reasoning, including OpenAI's o-series, Claude with extended thinking enabled, and Gemini 2.5/3.x, DeepSeek R1 distill variants, and Qwen3 thinking, charge for internal reasoning tokens that the user never sees. A short 200-token visible response from o3 can include more than 2,000 billed reasoning tokens.
  • Data residency premium: Data residency can add a premium for compliance-sensitive workloads. Anthropic charges a 1.1x multiplier for US-only inference on Opus 4.6 and later models. OpenAI charges 10 percent on regional data residency endpoints. These costs matter primarily for compliance-sensitive verticals.

What does it actually cost to run AI APIs in production?

Production AI costs vary widely by workload, model choice, token volume, caching, and routing. The examples below show what real monthly spend can look like across three common use cases, from simple chatbots to agentic coding systems.

A customer support chatbot

A customer support chat handles 10,000 user messages per day. Each message averages 500 input tokens and 300 output tokens. Based on those volumes, the estimated monthly cost would be:

  • Claude Haiku 4.5 at $1 input / $5 output: about $600 per month
  • GPT-5-mini at $0.125 input / $1 output: about $128 per month
  • Gemini 2.5 Flash at $0.30 input / $2.50 output: about $315 per month
  • Qwen3-235B-A22B at $0.09 input / $0.10 output: about $36 per month

If the chatbot uses a 2,000-token system prompt that is cached across requests, costs can drop by roughly 30% to 50%.

The bill can fall even further with routing. For example, 70% of simple queries such as FAQs, account lookups, and status checks can be routed to cheaper models like Flash-Lite or nano. That can reduce the remaining cost by another 60%.

A RAG pipeline

A B2B knowledge base serves 50,000 monthly queries across a B2B knowledge base with 5 million indexed documents. Each document averages 500 tokens. Indexing all 5 million documents with OpenAI text-embedding-3-small costs about $50 as a one-time expense.

Ongoing query embedding costs are much lower. Embedding 50,000 monthly queries costs about $0.50 per month. The main cost comes from generation. Each query averages 4,000 input tokens, including retrieved context and the prompt, plus 800 output tokens.

At standard rates, this generation step costs about $640 per month on Sonnet 4.6. With aggressive prompt caching on the retrieved context, it drops to about $190 per month. The same workload on DeepSeek V3.2 costs about $60 per month without caching.

Vector storage and retrieval infrastructure adds another $50 to $300 per month, depending on the provider and scale. Relevant providers are commonly listed in G2’s AI Search & Retrieval Infrastructure Platforms category.

An autonomous agent

A code-generation agent runs 1,000 tasks per month.  Each task averages 50,000 input tokens and 15,000 output tokens. The input tokens include codebase context and tool results. The output tokens come from multi-step tool loops.

On Claude Opus 4.8 at $5 input / $25 output, the standard monthly cost is about $625. Prompt caching can reduce this significantly. If the codebase context is cached, cache hits are billed at $0.50 per million tokens instead of $5 per million tokens. With caching, the effective monthly cost drops to roughly $200.

On Qwen3 Coder at $0.22 input / $0.90 output, the same workload costs about $25 per month at standard rates. Without caching, costs can rise quickly. A long-running Opus agent with many tool calls can reach $1,500 to $3,000 per month.

This is why agent observability and routing tools matter. Platforms such as LangChain, AWS Bedrock, and IBM WatsonX, which appear in G2’s Large Language Model Operationalization (LLMOps) category, help teams monitor, route, and control this kind of cost volatility.

How to reduce AI API costs

AI API costs fall fastest when you route, cache, batch, and limit token usage strategically. Here are five techniques in order of impact:

  • Route simple requests to cheaper models:  Send the cheapest 70 percent of traffic to a low-tier model, such as Haiku 4.5, Gemini 2.5 Flash-Lite, GPT-4.1 nano, DeepSeek V3.2, or Qwen3-235B. Reserve the flagship tier for the smaller share of requests that genuinely need it. This single change typically cuts the bill by 60 to 80 percent. AI gateway products can handle the routing logic.
  • Cache repeated prompts and context:  If you reuse a system prompt, document, or conversation history across multiple requests, enable caching. The cached portion costs 10 percent of the base input rate on Anthropic and Google, and as low as 1 percent on DeepSeek V4 Flash. RAG pipelines and agents benefit the most from this approach.
  • Batch work that does not need instant responses:  Anything that does not need a synchronous response, such as nightly summarization, classification sweeps, evaluation runs, or content generation queues, should go through the batch API at 50 percent off. The discount stacks with caching.
  • Limit output tokens and response length:  Output tokens cost 2x to 8x more than input tokens. Capping max_tokens and using structured output formats (JSON schema, enums) directly reduces output volume more than any input-side optimization.
  • Use smaller models when quality is good enough: A 2026-era 70B-parameter open model, such as Llama 4 Maverick on Groq, Mistral Large 3, or Qwen3-235B-A22B, can handle a surprising share of production tasks at one-tenth to one-fiftieth the cost of GPT-5.4 or Claude Sonnet 4.6. The price-to-quality gap between mid-tier and frontier models has compressed substantially in the past year.

For organizations running AI at meaningful scale, dedicated cost-observability tooling (Helicone, TrueFoundry, products in the LLMOps category) typically pays for itself within the first month.

FAQs about AI APIs pricing

Got more questions? We’ve got you covered.

Q1. What is the cheapest AI API in 2026?

Liquid LFM2-8B, priced at $0.01 input and $0.02 output per million tokens, is the cheapest production-grade model. DeepSeek-Chat, at $0.014/$0.028, is the cheapest option from a frontier-capable provider. Alibaba Qwen3-235B-A22B, at $0.09/$0.10, is the cheapest near-frontier model with a 262K context window.

Q2. What is the most expensive AI API?

GPT-5.5 Pro, priced at $30 input and $180 output per million tokens, is the most expensive standard model. Claude Fable 5, where available, costs $10/$50. For most flagship-tier work, Claude Opus 4.7/4.8 and GPT-5.4 fall in the $2.50–$5 input range.

Q3. How are AI APIs priced?

AI APIs are priced per token. A token is roughly four English characters, or about 0.75 words. Input tokens, which are the tokens you send, and output tokens, which are the tokens the model returns, are billed at separate rates. Output tokens are typically 2 to 8 times more expensive than input tokens.

Q4. How much does GPT-5 cost?

GPT-5.5 costs $5 input and $30 output per million tokens. GPT-5.4 costs $2.50/$15, while GPT-5-mini is priced at $0.125/$1.00. GPT-5-nano costs $0.05/$0.40. Cached input across the GPT-5 family is 90 percent cheaper than standard input.

Q5. How much does the Claude API cost?

Claude Opus 4.7 and Opus 4.8 cost $5 input and $25 output per million tokens. Claude Sonnet 4.6 is priced at $3/$15, while Claude Haiku 4.5 costs $1/$5. All current Claude models include a 1M-token context window at standard rates.

Q6. How much does Gemini API cost?

Gemini 3.1 Pro costs $2 input and $12 output per million tokens. Gemini 3.5 Flash is priced at $1.50/$9, while Gemini 2.5 Flash-Lite costs $0.10/$0.40. Long-context requests above 200K tokens on Pro models increase to roughly 2x input rates.

Q7. How much does AWS Bedrock cost?

Bedrock pricing varies by model. AWS Nova Micro costs $0.035/$0.14, Nova Lite costs $0.06/$0.24, and Nova Pro costs $0.80/$3.20. Third-party models on Bedrock, including Claude, Llama, and Mistral, follow each provider’s pricing, with the addition of a marketplace billing structure, such as CCU conversion at $0.01 per CCU for Anthropic.

Q8. How much does Alibaba Qwen API cost?

Qwen3-235B-A22B costs $0.09 input and $0.10 output per million tokens with a 262K context window. Qwen3 Coder is priced at $0.22/$0.90, while Qwen3.5 Flash costs $0.065/$0.26. Qwen offers the deepest open-weight model catalog of any single provider.

Q9. Is there a free AI API?

Yes. Google AI Studio provides the most generous free tier, with 5 to 15 requests per minute and up to 1,000 daily requests across six Gemini models. OpenAI, Anthropic, and most other providers offer time-limited free credits, typically ranging from $5 to $200, for new accounts, but they do not provide a permanent free production tier. Groq offers a generous free developer tier on open-weight models.

Q10. How much does it cost to build a chatbot with AI APIs?

A chatbot handling 10,000 messages per day costs roughly $36 per month on Qwen3-235B-A22B, $128 on GPT-5-mini, $315 on Gemini 2.5 Flash, or $600 on Claude Haiku 4.5 at standard rates. Prompt caching and model routing typically reduce these figures by 60 to 80 percent.

Q11. Should I use a single API or multiple?

Multiple APIs are the better choice in almost every production case. A tiered routing strategy, where a cheap model handles 70 percent of traffic, a mid-tier model handles 20 percent, and a flagship model handles 10 percent, can reduce average per-query cost by 60 to 80 percent compared with routing everything through one premium model. AI gateways such as Kong Konnect, liteLLM, Helicone, TrueFoundry, and KrakenD handle this routing.

Q12. Are AI API prices going up or down? 

AI API prices are going down. List prices across OpenAI, Anthropic, Google, and DeepSeek dropped roughly 60 to 80 percent between early 2025 and mid-2026. New flagship launches now routinely ship at the same or lower price than the previous flagship. Continued downward pressure is likely as model efficiency improves and provider competition intensifies.

Make the AI API spend predictable

AI API pricing is becoming cheaper, broader, and more competitive, but the smartest buying decision is no longer about picking the model with the lowest listed rate. The real advantage comes from understanding how each workload behaves: how much context it needs, how much output it generates, whether responses can be cached or batched, and when a cheaper model is good enough.

Teams that treat AI infrastructure like a routing and optimization problem, not a single-vendor purchase, can control costs without sacrificing quality, especially as model choice, modality pricing, and hidden usage fees keep expanding. In practice, the best AI API strategy is the one that turns pricing complexity into operational leverage.

Compare top-rated LLM platforms on G2 and find the best fit for your team with real user reviews, side-by-side insights, and trusted buyer guidance.