What Does an AI Customer Service Assistant Cost
Token Prices at the End of 2026
- AI In Business
Ten thousand customer service conversations cost roughly 37 dollars in tokens.
The figure comes from Anthropic’s own pricing documentation, and it is the number that goes missing from the slide when an AI assistant is demonstrated to a board. Thirty-seven dollars sounds too small to be credible, and on its own it misleads: the same ten thousand conversations can cost forty times more once they run in Hungarian, carry a knowledge base, and hand the hard cases to a person. Here is what tokens cost in September 2026, how those prices move, and how to build a cost case that survives the first invoice.
The question that ends the demo
At our September executive breakfast in Budapest, a division head from a Hungarian telecommunications group made a request that stayed with me. When you present a customer service assistant that reduces the need for human capacity, he said, put the AI cost and the token consumption on the table beside it. Show both pans of the scale.
It is a fair request, and a harder one than it sounds. The payroll side is easy to model, because it is a known number that changes once a year. The AI side has a list price that changes several times a year, a consumption pattern nobody can predict before the first pilot, and at least six multipliers that live outside the vendor price table.
How the meter runs
The three major providers meter the same four items: input tokens, output tokens, cached input, and cache writes. The differences between them show up in the modifiers stacked on top.
Input tokens cover everything sent to the model: the system prompt, the tool definitions, the retrieved documentation, and the conversation so far. Output tokens cover everything the model writes, including the reasoning it performs before it answers. Output is the expensive direction, running at five or six times the input rate across the current lineup.
Cached input is the biggest discount in customer service because a support assistant sends the same system prompt and product documentation with every request. OpenAI, Anthropic, and Google all bill a repeated prompt prefix at 10% of the standard input rate. Anthropic charges 1.25 times the input rate to write a five-minute cache, so caching pays for itself after a single re-read.
Then come the modifiers. A prompt above the provider’s long context threshold moves to a separate, higher meter. Asynchronous work through a Batch API is half price at all three providers. Server-side tools carry their own line: web search costs 10 dollars per 1,000 searches at both OpenAI and Anthropic. And pinning inference to a geography, which most Hungarian enterprises will want for data residency, adds roughly 10% everywhere. Anthropic applies a 1.1 multiplier for region-pinned inference, Google charges 10% more on non-global endpoints, and OpenAI adds a 10% uplift for regional processing on models released since 5 March 2026.
What tokens cost in September 2026
List prices in US dollars per million tokens, standard tier, short context, taken from the providers’ own pricing pages on 18 September 2026.
| Model | Input | Cached input | Output |
|---|---|---|---|
| GPT-5.6 Luna (OpenAI) | 0.20 | 0.02 | 1.20 |
| Gemini 3.1 Flash-Lite (Google) | 0.25 | 0.025 | 1.50 |
| Gemini 3.6 Flash (Google) | 0.75 | 0.075 | 3.75 |
| Claude Haiku 4.5 (Anthropic) | 1.00 | 0.10 | 5.00 |
| Claude Sonnet 5 (Anthropic) | 2.00 | 0.20 | 10.00 |
| Gemini 3.1 Pro (Google) | 2.00 | 0.20 | 12.00 |
| GPT-5.6 Sol (OpenAI) | 4.00 | 0.40 | 20.00 |
| Claude Opus 5 (Anthropic) | 5.00 | 0.50 | 25.00 |
| GPT-6 Astra (OpenAI) | 10.00 | 1.00 | 50.00 |
That is a fifty-times spread on input and roughly forty times on output between the cheapest production tier and the frontier. For customer service, the lower tiers carry the work, because the job is classification, retrieval, and short, accurate answers inside a controlled domain, and that sits well within their range.
One conversation, priced
Anthropic publishes a worked example for exactly this use case: an average support conversation of about 3,700 tokens, run on Claude Haiku 4.5, giving roughly 37 dollars per 10,000 tickets. That is about 0.4 US cents per conversation. At 20,000 conversations a month, it comes to around 74 dollars, a rounding error against a single full-time agent in any European market.
If the meeting stops there, the business case is built on sand. That figure assumes English text, no retrieval, no tool calls, one model pass, and no escalation. Change any one of those and the number moves.
Six things that move the number
- Language. Hungarian is expensive to tokenize. Tokenizers are trained mostly on English text, so Hungarian words break into more pieces and the same document consumes markedly more tokens. We covered the mechanics in Why LLMs Are Weaker in Hungarian. For cost modeling, the rule is simple: measure your own corpus, and treat every English benchmark as the best case.
The difference is not only about token count. The same models can also perform very differently in Hungarian than in English. We examine why in Why LLMs Are Weaker in Hungarian →
- Retrieval. A useful support assistant answers from your documentation, so every reply carries several thousand tokens of retrieved context on the input side. This is where prompt caching earns its place, since the stable part of that context bills at a tenth of the rate.
- Reasoning and agentic loops. Gartner puts the multiplier at 5 to 30 times more tokens per task for agentic models compared with a standard chatbot query. An assistant that checks an order, reads a contract, and drafts a reply is doing agentic work, and it deserves to be budgeted that way.
- Tokenizer changes. Anthropic notes that Claude 4.7 and later models use a newer tokenizer that produces approximately 30% more tokens for the same text. A published rate cut can still raise your monthly bill, which is why the model you measure must be the model you deploy.
- Voice. A telephone assistant sits in a different price class from a chat widget. OpenAI bills real-time audio at 32 dollars per million input tokens and 64 dollars per million output tokens, with a mini tier at 10 and 20 dollars, and GPT-Live voice sessions at 5 US cents per minute. Google prices live transcription at an effective 0.89 US cents per audio minute. Voice deployments need their own metering model.
- Escalation. Every conversation the assistant fails to contain still costs its tokens, and then costs an agent as well. Containment rate is the most important variable in the whole calculation, and it is the one number no vendor can hand you in advance.
Prices move, and some carry an expiry date
Two rows in the table above carry a published end date, and a third recent move shows how quickly the rest can shift.
- Google’s promotional rate of 0.75 dollars input and 3.75 dollars output on Gemini 3.6, 3.7, and 3.8 Flash runs through 31 December 2026. Standard pricing of 1.50 and 7.50 dollars applies from 1 January 2027, doubling the rate on the first day of the new budget year.
- OpenAI lists the 4 dollar input rate on GPT-5.6 Sol as promotional at least through 21 November 2026.
- OpenAI cut GPT-5.6 Terra by 20% and Luna by 80% on 30 July 2026, three weeks after the family became generally available.
Movement can run the other way too, and it can be canceled. Anthropic had scheduled Claude Sonnet 5 to rise from 2 and 10 dollars to 3 and 15 dollars on 1 September 2026, then confirmed the increase would not happen and made the launch price permanent. Anyone who budgeted for that step-up is carrying spare room in the 2026 plan.
The practical conclusion for a 2027 budget: model the post-promotion list price, treat today’s rate as a temporary discount, and keep model choice in configuration, where a switch costs an afternoon.
Cheaper tokens, larger bills
The long-run direction is settled. Epoch AI measured how fast the price of reaching a fixed capability level has fallen and found declines ranging from 9 to 900 times per year depending on the benchmark. The Price of Progress, a study built on benchmark-level price history, puts the frontier figure at 5 to 10 times per year and isolates algorithmic efficiency progress at around 3 times per year. The two methods land on different magnitudes and agree on the direction. Gartner forecasts that by 2030, inference on a one-trillion-parameter model will cost providers over 90% less than it did in 2025.
Now the counterweight. Gartner is explicit that those savings reach enterprise customers only in part, and that frontier intelligence will demand far more tokens than today’s mainstream applications. As consumption rises faster than unit prices fall, total inference cost is expected to increase. Will Sommer, the Gartner Sr. Director Analyst behind the forecast, describes the trap as confusing the deflation of commodity tokens with “the democratization of frontier reasoning”.
The market data follows that logic. Gartner’s July 2026 forecast puts worldwide end-user spending on AI models and platforms at 64 billion dollars for 2026, up 63.4% from 39 billion in 2025. Unit prices are collapsing while invoices grow, and both statements are true at once.
What belongs on the other side of the scale
A defensible cost case for a customer service assistant has seven lines, and it starts with a number only you can measure.
- Tokens per conversation, measured on your own content, in Hungarian, with your retrieval configuration in place. Run 200 real conversations through a pilot and count. Every other number depends on this one.
- Containment rate. Price contained conversations and escalated conversations separately, because an escalated conversation costs tokens and then costs an agent.
- Model routing. Send classification, intent detection, and short factual answers to the cheap tier, and reserve the frontier models for the small share of cases that need them. The spread in the table above is fifty times.
- Caching and batching. Cache the system prompt and the stable knowledge base at 10% of the input rate, and push overnight work through the Batch API at half price.
- Spend controls. Set hard monthly caps in the API console before anyone holds a key. Consumption pricing has no natural ceiling, and the stories of budgets emptied in a week are all stories about missing caps.
- The 2027 list price, with today’s promotional rate treated as a temporary discount that ends on a date you can already read.
- The costs that dwarf inference: integration with your CRM and telephony, data preparation, evaluation, monitoring, and the people who supervise the assistant. In most projects we see, tokens are the smallest line item in the first year.
Inference is only one part of the real cost. Integration, internal time, testing and ongoing operation often sit outside the vendor price table. We break down those hidden costs in What Does an AI Project Really Cost? →
The honest answer
The division head at that breakfast asked the right question, and the answer has two halves. Token prices in September 2026 are low enough that inference is rarely what decides a customer service business case. Token consumption is volatile enough that an unmeasured deployment can still surprise a CFO in the second quarter. Both halves belong on the slide, and the second one is only knowable by measuring your own traffic, in your own language, on your own documentation.
So the next time a vendor shows you the headcount saved, ask for the token model beside it. If the answer is a single number with no containment rate, no language multiplier, and no expiry date on the price, the estimate isn’t built yet.
Put measured numbers on both pans of the scale
If your organization is weighing an AI customer service deployment and the cost side of the case is still an estimate, a structured assessment is the fastest way to replace it with measured figures. The AI Opportunity Check is a focused, fixed-fee review at 1,450 euros: one workshop, a written summary, and a clear answer on whether to start a pilot now. The AI Compass Audit is the 4-week, fixed-fee version for organizations mapping several use cases at once, with success criteria and risks attached to each.
Sources
Anthropic. (2026). Claude Platform Pricing. Source of Claude model rates, prompt caching, batch pricing, tokenizer changes, and the customer support cost example. Read article →
Epoch AI. (2026). LLM Inference Price Trends. Source of historical declines in inference prices at fixed capability levels. Read article →
Gartner. (2026, March 25). Gartner Predicts Inference Costs Will Fall More Than 90% by 2030. Source of the inference cost forecast and the 5–30× token multiplier for agentic AI. Read article →
Gartner. (2026, July 20). Gartner Forecasts Worldwide AI Platforms and Models Market to Grow 63% in 2026. Source of the global AI models and platforms spending forecast. Read article →
Google Cloud. (2026). Agent Platform Generative AI Pricing. Source of Gemini pricing, promotional rates, regional processing premiums, and transcription costs. Read article →
Omnit. (2026). Why LLMs Are Weaker in Hungarian. Source of Hungarian tokenization behavior and its effect on AI cost and context usage. Read article →
OpenAI. (2026). API Pricing. Source of GPT model rates, cached input, long-context pricing, real-time audio, and web search costs. Read article →
The Price of Progress. (2025). Source of benchmark-level AI price declines and improvements in algorithmic efficiency. Read article →

Lajos Fehér
Lajos Fehér is an IT expert with nearly 30 years of experience in database development, particularly Oracle-based systems, as well as in data migration projects and the design of systems requiring high availability and scalability. In recent years, his work has expanded to include AI-based solutions, with a focus on building systems that deliver measurable business value.
Related posts

Turning Ambition into Real, Scalable Results

Why So Many AI Projects Stall — and How to Finally Move Beyond the Pilot Phase
Are you sure AI is the right next step?
We help uncover the real opportunities, limitations, and realistic next steps.


