AI Was Supposed to Get Cheaper, So Why Are Credits and Usage Getting More Expensive?
AI is becoming cheaper and many AI bills are increasing. Both statements can be true.
The price of achieving yesterday’s level of AI capability is falling. However, people are no longer running yesterday’s workloads. They are using deeper reasoning, longer conversations, connected tools, autonomous agents, high-resolution media and repeated generations.
At the same time, “credits” often hide what is actually being consumed. When a provider changes its models, limits or credit conversion, the same workflow may suddenly burn through a balance much faster.
The practical conclusion is simple: businesses should treat AI pricing like a variable wholesale cost, not a permanent utility tariff.
The short answer: unit cost and total spending are different
The claim that AI is simply becoming more expensive is misleading.
Stanford’s 2025 AI Index estimated that the cost of running a system at approximately GPT-3.5 performance fell more than 280-fold between November 2022 and October 2024. Epoch AI’s 2026 trend data estimates that AI-chip performance per dollar has improved by roughly 49% annually since 2023.
But Epoch also estimates that the cost of training frontier language models has grown by approximately 3.5 times per year since 2020. Providers are using efficiency improvements to build more capable systems, not merely to make existing workloads cheaper.
This is similar to mobile data becoming cheaper per gigabyte while people consume far more data. The unit becomes cheaper, but the total invoice does not necessarily fall.
A credit is not a stable unit of AI
A token has a technical meaning, although tokenization can still vary between models. A credit does not have a universal meaning.
For example, OpenAI’s current Business and Enterprise credit rate card explains that usage may be measured as:
- A fixed number of credits per message, task, generation or connected minute.
- Credits per million input, cached-input and output tokens.
- A variable amount based on the model, task complexity, context, output length, automations and fast mode.
Two services can therefore offer “500 credits” while delivering radically different quantities of useful work.
Even token prices require careful comparison. Anthropic says Claude Sonnet 5 produces approximately 30% more tokens for the same text than Sonnet 4.6. Its per-token price is lower, but the cost of an equivalent request does not fall by the full headline percentage.
The metric businesses should monitor is not cost per credit. It is:
Total cost per successful, usable business result, including retries, tools, failed attempts and human review.
Why credits disappear faster
1. Models now reason before answering
A conventional chatbot may produce one short response. A reasoning model can generate internal reasoning tokens before delivering its final answer.
Anthropic’s extended-thinking documentation allows developers to set reasoning budgets, with larger budgets potentially improving difficult results while consuming more output capacity.
A better answer may be worth the additional cost, but businesses should not send every classification, summary or basic support question to their most expensive reasoning configuration.
2. Agents turn one instruction into many operations
“Research these competitors and produce a report” may involve:
- Planning the task.
- Running multiple searches.
- Opening sources.
- Reading documents.
- Calling external tools.
- Revising the analysis.
- Generating the report.
- Checking the result.
The user sees one request. The provider may execute dozens of model and tool operations.
OpenAI’s rate card explicitly states that agentic usage varies with task size, automations, model choice and concurrent instances. Its API pricing also separates token costs from services such as web search, file search and execution containers.
3. Context keeps growing
Long conversations, uploaded documents, website knowledge and previous agent results can be sent back to the model repeatedly.
Prompt caching can reduce this cost, but only when the application and provider support it correctly. Without context management, every additional document or conversation turn can increase the amount processed during later calls.
4. Media workloads are heavier
Text generation is not comparable with generating or analysing high-resolution images, audio or video. Longer duration, higher resolution, additional reference images and premium generation settings generally require more processing.
Creative work also has a hidden multiplier: the first generation is rarely the finished advertisement. If a team produces ten versions to select one usable result, the relevant cost is the total cost of all ten.
5. Speed and priority carry a premium
Providers must decide which users receive limited accelerator capacity first. Fast, priority and high-reasoning modes can therefore have different prices or burn rates.
Paying extra may be rational when a result saves an employee several hours. It is wasteful when the workflow does not require immediate completion.
6. Promotions and credit mappings change
Current pricing pages demonstrate that AI prices can move in both directions.
Google’s Agent Platform pricing currently lists introductory Gemini Flash pricing through 31 December 2026, followed by higher standard pricing from 1 January 2027.
Conversely, OpenAI reduced GPT-5.6 model and credit pricing during July and August 2026, while Anthropic made Sonnet 5’s lower introductory pricing permanent.
The lesson is not that prices always rise. It is that promotional prices, model economics and credit conversions are not permanent assumptions.
Why providers still face enormous infrastructure costs
Efficiency improvements do not eliminate the physical cost of AI infrastructure.
The International Energy Agency’s 2026 analysis reports that capital expenditure by five major technology companies exceeded $400 billion in 2025 and was expected to increase by another 75% in 2026. It also estimates that:
- Overall data-centre electricity consumption grew 17% during 2025.
- Electricity use by AI-focused data centres grew approximately 50%.
- Global data-centre consumption could rise from 485 TWh in 2025 to around 950 TWh in 2030.
Meanwhile, the 2026 Stanford AI Index estimates that global AI compute capacity has grown approximately 3.3 times annually since 2022. It also notes that, once a model is deployed at scale, its cumulative inference energy can exceed the one-time training energy within months.
Providers are not only paying for individual answers. They are funding chips, networking, cooling, electricity, data centres, model research, safety systems and spare capacity for demand spikes.
It is also reasonable to infer that commercial pricing is designed to segment customers, control demand and recover investment—not simply pass through the electricity cost of each request. Most providers do not publicly disclose enough information to separate these factors precisely.
Is changing AI pricing fair?
A paying customer is justified in being frustrated when they build a workflow and client price around one cost structure, only for the underlying economics to change.
However, purchasing a subscription normally provides access under the current terms. It does not usually guarantee that the price, included usage or model mix will remain unchanged forever.
A responsible pricing change should include:
- Clear advance notice.
- A measurable explanation of the new credit conversion.
- Usage records that customers can audit.
- Prospective rather than retroactive application.
- Reasonable migration or cancellation options.
- Transparent treatment of failed generations and retries.
A change becomes difficult to defend when a provider keeps the same advertised subscription price but quietly reduces effective usage, alters credit burn without explanation or gives businesses too little time to adjust client commitments.
Notice periods also vary by product and contract. OpenAI’s current European consumer terms state that subscription price increases receive at least 30 days’ notice and apply at renewal. Its business services agreement says pricing-page changes become effective 14 days after posting.
These are OpenAI examples, not universal industry rules. Whether a particular change is legally permitted depends on the provider agreement, the customer’s jurisdiction and how the change was communicated.
What happens when your agency economics change six months later?
Consider this simplified monthly service:
This is an illustrative example, not an industry benchmark. The provider increase is €120, but the agency’s contribution falls by approximately 22.6%.
The real danger is combining volatile upstream costs with:
- A fixed long-term client price.
- Unlimited usage promises.
- No overage policy.
- No pricing-review clause.
- An application that depends entirely on one model or platform.
The agency has effectively promised stable retail pricing for an unstable wholesale input.
How to price AI-powered client services safely
Measure the complete workflow
Track model tokens, tool calls, media generations, retries, failures, storage and human-review time. Average cost per prompt is less useful than cost per approved output, qualified lead, completed report or resolved customer request.
Stress-test the economics
Calculate the margin at current AI cost and again at several higher-cost scenarios. The appropriate stress multiplier depends on the provider, workload and contract length.
A practical formula is:
Minimum client price = delivery cost + stressed AI usage cost + overhead allocation + target profit
Do not use current promotional pricing as the only scenario.
Sell an allowance, not undefined unlimited usage
Include a measurable amount of usage in the base fee. Charge an overage, upgrade the client to a higher tier or reduce service speed after that allowance is reached.
Describe the allowance in a client-relevant unit such as completed videos, qualified calls, documents processed or conversations handled, not invisible provider credits.
Add a pricing-adjustment clause
Longer contracts should permit a review when third-party platform costs, taxes or required infrastructure change materially. Specify the review frequency, threshold, notice period and customer options.
Have locally qualified counsel review the clause before relying on it.
Build provider flexibility
Keep prompts, customer data, workflow logic and business rules separate from a single model integration where practical. Maintain at least one tested alternative for important services.
Multi-provider architecture adds engineering effort, so it is not automatically worthwhile for every small workflow. It becomes more valuable as client revenue and switching risk grow.
Route work by difficulty
Use smaller, cheaper models for extraction, classification, formatting and routine support. Reserve frontier reasoning for tasks where it materially improves the result.
Caching reusable context, limiting unnecessary output, summarising old conversations and controlling agent steps can also reduce waste without reducing customer value.
Monitor before the invoice arrives
Set provider budgets, alerts and per-client usage dashboards. Review unit economics at least as often as client contracts can be repriced.
Will AI get cheaper or more expensive?
The most defensible forecast is a split outcome.
Equivalent intelligence will probably keep getting cheaper. Hardware performance per dollar is improving, algorithms are becoming more efficient, smaller models are becoming more capable, and provider competition is pushing down prices.
Frontier AI will remain expensive. Advanced reasoning, real-time systems, multi-agent work and high-quality video can use any efficiency gains to perform more ambitious tasks rather than simply lowering the invoice.
Total business spending may continue increasing. When each unit becomes cheaper, businesses often find more uses for it. This is a reasonable economic inference, not a guaranteed forecast.
Open alternatives should strengthen competition, but terminology matters. The Open Source Initiative distinguishes fully open-source AI from open-weight models that release model parameters without the complete training code or data.
Open-weight deployment can reduce dependency and become economical at sufficient volume, but it is not free. The business assumes hosting, engineering, security, monitoring, licensing and reliability costs previously handled by the API provider.
The winning strategy is therefore unlikely to be “always use the cheapest provider.” It will be to route each workload to the least expensive option that reliably achieves the required result.
Key takeaways
- AI is generally becoming cheaper at a fixed capability level.
- New workloads are consuming more compute than older chatbot tasks.
- Credits hide cost unless they map transparently to real usage.
- Never price a client service on the assumption that today’s promotional economics are permanent.
- Protect margins with allowances, monitoring, stress testing, adjustment clauses and provider flexibility.
Frequently asked questions
Are AI credits genuinely becoming more expensive?
Sometimes but not universally. A platform may increase its price or credit burn, while another model becomes cheaper. Compare the cost of the complete successful workflow rather than credit quantities alone.
Why does the same monthly subscription feel more restrictive?
The service may have introduced more expensive models or features, altered usage limits, changed the credit conversion, or moved activity into a shared usage pool.
Should an agency pass every provider increase to its clients?
Not automatically. Small fluctuations should normally be covered by the agency’s margin. Material and sustained changes may justify repricing if the contract permits it and the client receives adequate notice.
How much margin should an AI service include?
There is no reliable universal percentage. Base it on measured usage volatility, support labour, failure rates, contract length and switching risk. Test whether the service remains worthwhile when AI costs rise substantially.
Will open-weight models solve provider dependency?
They can reduce it, especially at scale, but they replace API dependency with infrastructure and operational responsibility. Licensing and model suitability must also be reviewed.
What is the most important cost metric?
Cost per successful business outcome. This includes every model call, generation, tool, retry, failure and human correction required to deliver an acceptable result.
Conclusion
AI is likely to become more capable and more efficient. That does not mean every subscription, credit or client workflow will become cheaper.
Businesses should assume that models, limits and pricing will change. Build sufficient margin, measure real usage, avoid unnecessary provider lock-in and price AI-powered services with enough flexibility to survive the next change.