AI is becoming a bargain hunter's market, with a few luxury models on top
Which part of your AI spending is earning its keep? This article from The Register breaks down the split between cheap commodity inference and rising frontier model prices, along with usage data showing where extra tokens stop improving output. Read the article for a sharper view of AI economics.
Why are AI token prices moving in two opposite directions?
The AI market is effectively splitting into two tiers: commodity inference and frontier models.
On the commodity side, prices have dropped sharply. For example, GPT-4-class output that cost about $20 per million tokens in late 2022 is now closer to $0.40 for equivalent capability — roughly a 55x decline in under four years, based on Introl’s December 2025 unit-economics analysis. When DeepSeek launched its R1 reasoning model at $0.55 per million input tokens and $2.19 per million output tokens, versus OpenAI’s o1-preview at $15 input and $60 output, the market effectively “repriced overnight” on the back of a ~97% discount.
At the same time, prices for cutting-edge frontier models have risen. OpenAI reportedly doubled GPT-5.5 pricing to $5 input and $30 output per million tokens. Google’s Gemini Flash 3.5 arrived at 3–6x the cost of the model it replaced. Anthropic’s newer models, like Claude Sonnet 5 and its Mythos and Fable lines, also sit at the higher end, especially once you factor in that some models use more tokens to reach the same result.
The net effect: basic inference is becoming a low-margin, almost commodity service, while the most capable frontier models are priced as premium, specialized tools. Buyers now have to decide when they truly need frontier performance and when a cheaper, near-parity model is “good enough.”
How are AI pricing changes impacting company budgets and productivity?
Many organizations are discovering that AI spend can grow faster than the value it creates if it isn’t managed carefully.
According to Ameya Kanitkar, CTO of AI measurement platform Larridin, AI costs were initially modest — often $20–$100 per month per LLM subscription. Around early 2025, as vendors pushed for more usage and models became capable of longer, more complex “agentic” tasks, costs rose sharply. Larridin has seen AI costs increase by about 10x between January and mid-year in some engineering operations.
In concrete terms, some companies are now spending 10–20% of labor cost on tokens. For a software engineer earning $200,000 annually, that can mean $2,000–$4,000 per month in AI token spend alone.
However, more spend does not automatically mean more output. Larridin’s data shows:
- 15–30% of AI users account for over 50% of total AI spend.
- Beyond an inflection point at about 35–40% of client AI spending, additional token burn no longer correlates with higher developer productivity.
Using that inflection point as a soft cap per employee, Kanitkar reports that companies can cut AI costs by around 40% without changing tools or workflows — simply by limiting overuse.
In parallel, pricing models are shifting. Anthropic, for example, has moved corporate customers from per-seat to metered pricing and tightened the permitted uses of subsidized subscription plans. This encourages more deliberate usage and makes cost management a core part of AI strategy.
How can we optimize our AI stack for cost without sacrificing capability?
Companies are starting to rethink their AI stack design to balance cost, capability, and reliability. Several practical levers are emerging:
- Use multiple models instead of a single default
Larridin is seeing about 75% of companies using more than one model. A common pattern is:
- Use lower-cost open-weight or commodity models for routine tasks.
- Reserve premium frontier models for complex reasoning, high-stakes decisions, or customer-facing experiences.
Larridin’s data shows that, despite higher prices, enterprises still direct almost half of their AI spend to Anthropic’s Opus because it handles complex engineering and reasoning tasks well. That suggests a tiered approach: pay for top-tier performance where it truly matters.
- Leverage open-weight models as a cost lever
Open-weight models like Kimi 2.6/2.7 and GLM 5.2 are now “almost at parity” with Anthropic’s Opus 4.7/4.8 in many scenarios, according to Kanitkar. They are:
- ~10x cheaper in theory, and about 5x cheaper in practice.
- Sometimes slower and more token-hungry, but with a low per-token price that keeps total cost down.
These models are particularly attractive for internal engineering tasks, batch processing, and experimentation.
- Set data-driven token budgets
Instead of arbitrary limits, use productivity data to define thresholds. Larridin’s analysis found that beyond 35–40% of total AI spend, extra token usage stops improving developer output. Using that as a per-user or per-team budget can reduce AI costs by roughly 40% without changing tools, simply by curbing low-value usage.
In short, a cost-conscious AI strategy today typically involves a mix of open-weight and frontier models, clear usage policies, and ongoing measurement of how token spend maps to real productivity gains.
.jpg)
AI is becoming a bargain hunter's market, with a few luxury models on top
published by Aksara United Management Sdn Bhd

As an ICT Solution Provider, We at AKSARA UNITED MANAGEMENT SDN BHD consistently strive to provide excellent service to our clients in delivering Management and Personal Development Services, Professional Best Practice , IT Technical Consulting & services, Business Application Development & Government Incentive Schemes.
Our expertise range from Microsoft Technologies, Cisco Networking, Amazon Website Services (AWS), Comptia Certification, JAVA, Sun Solaris Linux/Unix, Oracle, Security , Programming Languages, Web Development. Mobile Content Development , Digital Marketing Courses, and other cutting edge technologies.
AKSARA UNITED MANAGEMENT SDN BHD training centre is located at the heart of Kuala Lumpur with a well equipped state of the art IT equipment which allows our clients to have maximum hands-on practice. Our classrooms are also comfortably furnished with multi-platform, networked computers, integrated softwares & fully functional data centre with multiple arrays of real servers for our participants in order to expose themselves in relevant industries. In our endeavour to ensure all attendees fully understand and are able to use the new knowledge acquired, post class support or courses attended is made available via e-learning assistance & mentoring session.