
An internal Amazon project burned through 1.8 million dollars using Claude, Anthropic's AI model, for a result that never shipped. The episode, reported on July 30, 2026 by the Financial Times, illustrates a simple risk: without AI cost governance, spending can spiral fast, even inside a company as technically sophisticated as Amazon. For any small or mid-size business scaling up AI agents, the lesson is worth taking seriously before the surprise bill lands, not after.
In brief
- An internal Amazon project meant to automatically match authors to book listings cost 1.8 million dollars in Claude usage, 860% over its original budget, and was never shipped, according to the Financial Times (July 30, 2026, reporter Rafe Rosner-Uddin).
- Two other internal projects overran in similar ways: roughly 541,000 dollars in unplanned costs on a financial auditing tool, and 134,000 dollars on a logistics system.
- These overruns stayed invisible for five months before anyone caught them, for lack of centralized AI cost tracking.
- Amazon had built an internal leaderboard ranking employees by AI token consumption, nicknamed KiroRank, which teams gamed by running pointless tasks to inflate their scores ("tokenmaxxing"). The company scrapped it.
- Amazon senior vice president Dave Treadwell told staff in a memo, "please don't use AI just for the sake of using AI." The company now tracks value delivered instead of raw token volume.
What happened at Amazon
According to several internal sources cited by the Financial Times, Amazon engineers presented examples in a staff meeting of generative AI projects whose costs had run completely out of control. The most striking case involved a project meant to automatically match authors to book listings on the retail platform. Using Claude Sonnet for this task, the team spent 1.8 million dollars, or 860% above the original budget. The project was never deployed.
Two other examples cited show this was not an isolated incident: an internal financial auditing tool generated roughly 541,000 dollars in unplanned costs, and a system meant to shorten delivery times went 134,000 dollars over budget. In each case, the overrun stayed invisible for months, for lack of a consolidated dashboard tracking real-time consumption. One employee quoted by the Financial Times summed it up: "it's difficult to figure out how much anything costs."
The problem behind the problem: measuring the wrong metric
The episode is not just a budget-tracking failure. It also reveals an incentive problem. To speed up internal AI adoption, Amazon had built a leaderboard, KiroRank, that tracked how many tokens each developer consumed on its internal tools (Kiro, MeshClaw). The more AI a staff member used, the higher they ranked, regardless of the actual value produced.
The result: some teams ran pointless tasks simply to boost their score, a practice employees internally nicknamed "tokenmaxxing." That pressure mechanically inflated the bill, with no link to real productivity. Amazon eventually scrapped the leaderboard. In an internal memo, senior vice president Dave Treadwell acknowledged the tool was built with "good intentions" but asked teams to "please don't use AI just for the sake of using AI." The company says it now tracks a different metric, "normalised deployments," meant to reflect productive AI usage rather than raw token volume.
Key takeaway
A metric that rewards AI usage volume, with no link to value produced, mechanically pushes toward overconsumption. That is true at a cloud giant like Amazon, and just as true at a ten-person business.
Why this should matter to a small or mid-size business
A small or mid-size business obviously does not handle Amazon's volumes. But the mechanism is identical, at any scale: modern AI agents and coding assistants (Claude, ChatGPT, Cursor, Copilot) are billed by usage, by the token or the request, not at a fixed flat rate. An internal project handed to an AI agent with no planned budget and no checkpoint can, in theory, run in circles for weeks on a poorly scoped task, with nobody noticing until the monthly bill arrives.
The risk grows with agent autonomy. A coding assistant that reruns tests on its own, regenerates content, or explores multiple paths to solve a problem can consume far more tokens than an equivalent manual task, without that being visible task by task. That is exactly the scenario in Amazon's 1.8-million-dollar project: a task seen as minor, left with no cap and no interim review for five months.
| Amazon project | Overrun observed | Root cause identified |
|---|---|---|
| Author-to-listing matching (Claude) | $1.8M, 860% over budget | No cap, no review for 5 months |
| Internal financial auditing tool | $541,000 unplanned | AI usage costs not budgeted for |
| Delivery-time reduction system | $134,000 unplanned | No consolidated tracking |
An AI cost governance method, sized for a small business
You don't need a dedicated finance department to avoid this kind of overrun. The principles Amazon now applies, after the fact, can be put in place upfront in a small or mid-size business.
Set a budget before launching an AI project
Turn on consumption alerts
Name someone responsible for the monthly review
Avoid pure volume metrics
Measure value delivered, not tokens consumed
This method echoes an old IT management rule: what isn't measured tends to escape control. Agentic AI, billed by the token and able to act autonomously over long sequences, makes that rule more urgent than ever.
FAQ
Why did Amazon spend $1.8 million on an AI project that never shipped?
A team assigned Claude, Anthropic's AI, to automatically match authors to book listings, with no spending cap and no interim review. The overrun stayed invisible for five months, reaching 860% of the original budget, according to the Financial Times (July 30, 2026).
What was the "tokenmaxxing" revealed at Amazon?
It's the internal nickname for running pointless AI tasks to artificially inflate one's score on KiroRank, an internal leaderboard ranking employees by token consumption. Amazon scrapped the leaderboard after finding it encouraged overconsumption rather than productivity.
Is a small or mid-size business really at risk of this kind of overrun?
Yes, as soon as it uses autonomous AI agents billed by usage, for code, content generation, or research. Without a planned budget and monthly tracking, a poorly scoped task can consume far more tokens than expected, with nobody noticing until the bill arrives.
How can a small business avoid an overrun like Amazon's?
By setting a budget before launching agentic AI use, turning on consumption alerts, naming someone responsible for monthly cost review, and measuring value delivered rather than raw token volume.
Conclusion
The Amazon episode is not an isolated case of a poorly organized big company: it's an early warning for any organization adopting agentic AI without tracking its cost. A small business doesn't need Amazon's volumes to feel the same mechanism, a poorly scoped project, left with no cap and no review, can inflate the bill much faster than expected. The good news: the fixes are simple, a budget per project, alerts, a regular review, and a value metric instead of a volume metric.
To go further, browse our other AI resources for business leaders and see real company case studies that have structured their AI usage.


