Uber says AI token efficiency is replacing "tokenmaxxing"

After exhausting its 2026 AI budget within months, Uber says broader use of frontier tools is now lowering cost per token through prompt caching, model changes and tighter usage tracking.

Summary

Uber says it is moving away from "tokenmaxxing" after exhausting its entire 2026 artificial intelligence budget within the first few months of the year, and now argues that wider use of frontier AI tools can reduce unit costs rather than simply inflate spending. Chief Technology Officer Praveen Neppalli Naga said Uber had pushed employees to use AI coding tools heavily, particularly Anthropic's Claude Code, and even created internal leaderboards ranking software engineers by usage before reworking the strategy once costs rose too quickly. On Wednesday, he said Uber had quadrupled the number of employees using frontier AI tools while lowering cost per token by improving prompt caching, adjusting default model settings, evaluating new models for efficiency and giving engineers visibility into their AI usage and hourly costs. The shift comes as companies face growing pressure to show returns on large AI investments. Uber President and Chief Operating Officer Andrew Macdonald said in May it was still difficult to draw a direct line between AI use and materially higher output of useful consumer features, while Jim Reid, global head of macro and thematic research at the Deutsche Bank Research Institute, warned last month that broad productivity gains from AI may still be years away. Uber's lower token costs may also run into Jevons paradox: the price of a single token has dropped more than 90% since 2023, but spending on large language models has doubled since late last year, and Bain and Co. said token costs halved from December 2024 to 2025 while tokens consumed surged 450% as companies upgraded AI capabilities.

Terms & Concepts
  • tokenmaxxing: A corporate AI usage pattern that rewards heavy token consumption without clear focus on efficiency or returns.
  • prompt caching: A technique that stores and reuses repeated prompt data to reduce the cost of AI processing.
  • Jevons paradox: An economic effect in which efficiency gains lower unit costs but can increase total usage and spending.