Brian Armstrong said most employees never hit usage limits as Coinbase lowered costs through better defaults, routing and caching, with open-source models now set as default options.
Coinbase reduced AI spending by nearly 50% even as token usage grew exponentially, Brian Armstrong said, pointing to efficiency gains rather than tighter restrictions. He said the savings came from improving default model choices, routing requests to lower-cost options, and expanding caching (reusing prior model responses), instead of relying on caps or alerting systems to curb demand. Armstrong added that 91% of employees never reached usage limits, while open-source models such as Zhipu's GLM 5.2 and Moonshot AI's Kimi 2.7 have become default options through an LLM gateway (system that routes AI model requests). He also said LibreChat cache hit rates climbed from 5% to 60%, suggesting a much larger share of prompts could be answered from stored results rather than fresh model calls.