This week Tesla started capping employee AI token spend at $200 per week, with overages requiring manager sign-off, according to an internal memo reported by The Information. Earlier this year, Uber reportedly burned through its entire multi-billion-dollar annual AI budget in about four months after rolling AI coding tools out to thousands of engineers without per-user guardrails. GitHub moved Copilot toward usage-based billing after flat-fee plans turned into unlimited liability at scale. The AI cost story of 2026 has a shape now, and mid-market companies should read it carefully — because the same dynamics hit smaller budgets harder.
Why AI spend runs away
Metered billing meets enthusiasm. Token pricing means the bill scales with usage, and usage scales with success. The better your rollout goes, the faster the meter spins — the opposite of most software, where adoption is free after the license.
Agents multiplied the meter. A person asks a question and reads the answer. An agent loops: it plans, calls tools, reads results, retries, and reasons over growing context — consuming tokens at every step, sometimes for hours, sometimes overnight. One ambitious automation can out-spend a whole department of chat users.
Usage targets made it worse. Companies that set AI adoption quotas discovered employees and systems optimizing for measured volume rather than value — the phenomenon the industry took to calling "tokenmaxxing." When the metric is consumption, consumption is what you get.
What the caps get right — and miss
A spend cap is a tourniquet: correct in an emergency, not a governance model. The $200 ceiling stops the bleeding but treats every user and workload identically — throttling the engineer whose agent produces real return along with the one whose scripts run in circles. The durable fix is knowing which spend produces value, and that requires instrumentation most companies never installed: per-user and per-workflow attribution, cost-per-outcome tracking, and someone who reviews it monthly with authority to act.
The mid-market translation
You likely aren't spending Uber money — but proportionally, an unmonitored AI bill can distort a mid-market IT budget faster than it distorts a Fortune 500's. The playbook that works at any size: budget per workflow, not per company (a number attached to a use case is manageable; a blob is not); guardrails per user from day one (limits, alerts, and approval paths before the surprise invoice, not after); route workloads to cost-appropriate models (frontier models for frontier problems; smaller or local models for the repetitive volume work where they're nearly free); and a monthly ROI report someone signs — spend, usage, outcomes, and the recommendation, in writing.
That last item is literally a deliverable of Managed AI Operations — the monthly report before any renewal conversation — and it's why the practice exists: AI without cost accountability isn't a capability, it's a liability with good demos. Seat-level discipline is covered in why AI isn't for every employee. Or book a briefing before your Q3 invoice writes this post for you.
