Who Gets How Much AI? Three Answers, Not One

Every company that clamped down on AI spending in 2026 did it in the same order — and got the question wrong. Notes on allocating AI budget across roles.

aimanagementengineeringcost

Everyone Made the Same Mistake in the Same Order

In April 2026, Uber’s CTO told the company it had burned through its AI coding tool budget four months into the year. By June, Uber had capped spending at $1,500 per employee per tool — exceedable with approval — with an internal dashboard showing everyone their own usage. Microsoft cancelled most Claude Code licences across its Experiences + Devices division (Windows, Microsoft 365, Outlook, Teams, Surface) effective June 30, six months after rolling the tool out to roughly 5,000 engineers and watching adoption climb past 84%. Amazon shut down KiroRank, an internal leaderboard ranking developers by token consumption, after engineers started running pointless jobs to climb it. An SVP told staff: “Please don’t use AI just for the sake of using AI.”

Read those in sequence and the pattern is uncomfortable. Every one of these companies did the same four things in the same order:

  1. Distribute access widely
  2. Incentivize usage — leaderboards, weekly targets, adoption percentages in performance reviews
  3. Burn the budget
  4. Clamp down

Not one of them appears to have declared an allocation principle first. The caps are all reactive. Which means the interesting question isn’t “what number did Uber pick.” It’s “what would it look like to decide this on purpose?”

The Usage Metric Was Never the Point

Step 2 is where the damage happened. Ranking people by token consumption is Goodhart’s Law with a fresh coat of paint: when a measure becomes a target, it stops measuring. It is the same mistake as counting lines of code, and it fails the same way — by rewarding volume.

The measured effect size makes this worse. DX’s research across 184 companies and 38,880 developers puts median AI time savings at about 3 hours 45 minutes per developer per week, real productivity gains at 5–15%, and median PR throughput up around 7.8%. Those are good numbers. They are not the 3–10x that vendor marketing implies. A budget built on a 10x assumption doesn’t merely overshoot — it makes every later conversation about cost feel like a failure, when the tool was working about as well as tools work.

The Cost Driver Isn’t Workload

Here is the part I found most useful, and it comes from vendor documentation rather than any think piece. Anthropic’s Claude Code cost guidance names two specific causes of unexpectedly high API bills: long sessions that were never cleared, and leaving the most expensive model as the default.

The mechanism is worth understanding. Claude Code sends the full conversation with every request, and each tool use sends another request carrying that batch of results. Prompt caching means the history is re-read at the cached rate rather than full price — but it is still read. So a one-line question in a session that has been open all day draws usage for the whole conversation. And the first message after a break longer than the cache lifetime misses the cache entirely and reprocesses everything.

If the dominant cost driver is a habit rather than a workload, teaching the habit beats tuning the budget. That reframes the exercise. Before arguing about whether someone deserves $300 or $600 a month, find out whether they clear between tasks and whether they’re running a frontier model to rename variables.

The same docs note that agent teams — multiple parallel instances — use roughly 7x the tokens of a standard session when teammates run in plan mode. Worth knowing before you read a spike as someone’s productivity.

One Name, Three Different Problems

This is what most discussions miss. “AI cost allocation” is one phrase covering at least three structurally unrelated problems, and they have different — sometimes opposite — solutions.

Type A: metered, high-variance. Engineers, data scientists, and any product manager who has started building prototypes. One person’s monthly spend swings by hundreds of dollars, driven not by how much work they did but by how long they let an agent run autonomously. Anthropic’s published enterprise average is about $13 per developer per active day and $150–250 per month, with 90% of users staying under $30 per active day — heavy agentic users land far above that. This is the only category that genuinely needs per-person caps, approval flows, and spend visibility.

Type B: seat-licensed, stable. Sales, customer success, PM, management, HR, finance, field roles. Cost per person is roughly flat, because the work is short interactions rather than long autonomous runs. Designing a cap here is building a control for a problem that doesn’t exist.

Type C: outcome-priced. Front-line customer support. In 2026 this category moved off seats entirely. Intercom’s Fin bills $0.99 per outcome; Zendesk introduced roughly $1.50 (committed) to $2 (pay-as-you-go) per automated resolution and expanded the model at Relate in May. There is no per-person allocation to design. The decision is unit cost times volume versus loaded human cost times volume.

The practical consequence: only Type A needs the machinery. Imposing the same request-and-approve process company-wide taxes the roles it can’t help, and signals distrust for no return.

Engineers and Sales Have Opposite Cost Problems

Within Type B, sales deserves separate treatment, because its cost problem is structurally inverted from engineering’s.

Engineering’s problem is metered runaway: consumption exceeds expectation, so the fix is caps, visibility, and routing cheap work to cheap models. Sales’ problem is stacked seat licences. Reviewers of the Gong-plus-Clari combination report stacks clearing $500 per user per month at a couple hundred reps — and no amount of usage visibility reduces that, because it is already contracted. The fix is a licence audit and consolidation.

Same phrase, opposite prescription. Write one company-wide rule and you will necessarily get one of them wrong.

Inside and field sales also split. Inside sales is the one role where output attributes cleanly to an individual — sends, reply rate, meetings booked, revenue closed — which makes it the only place outcome-linked allocation can honestly work. It is also where the sales version of tokenmaxxing lurks: generate more, send more, watch reply rates collapse. Field sales, by contrast, spends most of its hours travelling and in rooms with people, and has structurally little time to consume AI at all. Give field reps a desk worker’s budget and it simply goes unused.

When the Vendor Defines the Outcome

Outcome pricing looks like it dissolves the gaming problem. It relocates it.

Read the two vendors’ definitions side by side. Fin’s pricing page states that resolutions, procedure handoffs, and disqualifications are each $0.99, and defines a billable resolution as no further help being requested after Fin’s last answer. That is a passive definition — a customer who gives up and leaves satisfies it. Note also that handoffs bill, meaning a conversation Fin escalated to a human still costs $0.99. Zendesk went the other direction, announcing that billed resolutions are verified twice: once by the agent completing the interaction, and again by a separate AI evaluation model.

Two implications. First, estimating spend as “resolution rate × unit price” understates it wherever non-resolution events are billable; the right shape is unit price × billable events. Second, the audit target has moved. Under seat-and-token pricing you audit whether usage metrics are being gamed. Under outcome pricing you audit whether the vendor’s definition of success matches yours — which means sampling conversations and having a human judge them, early.

The Research Points Against the Intuition

The intuitive allocation gives more budget to more senior, higher-paid people. The best available evidence points the other way.

Brynjolfsson, Li and Raymond studied 5,179 customer support agents receiving a generative AI assistant. Productivity — issues resolved per hour — rose 14% on average, but the gain was concentrated: +34% for novice and low-skilled workers, with minimal effect on the experienced and highly skilled. The mechanism they describe is the model propagating the practices of the best workers, letting newer ones move down the experience curve faster.

Dell’Acqua and colleagues ran a field experiment with 758 BCG consultants. On 18 tasks inside the model’s capability frontier, AI users completed 12.2% more tasks 25.1% faster with better quality. On a complex managerial task deliberately chosen to sit outside it, they were 19 percentage points less likely to reach the correct answer.

Put together: the people who gain most are the least experienced, and the thing that protects you at the frontier’s edge is expertise. So “more for juniors, protected review time for seniors” is defensible — but only where the review actually happens. Hand a large budget to someone who cannot yet tell when the model is confidently wrong and you have bought volume, not output.

Two honest caveats. Both studies compared having access against not having access; neither tested budget size. The step from “access helps novices” to “give novices bigger budgets” is an extrapolation. And Brookings’ framing is worth holding onto: today’s gains borrow expertise accumulated before these tools existed. If juniors never build that expertise, the effect may not survive the generation that had it.

What I’d Actually Do

  1. Sort roles into A, B and C before designing anything. Only A gets per-person machinery.
  2. Make consumption visible before capping it — including to the individual. Self-awareness moves the number on its own. Check what your contract actually exposes: routed through a cloud provider, the vendor’s own analytics may not reach you at all, and you would need OpenTelemetry or a gateway.
  3. Set caps as budgets, not targets. No leaderboards. No usage in performance reviews.
  4. Teach the habits — clear between tasks, match the model to the job — before adjusting anyone’s number.
  5. For Type B, audit contracts rather than usage.
  6. For Type C, audit the definition of the outcome.
  7. Measure cycle time, rework and satisfaction. Never lines of code or raw PR counts, which AI inflates automatically.
  8. Revisit quarterly. Annual budgets are how you find out in month four.

What Nobody Has Answered

The fairness question is still open, and I don’t think it has a technical answer. A Type A engineer might reasonably need $500 a month while a field sales rep needs $30. That is correct on the merits and reads as a perk gap. The only coherent position I can construct is to say out loud that fairness here means sufficiency for the work rather than equality of amount — and I could not find a company that has said it.

Individual ROI is unresolved too. Spend is measured precisely; time saved is mostly self-reported. Moving someone’s budget on a ratio with a precise denominator and a vague numerator is how you end up back at step 2.

One closing note on evidence. Uber’s $1,500, Microsoft’s cancellation and Walmart’s cap are press reporting, not company statements: the scope, the exceptions and what happened next are unknown. Neither Uber nor Microsoft has published whether output fell after they clamped down — which is, of course, the one number that would settle the argument.

← Back to Notes