There’s a number in the top-right of my screen right now, sitting in the macOS menu bar next to the Wi-Fi icon and the clock. It says $198.24 · 211.25 MTok. That’s how much I’m projected to have spent on LLMs this week, and how many tokens that bought, updating to the cent while I work. I wanted visibility for myself, so I added this cool little menu-bar tool that gives me real-time statistics about my Claude usage, using Raycast’s ccusage extension. People’s menu bars usually tell them the time and if they’re running out of battery soon. Mine also tells me what the week is costing.

I can see that number because that’s part of my job. I own finance and operations at a small cross-border AI company, and I’m the admin for our AI tools, so the spend passes in front of my eyes all day. The other day Anthropic’s console pulled up a warning that the org was about to hit its spend limit. When I looked at the usage, a colleague and I were among the heaviest users, both working on an internal project. A project with a deliverable where the numbers just had to be right, so we’d pointed Fable 5 at it and let it work – multiple agents in parallel. That was expensive, but it was deliberate spend on a job that doesn’t allow made-up figures.

When my finance peers talk about the cost of AI, they usually mean COGS: what the customer’s usage costs us to serve, and whether we pass it through or eat it. That number shows up in gross margin, and investors ask about it. Internal AI spend is different. It covers what the team spends while building the product, and it still sits outside most reporting. Usually it appears as a credit-card charge someone reconciles at month-end, after the work is done and the context is gone.

Before deciding who should pay for a unit of AI work, a company needs to know what that unit costs. The cap stories all start in the same place: the company can see the spend, but can’t explain it.

What the cap stories have in common

Take Uber, because it’s the cleanest example and it’s been discussed widely. They rolled Claude Code out to roughly 5,000 engineers starting in late 2025, and by April 2026 they’d burned through their entire 2026 AI budget, in about four months (Forbes, 2026). Their response was to cap employees at $1,500 a month per tool, tracked on an internal dashboard, exceedable with permission (TechCrunch, 2026). Their COO, Andrew Macdonald, said it’s “very hard to draw a line between one of those stats and ‘Okay now we’re actually producing like 25% more useful consumer features’” (Fortune, 2026).

Amazon, Walmart, Cisco, and Meta have also introduced restrictions, from token caps to cheaper-model defaults. Walmart capped tokens on its internal build-with-AI platform after usage took off. Amazon warned staff to stop using AI “just for the sake of using AI.” One widely quoted line, “we created a monster,” came from an unnamed software-company CIO whose bill jumped sevenfold in a single day when the vendor switched it from flat to token-based pricing (Financial Times, 2026).

In each case, adoption moved faster than the reporting around it. The usage data arrived with names, models, and tokens, but no useful account of what the work produced. A cap limits the number. It does not explain where the number comes from.

At Uber’s scale, I can understand the cap. Thousands of engineers can trigger workflows that fan one prompt out into an orchestrator, retrievers, and multiple model calls. The important cost may sit several layers below the person who started the run, inside a call graph nobody is watching live (Vantage, 2026).

A per-person limit buys time while the company figures out which uses are worth paying for. It also risks penalizing work that creates value, because usage gets measured before outcomes do. Uber’s adoption reportedly went from a third of engineers to most of them in about a month, and big parts of committed code were originating from those tools (Forbes, 2026). The spend needs to be tied to specific outcomes before anyone knows what to cut.

Why the bill grows while prices fall

Per-token prices have been falling across commodity and mid-tier models. The bill can still rise because teams use cheaper tokens more often, and because agentic workflows consume more tokens per task than a one-prompt answer. Frontier reasoning models add another source of growth through a higher price per token. People are running out of credits a few days into the new month. I’ve been there.

For an AI company, inference is one of the biggest lines in the P&L. AI gross margins are running around 52% in 2026, up from 45% the year before, with inference eating roughly 23% of revenue at scaling-stage B2B companies (ICONIQ via SaaStr, 2026). a16z made the same broader point back in 2020: AI businesses run 50 to 60% gross margins against software’s 60 to 80% because inference is a variable, per-use COGS rather than a fixed cost you amortize (a16z, 2020). Roughly $230,000 of every million in AI revenue slips out as inference before people get paid.

At traditional SaaS companies, separating customer usage from internal usage is mostly a bookkeeping exercise. At an AI company, customer-driven inference moves with usage. Mix the two together and the margin you report becomes an estimate.

The missing view

I haven’t figured it out entirely yet. My company is working through a similar challenge at a much smaller scale. We already track customer-driven AI spend in a dashboard. The internal usage view still needs work. I know the number because I run the operating backbone and own the tools. We can break usage down by product and model. Getting that in front of the people spending it is the part I’m still working on.

The useful version should run on a schedule. My self-built finance KPI dashboard already handles the monthly bank reconciliation automatically. Internal AI spend should work the same way: pull the provider data, reconcile it, and update a shared view showing spend by person, product, and model. I’m exploring how to automate that from the existing admin panels.

The FinOps Foundation puts visibility before allocation, optimization, and governance. Ninety-eight percent of FinOps teams now say they manage AI spend, up from 31% two years earlier (FinOps Foundation, 2026). The survey says companies are managing the spend. It says nothing about whether the people creating it can see it.

I spend company money like it’s my own. That’s a side effect of more than a decade in finance. I watch cash leave the bank account. The person using an AI tool usually sees the result, not the cost of the attempt.

A cap tells someone when they’ve run out of credits. It doesn’t show what the spend produced, whether the run was wasteful, or whether it supported the one deliverable that had to be right. We’ve started with the low-tech version, which is mostly me telling the team when the spend is climbing. We’re still learning what’s driving the extra tokens. The models have changed, and so have the ways people use them. I want to know how much of the increase comes from each before telling people where to cut. I’d rather share the numbers than tell a team of adults they’ve hit their allowance and have to wait for the limit to reset.

What I still don’t know

There’s a second thing I’m curious to check. I’d call it a guess rather than a fact, because it’s based on what I’ve noticed and I haven’t checked the data. I’d guess people who came to these tools without an engineering background run up more cost per result.

My evidence is embarrassingly small, one comparison, me against a friend who writes code for a living. Especially in my early AI-days, just about when OpenAI released their first ChatGPT model, I noticed that when I wanted the model to change something, I’d send it a screenshot and describe what I wanted in a paragraph, because I didn’t always have the technical jargon for, e.g., the specific area of a website. He changed the same thing with three keywords. His prompts were shorter, his context was tighter, and he didn’t make the model re-read the whole thread to remember what it was doing. Mine were longer and sloppier and cost more for the same result.

If that holds beyond the two of us, the non-engineering part of the bill may be larger than anyone expects. That’s a hypothesis, not a real number. I can’t test it until the data is aggregated and reconcilable.

Most of the cost-lowering moves are very basic. Smaller, sharper prompts. Fewer output tokens; I ask for caveman-style answers. Don’t force the model to re-read an entire conversation it already knows. Don’t switch models mid-task, because that triggers a full context re-read. Burn your provider credits before you add cash by pointing an API key at them first.

We’ve started working through those as well. I don’t know which changes matter yet, because we don’t have a shared view to compare them against.

Customer pricing comes after internal accounting. Before deciding what to charge, we need to separate customer-driven inference from what our own team burns. The companies capping spend are trying to manage a number they can’t yet explain. I’d rather spend a few more tokens making the internal view run automatically than apply a hard cap. Might share how it goes.