Episode 45 – The Meter Is Running

1

Twenty dollars for a simple drafting query. Twenty thousand to have agents review 100,000 contracts. Same platform, same contract; the only variable is how many tokens the work burns. The numbers come from Gabe Pereyra, co-founder of Harvey, who adds that the first $10 million AI bills are a matter of time. Legora has already moved: since June, its most advanced agent is priced by consumption, with live dashboards and spending thresholds. The seat licence had a good run. The meter is running now.

Three quarters of a word

A token is the unit a language model actually works in: text chopped into fragments, roughly three quarters of an English word each. Everything is counted. The question you type, the data room you attach, the answer, and with reasoning models the thinking in between, billed even though you never read it. Tokens are the kilowatt-hours of AI: nobody buys electricity by the lamp, and some vendors have decided to stop selling intelligence by the seat. For law firms and legal departments the switch rewrites the budgeting question. A seat licence was predictable: multiply lawyers by price, sign, forget. A meter asks how much work you will run through the machine, and almost nobody holds the data to answer; Harvey says token consumption on its platform grew 14x in six months. The profession has seen this film. Legal research began with metered pricing, per search and per minute, and lawyers learned to dread the terminal. Flat plans arrived to cure the anxiety. Legal AI is now running the film in reverse.

Ask Uber how it ends

Uber shows what happens when the meter meets enthusiasm. Through early 2026 the company urged employees to use AI as much as possible, with internal leaderboards ranking teams by usage. In April, its chief technology officer admitted the entire annual budget for AI coding tools was gone; in June came the caps, $1,500 per employee per month, per tool, on Claude Code and Cursor. And the value? Well, not so easy to find…The trap is structural, because list prices for token keep falling while bills keep rising: reasoning models spend tokens deliberating before they answer, agents turn one instruction into hundreds of billed steps, and cheaper units invite heavier use, the old Jevons paradox. Much of what the tokens buy is waste. Up to 40% of employees’ work may be “workslop”, AI output dressed up as work but too hollow to use. Seat pricing hid all of it inside a flat fee. The meter gives waste a line item, and that is the most useful thing about it. At the end of the day, the real price of a token is the unit price times your organisation’s behaviour, and behaviour is the part nobody manages.

Before the meter starts

Consumption pricing is arriving whether organisations are ready or not. Some firms have moved early and negotiated multi-year plans at locked terms, a sensible shock absorber. Treat it as a seatbelt rather than a guarantee: GPT, Claude and Gemini change models and price lists several times a year, the legal platforms running on them rebuild on whatever ships next, and a three-year contract freezes the fee while everything the fee buys keeps moving. Five moves worth making now, while the contracts are still on the table:

  • Inventory. List every AI tool the organisation pays for and its pricing model: seats, credits, tokens. Include the personal subscriptions people quietly expense; that is where consumption habits are forming.
  • Baseline. Measure current usage before signing anything metered. Thresholds, caps and volume discounts may all be negotiable, but you cannot negotiate around a number you do not know.
  • Training. Teach people which tasks deserve an agent, which fit a chat window, and which stay human. An untrained user was a quality risk; under a meter, an untrained user is a budget risk too.
  • Guardrails. Set spending alerts per team and per matter, and name the person who reads the dashboard every week. Uber introduced caps after the money was gone. The order matters.
  • Attribution. Decide now whether AI is overhead or a matter cost passed to clients, and how that line will read on an invoice. Clients already ask how AI was used on their work; what it cost is the next question.

Our take

Consumption pricing is more honest than the seat licence it replaces: it prices work rather than access, and it exposes habits a flat fee kept invisible. Honesty cuts both ways. Vendors who meter their revenue should expect customers to meter the value, task by task, matter by matter. Tokens will do to legal AI what the timesheet did to legal work: force everyone to say what the tokens were for. Organisations that measure, train and set rules before the switch will negotiate from strength. The rest will learn the real price of a token the way Uber did, at the end of the money.

Do you know what your organisation’s AI actually costs per matter? At Better Ipsum, we help law firms and legal teams get ready for consumption pricing: usage audits, AI literacy programs, spending guardrails and client billing language. Contact us for a consultation

Share:

Subscribe Our Newsletter