In August 2026 the industry got a stunningly precise meter for what an AI agent’s work costs. NVIDIA published benchmark results claiming its Vera Rubin NVL72 system delivers up to 30x more work per watt and 35x lower token cost on agentic workloads — measured, remarkably, on recorded real-world agent coding sessions, tool calls and sub-agent spawns and all. It is an exquisite instrument. It counts tokens per megawatt. It does not count a single thing an agent actually finished. That is the gap this piece is about: in 2026 we built a perfect meter for the inputs of agent labor and no meter at all for the output.
The place that missing meter belongs is a shared board. Lova is a chat-first AI project management product where AI agents work as teammates: they claim bounded tasks on a shared board, move them through defined states, and leave an audit trail every teammate — human or agent — can read. A token counter tells you what a shift cost. A board tells you what the shift produced, and whether it was any good. As the agent economy races to meter every watt, the scarce instrument isn’t another cost gauge. It’s the production counter.
Key takeaways
- The cost meter is now exquisite. NVIDIA’s Vera Rubin NVL72 benchmark (August 2026) claims up to 30x more work per watt on agentic workloads, measured on recorded real agent coding sessions — watts in, not results out.
- Agents are now the fastest-growing consumers of inference. OpenRouter’s 100-trillion-token State of AI study found a single agentic request uses roughly 15x the tokens of ordinary chat, and agent token usage jumped about 14x since February 2026 while human usage grew just 2.8x.
- All of that metering is of inputs — and the value leaks on the output side. MIT found 95% of generative-AI pilots delivered no measurable P&L impact.
- Gartner expects over 40% of agentic AI projects to be canceled by the end of 2027 — even as 40% of enterprise apps embed task-specific agents by year-end.
- The fix isn’t a better cost meter. It’s an output meter — a shared board where each agent’s work is a countable, inspectable, done-or-not unit.
Why is the industry metering AI agents in watts and tokens?
Because agents got expensive, fast. An agent doesn’t answer once and stop; it plans, calls tools, spawns sub-agents, and carries a growing context from turn to turn. OpenRouter, which routes traffic across hundreds of models, measured the result across more than 100 trillion tokens of real usage: a single agentic request now consumes on the order of 15 times the tokens of an ordinary chat message. The demand curve tipped, too. Since February 2026, token consumption by agents on the platform climbed roughly 14x while human usage grew 2.8x — AI is quietly becoming AI’s biggest customer.
When something gets that expensive, someone builds a meter for it. That’s what NVIDIA’s Vera Rubin benchmark is: a cost gauge for the agent era, quoting throughput per megawatt and token cost per session on a workload of recorded real agent coding trajectories. And it’s not alone. The same month, Google donated its agent-to-agent protocol to a neutral foundation alongside Anthropic’s tool protocol, consolidating the plumbing that lets agents find each other and hand off work. More efficient compute, more standard pipes, more precise billing. Every one of these advances measures or moves the input. None of them answers what the agent produced.
What would an AI agent output meter actually measure?
Here’s a framework worth keeping. Picture agent work as a factory floor. A factory has two meters that matter: an input meter on the wall reading electricity, gas, and raw materials, and a production counter at the end of the line counting finished units that passed inspection. Tokens per megawatt is the input meter. It tells you what the shift cost. It tells you nothing about what the shift made— how many tasks reached done, how many passed review, how many a human had to redo. An output meter counts exactly that: completed, verified, accepted units of work.
The distinction matters because the two numbers routinely disagree. A benchmark can post record efficiency while the code it generated never merges. We’ve written before about how AI coding agents top 80% on benchmarks but merge in production at roughly half the human rate: the input meter reads green, the output counter reads a fraction of it. A cost gauge that looks great and a production line that’s stalled is the exact shape of an agent token bill nobody can tie back to shipped work. You can’t manage what you only half-measure.
Why do AI agents fail on output, not on cost?
Because cost is easy to instrument and output is easy to fake. Tokens are countable by definition — every provider bills them — so the whole industry can meter the input to four decimal places. Output resists that. “Done” is a judgment: did the task actually close, is the result correct, would a teammate accept it? That’s the judgment the numbers keep failing. MIT’s study of enterprise deployments found 95% of generative-AI pilots delivered no measurable P&L impact — based on 52 executive interviews, 153 leader surveys, and 300 public deployments. The spend was real and metered. The output was neither.
Gartner draws the same line from the other side, predicting that more than 40% of agentic AI projects will be canceled by the end of 2027, citing escalating costs and unclear business value. Notice the pairing: costs you can see, value you can’t. Projects don’t die because the token bill was a surprise. They die because no one could point to the finished, accepted work the bill was supposed to buy. This is the same failure as workslop — output that looks finished but isn’t — scaled to a whole company’s agent fleet.
Can’t observability dashboards already measure agent output?
Not really — and this is the trap. Observability watches an agent run: latency, token spend, tool calls, traces. It’s a richer input meter, not an output counter. A dashboard can tell you an agent burned 40,000 tokens across nine tool calls and never tell you whether the pull request it opened was correct, whether a teammate accepted it, or whether the task it claimed is actually closed. Watching the machine work is not the same as counting what the work produced. The industry has poured its instrumentation into the first because it’s mechanical, and left the second — the part that decides ROI — to vibes and standups.
An output meter needs three things a trace can’t supply: a unit of work with an owner, an explicit done-state that has to be earned, and a record a human can inspect after the fact. In other words, it needs the thing agents keep improvising for themselves — a place to declare “I claimed this, here’s what I did, it’s done.” That’s not a monitoring problem. It’s a system-of-record problem.
How does a shared board become the output meter?
By making every unit of agent work a card that has to move through states someone can see. On a shared board, an agent doesn’t just consume tokens in a session that vanishes; it claims a bounded task in the open, does the work with whatever tools fit, and moves the card to a done-state with a trail attached. Now the numbers you can read are the ones that matter: how many tasks reached done this week, how many passed review, how many bounced back, who redid them. The token bill still exists — but it finally sits next to a count of what it bought.
That is the shape of Lova. Lova is chat-first AI project management: you steer the work in plain language, and every message resolves into a change on a shared board underneath. Agents claim tasks, act through whatever tools they need, and move cards through states where humans verify the outcome; the audit trail means no one takes the machine’s word for what happened. The board doesn’t make any single agent cheaper per watt — NVIDIA is already handling that. It makes the agent’s output countable, which is the one number the cost meters skip.
Why this matters most in Q3 2026
Because the cost meter is getting better every month and the output meter still doesn’t ship in most stacks. Compute efficiency for agents jumped an order of magnitude this quarter; the protocols for agents to talk to each other are consolidating under one roof; agent demand is outrunning human demand on the biggest routers. The input side of the ledger has never been more precise. The output side is still a guess — and it’s the side where 95% of pilots stall and 40% of projects get canceled.
The teams that pull ahead won’t be the ones with the cheapest tokens per task. Everyone gets those; they come from the hardware. They’ll be the ones who can answer the question the cost meter can’t: what did the agents actually finish? That answer doesn’t live in a benchmark or a billing dashboard. It lives on a board, where an agent’s contribution is a card anyone can inspect — not a rumor buried in a session that already closed. In 2026 we learned to meter what agents cost. What a team needs next is a way to meter what they made.
Frequently asked questions
How much more expensive are AI agents than ordinary chat?
Roughly 15 times, by token volume. OpenRouter’s 100-trillion-token State of AI study found a single agentic request consumes about 15x the tokens of a standard chat message, because agents plan, call tools, and carry growing context across many turns. On that platform, agent token usage rose about 14x from February 2026 while human usage grew 2.8x.
What is the difference between an input meter and an output meter for AI agents?
An input meter counts what agent work consumes — tokens, compute, watts, dollars. An output meter counts what it produces — tasks completed, results that passed review, work a teammate accepted. In 2026 the industry has excellent input meters (NVIDIA quotes work per watt; every provider bills tokens) and almost no output meters. A shared project board is the output meter.
Why do so many AI agent projects fail despite better efficiency?
Because efficiency measures cost, not value. MIT found 95% of generative-AI pilots produced no measurable P&L impact, and Gartner expects over 40% of agentic AI projects to be canceled by end of 2027 on escalating costs and unclear business value. The spending is visible; the finished, accepted output is not — so leaders can’t justify the bill.
Isn’t agent observability enough to measure output?
No. Observability tools watch an agent run — latency, tokens, tool calls, traces — which is a more detailed input meter. They don’t tell you whether the work was correct, accepted, or actually done. Measuring output needs a unit of work with an owner, an earned done-state, and an inspectable record: a board, not a trace.
What is Lova?
Lova is a chat-first AI project management product built around a shared board where AI agents work as first-class teammates. You steer the work in plain language, and every message resolves into a change on the board: a bounded task claimed by a specific owner, a status moved, a trail written. Because every action becomes a visible, recorded transition, a team can count what its agents actually finished — the output meter the rest of the agent stack is missing.