A serendipitous seat at a fireside chat with Boris Cherny, the creator of Claude Code, led to the clearest read I have on why the AI bubble pops: we have handed the word “return” to the people who build the tools, and they are getting it wrong.
I was a few rows back at Meta’s @Scale conference, watching Jesse Chen interview Boris Cherny. The room was full of builders and the energy was real. What struck me first was how often Boris came back to ROI on his own. He was not dodging the money question, he was leading with it. Talking about why he runs the most expensive model on everything, he put it plainly:
it actually maybe comes back to ROI.
He is right that it does. Then he showed the room what he means by it. He asked for a show of hands: who has stopped writing code by hand entirely? A lot of hands went up. He nodded, and gave the number behind it.
80 to 90% of the code is written by [Claude Code] on average.
For a growing number of teams, he added, it is all of it. The room loved that number. And somewhere in the applause, the thing clicked.
Every metric he reached for was an engineering metric. Code velocity. Tokens consumed. The ratio of agent-written code to human-written code. These are real, and they are the right metrics if you build and sell the tool. They are a beautiful product scorecard on usage and utility. They are not a business scorecard about the return a shareholder cares about. The gap between those two things is the whole story, and right now there are billions of dollars of AI spend sitting on the wrong side of it.
The math a board actually runs
Walk it the way a CFO has to. You spend a dollar on AI agents to build your code and it replaces a dollar of human engineering. Same output, swapped sources. Your return is zero. Nothing about your cost structure changed, nothing about your revenue moved. For the spend to earn its place it has to do one of two things: take out more cost than it adds, or bring in revenue the old setup never could.
That sounds obvious written down. It is shockingly easy to lose in a room full of green checkmarks. The faster, cheaper code feels like the win. But faster code is the thing you bought, not the thing you were buying it for. What you were buying it for shows up later, somewhere a finance team can actually see it.
Every tool was a marvel. Every return had a condition.
Boris walked through the tools, and each one is a small marvel. Take Cowork, which he described simply:
Cowork is Claude Code for non-engineers. It’s just Claude Code under the hood.
He uses it for everything, including his own travel. It reads his calendar, finds the trips, and books the flights and hotels before he asks. But run it through a company’s eyes, not Boris’. Does an agent book that travel cheaper than Boris would have booked it himself? Maybe a little, or maybe it costs more once you count the compute it burns to do the job. Does taking travel admin off his plate convert into revenue? Only if the hours it frees go somewhere that earns. And here is the part that actually matters: even if those freed hours make Claude Code a little better, or ship the next release a day or two sooner, does that bring in one more paying customer? I would have a hard time buying that argument, and I do not think a CFO would either. Claude Cowork feels like magic. The return is a maybe.
Same with code review. Anthropic ships a product called Claude Code Review, the same automated reviewer it runs internally on every pull request. He made a point of how expensive it is, on purpose, burning a large pile of tokens to catch nearly every bug before a human ever sees the pull request.
I’m not looking for bugs anymore.
Wonderful, and still conditional. It only pays if catching those bugs ships revenue sooner, or cuts rework cost by more than the token bill. Then came the showpiece, not another tool but an applied use case. His continuous integration was slow, so he pointed Claude at it with a dynamic workload and let it run. A few hours and a few million tokens later, it returned four pull requests that cut build times in half.
this is work that would have taken days or weeks or months in the past.
Days of senior engineering work, gone. Beautiful. And faster CI is internal speed. It becomes return only when the capacity it frees ships something a customer pays for, or lets the same work happen with a smaller team.
The sharpest moment was Boris reasoning about his own spend. He uses the most expensive model on everything, on purpose, and walked through why:
if you think about ROI, there’s probably like a 50% chance to reduce the investment but probably like a 1000% opportunity… to increase the return. So I would focus almost all your effort on increasing returns.
He is right. He is also, without quite saying it, describing the exact job his customers have and his metrics do not measure. So increase the return. But what, exactly, is the return?
More code is not more business
Hold that question against the data, because this is where one room becomes the whole industry. The headline measure of engineering output is velocity, and velocity is exactly where AI delivers. Google’s DORA program, drawing on roughly 39,000 practitioners, found that developers feel meaningfully more productive with AI. Three in four say so. Then DORA measured the system those developers work in. Every 25 percent jump in AI adoption lined up with delivery throughput falling about 1.5 percent and stability falling about 7.2 percent. More code, reaching users no faster, breaking slightly more often.
Faros AI put telemetry on more than 10,000 developers across 1,255 teams and found the same thing from the other side. Teams with heavy AI adoption merged 98 percent more pull requests and completed 21 percent more tasks per developer. Real output, no argument. Then it checked the delivery metrics those numbers are supposed to drive, and found no significant correlation between AI adoption and any improvement at the company level. The gains did not vanish. They got absorbed: review time on those same teams climbed 91 percent. Faros named it the AI productivity paradox. Its 2026 follow-up across 22,000 developers made the absorption explicit: throughput up, with epics completed per developer up 66 percent, while incidents per pull request more than tripled and median review time grew roughly fivefold.
The individual gauge moves. The business gauge does not. That is the engineering ROI story, and it is the story being sold to boards as the return.
Where return actually comes from
So where does return come from, if not from shipping more code, faster? Decades of transformation work, long before AI, already answered this. The companies that create real value do not choose between cost and growth. They run both. BCG calls it the dual mandate and warns that roughly a third of companies let costs outrun revenue, quietly eroding the profitability they were chasing. McKinsey’s transformation research is blunter: in the rare transformations that throw off several times the value of the rest, more than half of that value comes from top-line moves, not cost cuts. Cut-only programs do not sustain. And across the board, up to 70 percent of transformations fail to reach their full potential. The prize is the one that moves both lines at once. Profitability, not efficiency. The hard one.
Now lay AI on top of that, and the mismatch is almost funny: eighty percent of companies aim it at efficiency, the weakest lever. Only the six percent who also chase growth and redesign the workflow capture real enterprise value; everyone else, MIT found across 300 deployments, sees no measurable P&L impact at all.
That is the burst, in slow motion. Not a crash, a bill coming due. Companies are funding AI on a return their engineers defined, and the day a board asks where it lands on a real scorecard, most of that spend has no answer.
The one question I now ask
So here is the blueprint I run when considering any AI investment, mine or a client’s. One question, asked of every tool: where does this show up, and on whose scorecard?
Run code velocity through it. Faster releases are not return on their own. They are return if the sales motion can absorb them. Ship three features instead of one, and if that closes deals that were stuck waiting on those features, that is top-line you can trace. Open a market you could not sell into before, and that is a new revenue line, traceable. If neither happens, you bought speed and drove it into a wall.
Run headcount through it. Replacing engineers with agents is only return if the net cost actually falls, not if a dollar of agent stands in for a dollar of human.
Serendipity to epiphany
I left that fireside chat more bullish on the tools than when I walked in. Claude, in all its forms, is the single largest AI tool in my own kit, and I am not giving it up.
What changed was not my read on the tools. It was my read on the root cause of the bubble. It is not the wild compute costs. It is not Nvidia investing in the AI companies that turn around and buy its chips. It is not even the scores of companies built on someone else’s model with no IP of their own. The real cause sits one level closer to the work. We have handed the definition of ROI to our builders and our engineers, the people creating the spend, and they are getting the “return” variable wrong.
That was the serendipity. I sat there listening to a brilliant engineer insist, correctly, that ROI should drive AI adoption, and I had spent years on the other side of that equation: building board scorecards, leading product to yield financial return for investors. From that seat, the gap was obvious. What he was calling return is not how a business measures return. Serendipity to epiphany, one row from the stage.
I should add, for the record: knowing all of this, I will still probably try to get in on the Anthropic IPO.
Funding AI on a return your engineers defined, and not sure it lands on a real scorecard? Let’s talk.


