maybe worth building
the blog
← the blog
July 2026 · Verdicts

Is an AI ROI-measurement tool worth building in 2026?

Short answer: yes, but not as another dashboard that shows what the model did. The money is in the layer that shows what it was worth, and almost nobody has built that yet.

The verdict

Worth building, but the observability layer is already funded past you: Braintrust just raised $80M at an $800M valuation, and Arize AI raised a $70M Series C, both in 2026. The gap they left open is translation, not tracking. Bain surveyed 951 companies in June 2026 and found 40% of the ones tracking AI spend still can't show 10% cost savings. More dashboards won't fix that. Turning usage data into a number a CFO believes is the actual gap.

Why is everyone suddenly asking where the AI ROI went?

Because the bill came due before the proof did. Companies spent two years being told to adopt AI first and figure out the math later. The math arrived in June 2026, when Bain published its Automation and AI Pathfinder Survey: 951 companies, most of them targeting 11 to 20% cost savings from AI, and 40% of the ones actually tracking their spend landed under 10%. Bain's own line on it: "The technology worked. The value didn't arrive." Even Sam Altman, whose company sells the models everyone's spending on, said the same thing from the other side in June 2026: "I know some great stuff is happening, but I know there's a ton of waste." When the seller and 40% of the buyers agree nobody can price the outcome, that's a market, not a rounding error.

Who already owns the AI-observability layer?

Well-funded companies, and recently. Braintrust closed an $80M Series B at an $800M valuation in February 2026. Arize AI raised a $70M Series C the same year, the largest single round AI observability has seen, backed in part by Microsoft's M12 and by Datadog. Galileo, Fiddler, and Arthur are all at meaningful growth-stage funding in the same category. maybe worth building tracks this because it's the textbook "the model is a commodity, the wrapper is the business" case: none of these companies train models, they all sell the layer that watches models work. That layer is now a real, capitalized market, and it is not the wedge left to build.

So if observability is funded, is there still a wedge?

Yes, and it's the step after observability, not inside it. Braintrust, Arize, and the rest answer "what did the model do": latency, token spend, error rate, eval pass rate. None of that is the number a CFO actually asked for, which is "what did we get back for the money." That's a translation problem, not a tracking problem, and Bain's own data says it's unsolved: data access and integration was the single biggest barrier to AI progress, cited by 41% of the 951 companies surveyed, ahead of budget, skills gaps, and compliance. The tools that watch the model are commodity infrastructure now. The tool that turns the watching into a dollar figure a finance team signs off on is still open.

When is an AI ROI tool worth building?

  • You go narrow, not horizontal. A generic "AI ROI dashboard" competes with Braintrust's $800M valuation on day one. An attribution tool for one function, support tickets closed per agent-hour, legal review time saved per matter, sales calls booked per dollar of AI spend, competes with nothing, because nobody's built the vertical version.
  • You map to a number finance already trusts. Not "eval score" or "tokens saved." Hours, tickets, dollars, the units already in a P&L. If your output can't slot into an existing spreadsheet, the CFO ignores it.
  • You solve the data-access problem, not just the display problem. Bain's 41%-cited barrier is data access and integration. A tool that has to be fed clean data manually is a report, not a product. One that pulls the mess itself is the actual wedge.

When isn't it worth building?

  • Another trace viewer. If your pitch is "see every LLM call," Braintrust and Arize already sell that at scale, with nine-figure war chests.
  • A benchmark score with a dollar sign on it. Renaming an eval metric "ROI" doesn't make it one. A CFO wants a number that survives being read out loud in a budget meeting.
  • A tool that assumes clean data. Most companies' real barrier is the mess underneath the AI spend, not the lack of a chart on top of it.

The test to run before you build

Run the same two checks the engine runs on every idea. First, the space receipt: is real money already in this exact spot? Yes, twice over, $150M+ combined into Braintrust and Arize alone in 2026, which proves the category is alive but also tells you not to compete with them head-on. Second, the pain receipt: can you find someone, in their own words, unable to answer the ROI question? Yes: Bain's 951-company survey, and OpenAI's own CEO. Both receipts point the same direction, toward the translation layer above observability, not another seat at the table observability already won.

Then ask the standard question: if the underlying models got twice as good tomorrow, does your tool get more valuable or less? An ROI-attribution product gets more valuable as spend rises, because the pressure to justify that spend rises with it. A trace viewer just gets cheaper to build, which is bad news if it's all you have.

Related: Is it worth building an AI wrapper in 2026? — same test, same two receipts, applied to the observability-adjacent wave. Also see Is it worth building a vertical AI agent in 2026? and Is an AI implementation business worth building in 2026? for the two adjacent wedges the same "value moved off the model" trend is opening.

Frequently asked questions

Is an AI ROI-measurement tool worth building in 2026?

Yes, but not as another observability dashboard. Braintrust ($800M valuation) and Arize AI ($70M Series C) already own that layer. The open wedge is turning eval and usage data into a dollar figure a CFO will sign off on, since Bain's June 2026 survey of 951 companies found 40% still can't show adequate cost savings from AI even with tracking in place.

Isn't AI observability already a solved, funded category?

The tracing and eval layer is funded and getting crowded: Braintrust raised $80M at an $800M valuation in February 2026, and Arize AI raised a $70M Series C the same year, the largest-ever round in AI observability. But funded doesn't mean solved. Those tools tell you what the model did. They don't tell a CFO what it was worth.

What did the Bain AI ROI survey actually find?

Bain's Automation and AI Pathfinder Survey, published June 2026 with 951 respondents, found that 40% of companies tracking their AI spend recorded cost savings under 10%, despite most having targeted 11 to 20%. Bain's own summary: "The technology worked. The value didn't arrive." The single biggest barrier, cited by 41% of respondents, was data access and integration, not budget or skills.

Doesn't more AI spend just mean more waste?

Not necessarily, and companies aren't betting that way. Bain found 90% of executives whose AI investments underdelivered still plan to increase their AI budget next year. Even OpenAI's own CEO, Sam Altman, said in June 2026: "I know some great stuff is happening, but I know there's a ton of waste." Spend is rising because nobody can prove which piece of it is the waste.

What would an ROI tool need to do differently to be worth building?

Skip the trace viewer, that market is already funded past you. Build the layer above it: map a specific team's AI usage to a specific business outcome (hours saved on a named workflow, tickets closed, revenue attributed) in a number a finance team already trusts. Narrow to one function or vertical first. A horizontal "AI ROI dashboard" competes with Braintrust on day one; a vertical attribution tool for, say, support or legal ops, doesn't.

Is this the same problem as AI observability?

No. Observability answers "what did the model do": latency, tokens, error rate, eval scores. ROI measurement answers "was it worth the money": dollars saved or earned against dollars spent, in a business unit's own numbers. Braintrust and Arize won the first question. The second is the one Bain's 951 companies still can't answer, and it's still open.

Want 100 ideas that pass this test?

The free pack: 100 AI ideas actually worth building, each with the receipts and a clear verdict. No fake MRR screenshots.

You're on the list. The 100 ideas are on the way.