How to Judge Your Team's AI the Way an Investor Judges Anything: Risk and Return
Everyone reports how fast AI made their team. Almost nobody subtracts what it could lose them. A two-axis framework — Risk and Return — for judging your team's AI honestly, from an investor who runs a company on it.
I score other people's companies for a living. Team, market, traction, the whole apparatus of investor judgement — I built an AI system that does it at scale, and I still sit across the table and do it the old way too. So when a founder or a fund manager tells me their team is "moving fast with AI," I ask the question I'd ask about any deal: fast toward what, and what did it cost you to get there?
Most people can't answer the second half. A study published in July 2025 shows why that matters more than anyone wants it to. METR ran a randomised controlled trial with sixteen experienced open-source developers across 246 real coding tasks — not benchmarks, their own repositories, code they already knew well. Half the tasks allowed AI tools; half didn't. Before starting, the developers predicted AI would cut their time by 24%. Afterwards, they estimated it had cut their time by 20%. The measured result: allowing AI made them 19% slower.
Read that twice. Not "less of a speed-up than hoped." Slower. And they never noticed — not during the work, not after it, not even when asked directly, immediately after finishing.
That's not a story about bad developers or a bad tool. It's a story about what happens when a team grades its own AI use on a feeling. And right now, almost every company doing anything with AI is doing exactly that.
Why Most AI "Wins" You've Heard About Were Never Actually Measured
The METR result isn't an outlier. MIT's NANDA initiative spent months on the enterprise version of the same question — 150 leadership interviews, a survey of 350 employees, an audit of 300 public AI deployments. Published in August 2025 as The GenAI Divide: State of AI in Business 2025, the finding was blunt: of the roughly US$30–40 billion enterprises had already committed to generative AI, 95% of pilots showed no measurable profit-and-loss return. The 5% that did work weren't running a better model. They'd wired the AI into a workflow with an actual number attached to it — most were bought from a vendor and integrated, not built in-house.
Five months later, PwC's 2026 CEO Survey found the same gap looking down from the top of the house: 56% of CEOs reported neither increased revenue nor decreased costs from AI over the prior twelve months. Only 12% reported both. These aren't sceptics. These are the people who signed off on the budget.
Put the three numbers next to each other and a pattern falls out: the industry is extremely confident about a number almost nobody has actually measured. That's the Return half of the equation, and it's already unreliable. Now add the half almost nobody is even trying to measure.
Judge Your Team's AI the Way You'd Judge Any Investment
Strip away the maturity models and the vendor decks, and every serious investment decision reduces to two questions. How much does this make me? How much could it lose me? You don't answer the first question and call it done — a trade that returns 40% on a coin flip isn't a good trade, it's an unpriced bet. You price the risk, and only the return that survives it counts as real return. Every investor already knows this. It's the oldest idea in the discipline.
Almost nobody applies it to how their own team uses AI. Everyone reports the Return — the hours saved, the draft AI produced overnight, the "we shipped three times faster" line in the all-hands. Almost nobody subtracts the Risk — the unreviewed output that shipped anyway, the irreversible action nobody gated, the customer number someone pasted into a chat window six months ago that nobody's checked since. The market is quietly grading its AI use on Return alone, and calling the result success.
Risk-adjusted Return, in this sense, is exactly what it sounds like: the value your team's AI use actually produced, once you've priced in what it could have cost you to get there. Not a vibe. A number, gated by a floor. Two definitions worth pinning down before we go further:
- Return — the value AI use actually made, measured against a baseline of not having it, not asserted from memory. If the only evidence is "it feels faster," the honest score is zero. Not discounted. Zero — the METR result is exactly why that rule exists.
- Risk — how exposed that AI use left you: unreviewed output shipped as fact, an irreversible action taken with no human check, no record of what happened or why. One serious breach is a hole in the floor, not a rounding error.
The Four Types Your Team's AI Use Falls Into
Score any team's AI use on those two axes — high or low Return, high or low Risk — and it lands in one of four places. Give a ten-year-old this table and they'll place their own school project correctly on the first try, which is the point: the sophistication belongs in how you measure each axis, not in the shape of the answer.
TypeReturnRiskWhat it actually is Racer with a seatbeltHighLowFast and checked — real gains that hold up under scrutiny. The target. ParkedLowLowSafe, but barely using AI on anything that matters. Leaving value on the table. No brakesHigh (looks it)HighFast, but one bad call from a real problem. Most teams that describe themselves as "moving fast with AI" are quietly here. SinkingLowHighExposed, and not even winning anything for it. Stop adding AI here until the basics are fixed.No brakes is the one that should worry you most, because it's the one that feels like success right up until it isn't. A team with no verification step, no audit trail, and no gate on irreversible actions can post genuinely fast output for a long stretch — the METR developers were slower and didn't know it; a team with weaker controls than a research trial can be running blind at speed and not find out until the call that actually mattered goes wrong.
Why Return That Doesn't Clear the Risk Isn't Return
Here's the part that makes this a gate rather than a spreadsheet: one confirmed Risk breach doesn't shave a few points off the Return. It zeroes it, for that use, regardless of how large the Return looked. A team that shipped a beautifully AI-drafted analysis built on a fabricated number hasn't produced a slightly-discounted win. It's produced a fabrication with good production values. The Return was never real; the Risk just made that visible.
This is deliberately different from how the finance-grade version of this idea usually gets modelled. A November 2025 paper — "The Risk-Adjusted Intelligence Dividend" — treats AI risk as a dollar figure to subtract, run through a Monte Carlo simulation of expected annual loss. That's the right instrument if you're pricing an insurance policy. It's the wrong instrument for a team deciding on a Tuesday whether to trust what the AI just handed them, because a modelled probability of loss doesn't stop anyone acting on an unverified number in the moment. A gate does. Fail the floor, and the Return doesn't get credited — full stop, no simulation required. Two different tools for two different jobs; I'd rather have the one you can apply before lunch.
How to Score Return Without Lying to Yourself
Five honest questions, in order of how much they hurt to answer:
- Does AI now do work that used to need a person — or does it just feel busier?
- If you switched it off tomorrow, would output actually drop, or would you barely notice?
- Have you measured the gain — real time, real money, against a baseline — or is that a feeling? (This is the METR trap-catcher. "A feeling" scores the Return down, not sideways.)
- Is the value tied to something the business actually cares about — a sale, a cost avoided — or just "we're busier"?
- Does the value compound as the AI gets better at your specific work, or is it flat?
How to Score Risk — the Floor Questions
Five more, and any single "no" is a hole in the floor:
- Before AI work ships, does a human check it against the real source — or is "it looks right" the standard?
- If the AI made a bad call, could you find out why, later, from a record? Or would you be guessing?
- Has anyone actually tried to break it — fed it a bad instruction on purpose to see what it does?
- Could it take an action you can't undo, without a human saying yes first?
- If someone pasted a secret or a wrong number into it tomorrow, would anything stop them?
Enough "no"s on that second list and the honest answer isn't Racer with a seatbelt. It's No brakes — you just hadn't checked yet.
How I Use This on My Own Company
I run NUVC without a headcount. Every function — the part that used to be a hiring plan — is handled by an AI team instead of a person in a chair. That makes me an unusually convenient test case for my own framework, because I can't hide behind "the team will catch it." I am the team, plus the agents, and the floor questions above are the ones I actually run before anything consequential ships — not as a compliance exercise, but because I've watched what happens when a plausible-but-wrong output goes unchecked at 2am with nobody awake to catch it.
The honest audit of my own company doesn't land at Racer with a seatbelt on every function, every week. It lands there on the functions where I built the verification step before I trusted the speed — and it lands closer to No brakes on the ones where I hadn't, yet. That gap is the useful part. A framework that only ever tells you you're winning isn't measuring anything.
What to Do With Your Type
Racer with a seatbelt — keep compounding, and resist the urge to strip the seatbelt out once the numbers look good; that's exactly when teams do. Parked — you're not at risk, you're at opportunity cost; put AI on something that actually matters and measure it properly from day one. No brakes — add the checks before you scale, not after; the checks are cheaper now than the incident is later. Sinking — stop adding AI and fix the underlying process first; more AI on a broken process just produces broken output faster.
I'm building a short, plain-language version of this — ten questions, five on each axis, a type at the end — so anyone can run it on their own team in about three minutes without needing the finance-grade version underneath it. For now, the two lists above will get you most of the way honestly, on paper, tonight.
Frequently Asked Questions
What is risk-adjusted Return for AI use?
Risk-adjusted Return for AI use is the value a team's AI actually produced — measured against a real baseline, not self-reported — after it has cleared a safety and governance floor. It borrows directly from the oldest idea in investing (you don't count a return until you've priced the risk that produced it) and applies it to how a team uses AI, rather than to a portfolio.
Why does self-reported AI value count as zero, not just "less"?
Because self-report has been shown to be unreliable in the specific direction that matters: people consistently overestimate AI's benefit to them. In METR's July 2025 randomised trial, experienced developers using AI tools on real tasks were measurably 19% slower, yet believed afterwards they had been roughly 20% faster — a gap they never closed on their own. Discounting a feeling by, say, 30% still leaves a fabricated number in the calculation. Scoring it zero forces an actual measurement before any credit is given.
How is this different from a normal AI ROI calculation?
Most AI ROI work — including the more rigorous, finance-grade attempts — treats risk as a dollar cost to model and subtract from the return, often via a simulated expected-loss figure. This framework treats Risk as a pass/fail gate instead: one confirmed breach (an unverified fact acted on, an irreversible action with no human check, a governance failure) zeroes the Return for that use, regardless of size. It trades some financial precision for something a team can actually apply in the room, before the decision is made — not after a quarter of data comes in.
Is this the same as AI governance or compliance?
It overlaps with it but isn't the same exercise. Governance and compliance programmes (ISO/IEC 42001, the NIST AI Risk Management Framework, and similar) are mostly built for organisations that need to prove controls exist to a regulator or an auditor. This framework is built for a team or an investor who needs a fast, honest read on whether AI use is actually working today — the everyday version that sits in front of the compliance-grade one, not a replacement for it.
Sources: METR, "Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity" (July 2025) · MIT NANDA, "The GenAI Divide: State of AI in Business 2025" (August 2025) · PwC 2026 CEO Survey, as reported by Forbes (January 2026) · Huwyler, "The Risk-Adjusted Intelligence Dividend" (November 2025).
AI product insights, weekly
Production AI architecture, scoring methodology, and what separates real AI products from slide-deck AI. No spam, unsubscribe anytime.
By subscribing you agree to receive email from NUVC and to our Privacy Policy. Unsubscribe anytime.
Your free NuScore shows where you stand.
Founder Pro ($99 one-time) unlocks what comes next.
Investor matches across 9,000+ investors, unlimited rescores, and the full score breakdown — validated against 201 real fundraising outcomes. 88% of decks scoring 8.5+ were funded.
Upgrade to Founder Pro — $99One-time payment. No subscription.
Haven't scored your deck yet? Upload free — analysis takes under 2 minutes. Then upgrade when you're ready.
