The Mysteries We Unlocked: Field Notes from Building a Venture-Outcome Graph
Five practitioner lessons from building a ~6,900-company venture-outcome graph — identity is the real bottleneck, the dangerous errors are the plausible ones, survivorship hides in track records too, the founder is the unit not the company, and consensus isn't correctness.
This week we built something I'd wanted for years and kept underestimating: an outcome graph of roughly 6,900 accelerator companies — every Y Combinator batch, every Startmate cohort — with founders, investors, and exits linked, and each company joined to what actually happened next. Raised. Acquired. Went public. Or died.
I went in thinking the challenge was coverage. It wasn't. The mysteries that ate the week were all about truth — and they're worth writing down, because they're the ones no one warns you about when they tell you AI will just build your dataset for you.
1. Identity is the real bottleneck, not data
Most of the facts exist somewhere. Knowing who they belong to is the hard part. Startup names collide relentlessly across cities and industries; founder names merge two strangers into one phantom; an investor's name is a substring of a dozen others. The single most common way to corrupt a graph like this is to confidently attach a real event to the wrong entity. We caught a match that linked an accelerator to an unrelated fund on the strength of a two-character token. The fix is discipline, not cleverness — demand corroboration, and record a new entity rather than force a weak match.
2. The dangerous errors are the plausible ones
A pipeline reported that a company "raised $120k, investors including its accelerator." Technically true; catastrophically wrong to log as an outcome — because that money was the accelerator's entry cheque. The input dressed as a result. Label enough of those as successes and you teach the model that getting in equals winning, the exact circularity the whole project exists to break. Obvious garbage is easy to catch. It's the clean, well-formed, confidently-stated false fact that gets through — and an AI generates those all day. Most of the work was suspicion, not gathering.
3. Survivorship hides in the winners and in the investors
Everyone knows to distrust a dataset of only successful companies. Fewer notice the same trap one level up: if you only enrich the investors behind the exits, every investor you find has a flawless record, and the "exit rate" you compute is a pure artifact of where you looked. The count of a fund's winners is a real signal; the rate is a mirage until the losses beside them are linked in too. A track record without its denominator is a highlight reel.
4. The founder is the unit; the company is the vehicle
The most predictive pattern — the person on their second or third company — is invisible if you key on company names, and worse than invisible if you key on founder names, because names false-merge. Only a verified identity surfaces it. When it does, a small set of operators appears, carrying something no deck can hold. The company has the website and the cap table, so we treat it as the unit of analysis. But it's just the car the founder was driving that year.
5. Consensus is not correctness
Here's the one I'll only state in the general form, because the specific study is confidential: two evaluators trained on the same taste will agree with each other and can both be uncorrelated with what actually survives. A model pointed at what smart people picked will predict their picks accurately and still be blind to the outcome. It's calibrated to a proxy and confuses that for being right. The power law guarantees the gap — the companies that drive the returns are the ones consensus under-rates. The only referee that isn't a proxy is the outcome. Everything else is a room agreeing with a copy of itself.
What stays internal. This isn't a crystal ball; "dead" is a snapshot, deep-tech outcomes take a decade to resolve, and the machine-assembled links carry ordinary enrichment error we caught by hand where we could. And two of the most interesting things we found stay internal — a calibration study I can't detail, and engine mechanics that aren't for a blog post.
The meta-lesson travels, and it's the one I'd leave you with. Agents will assemble a dataset at a speed that still surprises me. Whether the dataset is true — right entity, real outcome, honest denominator, outcome not proxy — is not automatable. It's the entire remaining job, and it's the job that decides whether you built an intelligence graph or an elaborate way to launder your own assumptions back to yourself.
The orchestration was never the mystery. The truth was.
Frequently asked questions
What is the hardest part of building a startup dataset with AI?
Identity resolution, not data collection. Most facts exist somewhere; the hard part is knowing which company, founder, or investor they belong to. Names collide constantly, and the most common way to corrupt the dataset is to attach a real event to the wrong entity with confidence.
Why is an investor 'exit rate' often misleading?
Because it's usually survivorship. If you only enrich the investors behind companies that exited, every investor looks flawless and the computed exit rate is an artifact of where you looked. The count of wins is real; the rate needs the losses linked in too.
What does 'consensus is not correctness' mean in scoring?
Two evaluators trained on the same taste will agree with each other and can both be uncorrelated with what actually survives. A model that learns to imitate a group's picks inherits their internal consistency and their blind spots. The only referee that isn't a proxy is the real outcome.
AI product insights, weekly
Production AI architecture, scoring methodology, and what separates real AI products from slide-deck AI. No spam, unsubscribe anytime.
By subscribing you agree to receive email from NUVC and to our Privacy Policy. Unsubscribe anytime.
Your free NuScore shows where you stand.
Founder Pro ($99 one-time) unlocks what comes next.
Investor matches across 9,000+ investors, unlimited rescores, and the full score breakdown — validated against 201 real fundraising outcomes. 88% of decks scoring 8.5+ were funded.
Upgrade to Founder Pro — $99One-time payment. No subscription.
Haven't scored your deck yet? Upload free — analysis takes under 2 minutes. Then upgrade when you're ready.
