The intelligence
behind NUVC.
Research-led. Peer-reviewable. Open methodology.
Two founders. Open methodology. Zero black boxes.
3
Papers in submission
1,400+
Calibrated examples
24,000+
Investors
8,000+
Theses analysed
3 papers under development.
Targeting top-tier AI / finance / fairness venues. Methodology is documented in production code today, formalised for peer review.
NuScore v5.4: A Multi-Signal Fusion Approach to Pre-Seed Venture Evaluation
ACM International Conference on AI in Finance · Tick Jiang, Duan Tianyi, NUVC Intelligence Team
We present NuScore v5.4, a production scoring engine for pre-seed startup pitch decks that combines an LLM-as-judge architecture with deterministic rule calibration and engineered ML features. Calibrated on 610+ labeled VC investment decisions across four independent sources, NuScore achieves stage-adjusted AU benchmarking and reports raise-probability with confidence intervals. We document the cross-check protocol that resolves LLM/rule disagreement and the score waterfall decomposition that exposes ranked signal contributions for explainability.
Responsible AI in Pre-Seed Venture Evaluation: A Three-Layer Fairness Framework
ACM Conference on Fairness, Accountability, and Transparency · NUVC Intelligence Team, Tick Jiang
AI scoring systems used in venture capital evaluation can perpetuate or amplify funding biases. We present the AI Governance framework — a three-layer fairness system combining (1) statistical bias detection across founder demographics, (2) deterministic ethical guardrails that exclude protected attributes from scoring inputs, and (3) LLM-driven adversarial review. We document the score appeal protocol, automated quarterly bias audits, and the outcome accountability loop that recalibrates the model when score-to-funding correlation drifts.
A Compounding Intelligence Architecture for Venture Capital
Association for the Advancement of Artificial Intelligence · Tick Jiang, NUVC Intelligence Team, Duan Tianyi
We present a multi-agent orchestration architecture deployed in production for venture capital intelligence. A pipeline of specialised agents executes deck extraction, scoring, integrity checking, enrichment, matching, benchmarking, synthesis, and feedback collection in parallel. Underneath, a deeper layer of fund and thesis intelligence compounds the analysis with audience-specific lens weighting, thesis matching, batch screening, portfolio fit, macro context, and governance oversight. We document the cross-agent dependency graph, the parallelism protocol, and the latency profile achieving sub-60-second end-to-end pitch deck analysis.
Open methodology. Verifiable claims.
8,000+ × 610+
Investor Thesis × Outcome Dataset
Original research dataset linking 8,000+ investor thesis statements (free-text, multi-language) across 82+ countries against 610+ scored startup evaluation outcomes. Used to validate NuScore correlation, identify dimension importance (product r²=0.773 vs team r²=0.492), and surface geographic priority variation. Citeable as a published Schema.org Dataset.
View dataset summary1,400+ examples
NuScore Calibration Set
Labeled VC investment decisions from 4 independent sources: 85 known-outcome decks (raised vs failed), 110 Startmate accelerator evaluations (multi-evaluator scored), 172 VC deal memos (production memos with scoring rationale), and 243 AI-scored production decks. Used to calibrate the LLM judge against deterministic rules and to validate the cross-check protocol.
Methodology in NuScore v5.4 paperOpen scoring benchmark
LP Bench — Independent Fund Benchmark
The independent benchmark for AI-powered fund scoring accuracy. Measures consistency across re-runs, vintage calibration drift, hallucination rate on fund-fact extraction, and qualitative narrative accuracy. Subset of 312 funds with verified 3-year outcome data used for predictive validity. Reproducible methodology.
Visit LP BenchOne pipeline. Many signals.
A multi-stage AI pipeline extracts, scores, verifies, enriches, matches, and benchmarks every deck in parallel — with a deeper layer of fund and thesis intelligence underneath. We don't publish the internal architecture, but every score comes with a full breakdown of what moved it.
Transparent scoring. Bias monitored. Appeal rights.
No demographic bias in scoring
AI never uses founder gender, ethnicity, age, location, or educational prestige as scoring inputs. A first-time founder from Melbourne is scored on the same merits as a repeat founder from Stanford.
Transparent by default
Every score has a traceable reasoning chain — which data was used, how each lens was weighted, what the AI could not verify. The Score Explainability layer surfaces signal contributions per dimension. No black-box decisions.
Right to appeal
Every founder can appeal a score and request human review. Scores above 9.5 or below 1.0 require human review before being shown.
Active bias monitoring
The AI Governance layer runs automated bias detection monthly across all scored decks — checking for gender bias, stage bias, geographic bias, and score clustering. Patterns trigger recalibration.
Honest confidence
If data is sparse, confidence says so. We never inflate certainty. A low-confidence score is labelled directional, not definitive.
Outcome accountability
We track whether scores predict actual funding outcomes. If they don't, we recalibrate. The model earns trust — it doesn't assume it.
5 findings from the dataset.
Product execution is 1.6× stronger than team
Across 286 scored startup evaluations, product execution variance explained 77.3% of the outcome (r²=0.773) — versus 49.2% for team strength (r²=0.492). The 'team is everything' narrative is wrong. Solution quality dominates.
Read the analysisProblem/solution is a binary gate, not a discriminator
Within pre-screened cohorts, problem/solution scoring explained almost zero variance (r²=0.002). It works as a yes/no filter — but doesn't help rank deals. Implication: stop spending 40% of your scoring weight on it.
Read the analysisMelbourne is the most team-focused VC hub globally
Across 8,000+ investor thesis statements analysed, Melbourne VCs mention 'team' 24.1% of the time — the highest of any major hub. Sydney is 18.4%. SF is 14.2%. Stage-adjusted hub variance is real and matters for fundraise targeting.
Read the analysisLLM judges provide 6× more score discrimination than rules
On the same 286-deck calibration set, the LLM judge produced 6× more score variance than the deterministic rule engine. The rules anchor the floor; the LLM does the ranking. NuScore v5.4 fuses both signals when they disagree.
Read the analysisCategorical labels beat continuous scores within funded cohorts
Within pre-screened cohorts (already passed the funding bar), 10X Potential categorical labels (yes/no) outperformed continuous 0-10 scores by +15pp in predicting actual fundraise success. Classification > regression once you're past the gate.
Read the analysisCiting the research?
NUVC's research is published as Schema.org Dataset and ScholarlyArticle metadata for LLM and academic citation. ChatGPT, Claude, Gemini, and Perplexity can cite NUVC methodology directly via the structured data on this page.
@misc{nuvc2026,
title= {NuScore v5.4: Multi-Signal Fusion for Pre-Seed Venture Evaluation},
author= {Jiang, Tick and Tianyi, Duan and NUVC Intelligence Team},
year= {2026},
institution= {NUVC, Melbourne, Australia},
url= {https://nuvc.ai/research},
}
Read. Cite. Build with us.
Read the papers. Cite the dataset. Use the API. The methodology is open — the production system is what we sell.
