Skip to main content
Research

The intelligence
behind NUVC.

Research-led. Peer-reviewable. Open methodology.

Two founders. Open methodology. Zero black boxes.

3

Papers in submission

1,400+

Calibrated examples

24,000+

Investors

8,000+

Theses analysed

Publications

3 papers under development.

Targeting top-tier AI / finance / fairness venues. Methodology is documented in production code today, formalised for peer review.

ICAIF 2026·Draft · Submission Q3 2026

NuScore v5.4: A Multi-Signal Fusion Approach to Pre-Seed Venture Evaluation

ACM International Conference on AI in Finance · Tick Jiang, Duan Tianyi, NUVC Intelligence Team

We present NuScore v5.4, a production scoring engine for pre-seed startup pitch decks that combines an LLM-as-judge architecture with deterministic rule calibration and engineered ML features. Calibrated on 610+ labeled VC investment decisions across four independent sources, NuScore achieves stage-adjusted AU benchmarking and reports raise-probability with confidence intervals. We document the cross-check protocol that resolves LLM/rule disagreement and the score waterfall decomposition that exposes ranked signal contributions for explainability.

multi-signal fusionLLM as judgeVC evaluationcalibrationexplainability
FAccT 2027·Methodology in development

Responsible AI in Pre-Seed Venture Evaluation: A Three-Layer Fairness Framework

ACM Conference on Fairness, Accountability, and Transparency · NUVC Intelligence Team, Tick Jiang

AI scoring systems used in venture capital evaluation can perpetuate or amplify funding biases. We present the AI Governance framework — a three-layer fairness system combining (1) statistical bias detection across founder demographics, (2) deterministic ethical guardrails that exclude protected attributes from scoring inputs, and (3) LLM-driven adversarial review. We document the score appeal protocol, automated quarterly bias audits, and the outcome accountability loop that recalibrates the model when score-to-funding correlation drifts.

responsible AIfairnessventure evaluationbias auditappeal rights
AAAI 2027·Architecture documented

A Compounding Intelligence Architecture for Venture Capital

Association for the Advancement of Artificial Intelligence · Tick Jiang, NUVC Intelligence Team, Duan Tianyi

We present a multi-agent orchestration architecture deployed in production for venture capital intelligence. A pipeline of specialised agents executes deck extraction, scoring, integrity checking, enrichment, matching, benchmarking, synthesis, and feedback collection in parallel. Underneath, a deeper layer of fund and thesis intelligence compounds the analysis with audience-specific lens weighting, thesis matching, batch screening, portfolio fit, macro context, and governance oversight. We document the cross-agent dependency graph, the parallelism protocol, and the latency profile achieving sub-60-second end-to-end pitch deck analysis.

multi-agent systemsAI orchestrationventure capitalagent architecture
Datasets & Benchmarks

Open methodology. Verifiable claims.

8,000+ × 610+

Investor Thesis × Outcome Dataset

Original research dataset linking 8,000+ investor thesis statements (free-text, multi-language) across 82+ countries against 610+ scored startup evaluation outcomes. Used to validate NuScore correlation, identify dimension importance (product r²=0.773 vs team r²=0.492), and surface geographic priority variation. Citeable as a published Schema.org Dataset.

View dataset summary

1,400+ examples

NuScore Calibration Set

Labeled VC investment decisions from 4 independent sources: 85 known-outcome decks (raised vs failed), 110 Startmate accelerator evaluations (multi-evaluator scored), 172 VC deal memos (production memos with scoring rationale), and 243 AI-scored production decks. Used to calibrate the LLM judge against deterministic rules and to validate the cross-check protocol.

Methodology in NuScore v5.4 paper

Open scoring benchmark

LP Bench — Independent Fund Benchmark

The independent benchmark for AI-powered fund scoring accuracy. Measures consistency across re-runs, vintage calibration drift, hallucination rate on fund-fact extraction, and qualitative narrative accuracy. Subset of 312 funds with verified 3-year outcome data used for predictive validity. Reproducible methodology.

Visit LP Bench
The Architecture

One pipeline. Many signals.

A multi-stage AI pipeline extracts, scores, verifies, enriches, matches, and benchmarks every deck in parallel — with a deeper layer of fund and thesis intelligence underneath. We don't publish the internal architecture, but every score comes with a full breakdown of what moved it.

Responsible AI · AI Governance Layer

Transparent scoring. Bias monitored. Appeal rights.

No demographic bias in scoring

AI never uses founder gender, ethnicity, age, location, or educational prestige as scoring inputs. A first-time founder from Melbourne is scored on the same merits as a repeat founder from Stanford.

Transparent by default

Every score has a traceable reasoning chain — which data was used, how each lens was weighted, what the AI could not verify. The Score Explainability layer surfaces signal contributions per dimension. No black-box decisions.

Right to appeal

Every founder can appeal a score and request human review. Scores above 9.5 or below 1.0 require human review before being shown.

Active bias monitoring

The AI Governance layer runs automated bias detection monthly across all scored decks — checking for gender bias, stage bias, geographic bias, and score clustering. Patterns trigger recalibration.

Honest confidence

If data is sparse, confidence says so. We never inflate certainty. A low-confidence score is labelled directional, not definitive.

Outcome accountability

We track whether scores predict actual funding outcomes. If they don't, we recalibrate. The model earns trust — it doesn't assume it.

What We've Learned

5 findings from the dataset.

01

Product execution is 1.6× stronger than team

Across 286 scored startup evaluations, product execution variance explained 77.3% of the outcome (r²=0.773) — versus 49.2% for team strength (r²=0.492). The 'team is everything' narrative is wrong. Solution quality dominates.

Read the analysis
02

Problem/solution is a binary gate, not a discriminator

Within pre-screened cohorts, problem/solution scoring explained almost zero variance (r²=0.002). It works as a yes/no filter — but doesn't help rank deals. Implication: stop spending 40% of your scoring weight on it.

Read the analysis
03

Melbourne is the most team-focused VC hub globally

Across 8,000+ investor thesis statements analysed, Melbourne VCs mention 'team' 24.1% of the time — the highest of any major hub. Sydney is 18.4%. SF is 14.2%. Stage-adjusted hub variance is real and matters for fundraise targeting.

Read the analysis
04

LLM judges provide 6× more score discrimination than rules

On the same 286-deck calibration set, the LLM judge produced 6× more score variance than the deterministic rule engine. The rules anchor the floor; the LLM does the ranking. NuScore v5.4 fuses both signals when they disagree.

Read the analysis
05

Categorical labels beat continuous scores within funded cohorts

Within pre-screened cohorts (already passed the funding bar), 10X Potential categorical labels (yes/no) outperformed continuous 0-10 scores by +15pp in predicting actual fundraise success. Classification > regression once you're past the gate.

Read the analysis
Cite NUVC

Citing the research?

NUVC's research is published as Schema.org Dataset and ScholarlyArticle metadata for LLM and academic citation. ChatGPT, Claude, Gemini, and Perplexity can cite NUVC methodology directly via the structured data on this page.

@misc{nuvc2026,

title= {NuScore v5.4: Multi-Signal Fusion for Pre-Seed Venture Evaluation},

author= {Jiang, Tick and Tianyi, Duan and NUVC Intelligence Team},

year= {2026},

institution= {NUVC, Melbourne, Australia},

url= {https://nuvc.ai/research},

}

Read. Cite. Build with us.

Read the papers. Cite the dataset. Use the API. The methodology is open — the production system is what we sell.