Skip to content

Calibration Check™ · Public Methodology

Calibration Methodology: How CEOS measures its own predictive accuracy

CEOS produces two scores for every project: a pre-build prediction from the exploration loop, and an independent post-build evaluation after the prototype (or, for a standard Reality Check, the spec) is generated. The gap between them (Δ) is the signal we audit ourselves against. Below is the full dataset, open methodology, anonymized.

Self-calibration across 83 Reality Checks

83 runs

Self-calibrated across 83 Reality Checks: mean |delta| = 2.23. The post-build evaluation is doing real work, not rubber-stamping. The delta distribution below is what the calibration system actually does — how far the pre-build read sits from the independent post-build evaluation, across every Reality Check we have run. This measures the accuracy of our own method, not how many outside users we have.

Mean Δ
-2.18
Median |Δ|
2.00
Std dev
1.14
Min / Max
-5.5 / +1.0
% with |Δ| ≥ 1.5
78%

Delta distribution

Each bar shows how many Reality Checks fell into that absolute-delta range. Significant revisions on the right = the post-build evaluation is doing real work, not rubber-stamping.

[0, 0.5)
4
[0.5, 1.0)
6
[1.0, 1.5)
8
[1.5, 2.0)
16
[2.0, 3.0)
30
≥3.0
19

Vertical breakdown

Different verticals calibrate differently. Verticals with <10 projects are noisy; treat their stats as preliminary.

Calibration delta by industry vertical: project count, mean delta, and median absolute delta per vertical.
VerticalProjectsMean ΔMedian |Δ|
B2B SaaS17-1.91-2.20
Marketplace15-2.55-2.80
Climate / Energy10-2.04-2.20
AI Tool(noisy)8-1.56-1.60
FinTech(noisy)8-2.43-2.00
DeepTech(noisy)6-1.43-1.00
HealthTech(noisy)6-2.33-2.20
EdTech(noisy)4-1.80-1.80
Consumer / B2C(noisy)4-2.65-2.80
Hardware(noisy)3-2.87-2.80
Other(noisy)2-4.65-4.65

Top 10 largest |Δ|

Anonymized: only the vertical, delta, and flags are shown.

Top 10 projects by largest absolute calibration delta, anonymized to vertical, delta, and flags.
#VerticalΔ|Δ|Flags
1Other-5.55.5
2Marketplace-4.84.8simulated
3Hardware-4.04.0simulated
4FinTech-4.04.0simulated
5Other-3.83.8simulated
6B2B SaaS-3.83.8
7Climate / Energy-3.83.8simulated
8HealthTech-3.83.8simulated
9Marketplace-3.83.8simulated
10FinTech-3.83.8simulated

Methodology

Pre-build prediction: the best score across the feasibility loop, generated by the Strategist agent before any prototype work begins.

Post-build evaluation: an independent feasibility agent re-scores the project across 5 criteria after the prototype is built (or, for a standard Reality Check, after the spec is generated without a prototype).

Δ (calibration delta):final_score − preliminary_score. Negative = pre-build was too optimistic. Positive = pre-build underestimated.

What this number is (and isn't): every Reality Check here ran the identical pipeline, whether we submitted it ourselves while building or it came from a user. They are all real evaluations the algorithm scored, so we count them all. This is a measure of how well our own method calibrates, not a count of outside customers.

Simulated feasibility: when a project is run as a standard Reality Check (no prototype), the post-build score is generated by an LLM judge on the spec alone, not on a real prototype. These rows are flagged separately.

Download dataset as CSV →

Updated 09/08/2026, 02:42:45 · 83 total rows · open methodology