Calibration Check™ · Public Methodology
Calibration Methodology: How CEOS measures its own predictive accuracy
CEOS produces two scores for every project: a pre-build prediction from the exploration loop, and an independent post-build evaluation after the prototype (or, for a standard Reality Check, the spec) is generated. The gap between them (Δ) is the signal we audit ourselves against. Below is the full dataset, open methodology, anonymized.
Self-calibration across 83 Reality Checks
83 runsSelf-calibrated across 83 Reality Checks: mean |delta| = 2.23. The post-build evaluation is doing real work, not rubber-stamping. The delta distribution below is what the calibration system actually does — how far the pre-build read sits from the independent post-build evaluation, across every Reality Check we have run. This measures the accuracy of our own method, not how many outside users we have.
- Mean Δ
- -2.18
- Median |Δ|
- 2.00
- Std dev
- 1.14
- Min / Max
- -5.5 / +1.0
- % with |Δ| ≥ 1.5
- 78%
Delta distribution
Each bar shows how many Reality Checks fell into that absolute-delta range. Significant revisions on the right = the post-build evaluation is doing real work, not rubber-stamping.
Vertical breakdown
Different verticals calibrate differently. Verticals with <10 projects are noisy; treat their stats as preliminary.
| Vertical | Projects | Mean Δ | Median |Δ| |
|---|---|---|---|
| B2B SaaS | 17 | -1.91 | -2.20 |
| Marketplace | 15 | -2.55 | -2.80 |
| Climate / Energy | 10 | -2.04 | -2.20 |
| AI Tool(noisy) | 8 | -1.56 | -1.60 |
| FinTech(noisy) | 8 | -2.43 | -2.00 |
| DeepTech(noisy) | 6 | -1.43 | -1.00 |
| HealthTech(noisy) | 6 | -2.33 | -2.20 |
| EdTech(noisy) | 4 | -1.80 | -1.80 |
| Consumer / B2C(noisy) | 4 | -2.65 | -2.80 |
| Hardware(noisy) | 3 | -2.87 | -2.80 |
| Other(noisy) | 2 | -4.65 | -4.65 |
Top 10 largest |Δ|
Anonymized: only the vertical, delta, and flags are shown.
| # | Vertical | Δ | |Δ| | Flags |
|---|---|---|---|---|
| 1 | Other | -5.5 | 5.5 | |
| 2 | Marketplace | -4.8 | 4.8 | simulated |
| 3 | Hardware | -4.0 | 4.0 | simulated |
| 4 | FinTech | -4.0 | 4.0 | simulated |
| 5 | Other | -3.8 | 3.8 | simulated |
| 6 | B2B SaaS | -3.8 | 3.8 | |
| 7 | Climate / Energy | -3.8 | 3.8 | simulated |
| 8 | HealthTech | -3.8 | 3.8 | simulated |
| 9 | Marketplace | -3.8 | 3.8 | simulated |
| 10 | FinTech | -3.8 | 3.8 | simulated |
Methodology
Pre-build prediction: the best score across the feasibility loop, generated by the Strategist agent before any prototype work begins.
Post-build evaluation: an independent feasibility agent re-scores the project across 5 criteria after the prototype is built (or, for a standard Reality Check, after the spec is generated without a prototype).
Δ (calibration delta):final_score − preliminary_score. Negative = pre-build was too optimistic. Positive = pre-build underestimated.
What this number is (and isn't): every Reality Check here ran the identical pipeline, whether we submitted it ourselves while building or it came from a user. They are all real evaluations the algorithm scored, so we count them all. This is a measure of how well our own method calibrates, not a count of outside customers.
Simulated feasibility: when a project is run as a standard Reality Check (no prototype), the post-build score is generated by an LLM judge on the spec alone, not on a real prototype. These rows are flagged separately.
Updated 09/08/2026, 02:42:45 · 83 total rows · open methodology