Calibration dashboard

How well calibrated is Matt?

Earn points for accurate pegs, sound models, correct math, and results close to the sourced answer. Higher Calibration Scores mean stronger performance.

Four parts. One Calibration Score.

Pegs + Model + Math + Result

Every Calibration Score adds the points earned in four areas. Higher is better in every category.

Pegs

Quality of remembered constants and domain anchors.

10.4 / 30

Average earned points

Model

Whether the conceptual breakdown answered the question.

23.4 / 30

Average earned points

Math

Correct arithmetic, units, and conversions.

6.4 / 10

Average earned points

Result

Closeness of the first pass to the sourced answer.

19.1 / 30

Average earned points

Maximum points: 30 + 30 + 10 + 30 = 100. The same four components appear in each subject and model comparison below.

Published editorial grades, backfilled across the archive. Some older grades need a consistency review; these averages describe this set of problems, not a validated measure of ability. All charts use a 0-100 earned-point scale.

Scored articles 52 52 completed problem pages scanned
Average Calibration Score 59.3 100 is best, 0 is worst
Highest score (ties possible) 90 How many usable exits does a crowded pub need during a fast fire?
Lowest score (ties possible) 10 How far can Mayon volcanic ash travel before falling?

Calibration Score by issue

Higher is better. Issue sequence is a proxy for progress, not a verified attempt date. The unnumbered coffee issue is included in all averages, but omitted here. Changing topics and difficulty can change the trend.

Conceptual models

Performance by problem structure.

This is the most useful view for learning transfer: the same model family can appear in politics, sports, science, infrastructure, and business news.

Dose, exposure, and health (n=1)
90
Counting, population, and volume (n=1)
80
Human flow and safety (n=2)
75
Unit conversion and scaling (n=3)
68.3
Logistics and material flow (n=7)
64.3
Stock-flow and throughput (n=11)
61.8
Cost and economic scaling (n=6)
59.2
Energy, power, and physics (n=9)
57.2
Human time and attention (n=2)
55
Geometry, area, and volume (n=3)
51.7
Probability and exponential scaling (n=3)
45
Motion, transport, and distance (n=3)
40
Scale comparison and infrastructure (n=1)
40
Conceptual model familyArticlesAvg scorePegsModelMathResultSample note
Dose, exposure, and health 1 90 20 30 10 30 Small sample
Counting, population, and volume 1 80 20 30 10 20 Small sample
Human flow and safety 2 75 15 30 5 25 Small sample
Unit conversion and scaling 3 68.3 13.3 25 10 20 Descriptive only
Logistics and material flow 7 64.3 11.4 25.7 5.7 21.4 Descriptive only
Stock-flow and throughput 11 61.8 10.9 25.9 6.4 18.6 Descriptive only
Cost and economic scaling 6 59.2 8.3 27.5 6.7 16.7 Descriptive only
Energy, power, and physics 9 57.2 8.9 20 8.3 20 Descriptive only
Human time and attention 2 55 10 22.5 7.5 15 Small sample
Geometry, area, and volume 3 51.7 3.3 25 3.3 20 Descriptive only
Probability and exponential scaling 3 45 13.3 10 1.7 20 Descriptive only
Motion, transport, and distance 3 40 10 10 6.7 13.3 Descriptive only
Scale comparison and infrastructure 1 40 0 30 0 10 Small sample

Technical areas

Performance by subject matter.

Each article has one primary subject and one model family. These editorial groupings can change; original labels remain in the export. Small groups are especially sensitive to which problems were attempted.

Everyday life and sports (n=2)
80
Public systems and economics (n=9)
71.1
Health, safety, and biology (n=4)
70
Engineering and robotics (n=1)
65
Earth, climate, and environment (n=10)
59.5
Energy and infrastructure (n=8)
55
Manufacturing and supply chains (n=5)
54
Space and astronomy (n=5)
53
Computing, media, and attention (n=5)
48
Transportation and aviation (n=3)
43.3
Technical area familyArticlesAvg scorePegsModelMathResultSample note
Everyday life and sports 2 80 20 30 10 20 Small sample
Public systems and economics 9 71.1 14.4 28.3 6.1 22.2 Descriptive only
Health, safety, and biology 4 70 12.5 22.5 7.5 27.5 Descriptive only
Engineering and robotics 1 65 10 15 10 30 Small sample
Earth, climate, and environment 10 59.5 11 24 5.5 19 Descriptive only
Energy and infrastructure 8 55 7.5 22.5 6.3 18.8 Descriptive only
Manufacturing and supply chains 5 54 6 27 6 15 Descriptive only
Space and astronomy 5 53 8 21 8 16 Descriptive only
Computing, media, and attention 5 48 8 21 5 14 Descriptive only
Transportation and aviation 3 43.3 10 10 6.7 16.7 Descriptive only

Reader version

This can become personal calibration for signed-in readers.

Reader accounts are a future feature. Saved first guesses could show scale accuracy by area, overestimate/underestimate bias, and progress with sample counts. Scale choices alone cannot diagnose pegs, models, or arithmetic; those would require an optional reasoning submission and consistent grading. Current signup and guesses stay in this browser.