Each resolved record pairs a snapshot of what you expected with a snapshot of what you later recorded. The four cells are the shipped values of om_calibration_quadrant. They are never added together.
Why the four never become one number. Averaging these cells destroys the only thing the matrix knows: that a favourable outcome from unsound work is not a win, and an unfavourable outcome from sound work is not a mistake. A single calibration figure would score the two of them identically in the middle, which is the exact error this product exists to prevent. The grid is the finding.
A disclosure about the collapse. The stored assessments are finer than the grid — process is complete, partial or incomplete, and outcome is above, near or below. The quadrant collapses partial into sound and near into at-or-above, following what the API already writes. Open any cell to see the assessment underneath rather than the collapsed label.
When you record an expectation you also record how confident you were, one to five. This plots how often each of those levels later landed at or above what you had written down.
Two things this chart deliberately lacks. There is no forty-five degree reference line, because nothing in the product maps a one-to-five confidence step onto a probability — drawing that diagonal would invent a standard and then measure you against it. And the empty level is shown as an empty level, not as zero and not as a line passing through: observedRatePct comes back null from the service when a bucket has no records, and a curve that smooths over a null is inventing a fact about you.
Where a recorded risk cites a historical pattern, the pattern arrives with the thing that makes it readable. A base rate without its survivorship caveat is worse than no base rate, because it reads as precision.