Decision observability records every decision — its inputs, options, evidence, confidence, approver, action and outcome — and measures six things over time: decision quality, decision latency, confidence calibration, human intervention, compliance and economic impact.
The six metrics
| Metric | Definition | Why it matters |
|---|---|---|
| Decision quality | Share of decisions whose outcome met or beat the expected outcome | The real measure of value |
| Decision latency | Time from signal to executed action | Speed is often the biggest win |
| Confidence calibration | Whether 80%-confidence decisions succeed about 80% of the time | Trust and autonomy depend on it |
| Human intervention | Approval, override and escalation rates, with reasons | Shows where the model or policy is wrong |
| Compliance | Decisions made within policy, authority and envelope | Audit and regulatory readiness |
| Economic impact | Value protected or created vs baseline | The business case |
What a decision log must contain
- Decision ID and typeA stable identifier linking every downstream change.
- Inputs and context healthWhat was known, how fresh, what was missing.
- Options and scoresEvery option considered, including do nothing.
- Evidence and confidenceSources, confidence, dissenting evidence.
- AuthorityAutonomy level, approver, time, override reason.
- Action and outcomeWhat executed, when, and the measured result.
What a decision observability view shows
By decision type
Volume, quality, latency and value per decision type per week.
Calibration curve
Stated confidence vs actual success, with drift alerts.
Override explorer
Every human override with its reason, grouped by pattern.
Replay
Any decision reconstructed with the inputs and models of its day.
Review cadence
| Review | Frequency | Owner |
|---|---|---|
| Operational decision health | Weekly | Decision owner |
| Calibration and drift | Monthly | Model steward |
| Autonomy level changes | Quarterly or on incident | Owner + risk |
| Economic impact | Quarterly | Finance + owner |
Decision replay
Six months later, someone will ask why a decision was made. Replay reconstructs the decision with the inputs, model versions and policies of that day. It is the difference between an audit that takes an afternoon and one that takes a quarter — and it is the foundation for learning which decisions to automate next.
- Log every decision end to end.
- Track calibration, intervention and latency, not only accuracy.
- Overrides with reasons are the most valuable learning signal.
- Replay makes audits and learning possible.