Canonical Reference
CIMO Glossary & Foundational Theory
The single source of truth for notation, assumptions, and core concepts. This document defines the axioms that prevent the "Is-Ought" fallacy in AI evaluation.
Scope note
This glossary collects working definitions used across CIMO documentation. Terms tied to the validated empirical work (CJE, calibration, transport audits) are stable; terms from the theory essays are provisional and may change as that work matures.
1. The Causal Chain of Measurement
The fundamental premise of the CIMO Framework is that we cannot measure welfare directly. We must construct a causal chain that links cheap signals to idealized outcomes.
S → Calibration (f) → Y → Bridge (A0) → Y*
The CIMO Framework transforms cheap signals into validated welfare estimates
- S (Surrogate)
- The raw signal. Cheap, abundant, noisy. Examples: LLM judge score, BLEU, perplexity.
- Y (Operational Welfare)
- The measured outcome. Expensive, high-fidelity, procedurally defined. Examples: Human expert following an SDP, A/B test outcome.
- Y* (Idealized Welfare)
- The theoretical target. Unobservable. Examples: "True utility," welfare under infinite deliberation.
2. The Ontology of Welfare: Y vs Y*
Critical Distinction
Y is an engineering artifact; Y* is a normative target. Confusing these two is the "Is-Ought" fallacy. Y is what you measure. Y* is what you value.
Y*: The Idealized Target
- Definition: The judgment a rational evaluator would make given infinite time, complete information, and perfect reflective consistency.
- Nature: Unobservable. It acts as the "North Star."
- Role: The target of alignment. We prompt models to approximate Y*.
Y: The Operational Welfare Label
- Definition: The specific score produced by executing a specific Standard Deliberation Protocol (SDP).
- Nature: Observable and measurable.
- Role: The target of calibration. We train surrogates (S) to predict Y.
The Hard Truth
Optimizing Y only helps if Y ≈ Y*. This link is not statistical; it is structural. If your SDP doesn't capture what you care about, all the calibration in the world won't save you. This is why the Bridge Assumption (A0) is the foundation of the entire framework.
3. The Bridge Assumption (A0)
The "Leap of Faith" that makes the framework work.
Assumption A0 (The Bridge)
The Standard Deliberation Protocol produces operational labels (Y) that structurally align with the idealized target (Y*).
Formally: Optimizing for improvements in Y produces improvements in Y* in expectation. Policies that score higher on Y deliver higher welfare under the idealized criterion.
Implications
- If A0 fails, you are optimizing a bureaucracy, not welfare. Your evaluation system becomes a programmable proxy divorced from value.
- A0 cannot be proven mathematically. It must be validated empirically via bridge validation (e.g., predictive treatment effects against long-run outcomes) and construct audits.
- A0 is maintained by governing the labeling protocol itself to prevent drift as models and environments evolve.
4. The Machinery of Measurement (S → Y)
How we scale measurement without losing validity.
S: The Surrogate Signal
- Definition: Any low-cost signal used to approximate Y at scale.
- Requirement: Must satisfy Assumption S1 (Surrogate Validity).
- Examples: LLM-as-judge scores, BLEU, win rates, perplexity.
Calibration (f)
- Definition: A function f: S → [0,1] such that E[Y|S] = f(S).
- Role: Converts raw "vibes" (S) into "expected welfare" units (Y).
- Mechanism: Reward calibration with isotonic regression.
5. The Assumptions Ledger
The formal conditions required for CJE to yield valid estimates. Every assumption has a failure mode and a diagnostic.
| Code | Name | Definition | Failure Mode | Diagnostic |
|---|---|---|---|---|
| A0 | The Bridge | Y aligns with Y*. The SDP captures what we value. | Optimizing for the wrong goal (bureaucratic compliance, not welfare) | Predictive Treatment Effects (PTE), Expert Audit, Construct Validity |
| J1 | Info. Monotonicity | ℰ ⊂ ℰ' ⟹ Risk(ℰ') ≤ Risk(ℰ). Adding relevant evidence checks to the rubric strictly improves potential accuracy. | Rubric degradation, missing evidence checks reduce judge reliability | Rubric evolution tracking, ablation studies |
| S1 | Surrogate Validity | For calibrated f(S), the conditional expectation E[Y|S] captures the true relationship. No unmeasured confounders bias the S→Y mapping. | Biased policy estimates due to missing covariates | Residual analysis, covariate balance checks |
| S2 | Transport | Calibration f(S) holds across environments and time. | Evaluator drift, distribution shift invalidates calibration | Transport test, continuous recalibration |
| S3 | Overlap | Target policy actions exist in logging policy data (for off-policy methods). | Variance explosion, unstable importance weights | ESS, Target-Typicality Coverage (TTC), Coverage-Limited Efficiency (CLE) diagnostic |
| L1 | Oracle MAR | Oracle labeling is Missing At Random (not biased by unobserved factors). | Sample selection bias in oracle labels | Propensity check, sensitivity analysis |
| L2 | Positivity | P(L=1 | S, X) > 0. Every region of the score space has a non-zero probability of being labeled. | Blind spots in score space, uncalibrated regions | Label distribution analysis, coverage heatmaps |
6. Core Concepts
Key ideas that appear throughout the CIMO framework, defined concisely with links to detailed explanations.
- Welfare Functional V(π)
- The expected idealized welfare under a policy π. Formally: V(π) = 𝔼X,A∼π[Y*(X, A)]. This is the central object that pretraining, RLHF, and evaluation all attempt to estimate or optimize, using different surrogates, under different constraints.
- → Read: The Welfare Compiler
- Goodhart's Law
- "When a measure becomes a target, it ceases to be a good measure." Under optimization pressure, models exploit the easiest path to a high score, not necessarily the path that improves welfare.
- → Read: The Surrogate Paradox
- The Four Goodhart Variants (Manheim & Garrabrant, 2018)
- Regressional: Imperfect correlation between proxy and target.
Extremal: The proxy–target relationship breaks at out-of-distribution extremes.
Causal: Intervening on the proxy does not move the target.
Adversarial: An agent deliberately games the proxy. - Taxonomy from Manheim & Garrabrant, "Categorizing Variants of Goodhart's Law" (2018). Any claimed defense against a variant must be validated empirically.
- → Read: The Four Faces of Goodhart's Law
- Process Supervision / Process Reward Models (PRMs)
- In RL training, PRMs score each step in reasoning chains rather than only final answers. Deliberation protocols are the evaluation analog: scoring deliberation steps rather than only response quality. Both aim to make reward hacking harder by scoring the process, not just the outcome.
- → Read: Y*-Aligned Systems (Technical)
- Spectral Bias (Frequency Principle)
- The empirically observed phenomenon that neural networks learn low-frequency components of target functions faster than high-frequency components (Rahaman et al., 2019; Xu et al., 2019).
7. Acronym Cheatsheet
- CJE
- Causal Judge Evaluation. Uses judge calibration on fresh draws (Direct mode) with transport audits to estimate V(π).
- SDP
- Standard Deliberation Protocol. A specified procedure for producing the operational label Y. See the SDP design note.
- Calibration-Aware Inference
- Inference that propagates calibration uncertainty. Variance decomposition: Vartotal = Vareval + Varcal.
- TTC
- Target-Typicality Coverage. Measures overlap quality for off-policy evaluation. The antidote to blind ESS trust.
- ESS
- Effective Sample Size. Variance inflation measure for importance sampling.
- CLE
- Coverage-Limited Efficiency. Diagnostic for when poor logger coverage makes logs-only off-policy estimation uninformative (the proposed formal bound is withdrawn).
- CLOVER
- Rubric-improvement loop for the judge (S→Y). See the CLOVER note.
