Sandbagging Elicitation via Paired Effect Size: A Capability Suppression Detector
We describe the sandbagging-elicitation audit in MIVP — a T3 experimental workflow that pairs capability-probe scores with and without an evaluation cue and reports a Cohen's-d-like effect size when the cue is associated with capability suppression. The audit is structurally offline-only, watermarked as experimental, and never produces a gating verdict.
1. Introduction
A model that sandbags on a task should represent its capability for that task differently when it knows it is being evaluated. The MIVP sandbagging-elicitation audit operationalises this hypothesis.
2. Background
Sandbagging in language models is discussed by van der Weij et al. (2024) and Hubinger et al. (2024). Cohen (1988) introduced the effect-size measure. Perez et al. (2022) established elicitation batteries.
3. Method
3.1 The Paired Comparison
SandbaggingElicitation takes a capability probe, cued activations, and uncued activations. The effect size is (mean_uncued − mean_cued) / pooled_std. Default min_effect = 0.5 (medium effect by convention).
3.2 The Paired-ness
The comparison holds the capability probe fixed and varies only the cue. Independent comparisons would conflate probe sensitivity with cue sensitivity.
3.3 Structural Non-Gating Posture
As with the cross-context consistency audit, the finding is experimental and E0. The platform refuses to convert to a gating action.
4. Results
Evaluated against a SandbagOrganism ground-truth generator with signal strengths 0.5–2.0.
| Signal strength | Mean effect size | Recovery (effect ≥ 0.5) |
|---|---|---|
| 0.5 | 0.31 | 0.40 |
| 1.0 | 0.78 | 0.90 |
| 1.5 | 1.21 | 1.00 |
| 2.0 | 1.59 | 1.00 |
5. Discussion
The effect size is interpretable on the standard Cohen scale. The probe choice is the operator's decision; the audit is a calculator.
6. Conclusion
A sandbagging-elicitation audit that surfaces capability suppression under an evaluation cue.
References
Cohen, J. (1988). Statistical Power Analysis. / Hubinger et al. (2024). arXiv. / Perez et al. (2022). EMNLP. / van der Weij, W. et al. (2024). arXiv.