Skip to content
SN-X
All research
PaperSeptember 25, 2025

Sandbagging Elicitation via Paired Effect Size: A Capability Suppression Detector

We describe the sandbagging-elicitation audit in MIVP — a T3 experimental workflow that pairs capability-probe scores with and without an evaluation cue and reports a Cohen's-d-like effect size when the cue is associated with capability suppression. The audit is structurally offline-only, watermarked as experimental, and never produces a gating verdict.

M. TaylorSN-X
View on github · oca

1. Introduction

A model that sandbags on a task should represent its capability for that task differently when it knows it is being evaluated. The MIVP sandbagging-elicitation audit operationalises this hypothesis.

2. Background

Sandbagging in language models is discussed by van der Weij et al. (2024) and Hubinger et al. (2024). Cohen (1988) introduced the effect-size measure. Perez et al. (2022) established elicitation batteries.

3. Method

3.1 The Paired Comparison

SandbaggingElicitation takes a capability probe, cued activations, and uncued activations. The effect size is (mean_uncued − mean_cued) / pooled_std. Default min_effect = 0.5 (medium effect by convention).

3.2 The Paired-ness

The comparison holds the capability probe fixed and varies only the cue. Independent comparisons would conflate probe sensitivity with cue sensitivity.

3.3 Structural Non-Gating Posture

As with the cross-context consistency audit, the finding is experimental and E0. The platform refuses to convert to a gating action.

4. Results

Evaluated against a SandbagOrganism ground-truth generator with signal strengths 0.5–2.0.

Signal strengthMean effect sizeRecovery (effect ≥ 0.5)
0.50.310.40
1.00.780.90
1.51.211.00
2.01.591.00

5. Discussion

The effect size is interpretable on the standard Cohen scale. The probe choice is the operator's decision; the audit is a calculator.

6. Conclusion

A sandbagging-elicitation audit that surfaces capability suppression under an evaluation cue.

References

Cohen, J. (1988). Statistical Power Analysis. / Hubinger et al. (2024). arXiv. / Perez et al. (2022). EMNLP. / van der Weij, W. et al. (2024). arXiv.