Comparisons
32
Testing many metrics manufactures attractive p-values unless the full family is corrected and non-significance is kept separate from no effect.
Public · synthetic demoCherry-picking significant rows inflates false discoveries. Calling a non-significant result “no effect” ignores what the sample was too small to detect.
Family size: 32 Alpha: 0.05 Target power for MDE: 0.8 Headline arms: control n=4000 real-effect variant n=4000
The headline effect is 0.0265 with interval 0.0126–0.0404. Its raw p-value is 0.0002 and both corrected values are 0.0061.
| Comparison | Effect | Raw p | BH p | Verdict |
|---|---|---|---|---|
| headline real effect | 0.0265 | 0.0002 | 0.0061 | significant after correction |
| headline null variant | 0.0010 | 0.8796 | 0.8796 | not detected; MDE 0.0185 |
| metric_01 | 0.0800 | 0.0170 | 0.1356 | raw only |