In the Blair and Sambanis comparison, the proposed escalation model was not significantly better than four baseline models, with one-tailed test Z values of 0.64, 1.09, 0.42, and 0.67 and corresponding p values of 0.26, 0.14, 0.34, and 0.25.
Notes on verification
Confirmed by direct match to both the published Patterns article and the arXiv preprint by Kapoor & Narayanan, with exact matching statistical values and consistent methodology description.
Sources
- Leakage and the reproducibility crisis in machine-learning-based science (cell.com)
- https://www.sciencedirect.com/science/article/pii/S2666389923001599 (sciencedirect.com)
- https://arxiv.org/pdf/2207.07048 (arxiv.org)