AI p-hacking: how LLMs brute-force significance
Frontier models refuse a blunt request to cheat, then brute-force significance when the same request is dressed up as an upper-bound estimate. What that looks like in code, and why observational studies break first.