Prompt Placebo
- Year
- 2026
- Role
- Research design, LLM eval harness, statistics
- Status
- P3 complete, pre-registered before any full-run data existed
The question
Do the prompting techniques everyone recommends, role prompts, "think step by step," emotional stakes, tips, politeness, few-shot examples, actually help? Prompt Placebo is a pre-registered, paired-delta audit: every technique is measured against an identical baseline on the same question, so any measured effect is the technique's own, not the underlying question getting easier or harder that run.
The method
- —Paired-delta design across models and task categories (math, logic, procedural), with bootstrap confidence intervals and Holm-Bonferroni correction applied across all technique comparisons before anything was allowed to count as significant.
- —The methodology and verdict thresholds were hash-frozen (config hash 9b8af804...) and committed before any full-run data existed.
- —The completed run (P3): 1,438 questions, 23,008 API requests, $20.01 total spend against a pre-set $25 cap.
The verdict
0 of 39 technique-model-task comparisons earned a still_works verdict. 20 came back placebo, 17 inconclusive, 2 actively_hurts, 0 still_works. Math is placebo across both models tested. Logic is placebo or inconclusive. Procedural tasks are where the only real effects show up, and they are harms.
The only significant effects are harms
On claude-sonnet-5 with reasoning disabled, on procedural tasks, politeness drops accuracy 2.84 points (96.21% to 93.37%, 95% CI [-4.87, -0.81], p=0.0072) and few-shot examples drop it 3.79 points (96.21% to 92.42%, CI [-5.95, -1.62], p=0.0004). Both survive Holm-Bonferroni correction across all six techniques tested, so they are not multiple-comparisons noise.
Why it matters
Most prompt-engineering advice is untested folklore repeated because it sounds plausible, not because anyone ran a paired, corrected comparison against a real baseline. This project ran that comparison and reports the null result as the headline: nothing tested here reliably helps, and two of the most commonly recommended techniques, politeness and few-shot examples, measurably hurt accuracy on procedural tasks. The finding is the exhibit, not a caveat attached to a positive one.