How Long Should Your Experiment Be? The Answer Isn't Days
Two weeks on, two weeks off feels like a rigorous experiment. Statistically, it's one comparison — and any slow drift in your life is perfectly confounded with the thing you're testing. What actually buys you precision is the number of switches, not the number of days.
The Design That Feels Rigorous and Isn't
Here's the design almost everyone reaches for first. You want to know whether magnesium improves your sleep. So you take it every night for two weeks, then stop for two weeks, and compare the averages.
Twenty-eight days of disciplined logging. It feels serious. It feels like the kind of thing a scientist would do.
It is, statistically, a single comparison.
You have one magnesium period and one no-magnesium period. Every day inside a block is measuring roughly the same thing, because your sleep on Tuesday is highly correlated with your sleep on Wednesday. You did not collect 28 independent observations. You collected something closer to two.
And you have a worse problem than low precision. Whatever else drifted across those four weeks — a project deadline landing in week three, the weather turning, a new bedtime creeping in, a partner traveling — moved in one direction across the calendar. Your condition also moved in one direction across the calendar. The two are perfectly aligned. There is no statistical procedure that can pull them apart after the fact.
Precision Comes From Switches
The useful mental model: in a crossover experiment on yourself, the unit of evidence is not the day. It is the switch — a treatment stretch paired against an adjacent control stretch.
Hold the total at 28 days and vary only the block length:
| Block length | Blocks in 28 days | Paired comparisons |
|---|---|---|
| 14 days | 2 | 1 |
| 7 days | 4 | 2 |
| 3 days | ~9 | ~4 |
| 2 days | 14 | 7 |
Same effort. Same number of logged mornings. Seven times the comparisons at the bottom row versus the top.
Because day-to-day observations within a block are autocorrelated, the standard error of your effect estimate scales roughly with the number of paired comparisons, not the number of days. Shortening blocks from 14 days to 2 buys you something close to a √7 improvement in precision — about 2.6× tighter — for exactly the same 28 days of logging. This is the core insight behind the N-of-1 trial literature, and it's why the AHRQ's design guide for N-of-1 trials (Kravitz and Duan, 2014) emphasizes multiple treatment periods rather than long ones.
The drift problem improves just as much. With 2-day blocks, that week-three deadline hits your magnesium days and your no-magnesium days about equally. It becomes noise — annoying, but symmetric. It stops being bias.
The Constraint That Stops You Going Shorter
If short blocks are better, why not alternate every day?
Because of carryover. When you switch conditions, the previous condition has to have actually worn off, or your "control" days are quietly contaminated with residual treatment. Carryover is the one thing that genuinely requires long blocks, and it's set by the intervention's biology — not by your preference.
A workable rule: block length should be at least two to three times the intervention's washout time.
That produces very different designs depending on what you're testing:
- Caffeine timing. Caffeine's half-life is roughly 5–6 hours in most adults, so ~24 hours clears the large majority of a dose. Daily or 2-day alternation is fine. This is the ideal short-block experiment.
- Alcohol. Cleared in hours; sleep-architecture effects are largely same-night. 2-day blocks work well.
- Melatonin timing. Half-life under an hour. Same-day.
- Sleep or workout timing. No pharmacology at all — the "washout" is one night. Alternate freely.
- Vitamin D. Half-life measured in weeks. Short blocks are meaningless; the level in your blood barely moves. This needs long blocks or a different design entirely.
- Creatine. Muscle saturation takes weeks of daily loading and washes out over roughly a month (Hultman et al., 1996). A crossover here is close to impossible on any reasonable timeline. Run it as a single long before/after, and accept the weaker inference.
The mistake isn't picking short blocks or long blocks. It's picking 7 days because a week is a familiar unit of time, when the intervention you're testing washes out in six hours.
What About Weekly Rhythms?
There's one honest argument for 7-day blocks: if your outcome has a strong weekly cycle — you sleep differently on weekends, you train harder on weekdays — then a 7-day block guarantees every condition sees exactly one full week, balancing the day-of-week effect by construction.
This is real, but it's usually the expensive fix for a cheap problem. Two better options:
- Alternate on an odd cycle. 3-day blocks drift across the week, so both conditions eventually see every weekday. Balance emerges over the run instead of being enforced block by block.
- Track day-of-week as a covariate and adjust for it. If you're logging anyway, this costs nothing and preserves your short blocks.
Reserve 7-day blocks for interventions whose washout genuinely takes two or three days — not as a default.
Choosing Your Design in Three Questions
1. How long does it take to wash out? Hours → 2-day blocks. A day → 3-day blocks. Several days → 7-day blocks. Weeks → don't run a crossover.
2. How many switches can you get? Aim for at least 4 paired comparisons. Below that, a single unusual week can dominate your result and you'll have no way to know it did.
3. What's the smallest effect worth acting on? If you'd change your behavior for a 3-point sleep score difference but not a 1-point one, you need enough switches to resolve 3 points against your own night-to-night variability. Most people's sleep scores swing 8–12 points night to night, which is exactly why one comparison tells you almost nothing.
The Practical Upshot
The instinct that more days equals more rigor is half right. Days matter — but only through how many contrasts you extract from them. A 28-day experiment run as two long blocks is a nearly uninterpretable result. The same 28 days run as fourteen 2-day blocks is a genuinely informative one.
If you have been running experiments and finding that nothing ever reaches a clear conclusion, the block length is the first thing to check. It is the most common reason a well-intentioned personal experiment produces a shrug.
One caution before you shorten everything: faster designs reach a visible result sooner, which makes it far more tempting to stop the moment the number looks good. That's its own way of fooling yourself, and it's worth understanding before you start — we cover it here.
Try it: Take your current or next experiment and check the block length against the washout rule above. If you're on 7-day blocks for something that clears overnight, cut to 2 or 3 and keep the same end date. You'll finish with several times the evidence for the same effort.
References
- Kravitz, R. L., & Duan, N. (eds.) (2014). Design and Implementation of N-of-1 Trials: A User's Guide. Agency for Healthcare Research and Quality.
- Senn, S. (2002). Cross-over Trials in Clinical Research (2nd ed.). Wiley.
- Guyatt, G., et al. (1986). Determining optimal therapy — randomized trials in individual patients. New England Journal of Medicine, 314(14), 889–892.
- Hultman, E., et al. (1996). Muscle creatine loading in men. Journal of Applied Physiology, 81(1), 232–237.
- Schork, N. J. (2015). Personalized medicine: Time for one-person trials. Nature, 520, 609–611.