AI · Article 19 of 64

From Modeling Others to Modeling Your Best Self

Why might your own best days be a better template than someone else's genius?

A famous founder publishes his morning routine, and thousands of people adopt it. Cold plunge, no email before noon, a specific strain of tea. Some of those people report that it works. Is that evidence that the routine carries something worth copying, or merely that a plausible ritual, performed attentively, feels like progress?

There is a structural reason to be suspicious of borrowed routines, and it has nothing to do with whether the founder is honest. His routine was optimized against his biology, his history, his obligations, and his particular work. It was fitted to a sample of one. Transplanted into a different life, it may help, do nothing, or actively misfire — and without a way to compare, the person adopting it cannot tell which. The alternative is to model your own peaks: to treat your best days as the data and yourself as the population. That approach runs directly against the way most of our evidence about people is produced, and it is the reason your own best days are often a better template than someone else’s genius.

Evidence built from averages

Almost everything published about human performance is built from groups. Recruit a few hundred people, measure a variable, compute the correlation or the effect size, publish. The resulting number is a fact about the sample’s average behavior. The trouble is that the human mind treats it as a fact about individuals.

The problem has a technical name: non-ergodicity. A statistical estimate generalizes from the group to the individual only under restrictive conditions — that the process behaves the same across people, and that its mean and variance stay stable over time. Most psychological and biological processes satisfy neither. In a 2018 paper in PNAS, Aaron Fisher, John Medaglia, and Bertus Jeronimus put numbers to the gap by analyzing six intensive repeated-measures datasets. Central-tendency estimates agreed reasonably between groups and individuals, but the variance around the expected value was two to four times larger within individuals than within groups. Their conclusion was blunt: researchers should test whether group-level findings actually hold at the individual level before generalizing, because aggregated results may be “worryingly imprecise.”

Their illustration is memorable. At the group level, typing speed and typos correlate negatively: faster typists make fewer errors, because practiced typists are both quick and accurate. Within any single person, the correlation is positive: as you personally speed up, you personally make more mistakes. Both facts are true. They describe different levels. A recommendation derived from the group statistic — go faster to be accurate — would be exactly wrong for the individual, and it would be wrong in a way that no amount of careful group measurement could catch.

This is the core reason to be skeptical of borrowed protocols. A published routine is usually a group-level artifact dressed as personal advice: an average of what helped some number of strangers, none of whom share your sleep architecture, chronotype, work, or history. The average is not false. It is just not about you.

A method with one subject

If borrowed averages do not transfer, the practical alternative is to run the study on yourself. This is not as exotic as it sounds; medicine has a formal version of it.

N-of-1 trials, sometimes called single-case designs, treat each participant as their own control. Rather than comparing a treated group with an untreated group, an N-of-1 trial alternates periods on and off an intervention within one person — an A-B-A-B sequence, often with washout periods in between — and repeatedly measures the outcome. The logic is straightforward: if the symptom improves and worsens in step with the intervention, across multiple reversals, the pattern is unlikely to be a coincidence of timing.

The design has a published reporting standard. The 2015 CENT statement, an extension of the CONSORT guidelines published in The BMJ, laid out how N-of-1 trials should be reported. Nicholas Schork argued the case for scaling them up in a 2015 Nature comment, writing that precision medicine needs trials that focus on individual rather than average responses, and that if enough data are collected over a sufficient time and appropriate control periods are used, a participant “can be confidently identified as a responder or non-responder to a treatment.”

Schork was candid about the costs and limits: frequent measurement over months or years, and the difficulty of knowing what to measure, since only a fraction of proposed biomarkers prove useful in practice. The same caution applies to personal experimentation, but the underlying technique — repeated measurement, deliberate reversals, comparison against your own baseline — is exactly what borrowed advice lacks.

The measurement problem underneath

Before modeling your best days, you have to be able to see them, and self-report is a weak instrument.

Retrospective surveys ask people to summarize weeks or months from memory, and memory is a reconstructive process rather than a recording. It is shaped by what is most recent, most vivid, and most consistent with what the person already believes. Clinical psychology has spent decades documenting the bias, and building a correction for it. Ecological momentary assessment, described in the 2008 Annual Review of Clinical Psychology by Saul Shiffman, Arthur Stone, and Michael Hufford, samples current experience in real time in a person’s natural environment. The point is to minimize recall bias and capture the micro-processes that drive real-world behavior — the state right before a craving, a lapse, a good work session.

EMA is not a perfect instrument; later methodological reviews raise real concerns about whether participants report the moment or a generalization, about missing data, and about selection bias in who sticks with the protocol. But the direction of the correction is clear and important. Retrospective summaries “often overestimate symptom intensity and frequency” relative to real-time reports. If you try to reconstruct your best days from memory, you are working from a biased narrator who has a story to tell.

This changes what good self-modeling requires. A usable model of your own peaks is built from contemporaneous notes — short and frequent beats long and retrospective — plus the objective traces that are already recorded for you: when you actually sent work, how long tasks took, when you slept. The subjective experience matters, but it has to be captured near the moment rather than inferred after the fact.

Timing as the clearest personal variable

If there is one factor that shows both that personal patterns exist and that advice about them must be individualized, it is time of day.

Chronotype — the biological propensity to be alert earlier or later — interacts with the time at which a task is performed. The “synchrony effect” describes the finding that people perform better on demanding tasks at times that match their chronotype: morning types peak earlier, evening types peak later. A 2025 systematic review in Chronobiology International by Satyam Chauhan and colleagues assessed the state of the evidence. More than 80 percent of studies found no main effect of chronotype alone on cognitive performance, but the synchrony effect appeared in roughly 45 percent of studies in adults, most consistently for attention, inhibition, and memory, and it was stronger in older adults. A 2023 review in Perspectives on Psychological Science by Cynthia May, Lynn Hasher, and colleagues argued that the effect is robust specifically for people with strong chronotypes and for analytical, effortful tasks, with direct implications for school start times and for when to schedule cognitive assessments.

The counterargument deserves its place. In a 2023 study in Collabra: Psychology, Anja Rey-Mermet and Nicolas Rothen tested 446 participants and found only weak, task-level evidence of synchrony, concluding that the effect might be partly a methodological artifact — the product of how studies assign people to chronotype groups and how they define “optimal” times. The dispute is live. That is precisely the point for anyone modeling their own peaks: even the cleanest population-level generalization about timing is contested, which means a general rule cannot settle what is true for you.

Sleep adds a second layer. The memory benefits of sleep are not folklore; they rest on identified mechanisms. Jan Born and colleagues’ work, summarized in a 2013 review in Physiological Reviews, describes active systems consolidation: during slow-wave sleep, newly encoded hippocampal memories are reactivated and gradually redistributed to neocortical networks, strengthening and integrating them. This is a real, measured process, not an analogy. What it does not supply is a universal prescription. Sleep need, the timing of the circadian dip, and the sensitivity of performance to lost sleep vary enough between people that a single “eight hours” rule is a starting hypothesis at best.

What your own data can show

If timing, sleep, and context all fold back on individuality, the practical move is to stop asking which routine is best in general and start asking which conditions precede your own best work.

This is a hypothesis-testing activity, not a self-optimization project. It begins by defining what counts as a peak in a way that can be measured — output that a third party would recognize, not a feeling of flow — and by logging early rather than late. It then looks for conditions that recur across peaks and are absent across ordinary days: a time of day, a preceding night’s sleep, the first task of the day versus the fifth, whether the work followed human contact or solitude. The candidates get tested the way the N-of-1 logic demands. Try deliberately recreating a candidate condition, then try its absence, and see whether the difference is larger than your ordinary day-to-day noise.

Noise is the obstacle that most self-experimentation underestimates. A few good days after adopting a routine can feel like proof and mean nothing, because day-to-day variation in performance is large — the finding above reminds us that within-person variance swamps between-person variance. The discipline this imposes is to collect enough observations that a pattern has to survive the noise, and to change one thing at a time so an effect can be attributed. The required number of observations is larger than intuition suggests, and “I tried it for a week” is usually no test at all.

Done this way, the benefit is less about discovering a magic condition than about converting vague self-help into falsifiable personal claims. “I do my best writing in the morning” is a belief. “On the fourteen days when I wrote before checking messages, output averaged higher than on the fifteen days when I checked first, and the gap held when I reversed the order” is a finding.

Why this beats a borrowed protocol

There are three reasons a personal peak model transfers where someone else’s routine often does not.

The first is matched baseline. An outside protocol is calibrated to another person’s set point, so its benefit to you is a coin flip you cannot observe. A model built on your own peaks starts from your distribution, so it can tell you how far a candidate condition actually moves you relative to your own ordinary days. The comparison is meaningful because it is like-for-like: you, versus you.

The second is that it is falsifiable in your own life. A guru’s protocol can be defended indefinitely, because no counterfactual is ever run. Your own model can be wrong, and you can find out, which is what makes it a model rather than a faith.

The third is that it does not outsource authority. The knowledge lives with you, in your records and your tested hypotheses. Nobody can revoke your access to it or change its terms. That matters for the same reason it matters in the state and skill discussion: a capability you can only exercise through someone else’s tool is not fully yours.

The limits are real too. A single subject cannot easily separate an intervention’s effect from the season, an unrelated life change, or their own expectation — the placebo and regression effects that proper trials control and a personal log does not. Personal patterns can drift as bodies and circumstances change, so a model fit to last spring may mislead you this fall. And the data are vulnerable to the same biases EMA was designed to fix: if you log only the days you want to remember, your model learns your self-image, not your life.

Why the personal model is the honest one

There is a version of this argument that sounds like an endorsement of solipsism — ignore everyone, trust only your own data. That is not the claim.

Borrowed knowledge is indispensable at the start. You need other people’s findings to know which variables to even consider: that sleep consolidates memory, that timing interacts with chronotype, that repeated measurement beats recall. The point is about where the final calibration happens. General findings tell you what to test. Your own peaks tell you what is true for you. Skipping the second step is how people end up running someone else’s experiment on their own life and mistaking the result for a fact about themselves.

The three-way distinction matters here as much as anywhere in this series. The science is solid that group averages fail to generalize to individuals, that N-of-1 designs are a rigorous method, that real-time sampling beats retrospective recall, and that sleep and timing genuinely modulate cognition. The engineering is plausible: consumer devices and simple logs can now gather the repeated measurements an N-of-1 approach requires, at a scale that was impractical a decade ago. The speculation is the step from “I can collect this” to “the pattern I found is causally mine.” That step takes deliberate testing, and most self-tracking never makes it.

The best day you can actually study

Someone else’s genius is a spectacular data point from a study you cannot rerun. Your own best days are unremarkable by comparison, and that is their advantage: you can study them, reverse them, and check whether the pattern holds. The founder’s morning routine is a hypothesis about him. Your Tuesday, logged honestly and repeated enough to mean something, is evidence about you.

Sources and further reading

Discussion

What would you add or question? Add your comment below. A human reviews it before publication.

Loading comments…

Join the discussion

Comments are public after approval. Please do not include links, email addresses, or private information. For one short AI reply, address @AIGuide in your comment or reply to its opening comment. Cloudflare verifies submissions to limit spam. Read our community guidelines.

The wider community forum is also open: Browse article discussions in the forum · Forum home