MoreYears reader guide
How to read longevity evidence
A result can be exciting without being ready to guide a decision. This ladder helps separate biological possibility from evidence about what happens in people.
The evidence ladder
Five rungs, five different questions
Research does not always move neatly upward, and a study’s design does not guarantee its quality. The rungs describe what kind of conclusion a method can support—not whether a particular result is automatically trustworthy or important.
-
Cells and mechanisms
Biological plausibility
What it can tell usWhether a process can change in cells or tissues, and which biological pathways may be involved.
Central limitationA controlled cellular effect does not show that the same approach will be effective or safe in a whole organism.
-
Animal research
Whole-organism testing
What it can tell usHow an intervention behaves across organs and systems under controlled conditions that may be impossible in people.
Central limitationSpecies biology, doses, environments, and disease models can differ substantially from human aging and clinical use.
-
Human observational studies
Patterns in people
What it can tell usWhether an exposure, behavior, biomarker, or outcome tends to occur alongside another in real human populations.
Central limitationConfounding and reverse causation can create an association even when one factor does not cause the other.
-
Randomized human trials
Causal effects under trial conditions
What it can tell usRandom assignment can estimate whether an intervention caused a difference for the people, outcomes, and timeframe studied.
Central limitationA trial may be too small or short, use a surrogate endpoint, miss rare harms, or not generalize beyond its participants.
-
Replicated and synthesized human evidence
Consistency across studies
What it can tell usWhether findings hold up across independent studies, settings, populations, or a careful synthesis of the available evidence.
Central limitationA synthesis inherits weak studies, inconsistent methods, publication bias, and limits on personal applicability.
Across every rung
Study design is only the beginning
Two studies on the same rung can deserve very different weight. These questions help reveal whether a result is durable, meaningful, and relevant.
Replication and convergence
Has an independent team found a similar result? Confidence grows when different methods point in the same direction, not merely when one study is repeated in the same setting.
Effect size and uncertainty
Is the difference large enough to matter, and how wide is the plausible range around it? Statistical detection alone does not establish practical importance.
Population relevance
Age, health, sex, ancestry, baseline risk, and selection criteria shape whom a result describes. Evidence in one group may not transfer cleanly to another.
Endpoints
A biomarker, symptom, physical function, disease event, healthspan, and lifespan answer different questions. Movement in one should not be silently translated into another.
Duration
Short studies may detect an early signal while missing whether it lasts, whether benefits accumulate, or whether delayed harms emerge.
Conflicts and funding
Funding or commercial involvement does not automatically negate a result, but readers should know who designed, analyzed, and reported the work.
Peer review
Peer review adds scrutiny but is not a guarantee that methods are sound, reporting is complete, or later studies will agree.
Preprints
Preprints can share results quickly, but they have not completed journal peer review and may change materially before publication.