Evidence Standards

ABOUT & TRUST

A practical hierarchy for interpreting longevity claims without pretending every study answers the same question.

Longevity Research Review evaluates evidence according to the question being asked.

A mouse experiment can be excellent evidence for what happened in that mouse experiment. It is weak direct evidence that the same intervention extends healthy human life.

The study type has to match the claim.

Level 1 — Human outcome evidence

Highest weight is generally given to well-conducted human studies that measure outcomes relevant to people.

Examples include:

  • mortality;
  • major disease events;
  • disability;
  • physical or cognitive function;
  • quality of life;
  • clinically meaningful symptoms.

Randomization, adequate sample size, appropriate controls, sufficient follow-up and replication increase confidence.

Level 2 — Human intervention evidence using validated intermediate outcomes

Some questions cannot reasonably wait for lifespan endpoints.

Human trials using validated risk factors or clinically meaningful intermediate outcomes can still be highly informative.

The key question is whether the measured outcome is known to matter—not merely whether it moved.

Level 3 — Human observational evidence

Cohort, case-control and other observational studies can reveal associations across large populations and long periods.

Confidence depends on:

  • exposure measurement;
  • outcome measurement;
  • control of confounding;
  • reverse-causation risk;
  • selection bias;
  • consistency across populations;
  • dose-response patterns;
  • replication.

Observational evidence can be strong, but causal language requires care.

Level 4 — Mechanistic human evidence

Short physiological studies, pharmacokinetic work, tissue studies and biomarker experiments can explain what an intervention does biologically.

They can strengthen plausibility without proving long-term benefit.

Level 5 — Animal evidence

Animal studies are essential for understanding aging mechanisms and testing hypotheses.

Translation to humans is uncertain.

Species, strain, sex, age, housing, dosing and experimental conditions can strongly affect results.

Level 6 — Cell, molecular and computational evidence

Cell culture, molecular assays and computational models can identify mechanisms and generate hypotheses.

They sit far upstream from proof of a real-world human health benefit.

Reviews and meta-analyses

A meta-analysis is not automatically stronger than every individual study.

Its reliability depends on the quality, comparability and bias of the studies it combines.

A precise pooled estimate from weak or heterogeneous evidence can still produce a weak conclusion.

Guidelines and consensus statements

Professional guidelines can be useful because they synthesize evidence and clinical trade-offs.

We still consider:

  • date;
  • scope;
  • evidence-review method;
  • conflicts of interest;
  • target population;
  • whether the guideline answers the same question as the article.

Preprints

Preprints can be useful for tracking fast-moving research but have not completed journal peer review.

They should be labeled clearly and treated with additional caution.

Statistical significance

A p-value does not tell readers whether an effect is large, important or likely to replicate.

Where possible we consider:

  • absolute effect size;
  • confidence intervals;
  • baseline risk;
  • clinical importance;
  • multiplicity;
  • study power;
  • missing data.

Biomarkers

Biomarkers range from clinically validated measures to exploratory research signals.

A biomarker can be:

  • diagnostic;
  • prognostic;
  • predictive;
  • pharmacodynamic;
  • a risk factor;
  • a surrogate endpoint;
  • purely exploratory.

Those categories should not be blurred.

Replication

A single dramatic result is less persuasive than consistent findings across independent studies and populations.

Independent replication generally increases confidence more than repetition from the same laboratory, sponsor or dataset.

Safety evidence

Absence of a reported adverse event in a small short study does not establish long-term safety.

Safety assessment considers:

  • duration;
  • sample size;
  • dose;
  • vulnerable populations;
  • interactions;
  • monitoring;
  • post-market or real-world data where relevant.

Evidence language

Preferred wording should track confidence.

Examples:

Higher confidence: “reduces,” “improves,” “is associated with a clinically meaningful reduction” — when supported by strong evidence.

Moderate / conditional: “appears to,” “may improve,” “evidence suggests.”

Early / uncertain: “is being studied,” “is biologically plausible,” “preclinical evidence suggests,” “human benefit is not established.”

The purpose of the hierarchy is not to make research sound less interesting.

It is to keep the strength of the language proportional to the strength of the evidence.