How we rate evidence
Every article opens with an evidence label. It rates specific claims, not the paper as a whole, the scientists, or whether the topic matters. One paper can support one claim well and another poorly.
The four ratings
Several independent studies agree, or one large, well-run trial found a meaningful effect. Reasonable people would act on it.
Good evidence that hasn't been independently repeated yet, or consistent findings with a plausible mechanism. Worth knowing, not yet worth changing your life over.
Animal or cell studies, small studies, or a single association that hasn't ruled out its main alternative explanation. Interesting science; not yet a reliable guide for people.
Hypotheses, unreviewed preprints with extraordinary claims, or studies with serious design problems. We cover these mainly to explain why the headlines went too far.
Finding and popular claim
When the claim going around is bigger than what the paper found, the label rates both on separate lines, for example Finding: Early and Popular claim: Overstated. The lowest rating for a popular claim is called Overstated rather than Speculative, because the overreach belongs to whoever made it: news coverage, a press release, other scientists, or the authors' own earlier wording. The label says whose claim it was.
What earns each rating depends on the evidence
The four words mean the same thing everywhere, but what earns them depends on the kind of evidence behind a claim: studies in people, animal and lab work, field observation, or physical measurements and models. A single story can hold claims rated on different ladders.
Borderline ratings
Two readers rate every claim separately. When they agree and are both sure, the label shows that rating. When either leans toward the next tier, or they split by one tier, the label shows the lower tier and where it leans, for example "Early, close to Promising", and the article says what the call turns on. A borderline rating is a statement with error bars, not a hedge. If two readers ever split by two tiers, the claim is held and the rules get fixed.
Ratings can change
A rating reflects everything known when we publish, including later studies. If later work changes the picture, we say so and update the label, with a dated correction at the foot of the article.