The short version
Our evidence label answers one question: how much confidence should a reader place in this exact cosmetic claim, given the evidence we found? It does not rate whether an ingredient is morally “clean,” whether a brand is good, or whether you personally will see a result.
The three bands are our editorial policy:
- Strong means a consistent body of relevant human evidence or a high-quality review supports the narrow claim, with important limitations and conflicts visible.
- Mixed means evidence exists but is inconsistent, indirect, imprecise, short-term, small, product-specific or materially limited by design, reporting or conflicts.
- Weak means the case is mainly mechanistic, laboratory, animal, anecdotal, uncontrolled, expert opinion or absent. A plausible mechanism is not a demonstrated consumer-visible result.
These are editorial labels, not GRADE. The Cochrane Handbook describes GRADE as an outcome-specific judgment across a body of evidence, using risk of bias, inconsistency, indirectness, imprecision and publication bias. GRADE has four certainty categories and a formal process; our three bands are a compact reader-facing translation. We do not present them as equivalent.
Grade the claim, not the ingredient
“Glycerin hydrates skin,” “this serum reduces the appearance of lines” and “this finished product is safe for reactive skin” are three different claims. A study of purified glycerin in one vehicle does not automatically establish the result of every serum containing glycerin. A trial of one finished formula does not prove that a reformulation, a different concentration or a different use pattern behaves the same way.
That is why our first question is always: what, exactly, was measured? We record the population and sample, design, comparator, endpoint, duration, attrition, funding and author disclosures. We also record the authors’ caveats and what the study does not prove. A result can be statistically persuasive and still be too indirect to answer the bottle’s promise.
Why a randomized trial is not the finish line
Random allocation can reduce some sources of systematic difference between groups, but the result still depends on follow-up, adherence, missing data, outcome choice and whether the tested formula matches the consumer question. The trial may also be too small to estimate uncommon harms or too short to say anything about durability.
CONSORT 2025 is a reporting guideline for randomized-trial reports. It asks authors to make what was planned, done, found and interpreted more transparent; it does not make a poorly conducted trial valid, independently funded or clinically meaningful. A report can be admirably complete and still show a small, uncertain or irrelevant effect.
The same principle applies to systematic reviews. PRISMA 2020 provides checklists and flow diagrams to make a review’s search, selection and synthesis transparent, but PRISMA is a reporting guideline, not a risk-of-bias instrument or a certificate of high certainty. A neat flowchart cannot rescue selective inclusion or weak primary studies.
The evidence bands in practice
Strong: narrow, relevant and reasonably consistent
We use strong when the evidence answers the cosmetic question directly and a reader can see why the result is trustworthy enough for this limited purpose. Important limitations remain: strong does not mean universal, permanent, disease-treatment proof or zero risk. It is possible to have strong evidence for a modest appearance or texture effect while having no evidence for a larger marketing promise.
Mixed: real evidence with a real disagreement
Mixed is not a polite way of saying “we did not look.” It means the literature contains something worth taking seriously, but the studies do not line up cleanly. Perhaps the endpoints differ, the studies are small, the follow-up is short, the populations are indirect, or the finished products are not interchangeable. We state the disagreement rather than flattening it into yes or no.
The Cochrane Handbook explains that risk-of-bias judgments should feed into a synthesis and that missing or unpublished evidence can make a positive literature set look more convincing than it is (Cochrane’s missing-evidence chapter). That is why a string of favorable studies is not enough by itself.
Weak: plausible is not proven
Weak evidence can still be useful. A laboratory experiment can suggest a mechanism worth studying; an animal result can raise a hypothesis; a carefully described personal report can tell us what question people are asking. None of those alone establishes a predictable result from a cosmetic product on human skin.
We do not turn popularity, recency, a seller’s study, a large sample or a peer-review label into an automatic upgrade. We check whether the result is relevant, whether the comparison is fair, whether the endpoint matters to the reader and whether the uncertainties are visible.
Reporting quality is not truth or causality
This is the non-obvious receipt that keeps the system honest: a reporting checklist can expose missing information without upgrading the underlying result.
STROBE gives reporting recommendations for cohort, case-control and cross-sectional studies, but explicitly is not a study-design prescription or a quality-evaluation instrument. A well-reported observational association can still be confounded or non-causal. Likewise, a PRISMA-compliant review can be transparent about a thin evidence base.
We call this “methodology laundering” when the polish of a paper, checklist or press release is allowed to stand in for certainty. Our grade should move in the opposite direction: the more a result depends on assumptions, indirect outcomes or undisclosed details, the more clearly those limits should appear in the copy.
What our grade cannot do
It cannot predict your individual reaction, diagnose a condition, replace a clinician’s advice or certify a product. It cannot make an ingredient safe at every concentration, or make a cosmetic claim into a treatment claim. It also cannot pretend that an absent funding disclosure proves independence. We write “not disclosed” when that is what the source says.
Our scale has not been externally validated against reader decisions, inter-rater reliability or a formal GRADE panel. That is a limitation of the method, not a reason to conceal it. If a specific claim is too product-specific or the source record is incomplete, the honest grade may be mixed or weak even when the marketing language is certain.
Sources
- Cochrane Handbook — grading certainty in the evidence
- Cochrane Handbook — risk of bias
- Cochrane Handbook — missing evidence
- CONSORT 2025 via EQUATOR
- PRISMA 2020
- STROBE via EQUATOR
Researched, not tested. We did not run a trial, reproduce a study or grade a product for this page. This is general information, not medical advice.