Roll Call

How we grade

This page is the rulebook behind every ingredient sheet: what the labels mean, where the evidence comes from, and where we draw the line. The one-sentence version — we grade the strength of the evidence behind a claim, never the ingredient, and never the product.

What we grade

We grade claims. Not ingredients, not products. A claim is a specific statement that something does — or does not do — a specific thing: reduces migraine frequency, prevents the common cold, causes allergic sensitization. That statement is the thing being weighed, and the grade describes how strong the evidence behind it is.

One substance can carry several claims, and they can point in different directions. Riboflavin has good evidence behind correcting a deficiency and a separate, weaker case for reducing migraines at a much higher dose. Those are two claims, graded separately, and collapsing them into one verdict about riboflavin would lose the only part that matters — which claim, at which dose.

This is also why an ingredient never gets a “good” or “bad” label here. An ingredient is not a claim. It has no character to grade.

A claim has two parts, and they are separate

Every graded claim answers two questions independently. You will see the answers together on a card, joined by a dash — “Supports — established” — and they are two separate statements, not one.

First: which way does the evidence point?

Supports
The sources we cite point toward the claim being true.
Does not support
The sources we cite point toward it not being true.
Unresolved
The sources we cite have looked, and they do not agree.
Not graded
We have nothing we could retrieve and quote, so there is nothing here to grade. That describes what Roll Call holds, not what has been published.

Second: how strong is that evidence?

established
Strong, consistent evidence.
suggestive
Credible evidence, but the question is not settled.

Strength only exists where there is a finding to be confident about. Supports and does not support carry a strength. Unresolved and not graded carry none — there is no result there to be more or less sure of.

Keeping direction and strength apart lets us publish “does not support — established,” which is one of the most useful things a reader can be told: a popular claim, well studied, and the studies do not back it.

It is also why unresolved and not graded are separate. Unresolved is a finding: the sources we cite have looked and they disagree. Not graded is not a finding at all — it means we have nothing here we could grade against, which is a fact about what Roll Call holds rather than about what has been published. Merging them would let the second read as the first.

The direction is always about the claim next to it, never about the substance. A preservative reading “Supports — established” is not an endorsement of the preservative. The claim it sits on says the ingredient causes allergic sensitization, and the grade says the evidence for that is strong. The claim is the subject of the sentence, so we never show a direction without it.

The other label: the evidence grade

Alongside individual claims, a substance carries one overall evidence grade. It answers a broader question — how strong is the evidence around this substance at all — and it uses a five-step scale rather than the two strengths above. Both can appear on the same card, in similar-looking labels, so it is worth knowing which is which: a grade beginning “Evidence:” is about the substance, and a grade with a direction in front of it is about one specific claim.

What a grade means depends on what is being graded, and the same word does not say the same thing in all three places. On personal care and household products we are grading the evidence behind a concern. On supplements we are grading the evidence behind an effect, at the dose that was studied. So each grade is set out below in all three readings, rather than one of them standing in for the rest.

Evidence: established

Personal care Strong evidence supports a real concern at the exposures people actually get.

Household Strong evidence supports a real concern at the exposures household use involves.

Supplements Strong evidence supports this at the dose studied.

Evidence: suggestive

Personal care Some credible evidence exists, but the question is not settled.

Household Some credible evidence exists, but the question is not settled.

Supplements Some credible evidence exists, but the question is not settled.

Evidence: weak

Personal care The concern is discussed, but the evidence behind it is poor.

Household The concern is discussed, but the evidence behind it is poor.

Supplements The claim is made, but the evidence behind it is poor.

No credible concern

Personal care No credible concern at the concentrations used in cosmetics.

Household No credible concern at the exposures ordinary household use involves.

Supplements never carry this grade. It is a statement about concentrations, and inventing a “safe” grade for supplements is not something we will do.

Insufficient data

Personal care Too little has been studied to say either way.

Household Too little has been studied to say either way.

Supplements Too little has been studied to say either way.

Who was studied

A claim is only as general as the people it was tested on. We state the population the research enrolled, and it is frequently the reason two studies of the same substance appear to contradict each other: a nutrient corrects something in people who are deficient and does nothing measurable in people who already have enough. Both results are real. They are answers to different questions.

This is a fact about the study, not a guess about you. We do not infer whether a finding applies to any particular reader, and a card naming who was enrolled is not telling you that you are or are not one of them.

What was actually measured

A grade rests on measurements, and the kind matters. We say whether a claim was studied by a biomarker — something in a blood test moving — or by an outcome — something happening to a person. Both are real evidence. They are not the same evidence, and “established” on a biomarker is a weaker practical statement than “established” on an outcome.

Where we have not yet noted which one a claim rests on, the card says so. It does not default to the stronger reading.

Dose comes before the grade

A grade is meaningless without the dose it was earned at.

So the dose check runs first. We compare the amount on the label against the dose range used in the research behind that specific claim — and if the label’s amount falls outside that range, we do not show you a grade with a note underneath. We tell you the evidence exists at a dose this product does not provide, which is the honest description of that situation. A graded pill with a caveat below it invites you to read the grade and skip the caveat, and the caveat was the important half.

A recommended intake and a studied dose answer different questions, so we never substitute one for the other.

We also do not convert between units that have no honest conversion. International units have no fixed relationship to milligrams — the conversion differs by substance and by form, which matters most for vitamins A, D, and E. Colony-forming units are a count of organisms, not a mass. Percent Daily Value is a single figure aimed at the general population and is not the same number as a recommended intake for a specific group. Folate is defined in micrograms of dietary folate equivalents, which a label may or may not use. Where a label speaks in units we cannot honestly line up against the research, we say nothing rather than guess.

When there is no dose to compare

Some labels do not disclose one. A proprietary blend names its members and gives a single total, so no member’s individual amount is knowable. We say the dose is not disclosed, and we make no comparison, because none is possible.

The same applies to a panel line listing two forms of a substance under one amount. That number covers both forms and cannot be attributed to either, so no dose comparison is made and the card says why.

What every card shows, in every state

No card is empty. Even where we have nothing graded, a card shows:

  • The amount and form exactly as printed on the label. Not our tidied version.
  • The serving basis, stated plainly. Panels are per serving, and a serving is often two or three units — so the amount on the panel is frequently not the amount in one capsule.
  • The adult reference intake, with the population named — “adults 19+,” not a bare number floating free of who it applies to.
  • The upper limit, and what it covers. These are not always what they appear: niacin’s upper limit applies to supplemental forms specifically, not to niacin from food, and a limit quoted without that scope misleads.
  • A link to the source fact sheet.
  • The date its citations were last checked.

Where no upper limit has been set, that is its own statement, and it does not mean there is no limit. Thiamin, riboflavin, vitamin B12, biotin, and pantothenic acid have no established upper limit because the adverse-effect data was insufficient to set one. That is a gap in the research, not a finding about safety, and we say which it is.

The states you will see

Four of them match the four directions above. Two more are not about the evidence at all — they say what this card does and does not have:

  • No claims graded here yet — nothing on this substance carries a grade, and something here could. No implied all-clear.
  • Nothing here makes a claim to grade — this substance asserts nothing that a grade would describe. Different from the line above, and deliberately: one says a grade is outstanding, the other says none is coming, and water is the second.
  • A claim with no grade beside it — there is a claim, and the evidence behind it does not currently pass the checks above. The card says which: no sources cited, sources cited but none verified, or sources cited whose check has not cleared.

A substance with nothing graded carries no grade label of any kind — labelling it “insufficient data” would state a conclusion nobody reached. An absence of a grade is an absence, and the card says so rather than filling the space.

How an entry is checked

Every citation is resolved against the registry that issued it. The identifier is fetched and the title that comes back is compared to the title stored here. A citation that resolves to a different document, to an index page rather than a document, or that cannot be retrieved at all, is not verified — and the difference between those outcomes is recorded, because they call for different fixes.

Grades are set from the sources that survive that check. Where none survive, the entry shows what it cites and what that citation is currently worth, and no grade. A grade is not shown unless the evidence behind it would pass the checks above today — so a card can carry a claim, and its sources, and still show no grade beside it.

These checks are re-run. A document that moves, is rewritten, or stops resolving does not keep a grade standing on it: the next check finds the citation no longer holds and the grade comes down until it does. Nothing here is checked once.

What these checks cannot do is weigh a body of literature. They establish that a citation is real, that it is the document it claims to be, and that it is still there. They do not establish that it is the best available evidence, and this page will not describe them as though they did.

What counts as a source

Not every source weighs the same, and we do not pretend otherwise. At the top sit formal safety assessments from regulatory and expert review bodies — the Cosmetic Ingredient Review, the EU’s Scientific Committee on Consumer Safety, the FDA, and for supplements the NIH Office of Dietary Supplements — because they weigh the whole body of evidence at the exposures people actually get, not one headline study. Below those comes peer-reviewed research, ranked by how it was done: systematic reviews over single studies, human data over animal data, animal data over cells in a dish. Industry safety reviews and ingredient databases fill in the gaps, weighted accordingly and never mistaken for the last word.

A citation has to resolve to the actual document. A link to a database’s front page is not a source, however reputable the database.

Two hierarchies, because there are two kinds of question

Not everything on a label is a question about your skin. Does this preservative cause allergic sensitization and was the mica in this product mined with child labour are both claims, both gradeable, and both shown here — but they are not answered by the same kind of evidence, and ranking them on one list would be pretending otherwise.

So there are two hierarchies. Every graded claim says which one it was judged under, on the claim itself. A card can carry one of each: mica does.

Health and safety evidence — ranked by how well a study isolates an effect

  1. Formal assessments by regulatory and expert review bodies — CIR, the EU’s SCCS, the FDA, the NIH Office of Dietary Supplements. They weigh a whole body of evidence at the exposures people actually get.
  2. Systematic reviews and meta-analyses — many studies, weighed together.
  3. Randomised controlled trials — the design that isolates cause best.
  4. Observational studies in people.
  5. Animal studies.
  6. Cells in a dish.

Sourcing and supply-chain evidence — ranked by proximity and independence

The question here is not how well was this isolated but how close was the source to what happened, and how independent is it of the party being described. Sample size and randomisation do not apply, and no review body assesses supply chains at all.

  1. Government findings — investigative powers, and legal consequence for the subject.
  2. Independent investigative reporting — documents, named sources, people who went and looked.
  3. NGO investigations and third-party audits — frequently the closest to the site, with one significant limit: audit access is often granted by the party being audited.
  4. A company’s own account of its own supply chain.

A company’s statement about itself cannot establish a claim about itself. We will cite it, and we will tell you what it says — the company states X — but on its own it does not move a claim toward supported or unsupported. That is not a comment on any particular company. Nothing changes its own grade by asserting something about itself.

The list is where a source starts, not where it ends

In both hierarchies the type of source is a starting point that the specifics adjust. Who funded the work, and whether the people who did it declared a competing interest, can move a source below where its type would put it — a regulator that has gone quiet on an industry it oversees, or an audit commissioned by the company being audited, sits lower than its label suggests.

What we record about a source

A citation is not just a link. Several things about a study change how much it is worth, and we keep them on the record: who funded it, whether it was pre-registered, whether a registered trial went unreported, whether the measured outcome was switched after results were in, and whether the paper has been retracted, corrected, or had concerns raised.

These apply the same way to every source. Supplement-industry and advocacy-funded research is held to exactly the standard pharmaceutical research is held to here, and a study is not more trustworthy for having been funded by someone whose conclusions we like.

Unknown means nobody has checked yet. It is recorded as its own state, never as a blank and never as a pass — an unchecked study and one we have verified must not be treated alike.

For a supply-chain source, who commissioned it is most of what places it, so there the record is not optional: a sourcing claim cannot be published at all while any source behind it has no funding or conflicts recorded. That one is enforced by the database rather than by anyone remembering.

These details are not yet shown to you on the card. They are recorded, and they gate what can be published, and putting them in front of you is work we have not done — it is listed below with the rest of the gaps.

What we mean by context

An ingredient is never just an ingredient — it’s an ingredient at a concentration, in a formula, with a use pattern. A preservative that earns real caution in a lotion that sits on your skin all day can be a non-issue in a soap that’s rinsed off in thirty seconds. Concentration works the same way: a remarkable number of “linked to” headlines trace back to exposures hundreds or thousands of times higher than anything in a cosmetic. When a concern only exists at doses far above cosmetic use, we say so plainly — that’s the “sounds alarming, but” you’ll see on some sheets. Context isn’t a loophole. It’s the difference between reading the evidence and reading the headline.

Why there is no product score

We never roll ingredient grades up into a single number, colour, or badge for a product. Aggregate scores feel helpful and mislead almost by design: they treat a trace preservative and a base oil as equals, punish a long ingredient list for being long, and quietly convert “too little data” into “probably bad.” Worst of all, they hand your judgment to whoever tuned the formula. What we offer instead is the roll itself — every ingredient named, explained in plain language, with the evidence behind any claim graded honestly, including when it’s thin. The sheet is ours; the conclusion is yours.

Where the gaps are

This page describes how we grade. Here is where that grading does not reach yet.

  • Coverage is uneven, and stays uneven. There are thousands of substances on supplement and personal-care labels. We work down them by how often they turn up, so the things you are most likely to be holding are covered first. That is how we prioritise, not a queue we expect to empty.
  • The fields on this page get filled in by reading, not by code. Building a place to note what a study measured, who was enrolled, and at what dose does not note any of it. Where we have not read it, the card says so rather than leaving the space blank.
  • What we record about a source, you cannot see yet. Funding, pre-registration, retractions and the rest are stored and used, and no card displays them. Until it does, you are taking that part on trust, which is exactly the thing this page exists not to ask of you.
  • Most sources do not yet have their funding noted. Until a source is checked, it reads as unchecked.
  • Botanicals are the weakest area. A plant on a label is not one thing: the species, the part used, and how an extract was standardised all change what is actually in the capsule, and two products with identical label wording can differ enormously from each other and from whatever was studied. Milligrams alone do not settle it. Our coverage here is very thin.
  • Sourcing claims go stale faster than health claims. A finding about a mine or a plantation describes an arrangement that can change with a contract, and supply chains turn over far faster than biochemistry does.
  • Every card is current as of its date, and no later. Evidence moves.

We would rather list this than let a gap read as a finding.

What this is not

Roll Call does not check drug interactions, does not assess whether anything is safe for you, and does not personalise any of this. Claims printed on packaging — “supports immune health” and the like — are shown exactly as written and never graded true or false; they are not evaluated by the FDA, and grading them is not our job.

For anything about your own health, or how a supplement might interact with a medication: talk to your doctor or pharmacist.