Productive— faster every day
For your professionTeachersStudentsManagersMarketingDevelopersFreelancersParents

Tips & tricks · AI · Everywhere · ~hours of arguing and money spent on things with no evidence · 57 min read · in-depth guide, doing it ~3 h

Miracle supplement, or marketing? Verifying health claims with AI and PubMed

Last reviewed:

Illustration for: Miracle supplement, or marketing? Verifying health claims with AI and PubMed
In this article
  1. A typical scenario
  2. Why a shared headline isn’t a study
  3. What PubMed is, and what the connector can do
  4. The mental process: the same recipe as a company brand
  5. Phase 1: inventory — turning a claim into a question
  6. Phase 2: searching — finding the right studies, not the first study
  7. Phase 3: reading the abstract — AI as translator of technical language
  8. Phase 4: the hierarchy of evidence — why a meta-analysis beats a case report
  9. Phase 5: consensus, counterarguments — and how not to let AI agree with you
  10. Phase 6: the registry of verified claims — a living document instead of endless googling
  11. Case study: Věra's email, start to finish
  12. How to talk to loved ones who believe it
  13. When to leave it be and see a doctor
  14. Limits: what PubMed and AI won't give you
  15. Common mistakes
  16. The best tools
  17. What you get out of it
  18. Pro tip

Every family knows this. Grandma forwards an email about a miracle supplement “that doctors don’t want you to know about.” A friend at the pub declares that coffee destroys your heart, because “he read it somewhere.” An influencer with a million followers sells a detox cure while waving around a “study from Harvard.” And you’re standing in the middle of it all with a phone in your hand, doing what everyone does: googling the headline, finding five articles that contradict each other, and dealing with the exact same thing again a week later.

This guide offers a different path: instead of googling headlines, learn to go straight to the source — a database of real studies — and turn every check into a lasting record the family comes back to. Technically, this is made possible by the PubMed connector for Claude, which at the time of writing searches more than 36 million records of medical literature, for free. But the technology is the smaller half. The bigger half is the craft: knowing why a shared headline is not a study, turning a vague claim into an answerable question, reading an abstract with AI as your translator, placing a study on the hierarchy of evidence, and telling a scientific consensus apart from a single isolated paper. That’s exactly what you’ll learn here, step by step, with prompts to copy.

And right at the start, the most important boundary — one this text will deliberately keep returning to: this is a guide to media and scientific literacy, not a health advice line. Neither AI nor PubMed replace a doctor. A doctor decides on treatment, on starting and stopping medication, and on diagnoses — always, without exception. What this process gives you is something different, and also valuable: the ability to not fall for marketing, to tell strong evidence from weak, and to walk into the doctor’s office with better questions instead of a printed chain email. You can read it in phases; if you’re in a hurry, jump to Phase 1 and come back for the context.

A typical scenario

Jana has an eight-year-old son, a mother who’s retired, and a brother who runs a marathon in under four hours. It sounds like the setup to a joke, but it’s actually the setup to three situations that landed on her in the same week — and we’ll build the whole guide on them, because together they cover almost everything you’ll run into in ordinary life.

Situation one: an ad for a kids’ supplement. An ad showed up in Jana’s social feed for a syrup “for children’s immunity and focus,” with a photo of a smiling schoolkid and the phrase “clinically proven.” Her son is sick all the time — like every eight-year-old boy in November — and Jana feels exactly what the ad is banking on: what if it helped? It’s not cheap, but what’s that against a child’s health? A while ago she’d have bought it, or not, based on mood and a discussion on a parenting forum. Now she opens Claude with the PubMed connector, and in twenty minutes she knows: “clinically proven” rests on a single study of forty children, funded by the manufacturer, that measured something different from what the ad promises. That’s not proof the syrup is harmful — it’s proof that for that chunk of money she’d be buying hope, not effect. She decides for herself, informed; and if her son’s frequent illness itself worries her, that’s a question for the pediatrician, not for a supplement from an ad.

Situation two: a chain email at Sunday lunch. Jana’s mother, Věra, forwarded an email: “Scientists have found that [common food] causes cancer — share before they delete this!” At Sunday lunch she talks about it seriously and with real worry; she’s stopped buying that food and is urging the whole family to do the same. Jana’s first instinct is to roll her eyes and say “Mom, that’s a hoax.” But she knows that move never works — her mother just digs in and next time forwards the email to a cousin instead. So instead, that evening Jana runs the claim through PubMed, finds out what’s actually behind the headline (an observational study in rodents, at doses a thousand times higher than reality), and — more importantly — works out how to talk to her mother about it without condescension, using the bridge technique we’ll get to later. And she saves the claim in the registry, because this particular hoax comes back every couple of years wearing a new coat.

Situation three: a “study” from a fitness Instagram account. Jana’s brother Martin runs and follows accounts about sports nutrition. One of them claims a “study proved” that a specific supplement improves performance by 15 percent — and in the next post sells it with a discount code. Martin isn’t naive, but fifteen percent is fifteen percent. He asks Jana to show him how it’s done. Together they find the original paper: eight trained men, one week, measuring something only loosely related to running — and that 15 percent is a relative change in one out of five measured outcomes, with nothing significant on the rest. Martin doesn’t buy the supplement; instead he buys himself an extra hour of sleep, where the evidence holds up much better.

And here’s the part that matters most: next month there’ll be another ad, another email, another post. The difference between a family that relitigates every claim from scratch and a family that has a system is one markdown file — the registry of verified claims, where Jana logged all three checks: the claim, the verdict, the strength of the evidence, the date, links to the studies. When an aunt brings up the same detox at Christmas, Jana doesn’t open a search engine — she opens the registry. And when a big new study comes out a year later, one line gets updated. That’s exactly how a one-off “I’ll just google it” turns into a habit that sticks in a family.

Why a shared headline isn’t a study

Before we reach for the tools, you need six concepts. Not for a quiz — because each one is a specific point where a cautious scientific result turns into a bombastic headline. Anyone who knows these six spots will catch eighty percent of the problems before they even open PubMed. We’ll go through them with examples you’ll see on social media tonight.

A press release isn’t a paper

Between a study and your phone stands a chain: researchers write a paper → the university PR office turns it into a press release → a journalist turns the press release into an article → an editor puts a headline on the article → someone shares the headline with their own commentary. Every link in that chain has an incentive to sharpen the edge. Researchers write “suggests an association in the population studied,” the PR office writes “breakthrough discovery,” the journalist writes “scientists have found,” the headline writes “THIS causes cancer,” and the comment says “I always said so.” The original study can meanwhile be perfectly honest and cautious — you’re just reading the fifth link in the chain.

The practical takeaway: never treat a shared headline as information about the science — treat it as a tip that a study might exist somewhere. Your job is to find it, and that’s what the rest of this guide is for. By the way, the phrase “share before they delete this” is a reliable sign that no study exists at all; real studies never get deleted, unless they’re retracted for errors (and even that can be checked — we’ll show you how).

Relative versus absolute risk

The single most powerful trick in health headlines. “Consuming X raises the risk of disease Y by 50 percent” sounds terrifying. But if the baseline risk is 2 people out of 10,000 and it rises to 3 out of 10,000, that’s a 50 percent increase relatively — and an increase of one person out of ten thousand absolutely. Both numbers are mathematically true; only one of them helps you make a sensible decision. Headlines almost always report the relative change, because it’s the bigger number. The same trick works in reverse for promises: “40 percent better sleep onset” can mean falling asleep seven minutes sooner.

The question to internalize: how many people out of a hundred does this actually affect — before and after? When the abstract doesn’t give absolute numbers (and it often doesn’t), that’s itself informative: the authors, or the press release, chose the more dramatic number. AI is a great help with this conversion, because it pulls the numbers out of the abstract and recalculates them into “people out of a hundred” — you’ll see this in Phase 3.

Correlation isn’t causation

People who drink a lot of coffee show different health outcomes in some studies than people who don’t. Does that mean coffee is to blame? Not necessarily: coffee is drunk more often by people who work office jobs, smoke, sleep less, have a different income… Any of those things could be the real cause, with coffee just an innocent bystander. This is called a confounding variable, and it’s why observational studies (“we looked at who does what and compared outcomes”) can show an association but not a cause.

A different design tests causation: a randomized controlled trial (RCT), where participants are randomly split, with one group receiving the substance under study and the other a placebo. Random assignment ensures smokers, poor sleepers, and office workers alike are evenly spread across both groups — so any difference in outcome can be attributed to the substance. The gap between “is associated with” and “causes” is the single most common place where headlines lie without technically lying: the study found a correlation, the headline wrote causation.

Mice aren’t people

A huge share of “breakthrough” headlines happened in a petri dish or in rodents. That’s not a put-down — that’s how science starts, and there’s no other way to do it. But the road from “works on cells” through “works in mice” to “works and is safe in humans” takes years, and most candidates never finish it. The list of substances that kill cancer cells in a lab dish is long — bleach is on it too. On top of that comes the dose question: effects in animals are often achieved at doses that, scaled to a human, would mean eating kilograms of the food in question every day.

The question to internalize: who or what was this actually tested on? The answer is always in the abstract, usually right in the “methods” section. “In vitro” means cells in glass; “murine model” means mice. If the headline talks about people and the study is about mice, you’re done — not because the study is bad, but because the claim says something the study never claimed.

Sample size and follow-up length

A study with twelve participants can be an honest pilot study — and, at the same time, it’s nearly impossible to draw any conclusion from it that applies to you. Small samples produce wild results: flip a coin twelve times and you can easily land nine heads; flip it a thousand times and it settles near fifty-fifty. That’s why one dazzling result from a small sample so often vanishes when the study is repeated at scale. Follow-up length works the same way: a supplement tested for two weeks tells you nothing about what it does after a year of use — good or bad.

There’s no magic threshold for a “big enough” sample (it depends on the size of the effect), but as a rough guide: tens of participants is a pilot signal, hundreds is where it starts to get interesting, thousands is a solid foundation. And always ask how many people finished the study — if 200 people started and half dropped out, the result is calculated from the half who stuck around, and they’re often different people from the ones who left.

Conflicts of interest

A study on a supplement’s effect was paid for by the supplement’s manufacturer. That doesn’t automatically mean fraud — industry-funded research is common and often good quality. But it does mean a systematic tilt: manufacturer-funded studies reach favorable conclusions more often, partly because unfavorable results tend to quietly go unpublished. Serious journals therefore require “conflicts of interest” and “funding” sections — and you should always check them, or rather, have AI check them for you. The red flag isn’t “funded by the manufacturer” on its own; it’s the combination: manufacturer funding + small sample + a soft outcome + being sold three posts later with a discount code.

Here’s the guide’s first prompt — a quick “preflight check” on a headline, before you go looking for studies at all. It’s a triage tool: it tells you whether a claim is even worth twenty minutes with PubMed, or whether it falls apart at first glance.

Here's a health claim that reached me:
[paste the headline, chain-email text, or ad transcript]

Don't search for anything yet — just break the claim down like a
media analyst:
1. What exactly is being claimed? Rewrite it as one sober,
   emotion-free sentence.
2. What numbers does the claim contain — are they relative or
   absolute? If relative with no baseline, say what would be
   needed to fill that in.
3. Is it claiming an association, or a cause? Quote the words
   that give it away.
4. What marketing or hoax signals do you see in the text (urgency,
   "doctors don't want you to know," a discount code, anonymous
   "scientists")?
5. What would a study have to show to support this claim? Phrase
   it as a question we'll go verify next.
Be skeptical, but fair — don't write "hoax" until we've actually
checked something.

It returns a sober breakdown, and most importantly, point 5 — the question to verify, which we'll work with in Phase 1. Watch out for one failure mode: the model sometimes just "knows" whether a claim is true and writes a verdict with no sources. Don't accept that — a verdict with no traceable studies behind it is just an opinion, even if it's a language model's opinion. That's exactly why we move on to PubMed.

And the second prompt in this section — an exercise worth doing once with the whole family, kids included if they're old enough. Take a real health headline from today's news and have the whole dilution chain reconstructed for you:

Take this real health headline: [paste headline and link]
and run me a game of "science telephone" — reconstruct what the
individual steps between the study and the headline probably
looked like:
1. What the cautious conclusion in the original study probably
   said (in typical scientific language, with caveats)
2. What the university press release made of it
3. What the news article made of it
4. What the headline made of it, and what social media sharing
   made of it
At each step, flag exactly what got lost or sharpened.
At the end, write 3 questions I should ask myself before I share
a headline. This is a media-literacy exercise — clearly label the
invented steps as reconstruction, not as facts about this
particular study.

What's great about this exercise is that it works preventively too: once you've watched "suggests a mild association in men over 60" turn into "THIS is killing you," you never read headlines the same way again. For how to work with kids and AI more broadly — from bedtime stories to critical thinking — see the guide on AI in the family.

What PubMed is, and what the connector can do

PubMed is a public database of biomedical literature run by the U.S. National Library of Medicine. At the time of writing it holds more than 36 million records — from case reports to large meta-analyses — and it's free, with no registration, for anyone. It's not "some health website": it's the exact same database doctors and researchers search. When someone says "a study showed," if that study exists, there's a strong chance it has a record right here, with a unique identifier called a PMID.

One important clarification up front: PubMed is a catalog, not a full-text library. Every record includes an abstract (a structured summary of the study), but the full text is a different matter — it's free only for articles deposited in PubMed Central (PMC), which at the time of writing is roughly 8 million of those 36 million. The rest sits with publishers, often behind a paywall. That sounds like a major limitation, but for our purposes it's smaller than it looks: for judging "what did the study look at, in whom, and what did it find," the abstract is usually enough — and the review articles we'll care about most tend to be open more often.

Now, the connector. At the time of writing, there's an official PubMed connector for Claude — you'll find it in the connector directory (in claude.ai settings, under connectors), and it works on claude.ai, in team and enterprise accounts, and in Claude Code. Turning it on takes a few clicks and costs nothing extra — PubMed itself is a public service. What the connector can do, per its documentation at the time of writing:

  • Search articles by keyword, author, journal, and advanced syntax — Claude formulates the query for you, including the English technical terms you probably wouldn't have put together yourself.
  • Fetch metadata and abstracts: authors, journal, year, study type, abstract. That's the raw material for all of Phase 3.
  • Retrieve full text for articles that are in PubMed Central — that is, the openly available slice of the literature.
  • Find related articles — branch out from one study to others, which is the foundation of the consensus check in Phase 5.
  • Trace a citation: from a fragment like "Smith, 2021, some journal," find the actual record and its PMID. Invaluable when an influencer "cites a study" with no link.
  • Convert identifiers between PMID, PMC ID, and DOI — handy when you have a link from another site and want the PubMed record.

The exact shape of connectors keeps evolving, which is why I keep writing "at the time of writing" — before you start, check the connector directory for the current state. The general principle of how connectors work and how you hook them up (and why they're to AI what USB-C is to hardware) is explained in the overview of MCP connectors. And one privacy note that matters especially for health topics: your PubMed queries are literature searches, not medical records — but even so, don't put diagnoses and specific people's details into your prompts where it isn't necessary ("my mom, who has type 2 diabetes, takes drug X"). Ask generally ("what do studies say about X in seniors") and keep sensitive family details out of the chat — ideally in the spirit of the chapter on ethical and safe AI.

The first prompt with the connector is a connection test — a one-minute check that everything's working, and also a demonstration of what the output looks like:

You have the PubMed connector attached. Try it on a neutral query:
find the 3 most-cited systematic reviews on the effect of sleep
on immunity in adults from the last 10 years.
For each, return: title, authors, journal, year, PMID, and a
two-sentence summary of the abstract in English. Don't make
anything up — if the connector returns no results, tell me
plainly, including the error message.
At the end, briefly describe what query you sent to PubMed.

That last sentence is a habit worth having from day one: asking to see how the AI phrased its search. Partly you learn from it (you'll see the English terms and filters it used), and partly you'll catch it when a search fails — say, when the model searched too narrowly and came back with "nothing exists" on a topic that actually has hundreds of papers.

How PubMed search actually works — and why it's worth knowing

You don't need to learn PubMed's query language — that's what the connector is for; AI assembles the syntax for you. But it's worth understanding three principles, because they explain why AI searches the way it does, and when to step in.

The first principle: a controlled vocabulary. PubMed doesn't rely on plain keywords alone — records are manually tagged with a shared vocabulary of subject terms (called MeSH, Medical Subject Headings). Thanks to that, a search finds studies that use a synonym in their title too: searching under the tag for vitamin D also turns up papers that write "cholecalciferol." The practical takeaway: when the model's query listing uses terms you never typed, that isn't arbitrary — it's translating your question into the database's vocabulary. And conversely, when a search fails, it's worth saying "try it through the corresponding MeSH terms."

The second principle: filters by study type. PubMed can restrict results to meta-analyses, systematic reviews, RCTs, humans versus animals, and date ranges. This is exactly what the "top-down first" strategy from Phase 2 relies on — which is why you'll see phrasing like "systematic reviews from the last 10 years" in the prompts: those aren't just figures of speech, they're instructions that translate directly into filters.

The third principle: searching is iterative. The first query is almost never the right one — it either returns thousands of results (too broad) or none (too narrow). Professional search specialists tune a query over several rounds, and it works the same way with the connector — the difference is that the rounds take seconds. So always ask the AI to show how it searched, and to say how many results the query returned — "I found three studies" means something very different when the query returned three results total versus when it returned eight hundred and the model picked three.

A lasting setup: a verification project

Before we get into the phases, one five-minute investment that pays off on every check afterward. Set up a Project on claude.ai — call it something like "Claim verification" — and put the ground rules in its instructions, so you don't have to repeat them in every chat:

You're my assistant for verifying health and scientific claims.
Rules that apply in every conversation in this project:
1. You're not a doctor and you don't give treatment advice. If my
   question veers toward dosage, starting or stopping medication,
   or symptoms, stop and remind me that this belongs with a
   doctor.
2. Every claim about a study is backed by a PMID from the PubMed
   connector. A study with no traceable PMID doesn't exist.
3. Always state the study type (case report, observational, RCT,
   meta-analysis), the sample size, and who or what it studied
   (humans/animals/cells).
4. Actively look for counterarguments: for every conclusion,
   mention the strongest study or argument AGAINST it.
5. Distinguish "is associated with" from "causes," and convert
   relative numbers to absolute ones ("X people out of 100")
   whenever the data allows it.
6. Speak English, and explain technical terms the first time
   you use them.

This is also your first safeguard against sycophancy (point 4 — we'll come back to it in Phase 5) and against quietly sliding from literacy into "advice" (point 1). Later on you'll upload the registry of verified claims into this project too, so every new chat already knows what you've verified.

The mental process: the same recipe as a company brand

It might sound surprising, but the process we're learning here isn't a health-specific invention. It's the exact same mental pattern used elsewhere on this site by the guide to brand as a system — there, scattered materials get turned into a brand voice, a brochure, and a CRM; here, scattered headlines get turned into family knowledge. The steps are identical, and it's worth naming them, because you'll find them useful for the third and fourth topic you apply them to as well:

  1. A model scenario — you don't start with a tool, you start with a specific problem belonging to a specific person. Here: Jana and the syrup, Věra and the email, Martin and the supplement.
  2. Inventory — what do we actually have? There, the company's raw materials; here, the exact wording of the claim and its breakdown into a verifiable question (Phase 1).
  3. A single source of truth — there, brand voice in git; here, the primary literature on PubMed instead of the fifth retelling in the media (Phases 2–5).
  4. Derived outputs — there, the website and brochure read from the brand voice; here, the verdict, the answer for grandma, and the purchase decision are all derived from the same verification, not reinvented three times from scratch (Phase 6).
  5. Editability — there, an .af file instead of a generated image; here, the registry as a living document: a new study doesn't overwrite your memory, just one line in a file.
  6. Costs and limits, stated honestly — there, a chapter on what not to do in a chat; here, a whole section on paywalls, publication bias, and when to close PubMed and book a doctor's appointment.

If you've read the company guide, the next six phases will feel familiar. If not, no matter — but next time you build anything else with AI, remember the recipe stays the same: scenario, inventory, source of truth, derived outputs, editability, limits.

Phase 1: inventory — turning a claim into a question

The most common reason verification stalls isn't a shortage of studies — it's a vague claim. "Turmeric is healthy" can't be verified; it doesn't say for whom, for what, in what form, or compared against what. Researchers have a tool for exactly this problem, called PICO, and we'll borrow it because it's brilliantly simple: any health question can be broken down into Population (who?), Intervention (what's being done, or what are they exposed to?), Comparison (against what — placebo, nothing, a different substance?), and Outcome (what exactly is being measured?).

The difference is dramatic. "Does vitamin D help?" is unanswerable. "Does taking vitamin D in otherwise healthy adults in winter reduce the number of respiratory infections compared with placebo?" is a question the literature can actually answer, and it can be found. You don't need to memorize PICO as an acronym — just remember you're asking: who, what, compared to what, and what exactly is being measured.

Breaking a claim down with AI

In practice you don't do this breakdown by hand — it's the first thing you type into a new chat (inside the project from the previous section). Let's take Jana's syrup:

I'm verifying a claim from an ad for a kids' dietary supplement.
The ad says: [paste the ad text verbatim, e.g. "syrup with
extract X supports your child's immunity and focus — clinically
proven"].
Ingredients per the online store: [list the declared active
ingredients].

Break this down into verifiable questions using the PICO method
(population, intervention, comparison, outcome) — in plain
language, not academic:
1. How many separate claims does the ad actually contain?
   ("immunity" and "focus" are two different claims)
2. For each claim, write a one-sentence PICO question, plus
   English keywords for searching PubMed.
3. Flag any claim that's phrased so vaguely it can't be verified
   at all — and what the ad would have to say instead to make it
   verifiable.
Don't search yet, just prepare the questions. I'm a parent, not
a scientist — explain it without jargon.

A typical result: one ad sentence unpacks into three or four separate claims, and about half of them turn out to be legally bulletproof precisely because they don't claim anything concrete ("supports immunity" doesn't measure anything). That's a finding in its own right: the vaguer the promise, the smaller the chance there's anything behind it — concrete effects get sold concretely. Check that the suggested keywords match the declared substance, not the product's brand name; searching PubMed by a brand name almost never turns up anything, and the model sometimes misses that.

When a claim has no source at all

Věra's chain email differs from the ad in one way: it doesn't say "clinically proven," it says "scientists have found" — no names, no year, no journal. You can still work with that, just differently: instead of tracking down a specific study, you check whether any literature on the topic exists at all, and what it says.

I got a chain email with this claim: [paste the claim, e.g.
"scientists have found that reheating food in a microwave
destroys nutrients and creates carcinogens"]. The email cites no
source.

1. Break the claim into separate verifiable questions (PICO,
   plain language).
2. For each question, sketch what the world would look like if
   the claim were true: what would we expect to find in the
   literature? (large studies, official statements, reviews)
3. Prepare PubMed search queries — both to confirm it AND TO
   DISPROVE IT: also phrase a query that would surface studies
   showing the opposite.
4. Try to trace where the claim historically comes from — it's
   often a misreading of one specific old study. If you're not
   sure, say so.

Point 3 is a small thing with a big effect. If you only search for confirmation, you'll find it — there's at least one weak study "for" almost anything. If you search in both directions from the start, you get a picture instead of ammunition. We'll come back to this in Phase 5, because it's the single most important habit in this whole guide.

Martin's case: when there's a source, but it's hidden

Fitness Instagram is a strange genre: a "study" gets mentioned, sometimes even with a screenshot of a graph, but no link. This is where tracking down a citation comes in — the connector can find the actual record from fragments:

An influencer claims: [paste the claim, e.g. "new study:
supplement X improves endurance performance by 15%"]. The post
has a screenshot of a graph with a caption reading [transcribe
whatever's legible: author's last name, year, journal name,
anything at all].

1. Try using the PubMed connector to trace which study this is
   (search by author, year, and topic). Return candidates with
   PMIDs, and for each, note how well it matches the screenshot.
2. If you find the study, compare: what exactly does the abstract
   claim, versus what the post claims? List the differences.
3. If you can't find it, say so clearly — and explain what that
   means (maybe a preprint, maybe a conference poster, maybe it
   doesn't exist). Don't speculate about the results of a study
   you haven't actually seen.

Three possible outcomes, and all three are useful. The study exists and says what the post claims — good, you move on to assess it in Phase 3. The study exists and says something else — the most common case, and now you're done and you have your answer ready. The study can't be found — and it's fair to name it exactly that way: "a study nobody can find." Watch for one trap: the model sometimes tends to pick the "closest match" and present it as if it were the one. That's why we want candidates with a confidence level, not one overconfident answer.

The output of Phase 1 is one to five concrete questions with English keywords. It takes five to ten minutes and is the best-spent time in the whole process — feed the best database in the world a bad question, and all you'll find is noise.

Phase 2: searching — finding the right studies, not the first study

Now we put the connector to full use. The goal of Phase 2 isn't to find "some study" — that's exactly the mistake chain emails make. The goal is to figure out what the landscape of evidence around the question looks like: how much exists, of what kind, and whether summary papers already exist that have done the assembling for us.

Top-down first: reviews and meta-analyses

The golden rule of searching: don't start with individual studies, start with systematic reviews and meta-analyses. These are papers whose authors did exactly what you're now attempting — found all the relevant studies on a question, threw out the low-quality ones, and statistically summarized the rest. When a decent, recent meta-analysis exists for your question, ninety percent of the work is already done. When it doesn't, that's information too: the topic is either too new or too niche — and the one isolated study you stumble on stands alone, with nothing backing it up.

Use the PubMed connector. My question (from the PICO breakdown):
[paste the question, e.g. "does extract X reduce the number of
respiratory infections in school-age children compared with
placebo?"]
Keywords: [paste the English keywords from Phase 1]

Search in this order, and pause after each step:
1. Systematic reviews and meta-analyses on the question from the
   last 10 years. Return up to 5 of the most relevant: title,
   year, PMID, how many studies and participants they include,
   the main conclusion in one sentence.
2. If none exist, the largest randomized controlled trials.
3. If none of those exist either, tell me — and list what does
   exist on the topic (observational studies? animals?), so I can
   see what stage of understanding the topic is at.
For each result, state exactly how well it matches my question:
same population? same form of the substance? same measured
outcome? List the differences — "roughly the same" isn't good
enough for me.

The last paragraph of the prompt guards against the "almost the same study" trap. A classic example: you're asking about children, the study looked at adults; you're asking about a syrup with an extract, the study looked at the pure substance injected in a hospital. The model likes to hand over the "closest match" as if it were the answer — and a difference in population or delivery form can be the whole story. Ask for the differences to be spelled out explicitly, and decide for yourself whether they matter.

Read the map, not just the results

The second step of searching is short but changes your perspective: instead of a list of studies, have the terrain described to you.

Stay on my question. Before we dive into individual studies,
describe the landscape of evidence to me — like you're sketching
a map:
1. Roughly how much exists on this question? (order of magnitude:
   a handful of studies, dozens, hundreds?)
2. Which types dominate — cells/animals, observational, RCTs,
   reviews?
3. Do the results agree, or is the field divided? If divided,
   along what lines (older vs. newer, industry-funded vs.
   independent, different populations)?
4. Is there a visible pattern over time? (enthusiastic small
   studies early on, more sober large ones later — that's a
   common pattern)
Back every claim with a specific PMID you've found in step 1, or
go find more. Where you're not sure, say so.

Point 4 describes a pattern you'll see again and again, and it's worth knowing by name: the decline effect, a fading effect. The first small studies of a new supplement tend to be enthusiastic (small samples, publication bias, excited authors), and as larger, stricter studies come along, the effect tends to shrink — often close to zero. When the map shows you that shape, it's a strong signal that the enthusiastic headline is most likely citing an early point on that curve.

When a search finds nothing

"There's nothing on this" has two very different causes, and it's worth telling them apart. Either the query is bad — too narrow, wrong terms, a brand name instead of the substance. Or the literature genuinely doesn't exist. A widening prompt helps settle it:

The search for the question [question] returned almost nothing.
Before we conclude the literature doesn't exist, try widening it
systematically:
1. Rephrase the query 3 ways: broader terms, synonyms for the
   active substance (chemical name, common name, older
   literature's names), related outcomes (instead of "focus,"
   maybe cognitive tests).
2. Loosen the population: if there's nothing on children, what
   about adults? But tell me clearly that it's now a different
   population.
3. Check related articles around the little we did find.
If there's still nothing — write the conclusion: "the claim is
unsupported, because the literature on it is essentially
nonexistent," and explain why that's a different verdict from
"studies have proven it doesn't work."

That final distinction is subtle but honest: "there's no evidence it works" and "there's evidence it doesn't work" are two different sentences. They get written into the registry differently, and they get discussed with grandma differently. Marketing lives exactly in that gap — "it's never been disproven!" — and now you know how to name it.

By the way, for bigger questions where you'd rather have one thorough report with citations than a back-and-forth dialogue, deep research is worth using — a mode where AI works through dozens of sources on its own and returns a structured report. How to work with it is covered in the guide to deep research, using market research as the example; the principle is the same, just the sources here are studies. For everyday family verification, though, a dialogue with the connector is better — you see every step and learn as you go.

Phase 3: reading the abstract — AI as translator of technical language

You have five studies with PMIDs. Now you need to actually read them — and this is where most people give up, because an abstract looks like this: "In this double-blind, placebo-controlled trial, 120 participants were randomized…" Good news: the abstract is the most structured text in the world. It almost always has four parts — why the study was done (background), how (methods), what came out (results), and what the authors think about it (conclusions) — and AI can break it apart and translate it so that anyone who can read a recipe can understand it.

One rule up front: read the abstract for the methods and the results, not the conclusions. Conclusions are written by the authors, and authors are fond of their own studies; the phrasing in the conclusions tends to run a shade more optimistic than what the numbers in the results actually show. The press release then quotes the conclusions, and the headline sharpens it one more notch. You're going the other direction: toward the numbers.

A structured translation

The core working prompt of Phase 3 — you'll use it on every study you're taking seriously:

Use the PubMed connector to pull the abstract of study PMID
[number] (if the full text is in PubMed Central, pull that too).
Translate it into plain language, structured:

1. QUESTION: What was the study asking? (one sentence)
2. WHO: How many participants, what kind (age, health, country),
   humans/animals/cells? How many completed the study?
3. HOW: Study type (explain what that type means). What exactly
   did they receive or do, for how long, and what was it compared
   against?
4. WHAT CAME OUT: The main results, IN NUMBERS. Convert relative
   figures to absolute ones ("out of 100 people…") wherever the
   data allows. What was significant and what WASN'T — list the
   measured outcomes with no difference too.
5. CATCHES: Sample size, follow-up length, who funded the study,
   how conflicts of interest are disclosed, anything the authors
   themselves admit in the limitations.
6. IN ONE SENTENCE: What does this study say about my claim
   [claim] — and what does it not say about it?
Stick strictly to what's in the text. Where the abstract doesn't
cover something, write "the abstract doesn't say" — don't fill it
in from guesswork.

Two spots to stay alert on. First, point 4: "what wasn't significant" is essential. Studies often measure five things, one of them comes out significant, and that's all anyone talks about — this is called outcome cherry-picking, and it's usually visible from the abstract (the methods list what was measured, and the results go quiet on half of it). Second, the final instruction: without it, the model tends to quietly fill gaps in the abstract with general knowledge, and then you can't tell what came from this specific study and what came from the model's "memory." "The abstract doesn't say" is a full and important answer in its own right.

Numbers under the microscope

For a study a real decision rests on (Jana's syrup, Martin's supplement), it's worth going a layer deeper — this is exactly where you convert relative numbers to absolute ones, and ask whether the effect is even practically meaningful:

Stay on study PMID [number]. Now just the numbers, slowly and
plainly:
1. What was the result in the CONTROL group and what was it in
   the INTERVENTION group? Give both numbers side by side, not
   just the difference.
2. Convert the effect into an absolute expression: "Out of 100
   children who took the syrup, X avoided one extra infection" or
   similar. If that can't be calculated from the abstract, say so.
3. Statistical versus practical significance: even if the result
   is statistically significant, is the difference big enough
   that a person would actually notice it in everyday life?
   Compare it to something tangible.
4. Confidence interval, if given: what's the worst-case and
   best-case outcome the data are compatible with? Explain it to
   me on this specific case, not a textbook definition.
I'm a layperson — accompany every number with a sentence about
what it means for a "buy/don't buy" decision, but leave the
actual decision to me.

Point 3 is a gem no headline will ever give you. "Statistically significant" only means "this probably isn't random noise" — it says nothing about whether the effect is big. Shortening a cold by six hours can be statistically robust and practically worthless; with large enough samples, even a tiny difference comes out significant. The question "would I actually notice this?" brings the numbers back down to earth.

What the study doesn't say

The third prompt of Phase 3 is short, and it's my favorite question in this entire guide — because it's exactly in the gap between "what the study says" and "what's being claimed on top of it" that almost all health marketing lives:

I have a claim: [the original claim from the ad/email/Instagram
post]. I have a study, PMID [number], that we've gone through.

Put them side by side in two columns:
- What exactly the ad/post claims
- What exactly the study showed
Then list EVERY gap between them. Typically: a different
population (adults vs. children), a different dose or form, a
different measured outcome (a blood marker vs. an actual illness),
a different duration, an association presented as a cause, one
result cherry-picked out of many measured.
For each gap, note how big a problem it is: cosmetic difference /
substantial difference / the claim doesn't hold up.

For Jana's syrup, this prompt is where the whole verdict fell out: the study measured the level of one blood marker in forty adults over three weeks — the ad promises fewer illnesses and better focus in children. Between those two are four gaps, and any single one would have been enough. "Clinically proven" technically doesn't lie: some clinical study does exist. It just doesn't prove what a parent pictures under that sentence. This is exactly the literacy you're building here — and notice that none of it required any medical training, just the right questions in the right order.

Full text: when it's worth it

For most family checks, the abstract is enough. Reach for the full text (when it's available in PMC — the connector can retrieve it at the time of writing) in three cases: when money is on the line and the abstract is ambiguous; when you need a table of every measured outcome (to catch cherry-picking); and when you care about the limitations section, where authors admit their own weaknesses — it's often surprisingly candid, and only half a sentence of it survives into the abstract. You can ask with the same prompts, just with richer material to work from. And when there's no full text? Don't despair — honestly note "assessed from the abstract" in the registry; for a family-level verdict, that's almost always enough, it's just worth being aware of and recording.

A small glossary of scientific hedging

One more skill for Phase 3, small but quietly powerful: reading the level of confidence baked into the phrasing. Scientific English has a finely graded vocabulary of caution, and headline translations systematically flatten it — everything gets rendered as "scientists have found." Here's a translator for the most common phrases; AI preserves them when translating an abstract if you ask it to (and the project instructions from earlier already do):

  • "Is associated with" — the most common, and most commonly twisted, phrase. It means correlation, nothing more. The headline turns it into "causes," and right there you've lost the single most important piece of information in the abstract.
  • "May/might contribute to" — double caution: maybe, and at most a contributing factor. The authors themselves are saying this is a hypothesis.
  • "Suggests" — the data point in a direction but can't carry a firm conclusion. An honest word that press releases turn into "proves."
  • "In a murine model / in vitro" — mice, or cells, respectively. Whenever you see this, the whole claim is about the lab, not about people.
  • "Statistically significant" — "this probably isn't random noise." It doesn't say the effect is large or important; that's what the practical-significance question from the numbers prompt is for.
  • "No significant difference" — the study found no difference. Watch both directions here: it doesn't mean "proven that there's no difference" (the sample might have been small), but it's also not a minor detail to skip when citing the study — and that's exactly what happens all the time.
  • "Well tolerated" — a statement that participants could handle the substance, not that it works. Marketing loves to pass this off as proof of effectiveness.
  • "Further research is needed" — a ritual closing line in almost every study, but it does carry information: the authors themselves don't consider the question settled. When this sentence appears in a study that an ad cites as definitive proof, you've got a contradiction in black and white.

A good exercise to build the habit: on your next three checks, have these phrases highlighted and translated literally in the abstract translation. By the third study you'll start spotting them yourself — and along with them, the gap between what studies cautiously suggest and what gets waved around in arguments afterward.

Phase 4: the hierarchy of evidence — why a meta-analysis beats a case report

Not all studies are created equal, and science has a ladder for this, called the hierarchy of evidence. It's not dogma — a badly done meta-analysis is worse than an excellent cohort study — but as a first orientation it's reliable: the higher up the ladder, the harder it is to explain the result away as chance, bias, or the authors' wishful thinking. Here's the whole ladder, with a plain-language explanation:

LevelWhat it isIn plain termsWhat to do with it
Cells and animals (in vitro, mice)Lab experiments, on cell cultures or rodents"Something interesting happened in a dish/in mice"Science's starting line, not evidence for humans. A headline about people + a study on mice = end of verification
Case reportA description of one remarkable patient, or a handful of cases"This happened to one person, and it was worth writing down"A curiosity and a hypothesis, never evidence. "It worked for my neighbor" in a lab coat
Cross-sectional studyA snapshot of a population at one point in time: who's doing what and how they're doing"We photographed a crowd and are looking for patterns"Shows an association, doesn't tell you what came first — whether sick people drink tea, or tea drinkers get sick
Case-control studyComparing sick people with healthy ones and looking back at how they differed"A backward investigation through people's memories"Useful for rare diseases, but memory is unreliable and bias creeps in everywhere
Cohort studyFollowing a large group forward for years and recording what happens"A long-running reality show with thousands of participants"The strongest observational design; but it still shows associations, not causes
Randomized controlled trial (RCT)Random assignment to a group receiving the substance and one receiving placebo, ideally blinded"A fair head-to-head match against placebo under identical conditions"The first level that can actually say "causes." Strength depends on size and duration
Systematic review and meta-analysisA summary of all the good-quality studies on a question, often with combined statistics"A referee who watched every game, not just last night's"The strongest evidence available — provided the underlying studies are good quality and independent

Three notes that don't show up in the table but make the difference between using the ladder mechanically and using it well.

First: the ladder has to be read alongside the question. For "does X cause Y in humans?" an RCT is king. But you can't always run an RCT — nobody's going to randomize people into smoking for thirty years. For questions about long-term risk, large cohort studies are the best evidence we'll ever have, and it's fair to treat them that way. Conversely, for a dietary supplement that's easy to test against placebo, the absence of an RCT is telling: if the effect were real and manufacturers were confident in it, the trial would exist.

Second: a meta-analysis is only as good as the studies it summarizes. A meta-analysis of ten small, sloppy studies is a sloppy study with a fancy name — "garbage in, garbage out." So even with a meta-analysis, ask: how many studies, how many participants total, how good are the inputs (good reviews assess the quality of their inputs themselves and say so).

Third: one level up beats a pile of levels below. A thousand case reports don't add up to one RCT. This principle is the antidote to "but there are SO MANY studies!" — the question is never how many, but which kind.

Classifying with AI

In practice you don't apply the ladder by hand — it's a prompt that follows on from Phase 3:

For the studies we found on the claim [claim] (PMIDs: [list]),
build a hierarchy-of-evidence table:
| study | type | level on the ladder | sample | duration | population | funding |
Below the table, write:
1. What's the HIGHEST level of evidence this claim actually has?
   (one sentence)
2. What's missing from the ladder — and is that suspicious? (e.g.
   the substance has been sold for 20 years, would be easy to
   test against placebo, and no RCT exists)
3. If the claim were true at the strength it's being sold at, what
   would the evidence landscape look like? Compare that with
   reality.
For the study type, rely on the abstract and metadata, not
impression — if the type isn't clear, write "unclear from the
abstract."

Points 2 and 3 are a version of the classic "where's the catch" question: for things that genuinely work, the evidence landscape tends to fill in over time, moving upward (from mice to RCTs and reviews); for marketing claims, the same handful of low-level papers gets recycled for years. When a substance has been around for twenty years and the strongest evidence is still an observational study on a hundred people, that means something.

Quick calibration: how much evidence is "enough"

The last prompt of Phase 4 helps with the question everyone eventually asks: okay, so what? The answer depends on what's at stake — and this framework is worth having explicit:

Summarize the strength of the evidence for the claim [claim]
based on what we've found, then run it through three different
thresholds of demand:
1. "Interesting enough to keep an eye on" — a consistent signal
   from observational studies would be enough. Does it clear this?
2. "I'd spend money on it" — I'd want at least one solid RCT in a
   population similar to mine. Does it clear this?
3. "I'd change my behavior long-term because of it" — I'd want a
   review/meta-analysis or multiple independent RCTs. Does it
   clear this?
One sentence per threshold, yes or no and why. Reminder: this
isn't about treatment or medical advice — it's a calibration of
how strong the evidence is against what I'd actually do based on
it. The decision is mine, and anything health-related belongs in
consultation with a doctor.

This is also a subtle but important point: there's no single answer to "is this proven?" — there's an answer to "proven enough for what." Curiosity takes very little; your wallet takes more; changing your lifestyle takes a lot; and anything involving treatment isn't settled by any amount of research you do yourself — that threshold has a name, and it's "doctor." Once you internalize this framework, the endless media flip-flopping of "coffee is good for you/bad for you" stops being annoying: it's usually observational signals bouncing around threshold 1, presented by headlines as if they'd cleared threshold 3.

Phase 5: consensus, counterarguments — and how not to let AI agree with you

One study isn't science; science is a conversation among thousands of studies that correct and add to each other. A headline stands almost always on one paper — you want to know what the conversation as a whole says. And there's a trap lurking right inside the tool you're using: language models have a built-in tendency to agree. Ask "right, that detox is nonsense?" and you'll get eager agreement; ask the exact same question as "right, that detox works?" and the asker gets a more accommodating answer too. The model mirrors your framing — and during verification, that's fatal, because all you end up doing is confirming what you already thought, at greater expense.

The defense has three layers: neutral question phrasing (ask "what does the evidence say about X," not "right, X?"), explicitly requesting counterarguments (already baked into the project instructions, but ask for it separately on any verdict that matters), and separating roles — having the AI argue for once and argue against once, like in a courtroom. Here they all are, in action.

The opposing counsel

Before you settle on a verdict, have it challenged. Use this prompt on every check where the outcome matters to you — and deliberately on the verdict you're leaning toward:

For the claim [claim], based on the research so far, I'm leaning
toward the verdict: [your working verdict, e.g. "weak evidence,
mostly marketing"].

Now be opposing counsel. Your job is to take my verdict apart:
1. Use the PubMed connector to find the strongest studies that
   argue AGAINST my verdict. No straw men — the best the other
   side has. With PMIDs.
2. List the weaknesses in my research so far: what might I have
   overlooked, what population or framing did I fail to consider,
   where might I have slid into confirming my own opinion?
3. Under what circumstances would the person spreading the claim
   actually be right? Formulate the strongest honest version of
   their position (a steelman).
At the end, do NOT write a conciliatory "both sides have a point"
— write whether your counterarguments actually threaten my
verdict or not, and why.

That final instruction matters twice over. Without it, the model, after a sharp critique, often folds into a toothless "it's complicated" — and you need to know whether the critique sank the verdict, dented it, or bounced off. And if the opposing counsel genuinely finds a strong RCT you'd missed, that's the best possible outcome: the registry gets a better verdict than it would have had otherwise. Losing an argument with yourself is cheap; losing it in front of grandma is expensive.

Is this a consensus, or a lone outburst?

Second step: figure out where your findings stand relative to the rest of the field. The connector can find related work from any given study — exactly the tool for "does everyone else agree?":

Take the key study of our verification (PMID [number]) and map
its surroundings:
1. Find related articles and newer papers on the same question.
   Did anyone independent replicate the result? With PMIDs.
2. Are there studies asking the same question with the OPPOSITE
   result? Actively search for them — phrase the query so it
   would find them (e.g. terms like "no effect," "failed to
   replicate," and similar).
3. Do newer reviews cite this study? And how do they treat it —
   as support, or with reservations?
4. A verdict on its standing: is this study part of a mainstream
   (multiple independent teams, similar results), or a lone
   outburst (one group, nobody replicated it, reviews ignore it)?
If the findings are mixed, describe WHAT DIVIDES them — that's
usually more informative than a 3:2 scoreboard.

Point 2 deserves underlining: studies with a negative result are harder to publish and harder to find (publication bias — more on this in the limits section), so you need to search actively, in different words than the ones you'd use to find positive findings. And the practical translation of point 4: a reasonable decision is never built on one study — headlines are built on one study. When someone argues "but THIS study showed…," the right response isn't to argue about that one study — it's to ask what the others showed.

A background check: retractions and conflicts of interest

The last check before a verdict is quick, and occasionally dramatically worthwhile. Studies sometimes get retracted — for errors, and for outright fraud — and a retracted study can keep circulating online as "evidence" regardless; some of the most famous health hoaxes stand on exactly retracted papers. And the conflicts of interest we mentioned in the intro can be checked for a specific paper, directly:

For the key studies of our verification (PMIDs: [list]), run a
background check:
1. Has any of them been retracted, or does it carry a warning or
   correction (correction, expression of concern)? Check the
   record's metadata; if the connector can't tell, say so and
   suggest where I can check manually.
2. Who funded the studies, and what conflicts of interest do the
   authors disclose? Quote what's in the record; where it's
   missing, write "not disclosed."
3. Has the lead author's group published a suspiciously large
   number of similar papers on this topic? (one group can create
   the appearance of a consensus all by itself)
None of this disqualifies a study on its own — list the findings
neutrally and flag which ones are worth noting in the registry.

The phrasing "if the connector can't tell, say so" isn't there by accident. Connector capabilities keep evolving, and it's better to get an honest "I can't check this — look at the journal's website or Retraction Watch" than a confident answer pulled out of thin air. As a rule, the most useful sentence you can hear from AI during verification is "I don't know" — and every prompt in this guide is written so the model is allowed to say it.

That's the verification work done. You have the claim broken down, a map of the literature, translated key studies, their place on the evidence ladder, a stress-tested verdict, and a background check. All of it — once the process settles in after a few repetitions — takes twenty to forty minutes for an ordinary claim. What's left is the most valuable part: making sure that half hour doesn't vanish into your chat history.

Phase 6: the registry of verified claims — a living document instead of endless googling

This is where the guide pivots from "how to verify a claim" to "how to turn it into a habit." Without this phase, you've done one good one-off search — and in six months, when an aunt sends the same detox in new packaging, you start from zero, because the conclusions are buried somewhere in your chat history. With this phase, you get a registry of verified claims: a single markdown file that is the family's memory. It's the same "single source of truth" principle the company guide builds its brand voice on — the asset here just isn't tone of voice, it's what you took the trouble to find out.

Why a plain markdown file, and not a board, an app, or notes on your phone? For the same reasons as the brand voice: it's plain text that survives any tool; it can be versioned and shared; and above all, AI can read it — you upload it into a Project, and every new verification starts from what you already know. Where the file lives is up to you: a shared family drive, a notes app, even a git repository if that's comfortable for you. Only one thing matters: there's one version, and everyone knows where it is.

The structure of an entry

You'll adapt the format to your own taste, of course, but this one has proven itself — each claim is one section with fixed fields:

# Registry of verified claims — our family

## Syrup XY "for children's immunity and focus"
- **Claim:** the extract in the syrup reduces illness and
  improves focus in school-age children
- **Verdict:** unsupported (marketing runs ahead of the evidence)
- **Strength of evidence:** 1 small RCT in adults (n=40, 3 weeks,
  manufacturer-funded), measured a blood marker, not illness rate;
  no study at all on children or focus
- **Gaps between study and ad:** adults→children, blood marker→
  illness rate, 3 weeks→continuous use
- **Verified:** 2026-08-16 (from abstract; full text behind a
  paywall)
- **Sources:** PMID 12345678; also searched "extract X children
  respiratory" — no results
- **Reopen when:** an RCT in children is published, or a review
- **Note for the family:** harmless, but we'd be buying hope; talk
  to the pediatrician about our son's illness rate, not the
  online store

Notice three fields that aren't obvious. "Strength of evidence" isn't a grade, it's a sentence — a year from now, it'll bring the whole verification back to mind. "Reopen when" is what makes the registry a living document: the verdict isn't a permanent ruling, it's a state of knowledge as of a date, with a condition for when to revisit it. And "Note for the family" is the translation into practical language — this is the line you'll actually read out loud at Sunday lunch. Use a fixed scale for verdicts, so entries stay comparable; a five-point scale has proven itself: strongly supported by evidence — mixed/uncertain — weak evidence — unsupported — disproven. Plus a special category, "can't be assessed from what's available," for paywalls and missing literature.

Writing the entry with one prompt

Manually rewriting entries would kill the habit — so AI does the writing, as the closing step of a verification chat:

The verification is done. Create a registry entry in exactly this
markdown structure: [paste the template, or say you have it in
the project instructions].
Rules:
- Pick the verdict from our scale (strongly supported by evidence
  / mixed / weak evidence / unsupported / disproven / can't be
  assessed) and justify it in one sentence in the strength-of-
  evidence field.
- Put every PMID we worked with into the sources field, plus a
  brief note of what we searched for and did NOT find.
- Phrase the "reopen when" field concretely — what type of study
  or event would change the verdict.
- Write the note for the family in plain language, no jargon, no
  more than two sentences, no health advice — just what we know
  about the evidence.
Return clean markdown to paste into the file, nothing else.

Read the entry before you paste it in — it's your name behind the verdict, not the model's. The most common fix: the verdict tends to land a notch stricter or looser than your scale calls for, because the model interprets the scale in its own way; compare it against previous entries. Then upload the updated file into the project (replacing the old version, so two copies don't end up living side by side) — from that point on, every new chat knows what the family has already verified, and can answer "have we dealt with this before?"

Coming back to a topic: editability in practice

The same claim comes back a year later — a new product, the same extract. Without the registry: an hour of work, all over again. With the registry:

Our registry (you have it in the project) has an entry for
"Syrup XY." I've come across a new product with the same extract
and claim: [paste it].
1. Tell me what we found back then, and what verdict we reached —
   briefly, in the entry's own words.
2. Check via the PubMed connector what's been added to the
   literature since the verification date [date]. I'm mainly
   interested in studies that meet the "reopen when" condition
   from the entry.
3. If nothing substantial has been added, just suggest updating
   the review date. If something has been added, let's go through
   it again starting from Phase 3, and rewrite the entry.

This is the editability the whole registry exists for: a new study doesn't overwrite your head or trigger a panic — it changes one line in a file, with a date and a reason. And that goes double for the uncomfortable but honest case: when a big new study shows your old verdict was wrong. A registry where you can see "2025: weak evidence → 2026: supported, a large RCT came out" isn't a failure — it's exactly what honest learning looks like. Incidentally, this is where your family document differs most from chain emails: it's capable of changing its mind.

A routine: the registry watches itself

The last step turns the document into a system. Claude can run scheduled tasks — conversations that fire off on their own, on a schedule. Have the registry reviewed once a quarter:

Quarterly review of the claims registry (you have the file in the
project):
1. Go through every entry and, for each, check via the PubMed
   connector whether the "reopen when" condition has been met
   since the verification date.
2. Return a table: claim | verdict | unchanged / NEEDS REVIEW |
   why.
3. For items marked NEEDS REVIEW, add the PMIDs of the new papers
   and one sentence on what changes.
Don't edit the registry yourself — I edit the file after
reviewing. If everything's unchanged, one sentence and a date is
enough.

One last rule worth stating: a human writes into the registry, not an agent. That's not a technical necessity, it's a quality safeguard — verdicts nobody reads degrade. Five minutes over a quarterly table is the entire maintenance cost; compare that with the alternative, an endless "I don't remember what we found back then."

Case study: Věra's email, start to finish

We've gone through the phases one at a time; now let's put them together on the second model scenario, so you can see just how little time the whole thing actually costs. Věra's chain email claimed that a common food "provably causes cancer — scientists found this out, share before they delete it." Here's how Jana's evening went, with minutes noted at each step.

Minutes 0–3: the preflight check. The prompt from the opening section broke down the email text: emotionally loaded language, zero sources, the classic censorship appeal ("before they delete this") — and, most importantly, the claim conflates two different things: that a substance formed during a specific way of preparing the food is carcinogenic in the lab, and that eating the food normally causes cancer in people. Those are two different questions with very different answers; this exact kind of blending is what makes hoaxes such stubborn opponents — half of it is technically true.

Minutes 3–8: inventory and questions. The PICO breakdown (the Phase 1 prompt for a sourceless claim) returned two questions: is there evidence that substance X is carcinogenic — in whom, and at what doses? And: is eating the food in ordinary amounts associated with cancer rates in people? Plus English keywords for both.

Minutes 8–18: top-down search. The connector found lab and animal studies for the first question (high doses, an isolated substance) and, for the second — typical of long-running hoaxes — several large cohort studies and a meta-analysis that find no association with cancer in humans at normal consumption levels. The evidence map had a telling shape: busy down in the basement (cells, animals), quiet up top (large human studies). The headline had taken the basement and passed it off as the whole house.

Minutes 18–25: reading and hierarchy. The abstract translation of the meta-analysis confirmed a detail worth remembering: the animal studies used doses that, scaled to a human, would correspond to consumption levels completely outside reality. Not "a bit more" — orders of magnitude off. "The dose makes the poison" is an old saying, but the hierarchy table gave it numbers.

Minutes 25–30: opposing counsel and verdict. The opposing-counsel prompt (Phase 5) did its job: it found a real nuance — for one specific way of preparing the food, there's a recommendation not to overdo intake, based on the precautionary principle, not on proven risk. That belongs in the registry, because it's the rational kernel that Věra's email inflated into "causes cancer." Verdict: disproven, in the form the claim is made; rational kernel: caution around [the specific preparation method], with no evidence of risk at normal consumption levels. Registry entry written, "reopen when" field set to "a large prospective study with an opposing finding is published," done.

Thirty minutes — and the difference from googling isn't just the time: Jana now has an answer she can defend sentence by sentence, recorded so that two years from now, when the hoax comes back in new clothes, she'll find it in ten seconds — and she also has the rational kernel that lets her say to her mother at Sunday lunch not "that's nonsense," but "you know, there's actually a grain of truth in this? Let me show you what I found."

The abbreviated version for everyday use

Not every claim can support thirty minutes and six phases — and it's better to have an honest shortcut than no check at all. For low-stakes claims (no money on the line, no decision riding on it, you just want to know where you stand), the six phases can be compressed into a single prompt. You're consciously trading away depth — for anything that matters, go back to the full process:

Quick check with the PubMed connector — all in one step, but
honest. Claim: [paste the claim and where it's from].
1. Rewrite the claim as one verifiable question (who, what,
   compared to what, what outcome).
2. Find the highest available level of evidence: reviews and
   meta-analyses first, then RCTs, then the rest. Up to 5 papers,
   with PMIDs.
3. Summarize what they say — in numbers, absolute, including what
   wasn't significant.
4. Actively mention the strongest finding AGAINST whatever the
   conclusion is leaning toward.
5. A verdict on the scale: strongly supported by evidence / mixed
   / weak / unsupported / disproven / can't be assessed — plus one
   sentence why, and one sentence on what would change the
   verdict.
Explicitly flag what came from abstracts versus full texts, and
what's your judgment versus what's actually in the studies. Don't
claim anything without a PMID. And if the question touches
treatment or symptoms, stop and refer me to a doctor.

The output takes two minutes to read and is plenty for the "a friend at the pub claims" category. Just two disciplines apply: even a lightning check deserves a line in the registry (an abbreviated one is fine), and if the verdict comes out surprising or close, that's a signal to switch to the full process — the shortcut is for triage, not for arguments.

How to talk to loved ones who believe it

Now the harder part than the research. Jana has the facts about her mother's chain email — and facts are, surprisingly, not worth much in a conversation with someone you love. Research on persuasion and plain experience say the same thing: people don't change their minds when they lose an argument; they change it when they don't lose face. Whoever shows up with "Mom, that's a hoax, here are the studies" wins the exchange of arguments and loses the relationship and the point — next time, the emails just stop getting forwarded to them, not stop being shared at all.

A different approach works, called the bridge technique. It has three steps. First: common ground. Start with what you share — "I'm glad you care about your health; I don't want us eating anything harmful either." This isn't manipulative padding: it's true, and it turns the conversation from "me versus you" into "us versus the question." Second: acknowledging the rational kernel. Almost every hoax has a rational kernel — distrust of advertising, a bad experience with a dismissive doctor, a real historical case where "official sources" got it wrong. Acknowledge it out loud: "you're right that you can't trust everything manufacturers say — that's exactly why I looked into this more deeply." Third: a bridge to the evidence — as an invitation, not an execution. "I found out where this claim originally comes from, want to see? It's actually a bit of a detective story." You're offering a joint investigation and leaving the other person room to reach the conclusion themselves; that's the only conclusion that sticks.

And one extra principle, specifically for family: pick your battles. A harmless ritual grandma cares about, that costs nothing, might not need a registry verdict — it needs a nod. The registry is for claims that cost money, fear, or health. Whoever corrects everything ends up being heard by no one, right when it actually matters.

Preparing the conversation

AI helps twice here: preparing what to say, and rehearsing how it goes. First, the prep — and notice that the input is a registry entry, so you're not doing anything twice:

I'm about to talk to [my mom / a friend / my brother] about a
claim from our registry: [paste the entry]. This person believes
the claim, shares it, and it's tied for them to [health worry /
distrust of doctors / a good personal experience — describe what's
likely behind the belief].

Prepare me a conversation using the bridge technique:
1. Common ground: 2 sentences to open with — what I honestly share
   with them.
2. The rational kernel in their position: what's actually
   reasonable about their stance, and how to acknowledge it out
   loud, with no irony.
3. The bridge: how to offer what I found as a joint investigation
   — including one specific interesting detail from the
   verification that works like a story, not a lecture (e.g.
   where the claim actually comes from).
4. Three sentences to avoid, because they'll close the door.
5. A fallback plan: how to end the conversation gracefully if it
   gets stuck — the goal isn't to win today, it's to keep the door
   open.
Write it in my voice, no psychology textbook language.

I want to underline point 5: for deeply held beliefs, a realistic goal is "plant a doubt and stay someone they can still talk to about this." Beliefs held for years don't change over one lunch — and that's fine. Jana, in the end, didn't win her mother over on the studies: she showed her how she'd looked the claim up herself, and her mother now sometimes says "let's go look that up." That's the win. There isn't another kind.

A dry run

The second kind of help is practice — have AI play the other side before you try it for real:

Run through a practice conversation with me. You are [my mom, 68,
believes claim X because — describe the context; distrusts both
advertising and "official" information, is sharp, and hates being
lectured].
I'll try the bridge technique.
Rules:
- React realistically, not as a caricature: you have good
  counterarguments ("everyone finds what they want online,"
  "doctors don't know everything either").
- If I slip into lecturing, sarcasm, or overloading you with
  facts, react the way she would — then, outside the role, briefly
  tell me where I went off track.
- After at most 10 exchanges, end the conversation and give me
  feedback: what worked, what didn't, one thing to do differently
  next time.

Ten minutes of this simulation saves a lot of Sunday lunches. A typical first-round finding: you talk too long. In a real conversation you get a few sentences per turn, not a paragraph — and the simulation teaches you that before your mother does.

When to leave it be and see a doctor

This section is short on purpose — it's meant to be unmissable, not long. This whole guide operates in the territory of information: what studies say about claims from ads and chain emails. There's a hard line where that territory ends, and it's worth knowing by heart.

Beyond that line is everything about the specific health of a specific person. Symptoms bothering someone. Decisions about treatment. Starting, changing, or stopping any medication — even "just herbs," because supplements interact with medication. Diagnoses and interpreting them. Children's health and pregnancy, doubly so. In none of these situations is the answer more research — the answer is a doctor, who knows the context, carries the responsibility, and can actually examine the patient. AI has neither of those, and PubMed even less so; thirty minutes with the connector makes you a better-informed patient, not a doctor.

In practice that means two things. First, recognize the tipping point: the moment you catch yourself using research to put off a doctor's visit, or looking for the nerve to stop a medication, close the chat and book an appointment — that's not a failure of the method, that IS the method. Second, use the research correctly: as preparation for questions. "Doctor, I read that there are no studies of extract X in children — is there something you'd actually recommend for supporting immunity?" is a far better conversation than silence, and far better than a printed-out email. Doctors talk better with a well-prepared patient, not worse — the difference is whether you show up with questions or with ready-made conclusions.

And one more situation where it's time to leave it be: when the topic starts consuming you. Checking every meal against studies isn't literacy, it's anxiety with citations attached. The registry is meant to save the family's attention, not eat it up — when verification stops being an occasional tool and becomes a daily need, that itself is worth talking about, possibly with a professional.

Limits: what PubMed and AI won't give you

An honest guide has to say what the process can't do too — and not just for fairness's sake: whoever knows a tool's limits uses it better, and won't be caught off guard by an objection they've already heard.

Paywalls. Already mentioned, but it belongs here again: free full texts, at the time of writing, cover only part of the literature (PubMed Central); the rest sits behind subscriptions that make sense for institutions, not families. Practical consequences: you mostly work with abstracts (and honestly note that), you prefer openly available reviews, and you make peace with occasionally running into a study you'd like to read in full and can't. Don't turn to pirated libraries, and don't ask AI to "somehow get around" a paywall — among other reasons, because text "obtained some other way" may not match the final published version.

Publication bias. The literature you're searching isn't a neutral sample of reality: studies with a positive finding get published more readily than ones where "nothing came out." That means a summary of the literature is systematically more optimistic than reality — and for topics where a handful of small positive studies exist, it's fair to ask how many negative ones are sitting in a drawer somewhere. Good meta-analyses try to account for this (they test for asymmetry in findings), and you should account for it too, by calibrating: a small effect from small studies tends to be even smaller in reality.

English, and the limits of the database. PubMed is an English-centric database of biomedicine. Work published in other languages is often missing, or has only an English abstract — which barely matters for everyday health claims (major science today mostly publishes in English), but it does mean "nothing in PubMed" doesn't mean "nothing anywhere." And conversely: PubMed doesn't cover all knowledge — psychology, nutrition science at the border with agriculture, or sports science all have part of their literature elsewhere. For family purposes, PubMed is the best single choice; just don't treat it as the definition of what exists.

Fresh results and preprints. There's a lag between when a study is done and when it's indexed in the database; the freshest work circulates as preprints — versions that haven't gone through peer review. When AI comes across a preprint, treat it as "hasn't been checked yet": interesting, worth a registry entry, but with an explicit note and lower weight.

AI makes things up — still. The connector drastically shrinks the problem of hallucinated citations, because the model actually searches instead of recalling from memory. But it doesn't shrink it to zero: the model can still misreport a real study, blend two together, or mix in "general knowledge" while summarizing. The defensive rules are scattered through this guide; here they are together: always want every claim backed by a PMID; spot-check by hand (type the PMID into the PubMed website's search box and read the abstract yourself — always do this for key studies); and always want "what the study says" kept separate from "what you're adding on top." It's the same principle this site applies to verifying any fact from AI — just with zero tolerance when it comes to health.

Time and attention. Finally, the most honest limit of all: this process isn't free — it costs half an hour per claim, and a bit of discipline with the registry. Which is why the last rule of calibration is: don't verify everything. Verify what costs money, fear, or a decision. For the rest, the plain reflex "a headline isn't a study" is enough — and you'll pick that up naturally after a few rounds of checking anyway.

Common mistakes

  • Asking AI for a verdict instead of evidence. "Does collagen work?" gets you a confident essay with not a single verifiable source. The right question is "what do the studies say, and which ones" — and every answer should carry a PMID you can click through to.
  • Settling for the first study that confirms it. There's at least one weak study "for" almost anything. Skip the "find the strongest work against it" step (Phase 5), and it isn't verification — it's just a more expensive way of confirming your own opinion.
  • Skipping the claim breakdown. Search for "is turmeric healthy" and you'll drown. Five minutes with a PICO question (who, what, compared to what, what outcome) decides whether the rest of your half hour goes anywhere.
  • Reading conclusions instead of methods and numbers. Conclusions are written by authors about their own baby. The real content of a study is in who was studied, what was measured, and what came out in numbers — including what didn't.
  • Leaving the conclusions in your chat history. Without a registry entry, everything gets redone from scratch six months later. Verification ends with an entry — claim, verdict, strength of evidence, date, sources — or it never really happened.
  • Sliding from literacy into self-treatment. The moment research starts replacing a doctor's visit, or justifying changes to medication, it has stopped being literacy. The line from the second-to-last section always applies, to everyone in the family.

The best tools

  • The PubMed connector for Claude — the core of the whole workflow: searching studies, abstracts and metadata, full texts from PMC, related articles, citation tracing, and identifier conversion. Free, in the connector directory, at the time of writing.
  • Projects on claude.ai — persistent context: instructions with the verification ground rules (counterarguments, mandatory PMIDs, no medical advice) and an uploaded registry, so every chat already knows what you've verified.
  • The registry in markdown — a plain text file in a shared spot for the family. Zero technology, maximum durability; the one genuinely irreplaceable piece of the whole system, because it's yours.
  • Claude's scheduled tasks — a quarterly registry review: go through the "reopen when" fields, check for anything new in the literature, and return a table for a decision. AI watches, a human edits.
  • Deep research — for big topics where you'd rather have one thorough report with citations than a dialogue; the process is shown in the guide to deep research.
  • The PubMed website (pubmed.ncbi.nlm.nih.gov) — for spot checks: paste a PMID from the chat into the search box and see the abstract with your own eyes. Trust, but verify applies to the connector too.

What you get out of it

  • Time: instead of an hour of googling five contradictory articles, twenty to forty minutes of structured work — and for recurring topics, minutes, because the registry remembers. Family debates about "that study from the internet" shrink down to reading one entry.
  • Money: supplements and cures whose evidence tops out at mice, or a manufacturer's own study, stop flowing into the cart out of fear and hope. One "miracle" order you don't place pays for all the time spent on this guide.
  • Peace of mind: scary headlines lose their grip once you can find out what's behind them in twenty minutes — and what you usually find is an observational study with a tiny absolute risk. Family disagreements about hoaxes turn from arguments into joint investigations.
  • Quality: decisions — the purchases, and the ones you bring to a doctor — rest on the best available evidence instead of the last shared headline. And kids who pick up this process along the way gain a habit no school teaches quite this practically.

Pro tip

Once this process becomes second nature, turn it around: instead of waiting for a claim to arrive from outside, occasionally give your own head a "cleanup." Everyone carries health beliefs that came from somewhere once and never got checked — from "don't swim right after eating" to guaranteed family truths about vitamins. Sit down with your partner or your kids, write down ten things you believe about health without knowing why, and run them through Phases 1 through 6 like anything else. It tends to be the most fun use of this whole guide — and the registry suddenly fills up with claims the family actually lives by, not just the ones that happened to arrive by email.

And the closing rule that holds this whole piece together: a headline is a tip, a study is a lead, consensus is the answer — and for your own health, the answer is a doctor. Everything else in this guide is just the craft of getting from headline to consensus without getting lost along the way, and knowing the route next time.

Common questions

Does this replace my doctor?

No, and it shouldn’t even try. The guide teaches you to read studies and not fall for marketing — it helps you ask your doctor better questions, not answer them yourself instead of them. The moment it touches symptoms, treatment, dosage, or stopping medication, the decision belongs solely in the doctor’s office. Neither AI nor PubMed have your medical context or carry the responsibility.

Do I need to know English and understand statistics to pull this off?

No — that’s exactly what AI does for you. Abstracts are in English and written in technical language, but you have them translated and explained sentence by sentence, including what relative risk or sample size actually means. Your job isn’t to translate — it’s to ask: who was studied, what exactly was measured, and what the study doesn’t say.

What if the study is behind a paywall?

The abstract is always free on PubMed and is usually enough for a first assessment. At the time of writing, full texts are only available for articles in PubMed Central — roughly a quarter of the total. When you don’t have the full text, work with the abstract, look for review articles (they tend to be open more often), and honestly note in the registry that you only read the abstract.

Can I trust what AI tells me about studies?

Only with verification. Language models can make up a study — convincingly, with authors and a year attached. That’s exactly why the connector is the core of this guide: every claim needs a PMID (the PubMed identifier) you can click through to. A citation with no traceable PMID doesn’t belong in the registry, no matter how good it sounds.

How do I tell a strong study from a weak one?

Through the hierarchy of evidence: a case report is a curiosity, an observational study shows an association, a randomized controlled trial tests a cause, and a meta-analysis sums them all up. Add to that sample size, humans versus mice, and conflicts of interest. The article includes a table with a plain-language explanation of every level — and prompts that will do the classifying together with you.

What do I say to a grandparent who believes in the miracle supplement?

Not “that’s nonsense.” Use the bridge technique: start with what you have in common (you both want her to be healthy), acknowledge that she cares about her health, and only then offer what you found — as a joint investigation, not a correction. People don’t change their minds when they lose an argument; they change it when they don’t lose face.