research established

How to Read a Clinical Trial Like a Researcher

A practical guide to evaluating clinical trial papers, from sample size and blinding to outcomes and conflicts of interest, using psilocybin research as the example.

MMI Editorial July 7, 2026 13 min read

Most readers of clinical research are not researchers. They are journalists, patients, family members, policy makers, and curious people trying to make sense of headlines that often summarize complicated studies in misleading ways. The gap between what a study actually shows and what the coverage around it claims is, in many fields, considerable. In psychedelic research, that gap can be especially wide, which makes this a useful place to learn the skill.

This article is a practical guide to reading clinical trial papers. It is not a methodology textbook. It is the set of habits and questions that working clinicians, methodologists, and careful science journalists use to judge what a study can and cannot support. We use psilocybin trials as the running example because they are recent, accessible, and methodologically representative of the broader landscape, but the skills transfer directly to any medical or psychiatric research.

If you read all the way through and remember nothing else, remember this. In a clinical trial paper, the most important section is usually the Methods, not the Results. The Results tell you what happened. The Methods tell you what it means.

Golden Teacher Mushroom Anatomy

Start with the question

Before you read the paper, you should be able to say what question it is trying to answer. This sounds trivial. It is not. Many readers go straight to the abstract, absorb a result, and never circle back to ask whether the result actually answers the question they had in mind.

The questions a clinical trial can answer have a specific shape. In patients with condition X, does treatment Y produce better outcomes on measure Z than treatment W or placebo over time period T. Which patients, what treatment, what outcome, what comparator, what duration. That structure is the most any single trial can speak to. The questions a single trial usually cannot answer are the ones people most want answered. Is this treatment going to work for me. Should it be standard of care. What is the long-term safety profile across the general population. Those require accumulated evidence from many studies of different designs, and they rarely yield confident answers from one paper. So when you start, decide what you are hoping the paper can tell you, then check as you read whether the design can actually support that conclusion. Often it cannot, and noticing that is the most important takeaway.

Golden Teacher Mushroom in Natural Habitat

Read the methods first

Most people read papers in the order they are written, from abstract to discussion. Trained readers often jump straight to the methods after a quick look at the abstract, for a simple reason. Until you know how a study was conducted, the results do not have a defined meaning. Several methodological elements deserve close attention.

The population. Who was studied, with what inclusion and exclusion criteria, and what demographics. In psilocybin trials, common exclusions include a personal or family history of psychotic disorders, cardiovascular disease, pregnancy, and current substance use disorders. These are sensible safety choices, but they mean the trial cannot speak to anyone who would have been excluded. A depression trial that screened out psychotic vulnerability does not generalize to depression with comorbid psychotic features. That is not a flaw. It is a boundary on the claims.

The intervention. What exactly was administered, in what dose, on what schedule, in what setting. In psilocybin research the intervention is usually a bundled package of drug dose plus preparation plus structured dosing-day support plus integration follow-up. Treating it as the drug alone misrepresents what was tested.

The comparator. What the control group actually experienced. An untreated waitlist is not the same as an active drug. A tiny "placebo" dose of psilocybin is not the same as a true placebo. A niacin pill that produces flushing is not the same as an inert pill. The comparator defines what the intervention is measured against, and weak comparators inflate apparent effects.

Randomization. Were participants assigned to groups by chance or by some other procedure. Random assignment is the main tool for making the groups comparable before treatment, and non-randomized designs, including observational studies, open-label single-arm trials, and retrospective analyses, provide weaker evidence about cause and effect.

Blinding. Did the participants know which condition they were in, did the people administering the treatment know, and did the people assessing outcomes know. Each is a separate question, and the single, double, and triple-blind labels describe different patterns of who knows what. In psychedelic research, true blinding is essentially impossible at meaningful doses. The participant knows, and the staff usually know. This unblinding is a structural limit on what these trials can conclude.

The outcome measures. What was measured and how. A participant-reported scale is subject to expectancy effects, a clinician-rated measure is subject to rater bias unless the rater is blinded, and a biological outcome is more objective but possibly less clinically meaningful. The choice of outcome shapes the apparent results substantially.

Psilocybin Mushroom Comparison

The sample size conversation

A perennial question is whether the sample size was adequate, and the honest answer is that it depends on the size of the effect being investigated and the variability of the outcome. For very large effects in low-variability outcomes, small samples can be informative. For modest effects in noisy outcomes, very large samples may be needed. Most trials are statistically powered to detect an effect of a specified size with a specified probability, and the methods section usually reports a sample size calculation explaining what the trial was designed to detect.

In psilocybin trials, sample sizes have historically been small, often fewer than 50 participants and sometimes fewer than 20. Small samples have consequences. They allow only large effects to be detected, they inflate the apparent magnitude of the effects that are detected, they are more vulnerable to chance findings, and they limit subgroup analyses, because dividing a small sample quickly leaves too few participants to learn anything. None of this is a reason to dismiss small studies. Pilot and feasibility studies are deliberately small, designed to answer narrow questions about whether a larger study is worth running. The problem comes when a small pilot is interpreted, in the surrounding coverage, as if it had answered the broad questions only larger trials can address.

Blue Meanies Handheld Specimens

What the results actually say

When you finally reach the results, a few habits help. Look at the numbers, not just the verbal summary, because the text frames the data in line with the authors' interpretation while the tables and figures hold the data itself. Sometimes the data is stronger than the text suggests, sometimes weaker, and reading both lets you form your own view.

Pay attention to effect sizes, not just statistical significance. A significant result tells you the observed difference is unlikely to be due to chance given the sample. It does not tell you the difference is large or clinically meaningful. A trial with thousands of participants can detect a tiny difference with high significance, and that difference may still be too small to matter for any individual.

Look at the variability. Means and medians summarize the average response, while standard deviations, interquartile ranges, and the distribution tell you whether that average reflects a tight cluster or wide scatter. In psychiatric research, response is often highly variable, and the average can hide the fact that some participants improved dramatically and others not at all. Look at dropout rates and how missing data was handled, since people who leave a trial early may differ systematically from those who stay, and the choice between intention-to-treat and completer analysis substantially affects results. And look at the adverse events. Reports that dwell on efficacy and minimize difficulty are incomplete, and the language used for adverse events often reveals the authors' framing, since a phrase like "mild and transient" can quietly cover events the participants found significant.

Psilocybin Mushroom Cap Bruising Texture

The discussion section

The discussion is where authors interpret their results, place them against prior work, and acknowledge limitations. It is also where the most enthusiastic language tends to appear, so two questions are especially useful.

What do the authors themselves identify as limitations. Conscientious authors enumerate the limits of their work, and if the limitations are brief and dismissive, that is itself a signal, while a detailed and candid account makes the rest of the paper more credible. In the better psilocybin trials, the limitations sections are notably extensive, which is a feature of the field's strongest work. And do the conclusions actually follow from the data. Compare the specific claims in the discussion to the specific results in the tables, because the discussion sometimes reaches well beyond what the data can support, particularly in the closing paragraphs that gaze toward future implications. Trained readers learn to notice that gap.

Psilocybin Stem Bruising Close Up

Conflicts of interest and funding

Modern clinical research happens in a landscape where most large trials are funded by parties with a financial stake in the outcome. This is not, by itself, evidence of wrongdoing, but it requires attention. Look at who funded the trial, who employs the authors, and who holds intellectual property related to the intervention, usually disclosed at the end of the paper under headings like conflicts of interest or funding sources.

A trial funded by an industry sponsor is not automatically suspect, but across many fields of medicine, industry-funded trials are statistically more likely to report positive findings than independently funded ones, a pattern documented in several meta-analyses. Trials registered in advance, with pre-specified primary outcomes that match what the published paper reports, are more credible than trials whose outcomes seem to have been chosen after the data came in. In contemporary psilocybin research, a growing share of trials are sponsored by companies pursuing regulatory approval, which is normal for any drug development pathway. It does mean readers should pay particular attention to funding, pre-registration, and the consistency between reported and pre-specified outcomes.

Small and Large Sample Size Comparison

Putting it together

A good way to integrate what you have read is to write, in your own words, two sentences when you finish a paper. First, what the study actually showed, the narrow, specific finding, in the population studied, with the intervention tested, against the comparator used, on the outcome measured, at the time point reported. Second, what the study did not show, the larger questions you might have hoped it addressed but that its methodology cannot support. If you can write those two sentences without reaching for the abstract, you have understood the paper. If you cannot, you have not yet read it as carefully as it deserves.

Enigma Mushroom Bruised Specimen

A worked example

Consider a hypothetical psilocybin depression trial that finds a 50 percent response rate in 30 patients with treatment-resistant depression, against 25 percent in a tiny-dose "placebo" arm, at four weeks. The press release calls it an unprecedented breakthrough. A careful reader can quickly sort what it supports from what it does not.

It supports the existence of some signal of a larger effect from a higher dose than a lower dose, in a small sample of selected patients, over a short follow-up. That justifies further investigation. It does not support general clinical effectiveness, durability of effect, generalizability to less selected populations, or the absence of expectancy-driven inflation, since the small "placebo" dose was probably identifiable to participants. The press release's framing is enthusiastic in a way the trial cannot support. The careful reader notices, treats the trial as preliminary, and waits for replication in larger and better-controlled studies. This is not pedantry. It is the basic operating posture of evidence-based medicine. Most readers do not need to become experts. They need enough humility about single studies that the broader weight of evidence, over time, can shape their conclusions.

Tidal Wave Mushroom on Moonlit Forest Floor

A few habits

If you are not already in the habit of reading primary research, a few small practices build the skill. Pick a topic you care about and read three full trials on it, not abstracts but full papers including the methods, and after three you will start to see patterns in how trials in that field are designed and reported. When you read press coverage of a new trial, look up the original paper and compare what it says to what the coverage claims, an exercise that is consistently humbling. Subscribe to a methodology blog or newsletter that critically evaluates new research, because reading critiques alongside the trials sharpens your own reading. And notice when you are most tempted to overclaim, because the trials that produce the most exciting headlines are exactly the ones whose limitations get overlooked, and reading carefully under those conditions is most of what research literacy means in practice.

The evidence base for psilocybin in depression and other conditions is, at this writing, real but incomplete. Reading it carefully is how a serious reader stays in honest relationship with a developing field. The same skills, applied broadly, are how an informed citizen navigates the larger landscape of medical research without being misled by the loudest claims.

Frequently asked questions

Which section of a trial paper should I read first?
The Methods. Until you know who was studied, what was administered, what the comparator was, and how outcomes were measured, the results have no defined meaning. The Results tell you what happened, the Methods tell you what it means.
Does a small sample size mean a study is useless?
No. Pilot and feasibility studies are intentionally small and answer narrow questions about whether a larger trial is worth running. The problem is interpreting a small pilot as if it answered the broad questions only large trials can address.
What is the difference between statistical significance and effect size?
Significance tells you a difference is unlikely to be due to chance given the sample. Effect size tells you how large the difference is. A huge trial can find a tiny, statistically significant difference that is too small to matter for any individual patient.
Why is blinding such a problem in psychedelic trials?
At meaningful doses, participants and staff can tell who received the drug, so true blinding is essentially impossible. That unblinding can inflate measured benefit through expectancy, which is a structural limit on what these trials can conclude.
Does industry funding mean a trial is untrustworthy?
Not automatically, but it warrants attention. Industry-funded trials are statistically more likely to report positive results, so look at funding sources, whether the trial was pre-registered, and whether the reported outcomes match the pre-specified ones.