The early wave of modern psilocybin trials shared a quiet limitation that is easy to miss amid the excitement. The most striking results, the ones that made headlines, came from people facing life-threatening cancer, in the grip of a particular existential distress at the end of life. That work was genuinely moving and genuinely important, but it left an obvious question hanging. Did psilocybin help with depression as such, the ordinary, grinding major depressive disorder that affects millions of people who are not dying, or did it only work on the specific kind of dread that comes from staring down mortality? The 2021 trial led by Alan Davis, with Roland Griffiths and colleagues at Johns Hopkins, published in JAMA Psychiatry, set out to answer exactly that, by testing psilocybin for plain major depressive disorder.
This article reads that trial carefully, because it marks an important widening of the question. Moving from cancer-related distress to garden-variety major depression took the substance out of the specialized end-of-life context and pointed it at the far larger population that conventional antidepressants are meant to serve. The trial reported strong results, and it carried the familiar limits that travel with this whole body of work. It is also the study that anchors several other pieces in this series, since later follow-ups tracked its participants forward in time. Check the specifics against the original paper before publishing, since the figures and the design are where this story rests.

Why moving beyond cancer mattered
To see why this trial was a meaningful step rather than just another entry in the list, you have to understand what the cancer studies could and could not show. The landmark trials from Johns Hopkins and NYU in 2016 had tested psilocybin in people with life-threatening cancer, and they had found large, durable reductions in depression and anxiety. Those were real and important results. But the distress they targeted was a specific thing, existential distress, the depression and anxiety that arise in direct response to a terminal diagnosis and the confrontation with one's own death. It is a particular flavor of suffering, tied to a particular and extreme circumstance.
That specificity left a genuine gap in what the research had shown. Major depressive disorder, the common condition that conventional antidepressants are prescribed for, is not the same as existential distress in the face of dying. It afflicts people whose lives are not in immediate danger, it often has no clear external cause, and it is the actual target of most depression treatment. A reasonable skeptic could look at the cancer results and ask whether psilocybin was doing something specific to the dread of death, perhaps by shifting a person's relationship to mortality, rather than treating depression in the general sense. If that were true, the impressive cancer findings might not transfer at all to the much larger population of people with ordinary major depression, and the whole excited extrapolation from the cancer studies to depression broadly would rest on an untested leap.

So testing psilocybin in major depressive disorder specifically was the necessary next move, and it carried real stakes. A positive result would suggest that the effect was not confined to the special case of end-of-life distress but extended to depression as such, vastly expanding the potential relevance of the treatment. A null result would have been a serious check on the enthusiasm, hinting that the cancer findings were about something narrower than depression in general. The Davis trial walked straight into that question, taking the substance out of the cancer ward and into the territory of the condition that affects the most people.
It is worth pausing on the scale of what was at stake in that distinction, because it is easy to underestimate. Major depressive disorder is one of the most common and disabling conditions on earth, affecting hundreds of millions of people across every demographic, and it is the actual day-to-day business of most psychiatric treatment. Cancer-related existential distress, by contrast, however acute and important, affects a far smaller and more specific group. So the question of whether psilocybin's effect generalized from the second to the first was not a minor matter of scientific tidiness. It was the difference between a specialized intervention for a particular and tragic circumstance and a potential treatment for one of the largest mental health burdens humanity carries. The entire case for psilocybin as a major new approach to depression, rather than a niche tool for the dying, rested on the leap this trial was built to test.
There is a mechanistic angle to the question too, worth naming. One reading of the cancer results was that psilocybin worked by changing a person's relationship to death specifically, easing the particular terror of mortality through some shift in perspective that the experience can produce. If that were the mechanism, it might say little about ordinary depression, which often has no such clear existential object. Another reading was that psilocybin works on depression more generally, by loosening the rigid, self-reinforcing patterns of negative thought that characterize the disorder in all its forms, of which the cancer patients' dread was just one variety. The Davis trial was, in effect, a test between those two readings. A strong effect in ordinary depression would favor the second, more general mechanism, and that is roughly what it found, which is part of why the result mattered beyond its immediate numbers.

What they actually did
The study was a randomized, waitlist-controlled trial, and that design choice is worth understanding, including its weaknesses, because it shapes what the result can claim. It enrolled adults with major depressive disorder, ordinary depression rather than cancer-related distress, and randomly assigned them to one of two conditions. One group received psilocybin treatment promptly, two doses given within a structured program of psychological support, with preparation and integration in the now-standard pattern. The other group was placed on a waitlist, receiving the same treatment only after a delay, and serving in the meantime as the comparison against which the immediate-treatment group was measured.
The waitlist control is the design feature that deserves the most scrutiny, because it is both the study's practical solution and its central weakness. A waitlist group receives no active treatment during the comparison period, which means the trial is essentially comparing psilocybin-plus-therapy against nothing much at all for that stretch. This is a real limitation, and an important one to be honest about, because comparing an active, dramatic treatment against a waitlist stacks the deck. The waitlist group gets no engaging intervention, no attention, no expectation of improvement, so almost any difference between the groups could reflect those non-specific factors as much as the psilocybin itself. It is a weaker comparison than an active placebo or a head-to-head against another treatment would be, and it inflates the apparent effect. The trial used it presumably for practical and ethical reasons, but a careful reader has to weight the results accordingly.
The outcomes were tracked with established, validated depression rating scales, measured at intervals after treatment, and reported in the clinical terms that matter, response and remission rates as well as changes in symptom severity. By using recognized measures and reporting how many people actually got well rather than only how group averages shifted, the trial spoke in concrete clinical language about whether the treatment worked for ordinary major depression. Keep the waitlist design firmly in mind throughout, because it is the key qualification on everything the trial found.

What they found
The results were strong, strong enough to make the study one of the more cited in the field. The participants who received psilocybin showed large reductions in depression compared with the waitlist group, with effect sizes among the larger ones reported in depression research. This was not a subtle signal. The immediate-treatment group improved markedly, and substantial proportions of them met the clinical thresholds for response and for remission, meaning many did not just improve somewhat but dropped below the cutoff for clinical depression altogether. By the figures in the paper, a majority of the treated participants met response criteria and a large share reached remission at the post-treatment assessments, numbers worth verifying against the original but striking ones if they hold.
That the effect appeared in ordinary major depressive disorder, not just in cancer-related distress, was the headline contribution. It suggested that whatever psilocybin was doing, it was not confined to the specific existential suffering of the dying but extended to depression in the general sense, the condition that affects the broad population conventional antidepressants are meant to treat. This widened the potential relevance of the treatment enormously, taking it from a specialized end-of-life intervention to a candidate for one of the most common psychiatric conditions there is. The cancer findings, on this evidence, were not a special case after all but part of a broader effect on depression, which is a substantially bigger claim.
The speed and the clinical magnitude were both notable. The improvements appeared relatively quickly after treatment rather than building slowly over weeks, consistent with the rapid-onset pattern seen across psychedelic research, and the proportions reaching remission were high enough to be clinically meaningful rather than merely statistically detectable.
Within the limits of the design, this was a strongly positive result, the kind that justifiably drew attention and that motivated the larger and better-controlled trials that followed. The catch, and there is always a catch, is that the waitlist comparison means the raw size of the effect almost certainly overstates what a more rigorous comparator would show, a point that matters enormously for interpretation.

Why the waitlist control changes how to read it
It is worth slowing down on the waitlist issue, because it is the single most important thing to understand about this trial, and a clean example of why comparator choice matters so much in clinical research. The size of a measured treatment effect depends heavily on what you compare the treatment against, and a waitlist is about the weakest comparator there is. The waitlist group received no active treatment, no engaging experience, no therapeutic attention, and no reason to expect improvement during the comparison window. So when the psilocybin group improved dramatically and the waitlist group did not, the gap between them folds in not just the specific effect of the drug but everything else the treatment group received that the waitlist did not, the attention, the structured support, the expectation of help, the dramatic experience, all of it.
This means the large effect size, impressive as it looks, almost certainly overstates the specific contribution of the psilocybin. Some unknown portion of that gap reflects the non-specific benefits of being actively treated and cared for, which the waitlist group entirely lacked. A more demanding design, comparing psilocybin against an active placebo or against another real treatment, would shrink the apparent effect, because the comparison group would also be getting something engaging. The later Carhart-Harris trial that ran psilocybin head to head against escitalopram is a useful contrast here, since against a real antidepressant the difference was far less dramatic than against a waitlist, which tells you a great deal about how much the comparator shapes the headline.
None of this means the Davis result is fake or that psilocybin did nothing. It means the honest reading discounts the raw effect size substantially, treating the trial as strong evidence that psilocybin produces a real benefit in ordinary major depression while recognizing that the precise magnitude is inflated by the weak comparator. The trial answered the important question, does the effect extend beyond cancer distress to depression as such, with a credible yes. It did not, and could not, given its design, tell you how large the effect really is once the non-specific factors are stripped away. Holding both of those truths at once is what reading this trial well requires.
It is worth being clear about why the researchers might have chosen a waitlist despite its weakness, because the choice was not simply carelessness. Waitlist controls are common in psychotherapy research and in early-stage studies, partly for ethical reasons, since it can feel hard to justify giving seriously depressed people a known-inert placebo when you could instead promise them the real treatment after a delay, and partly for practical ones, since a waitlist is simpler and cheaper to run than a convincing active placebo. So the design reflects real constraints, not just a failure of rigor. But understanding why a weak comparator was used does not change what it does to the result, and a reader has to separate the sympathy for the researchers' situation from the clear-eyed discounting the design requires. A reasonable choice under constraints can still produce an inflated number, and both things are true here.
The broader lesson is one of the most portable in all of research literacy, which is that you can never interpret an effect size without asking what it was measured against. The same treatment can look spectacular against a waitlist, solid against an active placebo, and merely competitive against an established alternative, and all three pictures can come from honest trials of the identical intervention. The number alone is close to meaningless until you know the comparator. This trial, read alongside the later escitalopram head-to-head from the same group, is almost a controlled demonstration of that principle, since it shows the same drug, studied by the same researchers, producing very different apparent magnitudes depending entirely on what it was stacked up against. A reader who internalizes that will never again take a bare effect size at face value, which is a tool worth more than any single fact about psilocybin.

Where the study holds up
The strengths are real, beginning with the question it asked. Testing psilocybin in ordinary major depressive disorder, rather than only in the special case of cancer-related distress, was an important and necessary widening of the field, and the trial answered that question with a clear, strong signal. Demonstrating that the effect was not confined to existential distress but extended to depression in the general sense substantially expanded the potential relevance of the whole approach, and that contribution stands regardless of the comparator issue.
The trial was properly randomized, used established and validated depression measures, and reported its results in meaningful clinical terms, response and remission rates rather than only average score changes, so you can see how many people actually got well by recognized standards. The treatment followed the careful structured protocol that the field had developed, with preparation, supervised dosing, and integration. The population was a clinically relevant one, people with the common condition that depression treatment actually targets. And the study was published in JAMA Psychiatry, a leading peer-reviewed venue, and went on to anchor an important line of follow-up research tracking its participants forward in time, which speaks to its seriousness and influence. For what it set out to do, establish that psilocybin can produce a strong antidepressant effect in major depression, it delivered clearly.

Where it strains
The central weakness is the waitlist control, discussed at length above, and it cannot be overstated as the key qualification on the result. Because the comparison was against an untreated waitlist rather than an active comparator, the large effect size almost certainly overstates the specific contribution of the psilocybin, folding in the non-specific benefits of attention, expectation, and engagement that the waitlist group lacked. Any honest reading has to discount the raw magnitude substantially on these grounds, and the later head-to-head against escitalopram, which produced a much less dramatic difference, shows just how much the comparator was doing.
The other familiar limits apply too. The sample was modest, limiting generalization and the precision of the effect estimates. The selection issue is present as always, since these were people who volunteered for psilocybin treatment for their depression, an open and motivated group whose response may not reflect that of the broader depressed population. The blinding problem persists, since participants knew they were receiving psilocybin, which lets expectancy inflate the self-reported outcomes on top of the waitlist effect, compounding the two sources of inflation. And the follow-up reported in the original trial was relatively short, though later studies extended it, so the durability question was not answered by this study alone. None of these undoes the core finding that psilocybin produced a real antidepressant effect in major depression, but together they mean the trial is best read as strong evidence of a genuine effect of uncertain and probably overstated magnitude, awaiting the better-controlled trials that would pin the true size down.

What it justifies, and what it does not
A measured reading lands clearly. What the Davis 2021 trial justifies is the meaningful conclusion that psilocybin can produce a real antidepressant effect in ordinary major depressive disorder, not just in the special case of cancer-related existential distress, which substantially broadens the relevance of the treatment and was the key question the study set out to answer. It justifies the larger and better-controlled trials that followed, and it justifies treating major depression as a serious target for psilocybin research rather than assuming the effect was confined to end-of-life distress. By extending the finding beyond the cancer ward, it earned its place as an important step.
What it does not justify is taking the large effect size at face value, because the waitlist control almost certainly inflates it, and the true magnitude against a fair comparator is probably considerably smaller, as the later escitalopram head-to-head suggests. It does not justify claims that psilocybin is an established treatment for major depression, since a waitlist-controlled trial of this size, with the blinding problem layered on top, cannot establish that. And it does not justify extrapolating from this supported clinical context to casual use, since the effect was produced inside a structured therapeutic program. The trial answered an important question and answered it positively, while leaving the precise size of the effect, and the durability, for more rigorous work to determine.

Why this study matters
The Davis 2021 trial earns its place because it widened the central question of the field at exactly the right moment. The cancer trials had been moving and important, but they tested a specific kind of suffering, and the leap from end-of-life distress to depression in general was an untested assumption that the excited coverage had largely glossed over. By testing psilocybin in ordinary major depressive disorder and finding a strong effect, this study turned that assumption into evidence, establishing that the treatment's relevance extended to the common condition that affects millions rather than only to the dying. That is a genuine expansion of what the research had shown.
It is also instructive as a lesson in comparator choice, which is why it is worth reading even apart from its specific subject. The gap between this trial's dramatic waitlist-controlled effect and the far more modest difference the same group later found against escitalopram is a vivid demonstration of how much the choice of comparison shapes the headline, and a careful reader who absorbs that lesson carries it to every trial they encounter.
There is one more reason this particular trial repays attention, which is its role as a foundation for later work. Because its participants were tracked forward in time, it became the basis for follow-up studies examining how long the antidepressant effect lasted, including the year-long durability analysis that the same group published afterward. That makes the Davis trial not just a standalone result but a starting point, the original treatment whose downstream effects other studies went on to measure. Understanding it well is therefore useful beyond itself, since several other pieces of the depression evidence are built directly on top of it, and a weakness or strength in the original trial propagates forward into everything that extends it. The waitlist limitation, in particular, is worth carrying in mind when reading those follow-ups, since they inherit the selected, openly-treated sample that the original design produced.
The honest way to hold the study is as an important, strongly positive, but comparator-limited result, real evidence that psilocybin works against ordinary depression, paired with full awareness that the waitlist design inflates the apparent size and that the true effect is smaller and still being measured. It took psilocybin out of the cancer ward and into the broader world of depression, which is where the question always had to go, and it is the foundation on which a good deal of the subsequent depression research was built. That dual character, genuinely important and genuinely limited, is exactly what most real science looks like up close, and learning to see both at once is the whole point of reading trials carefully.