Ordinarily Well: The Case for Antidepressants

32

Washout

THERE WERE TWO such studies. One, published in 2010, found that antidepressants work only for severe depression. The other, appearing two years later, found that they work well for mild-to-moderate major depression. The first study received extensive publicity. The second—we will turn to it in the next chapter—received none.

Again, we need to know. With antidepressants, primary-care doctors do most of the prescribing. They see less severe illness. Should they back off? Or does the growth of depression treatment in general practice count as an achievement for public health?

If you have doubts, likely they can be traced to one publication, a meta-analysis by the Penn-Vanderbilt group who had shown that Paxil mutes depressed patients’ neuroticism. This team is highly accomplished. As befits the topic’s and the authors’ importance, their confirmation of the severity hypothesis received marquee publication, in JAMA. USA Today covered the report under the headline “Antidepressant Lift May Be All in Your Head.” The New York Times led with “Popular Drugs May Help Only Severe Depression.” When I spoke with my friend Alan’s neurologist, the concerns he expressed contained echoes of the JAMA paper.

To settle the severity debate, the Penn-Vanderbilt group conducted a patient-by-patient analysis of a carefully curated set of data from outside the FDA collection. They concluded that antidepressants’ benefits are “substantial” for those with severe depression but “minimal or non-existent, on average, in patients with mild or moderate symptoms.”

To evaluate that finding, we need to understand an additional element in the design of outcome trials. In most drug studies, after patients sign on they get placebo for a week. Participants who improve markedly are dropped, and the real drug trial begins.

The “washout” or “run-in” phase addresses any number of problems. For our purposes—seeing whether medicine helps in mild depression—the important consideration is that washouts may counteract a highly distorting tendency in outcome trials, baseline score inflation.

Let’s recall what happens. If raters have a sense of the minimum Hamilton score for admission to a study, and if they are under pressure to fill an enrollment quota, they will be inclined to tack on questionable Hamilton points. The boost will not be uniform. There’s no need to raise ratings in the very ill. Scores for least afflicted participants will be most inflated.

If drug companies demand a Hamilton of 17 for entry, they will see many patients with scores of 17, 18, and 19. The rest of the ratings distribution may be unremarkable. This distortion has been noted even in trials run at centers like UCLA and Harvard. When off-site raters, with no stake in the pace of enrollment, analyze tapes of admission interviews, they find patients to be much healthier than the on-site Hamilton scores suggest. According to off-site assessments, many patients admitted to drug studies simply are not depressed.

Baseline score inflation has various meanings. In my use, it stands for the bunching of scores near the entry level—low severity—of any trial. Arif Khan identified the pattern in fifty-one drug-company studies testing ten antidepressants. He concluded that raters “may arbitrarily assign minimally qualifying initial scores to borderline subjects to increase enrollment (baseline score inflation).” Similar distortions have been found in research on dementia and Tourette syndrome.

Without a run-in phase, the less ill patients whose scores have been inflated will go on to lose—that is, seem to lose—symptoms they never really had. Since placebo will “succeed” with these participants, even potent antidepressants will have trouble competing in the low range. Because of this distortion, trials without washouts will be especially likely to find a severity gradient—less efficacy in mild depression—even where (in the real world, when patients are judged accurately) there is none.

That said, the JAMA authors mistrusted placebo washouts. Use them, and you dismiss the best placebo responders, the ones who, on pills, get better immediately. Without run-ins, a protocol mimics what happens if, at a first meeting, a doctor prescribes an antidepressant. In their own clinical research, the Penn-Vanderbilt team ran trials without washouts.

The Penn-Vanderbilt researchers investigated the severity hypothesis by gathering studies like their own. They amassed more than two thousand citations of randomized medication trials for depression. They eliminated those with placebo washouts. They eliminated studies of dysthymia. They eliminated trials where they could not obtain patient-by-patient data.

Six trials met the criteria. Three tested imipramine, and three, Paxil.

To my reading the collection was problematic. In it, some patients with moderate depression had received low doses of medication. Most patients at the mild end of the spectrum did not even have major depression. So, yes, in this idiosyncratic batch of research, medication might work best in severe illness, but because the dosing and diagnosing were uneven, the question of what standard drug treatment would do in nonsevere major depression remained open. Even this material contained indications that it might work well.

Certainly, the three imipramine studies were mismatched, and in ways that will be familiar to us.

The most straightforward trial tested severe depression and aimed to get patients on 200 milligrams of imipramine, a substantial dose. They benefited.

Next came our constant companion, the NIMH collaborative study. Whether it should be considered a drug trial is a matter for discussion, but the Penn-Vanderbilt researchers worked hard to use the material well. To make results comparable to those in standard studies, they relied on outcomes midway through the experiment, at the eight-week mark, and they included data from all participants. One researcher assured me that the team had checked for “dose skewing”—did healthier patients receive less medication?—and found none. That analysis remains unpublished, and given Donald Klein’s concern (“It is quite possible the milder patients received ineffective, small doses”), it is hard not to wish for transparency about whether and when in the eight weeks patients with different levels of depression received plausible drug treatment.

But the real problem is with the third trial. Seemingly designed to showcase Saint-John’s-wort, the herbal remedy for mood disorder, it used imipramine as a comparator. The imipramine was given at 100 milligrams, which the study’s own authors acknowledged to be below the “recommended mean dosage for efficacy” and arguably “suboptimal.” Imipramine succeeded, just barely, as a benchmark—but only if you made a slight adjustment to the Hamilton scale.

The participants in the Saint-John’s-wort study had moderate levels of depression, so it was not surprising that the JAMA analysis would find poor outcomes in that range. By including results from patients on 100 milligrams of imipramine, the meta-analysis, which set out to test the severity hypothesis, seemed to have put a different question into play: whether depression responds best to full, recommended doses of medication.

The three Paxil studies only added complexity. Two were the Penn-Vanderbilt team’s own. In all, the JAMA meta-analysis involved 307 patients from Penn-Vanderbilt trials and 411 from the rest of the research universe. Criticizing review articles, Gene Glass had complained that authors who define “adequate design” narrowly end up highlighting their own work. Perhaps meta-analysis is not immune to that tendency.

One Penn-Vanderbilt team trial tested two psychotherapies for depression and used Paxil as a comparator. Despite high dropout rates in the Paxil arm (apparently volunteers wanted psychotherapy), the drug outperformed placebo, but with better efficacy in more severe cases. The other of the team’s trials involved only patients with severe depression, and Paxil had worked there, too, with a number needed to treat of 4. Those results were known in advance, prior to the development of the criteria for the JAMA meta-analysis. More symptomatic depression had a leg up before scores from any other doctors’ patients were added in.

The third Paxil trial, one of those quirky small studies that carry their own interest, had the greatest potential to cause trouble. Conducted at Dartmouth and other sites, it tested a psychotherapy focused on problem solving and included medication and placebo arms. Here’s the rub: the trial excluded all cases of major depression.

The Dartmouth team was investigating whether medication works for “minor depression.” To meet the diagnosis, patients needed to have three or four symptoms from the list that includes sadness, difficulty experiencing pleasure, insomnia, and so on. Major depression requires between five and nine symptoms.

Looking at all the patients in the Dartmouth trial—patients with three or four depressive symptoms—the researchers found good responses to Paxil. The number needed to treat ran between 4 and 5. Quite low-level depression (depression yet milder than the mildest major depression) responded to medication about as well as severe major depression does. The Paxil data seemed to disprove the severity hypothesis or cast it in grave doubt.

The story would end here but for the way that minor depression is defined.

“Major depression” encompasses both acute and chronic cases. You may be in your first, brief episode, or you may be in a prolonged episode preceded by many like it—either way, if you have the required five symptoms, you have the diagnosis.

Minor depression is different. If the disorder is chronic or recurrent, it’s called dysthymia—or was in the late 1990s, when the Dartmouth group did their diagnosing. In their protocol, dysthymia was considered first. If you had chronic low-level depression, you were dysthymic. If your low-level depression was not chronic, you had minor depression.

In drug trials, chronic depression tends not to respond to placebo, so drug effects show through. Most of the patients in the Dartmouth study had dysthymia. They did phenomenally well on Paxil. The remission rate—full recovery—was 80 percent, with a number needed to treat of 3. Chronic minor depression appears to be what Paxil treats best.

The JAMA analysis included only the remaining patients, those with (acute) minor depression. Paxil did not appear to work—but that’s because the dysthymic patients, for whom Paxil was curative, were missing.

Because in the JAMA sample severe depression was often chronic depression and mild depression generally was not, it was clear, before running any numbers, that the weakest antidepressant effects would appear in patients who began with low Hamilton scores. We would expect that trend even if antidepressants work equally well up and down the line—as, in fact, they appear to do when we approach minor depression the way we do major depression, including all illness, acute and chronic.

The focus on washouts begins to look like a fetish, one that comes at a cost: the bias that arises from reliance on trials with high dropout rates, uneven prescribing, and a lack of uniformity in diagnosis.

As for minor depression, the Dartmouth study contained hints that Paxil helps after all. Using ratings like those employed in quality-of-life studies, the researchers tracked overall mental health. On that measure, the most impaired patients—those for whom the minor depression caused serious difficulty in thought and action—made gains on medication relative to placebo, and at a statistically significant level.

For mildly depressed patients who were actually doing poorly, then, Paxil provided relief that the Hamilton scale failed to pick up. As an approach to brief, minor depression, watchful waiting might be a reasonable strategy; where suffering impairs function, Paxil might be the preferred next step. This clinically relevant information disappears in the maw of meta-analysis.

Where should doctors seek guidance?

The individual trials appear informative. In all but one, the antidepressant succeeded in its assigned role, as comparator or as primary treatment. (In the Saint-John’s-wort study, imipramine tested out better if you excluded one Hamilton item, “somatic anxiety”; apparently the raters had mistaken drug side effects, such as indigestion, for depression symptoms.)

Paxil is broadly effective. In severe and very severe depression, it sports a number needed to treat of 4. In dysthymia, 3. For patients impaired by brief, minor depression, Paxil brings quality-of-life benefits.

Meanwhile, imipramine soldiers on. Moderate doses treat grave depression. Even a suboptimal dose can serve, more or less, to validate a sample of patients with midrange illness.

That’s before we take into account baseline score inflation, which can create a gradient of results even when, in reality, antidepressants work equally well everywhere.

The JAMA meta-analysis is no smoking gun—hardly the sort of evidence that would convince a doctor that when his less depressed patients do well on antidepressants, he is seeing placebo effects. To me, the paper illustrates a different moral: not to rely on a meta-analysis until you’ve perused the constituent trials.

Here, their lessons run counter to the thrust of the overview paper. They are: Ask about impairment. Yes, patients with severe symptoms may do well on antidepressants. But so will patients with chronic mood disorders and patients (whatever the symptom count) hobbled by hard-to-measure effects of illness. Where it impacts quality of life, depression tends to respond to medication. If low doses don’t work, try full ones.



If you find an error or have any questions, please email us at admin@doctorlib.org. Thank you!