Perez & Brady's Principles and Practice of Radiation Oncology (Perez and Bradys Principles and Practice of Radiation Oncology), 6 Ed.

Chapter 14. Methodology of Clinical Trials

Yaacov Richard Lawrence, James J. Dignam, and Maria Werner-Wasik

In this chapter, we discuss the design, conduct, and analysis of oncology clinical trials, pointing out particular areas of interest to radiation oncology and reviewing some recent ideas in clinical trials and related research studies. This chapter provides a brief and essentially nontechnical sketch of the main concepts and current research areas, and we refer the reader to comprehensive texts on clinical trial conduct in oncology for further details. Excellent recent texts, such as Handbook of Statistics in Clinical Oncology, Clinical Trials in Oncology, and Oncology Clinical Trials: Successful Design, Conduct and Analysis provide the fundamentals, as well as up-to-date discussion of new challenges and active research in statistical methods for oncology clinical trials.13

Clinical trials enable physicians to advance medical care in a safe, scientific, and ethical manner. Formally defined, clinical trials are a set of procedures in medical research conducted to allow safety and efficacy data to be collected for health interventions.4 A more detailed definition for our purposes would describe a clinical trial as a prospective study that includes an active intervention, carried out in a well-defined patient cohort and producing interpretable information about the action of the intervention.5 Although in some ways similar to a well-designed laboratory experiment, the involvement of living human subjects demands adherence to strict ethical principles while also adding to the complexity of interpretation of the results. The past 60 years have witnessed an unprecedented appreciation of the importance of clinical trials and a consequent surge in the number of clinical trials performed. In the field of radiation oncology alone, according to a PubMed search, 376 clinical trials were published in 2010, of which 65 were phase III trials.

Over the course of less than a century, the evidence on which medicine is practiced has evolved from being entirely empirical to being a highly regulated scientific process based on vigorously designed clinical trials tightly overseen by numerous scientific and governmental agencies. Cardinal chapters in the history of clinical trial design and implementation include the following:

• 1747—James Lind’s work on the effect of citrus fruits in the prevention of scurvy among sailors in the Royal Navy

• 1863—Austin Flint’s use of a placebo group for comparison with an experimental treatment in the treatment of rheumatic fever

• 1947—Nuremberg Code, a result of the appreciation that much of the medical experimentation performed by physicians in Nazi Germany was both ethically wrong and scientifically uninterpretable

• 1948—The first double blind trial, performed by the British Medical Research Council to assess the value of streptomycin in the treatment of tuberculosis

• 1964—Declaration of Helsinki developed by the World Medical Association, ethical guidelines, which continue to be updated, for the performance of clinical trials

• 2000—Creation of the ClinicalTrials.gov Web site, a registry of clinical trials under the auspices of the National Institutes of Health (NIH)

The performance of high-quality cancer clinical trials involves the cooperation of multiple bodies, so-called stakeholders, including cancer patients and their families, physicians (who accrue the patients), their operating environment (academic institution or practice), the research team (who runs the trial on a day-to-day basis), sponsors (who oversee and fund the trial), independent monitors (to ensure the correct performance of the research team), government regulatory agencies, contract research organizations (CROs, which may carry out specific aspect of the trial such as auditing), and medical insurance companies. Modern clinical trials may require involvement of translational scientists, and clinical psychologists, as well as experts in quality of life, cost-effectiveness, and other disciplines. The complexity of clinical trials adds to the regulatory work involved. For instance, a multi-institutional federally sponsored clinical trial protocol will require the approval of at least three different research ethics oversight committees (institutional review boards [IRBs]) within the group coordinating the trial, at each institution that opens the trial to accrue patients, and at the sponsor.

Although many phase I and phase II trials are carried out by investigators within a single institution, many larger phase II and most phase III trials are generally multi-institutional. Thus, clinical trial investigator networks have emerged as essential players in the performance of large phase II and III clinical trials. The U.S. National Cancer Institute–sponsored Cancer Cooperative Group Program is an example,6 as are similar groups such as the European Organisation for Research and Treatment of Cancer (EORTC). One cooperative group, the Radiation Therapy Oncology Group (RTOG), founded by Simon Kramer in 1968, is dedicated to trials involving radiation therapy, has activated 460 protocols, and has accrued approximately 90,000 patients to its trials. Early studies sought to answer questions regarding radiation dose and fractionation. As cancer therapy became multimodal, the group addressed questions relating to combining systemic chemotherapy with radiation therapy. Most recent trials seek to combine targeted agents with contemporary radiation therapy techniques.6

Clinical trials are extremely expensive to perform, with costs continuing to rise as a result of both increased regulatory oversight and greater trial complexity. It has been estimated that implementation of the European Union’s Clinical Trials Directive (laws and regulations related to implementation of good clinical practice in the conduct of clinical trials) led to a doubling of the cost of running noncommercial cancer trials in the United Kingdom.7 Large phase III trials can cost in excess of $100 million; as a result, clinical trials are frequently financed by the pharmaceutical industry. An unfortunate consequence is that clinical trials are rarely performed on established generic drugs, where there is little commercial interest in establishing new indications. Conversely, the performance of rigorous clinical trials contributes significantly to the costs involved in the development of new pharmaceutical agents, which are subsequently reflected in the commercial pricing of the product.

TABLE 14.1 COHERENT FRAMEWORK FOR DETERMINING WHETHER CLINICAL RESEARCH IS ETHICAL

OVERVIEW OF ETHICAL CONSIDERATIONS

Medical ethics are based on the principles of autonomy (the patient’s right to refuse or choose treatment), beneficence (a practitioner should act in the best interest of the patient), nonmaleficence (first, do no harm), justice (fairness and equality), dignity, and honesty. Without due diligence, physicians may infringe on these principles when encouraging patient participation in clinical trials. Are the physicians confident that the proposed treatment is beneficial and not harmful? Are the potential subjects fully aware of the implications of participation? Do all segments of the population have equivalent chance to participate and receive potentially better treatment? Despite the universal acceptance of these principles, there have been numerous examples of grossly unethical research being performed in the Western world within living memory. Documents seeking to address these issues include the Nuremberg Code, the declaration of Helsinki, and the Belmont Report. Emanuel et al.8 have listed seven requirements that provide a systematic and coherent framework for determining whether clinical research is ethical (Table 14.1).

Ethical principles themselves and the creation of guidelines are insufficient to ensure the ethical conduct of medical research. Physicians within Nazi Germany performed atrocities despite the existence of German guidelines published in 1931,9 reflecting the need for legislation. The International Conference on Harmonisation of Technical Requirements for Registration of Pharmaceuticals for Human Use (ICH) brought together the regulatory authorities of Europe, Japan, and the United States and experts from the pharmaceutical industry to regulate scientific and technical aspects of pharmaceutical product registration. The ICH guidelines are legally binding in many countries (although not the United States) and are updated every few years, reflecting the increasing sophistication of the field. An example of a recent addition is the introduction of data and safety monitoring committees (DSMCs), which are independent groups of experts who monitor patient safety and treatment efficacy data in ongoing clinical trials. Another recent advance is the requirement for the registration of clinical trials at sites such as ClinicalTrials.gov. Such registration both improves transparency concerning what clinical trials have been and are being performed and empowers patients to find relevant clinical trials.

OVERVIEW OF TRADITIONAL CLINICAL TRIAL DEVELOPMENT PHASES

Traditionally, a new anticancer agent is tested in a three-step process, starting with a small dose finding trial (phase I), followed by a pilot efficacy trial (phase II), and culminating with a large comparative randomized trial (phase III). This development paradigm was created and established in the era of cytotoxic chemotherapies and continues to be used today in the era of often less-toxic (and possibly not dose-dependent) targeted therapies, with some adaptations and innovations that we discuss later. Here we review the traditional paradigm without technical details, which can be found in many excellent sources for clinical trial design and conduct.2,5

Phase I

The main objective of the phase I trial is to determine the maximum tolerable dose of an agent to be subsequently used in testing for efficacy. Although initially developed in the setting of drug testing, this concept has been adapted to test radiation alone and combined drug/radiation regimens. The basic conceptual approach is that of sequential dose increases, or escalation, in small patient cohorts until the treatment-related adverse event rate reaches a predetermined level or unexpected toxicity is seen. This stepwise testing in phase I trials determines what is known as the maximum tolerated dose (MTD), which is putatively the most effective level at which to evaluate efficacy.

The key parameters to be defined at the outset of a phase I trial include patient eligibility criteria, starting dose and schedule of dose escalation (which should frame the expected MTD), events comprising the adverse event/toxicity response and the expected MTD, and finally the escalation design plan.10 By far, the dominant design has been the so-called 3+3 approach, where cohorts of three patients are exposed to a given dose, and based on the outcomes in that cohort, either de-escalation, escalation, or additional enrollment takes place. There are a large number of other designs, one of which may be particularly suited for radiation therapy trials, as we discuss shortly. It should be appreciated that the MTD is a relative concept that can change over time. For example, hematologic toxicity may be less dose limiting today than it was before the development of bone marrow stimulators such as erythropoietin and filgrastim.

Phase II

The primary purpose of the phase II trial is to determine the response rate of the treatment, seeking early evidence of clinical activity. An important secondary purpose is to gather more robust adverse event information at the established dose. The primary efficacy end point of the phase II trial has traditionally been tumor response; however, duration of response, progression-free survival (PFS), and site-specific activity such as locoregional control are all relevant and increasingly used. Measures of patient survival are usually secondary end points in phase II trials because of the limited sample size and follow-up duration of these trials. In general, phase II trials usually are not designed to provide definitive evidence that the test treatment is superior to current options. In fact, phase II trials have traditionally been single-arm studies comparing against a benchmark historical response rate. Reliance on this nonconcurrent external control rate can be problematic.1113 Furthermore, patient selection factors can influence the results—for example, overall response rates in single-institution phase II trials can be significantly higher than in multicenter studies or subsequent phase III controlled trials.14 Randomized phase II trials have historically had a role in multiarm trials aimed at selecting the best treatment(s) to take forward for further testing.15 More recently, they have become favored as a means of providing more reliable pilot efficacy data.16 However, the preferred approach is changing as described later.

Phase III

Phase III trials are randomized comparisons between a new treatment regimen that has already shown promise in phase I/II trials and the current best standard of care (i.e., the control). Randomization offers a critical advantage over nonrandomized studies. Specifically, randomization balances the distribution of prognostic factors between treatment arms and ensures that treatment is assigned independent of these factors, thereby minimizing or eliminating these effects when comparing outcomes by treatment. When randomization is combined with treatment blinding of patients, researchers, or both, then even subjective outcomes can be assessed with minimal bias. Additionally, in multicenter studies, randomization can balance any systematic bias of the treating physicians or institutions.

Phase III trials can address one of several types of primary questions. For example, a study can be designed to determine if standard treatment is better than best supportive care. More often, phase III studies are designed to compare a new treatment with the current standard treatment. Phase III studies can also compare two or three different regimens with each other, as well as with standard treatment. Finally, a trial may be designed to demonstrate that a given treatment option is not worse than another by more than a tolerable margin. These “equivalence” trials, more accurately referred to as noninferiority trials, play an important role in development of less invasive or less burdensome treatment regimens.

The primary end point of a phase III trial is most typically overall survival (time to death from any cause); however, other important clinical end points such as disease-free survival (DFS) are increasingly justified. Because these trials aim to definitively demonstrate benefit with respect to these end points, the number of participants and follow-up period required for phase III trials is much longer than in phase II trials. Important secondary end points can include locoregional control and other site-specific failure end points, adverse event profiles, and quality of life measures.

Phase IV

Phase IV trials are also known as a postmarketing surveillance trials. These trials involve the safety surveillance of a drug after it receives regulatory approval for standard use. The safety surveillance is designed to detect any rare or long-term adverse effects over a much larger patient population and longer time period than was possible during the phase I–III clinical trials.

UNIQUE FEATURE OF CLINICAL TRIALS IN RADIATION ONCOLOGY

Therapeutic clinical trials in radiation oncology typically involve the introduction of new technologies (e.g., the use of stereotactic body radiation for a new indication) or more frequently the novel combination of radiation therapy with a systemic agent. There are unique challenges—both biologic and clinical—that characterize clinical trials in radiation oncology compared to those not involving radiation.

Response rate is frequently used in early-phase medical oncology trials to indicate activity; however, considering radiation therapy itself is highly effective at shrinking tumors, this end point is not useful in radiation trials. A more appropriate “activity” end point for radiation trials may be PFS, although this itself is often difficult to objectively assess. Modern imaging end points (such as fluorodeoxyglucose [FDG] uptake) show promise as early readouts of activity but still require vigorous validation for individual disease sites. Furthermore, efficacy and toxicity end points in radiation trials depend on multiple biologic factors, including size of the target, proximity of tumor to sensitive normal tissues, accuracy of target volume definition, degree of patient immobilization, dose of radiation, and fractionation scheme. Consequently, quality-assurance measures are an essential feature of radiation trials, especially in the multi-institutional setting.17 Inadequate quality assurance and lack of consistency in radiation delivery have led to the conclusions obtained from large, expensive clinical trials being questioned.1824 As a recent example, in RTOG 9704, which evaluated postoperative adjuvant chemoradiation treatment of pancreatic cancer, subtle protocol violations in target definition influenced both toxicity and survival.25

In medical oncology trials, adverse events typically occur during or within days of completing treatment. In contrast, toxicity following radiation therapy follows a biphasic course, early (within 3 months of starting treatment) and late (months to years later). Late toxicity is typically irreversible and hence important in determining the tolerability of an experimental treatment. Utilizing long-term toxicity as the primary end point in clinical trials is not practical; however, it nonetheless is imperative to collect and report robust information on long-term outcomes from radiation therapy trials. In fact, even in phase I radiation therapy trials, the follow-up period can be significantly longer than for those evaluating chemotherapy. A recent approach to dose escalation in phase I trials that considers late toxicities when deciding whether to advance to the next dosing level is discussed later.26

A further difference relates to the population studied. In medical oncology early-phase trials, participants have often received several lines of treatment and lack further therapeutic options. In contrast, patients in phase I radiation trials typically receive full-dose radiation treatment, and subjects may be treatment naïve. Furthermore, phase I trials in medical oncology are frequently “first in human” experience for the agent; toxicity is unpredictable and pharmacokinetic studies essential. Conversely, most multimodality phase I trials in radiation oncology involve systemic agents that have already been through extensive clinical testing; systemic toxicity is known and pharmacokinetic studies unnecessary. The purpose of the trial is to define the extent of local toxicity within the radiation field; consequently, radiation oncology phase I trials are organ specific. A recent study demonstrated that in reality, radiation phase I trials rarely utilize first in human agents, are associated with qualitatively predictable toxicity, and are comparatively safe.27 Examples of contemporary trial design in radiation oncology are provided in Table 14.2.

STATISTICAL ISSUES IN CLINICAL TRIAL DESIGN

Patient Population Definition and Stratification

A key issue in a clinical trial is a well-defined patient population to which the potential therapy applies. This is typically defined in terms of traditional disease characteristics reflecting putative prognosis, such as stage or its components. Increasingly, tumor pathology or marker features may be included. In any case, these must be unambiguously defined. Because factors defining eligibility are often critically related to prognosis, any randomized comparative study may use a stratified randomization approach to ensure equal representation of prognostic risk among treatment arms. Stratification factors need to be limited to a reasonable number because the total number of strata equals the product of the number of categories for each. For example, in a recently completed RTOG prostate cancer trial, stratification factors consisted of two levels of prostate-specific antigen (PSA) (<4 vs. 4 to 20), three cell differentiation categories (well, moderate, poor), and nodal status (N0 versus NX). The possible combinations of these variables create 2 × 3 × 2 = 12 strata within which treatments are to be balanced in allocation.

TABLE 14.2 EXAMPLES OF CLINICAL TRIAL DESIGNS FROM THE RADIATION THERAPY ONCOLOGY GROUPA

Randomization

Randomization is used differently in various phases of development but serves a similar purpose, which is to render treatment groups similar with respect to factors other than treatment that can influence outcomes. In phase I trials, there typically are not comparative groups, although there are situations where parallel cohorts of patients are being evaluated. Thus, randomization into cohorts ensures that these groups can be compared later for response biomarkers or other factors of interest. In phase II trials, randomization has been used in two similar but distinct ways. First, so-called selection designs have been used to help decide which of several potential candidate treatments to take forward to further definitive testing.15 In these trials, interest is not in statistically significant differences between treatments but rather is in the ability to nominally rank candidates in terms of best potential efficacy. It can be shown that this approach has high probability of identifying the most likely superior arm, although at the cost of false-positive findings, particularly if misused.28 A second and more recent role of randomization in phase II trials is to provide evidence, albeit at a less stringent criteria, that a test treatment is indeed promising.16 In phase III, randomization is critical for definitive unbiased evaluation.

With regard to implementation, randomization assignments can be simple or, more commonly, implemented using blocking or dynamic approaches with respect to balancing treatment arms by key factors, such as stratification variables mentioned earlier. Several proven methods are available.5

It is important to note that investigators must protect against practices that can erode or nullify the benefits of randomization. First, any breach of the random assignment process has an irreparable effect on the validity of the trial. Second, a large number (or differential number per treatment arm) of patient withdrawals can make the validity of the comparison suspect. Similarly, differential follow-up and consequently ascertainment of patient status between treatment arms can bias the treatment effect estimate. Third, bias in assessment of outcomes can have a major impact on the estimated treatment effect; thus, objective outcome measures and blinding of treatment assignment become important. Treatment assignment blinding is not feasible for radiotherapy and most chemotherapy regimens but can be used for many agents. In either case, and in particular for studies that cannot be blinded (among patients or caregivers), unambiguous, objectively defined end points are essential. In cases where determination of the end point involves possible observer subjectivity, such as when reading a diagnostic scan to determine disease progression, keeping assessors unaware of treatment assignment may be necessary.

End Points

In clinical trials, end points must be unambiguously defined, be assessable and reproducible, and reflect the action of the intervention. Typically, there is a single primary end point in a clinical trial; however, there may be numerous secondary end points.

Traditionally, in phase II cancer trials, treatment activity has been defined in terms of reduction in tumor burden. The most recent criteria for measuring activity are known as the Response Evaluation Criteria in Solid Tumors (RECIST).29 The criteria require the identification of target and nontarget lesions at baseline and their largest single dimensions. Categories of response are then defined—for example, complete response (CR), or disappearance of all target and nontarget lesions and no new lesions; partial response (PR), or 30% or greater decrease in the sum of the longest diameter of all target lesions, no progression of nontarget lesions, and no new lesions; and progressive disease (PD), or 20% or greater increase in target lesions, progression of nontarget lesions, or the occurrence of new lesions. A patient not satisfying either response or progression criteria is classified as having stable disease. Those achieving either a CR or PR are typically defined as objective responders, and the proportion of patients responding is then the primary end point of interest. Although widely used, there has long been concern that response defined this way is an inadequate substitute for more clinically relevant and objective end points such as survival time. In one study, fewer than 25% of agents that produced tumor response were eventually found to extend survival in comparative trials,30 whereas another suggested that tumor response is a reasonable surrogate for survival extension.31 Additional problems with the use of response rates in phase II trials include subjectivity and lack of reproducible assessments.32

As mentioned earlier, response is not as frequently used when radiation therapy is the test question, and in any case, other discrete binary end points can readily be used. For example, the proportion free from a given event (i.e., proportion alive, proportion recurrence-free, etc.) at a fixed time landmark such as 2 years is a common and straightforward end point.

A more informative end point that is used in many phase II and most phase III trials is the elapsed time from trial entry until occurrence of some event. The most straightforward of these is overall survival time, or time to death from any cause. This simple end point does not depend on adjudication of cause of death and its attendant complexities and naturally corrects for both favorable and unfavorable consequences of treatment. Although it can be verified or even ascertained from public records because of its simplicity, active follow-up per protocol remains of paramount importance. Cause-specific survival end points may also be considered; however, as mentioned, assigning cause of death is not simple, and one must account for “other cause” deaths and whether these have any relationship to treatment. In addition, whenever cause-specific deaths or other site-specific failure end points are used, methods for appropriately dealing with competing risks are required.33,34

Other commonly used time-to-event end points include DFS (time to recurrence or death from any cause) or PFS (time to disease progression, possibly determined via imaging or other assessments at regular intervals), although definitions of these are not standardized (e.g., see Hudis et al.35), and the specific failure events comprising given end points should be carefully specified. The main advantage of using DFS or PFS is the more rapid rate of events, leading to a smaller required sample size. In many cancer types, benefit with respect to these end points does not necessarily imply subsequent lengthened survival, although they may still represent clinical benefit for patients. For other disease settings (e.g., adjuvant therapy in colon cancer), DFS is a reliable and well-accepted primary end point that is strongly correlated with survival.36 This raises the topic of so-called surrogate end points, which are end points on which treatment benefits can be reliably measured. Various biomarkers and clinical end points have been studied and proposed as surrogate end points in clinical trials. Prentice37 specified criteria that a surrogate end point must fulfill if it is used to substitute for a clinical end point: the therapeutic intervention must exert benefit on both the surrogate and the clinical end point; the surrogate and clinical end point must be associated; and the effect of intervention on surrogate end point must mediate the clinical effect. An example of a widely studied surrogate marker is PSA, applied either as a static measure or as dynamic measures (PSA velocity; PSA doubling time; time to PSA nadir, particularly useful after radiation therapy of the intact prostate) to assess time to biochemical failure in patients with nonmetastatic adenocarcinoma of the prostate.38,39 This continues to be a developing area, and there remain many caveats and cautions regarding surrogate end points in clinical trials.40

Statistical Power and Sample Size

The overarching design consideration in clinical trials is to obtain sufficient information about an intervention so that a reliable decision can be made regarding its further development or use. In the classical (e.g., frequentist) statistical hypothesis testing paradigm, one sets up a null hypothesis of no treatment effect and an alternative hypothesis (which one hopes to validate) indicating a treatment effect. The type II or β error equals the probability that a statistical test fails to produce a decision in favor of a treatment effect when in fact the treatment is superior in the population. The complement of this probability (1 – β) is referred to as statistical power and equals the probability of correctly deciding in favor of a treatment benefit. Statistical power depends on the other principal parameters considered when planning the trial, specifically the probability of incorrectly finding in favor of a difference when none exists (type I or alpha error, usually set to 0.05 or 0.01 by convention), the a priori specification of a treatment effect that is considered both realistic and clinically material, and of course the sample size. It is imperative that trials be designed to achieve adequate statistical power; typically, 0.80 to 0.90 is desirable so as not to obtain equivocal findings concerning the potential worth of new treatments under consideration. Studies with low statistical power can cause delay or even abandonment of the development of promising treatments, as well as waste valuable resources, not least of which is the participation and goodwill of patients.41 In contrast, a “negative” trial that does not find the test treatment to be superior, if adequately powered, is informative in that resources can be directed into other more promising alternatives.

Thus, sample size to satisfy the power desired for the specified effect of interest is the key calculation in phase II and III clinical trials. (Phase I trials do not rely on hypothesis-driven sample size calculations, and the sample size derives from the specific design used.) The specific sample size calculation depends on the end point, and technical details will not be provided here; however, the two most common types of end points can be summarized as follows.

For a discrete binary end point in a phase II trial—for example, responded or did not respond—sample size calculations are straightforwardly performed using formulas for comparison of proportions. In a single-arm study, one aims to compare the observed response rate for the new agent to some historical response proportion, p0, or the response rate achievable with standard therapy in the target population. The main objective is to determine whether there is sufficient evidence to conclude that the response rate for the new regimen is greater than p0. We designate pA as a response rate which, if true, would be clinically material. We test the null hypothesis, H0 : p = p0, against the alternative hypothesis HA : p = pA. The values of p0 and pA and (and more importantly the difference), along with the sample size, will determine the power of the study. Note that to detect a small improvement (say, ≤10%) requires a large sample size. For example, to detect an improvement from a historical value of 20% to 30% with 85% power, more than 120 subjects are required. In addition, the value for both p0 and pA must be realistic; it is of little value to design and carry out a study to detect an effect size pA – p0 that is unlikely to be realized, simply because it is compatible with the number of patients that can be recruited.

For two-arm randomized phase II trials with discrete end points, the previous discussion is simply redefined in terms of two-sample comparisons of proportions, and the sample size is consequently much larger. Table 14.3 shows some sample size requirements for various response differences, illustrating the influence of the effect size and power on the number required. Finally, if the end point is a fixed time landmark, such as proportion event free or alive at 1 year, then the estimates of the proportions may be derived from survival analysis methods to appropriately account for losses to follow-up.

In many randomized phase II and nearly all phase III trials, the time from randomization until occurrence of the event is of principal interest rather than the event status at some fixed time landmark. In larger phase II and phase III trials, recruitment may take place over a lengthy interval, with each patient having a different follow-up duration, and the use of follow-up time per patient is more efficient than waiting until all patients have reached some fixed time. The treatment effect measure is then specified in terms of failure hazards, which can be thought of as failure rates per unit of time. Hypotheses are thus usually formulated in terms of the hazard ratio (HR) as H0 : λA/λB = HR = 1.0, where λA and λB are the hazards for treatments A and B, versus the alternative, HA : HR <1.0, for some value of the HR that represents a clinically important difference in outcomes. Under the assumption that this ratio is relatively constant over time, a given HR can be converted to an absolute difference between groups in proportions remaining event free at a specific follow-up time. For example, a new/standard HR equal to 0.75, or a 25% reduction in failure rate in the experimental group relative to the standard group, may translate into an absolute difference in the proportion of patients remaining free from the event between groups of 4.6% at 5 years, if the standard group 5-year survival percentage is 80% (Table 14.4).

From the specification of difference of interest or effect size, then, the sample size in terms of number of events required to detect this difference with desired statistical power and significance level is determined. Depending on the anticipated accrual rate and the prognosis (e.g., rapidity of failure events) in the control treatment group, the number of patients required can then be approximated. The number of events required depends strongly on the HR, becoming dramatically larger as the HR approaches 1.0 (Table 14.4). The number of patients required and total duration of the trial depend on the rate of patient accrual and the failure rate in the control group, both of which contribute to the determination of how rapidly the requisite events will be observed. The accrual rate is typically estimated from previous experience and may also involve querying investigators to project the accrual rate per unit of time. Similarly, the failure rate for patients under standard therapy is derived from available data. The final computations are straightforward but generally require computer programs,42 although under certain assumptions can be approximated.43 Sample size methods have been extended to take into account other factors that will influence power, such as patients withdrawing from treatment (dropout), switching from the assigned treatment to the other group (crossover), or deviating from protocol treatment (noncompliance).4446

TABLE 14.3 SAMPLE SIZE FOR A TWO-ARM COMPARATIVE (1:1 ALLOCATION) TRIAL WITH A BINARY END POINTA

TABLE 14.4 SAMPLE SIZE FOR A TWO-ARM COMP ARATIVE (1:1 ALOC ATION) TRIAL WITH A TIME-TO-EVENT END POINTA

Interim Analysis and Stopping Rules

Primarily for ethical considerations but also to make the best use of resources, interim analysis plans are used in all phases of clinical trials. These plans provide for early decision making in a trial regarding continuation, disclosure of findings, or modification of the trial while preserving integrity of the study with respect to power and type I error control described earlier. These methods are needed because with repeated hypothesis tests, the probability of at least one test resulting in an erroneous rejection of the null hypothesis increases.

Phase I trials have stopping rules that are integral to the design, in that termination of enrollment to a given dose is based on observed cumulative adverse event rates at a given time. We refer to a review of the designs for more details.1

Phase II trials more formally incorporate stopping rules, usually restricted to futility stopping, or discontinuation when results do not appear promising. For trials with discrete end points such as tumor response, this is accomplished through multistage study designs, whereby a cohort of patients is enrolled and assessed for response, and if a specific minimum response proportion is observed, the trial continues to full accrual or otherwise discontinues enrollment. The most commonly used designs are those proposed by Simon,47 although there are other similar approaches.2 In trials with a time-to-event end point, futility rules similar to those for phase III trials (discussed next) can be used to discontinue after a period of follow-up if results appear unpromising. Early stopping of single-arm phase II trials for extraordinary efficacy is unusual but certainly not prohibited.

Phase III trials use repeated testing strategies derived from an area of statistics known as group sequential methods.48 Briefly, the primary hypothesis is evaluated at predefined increments (typically three to five looks) of the total information (usually in the form of failure events) needed for definitive analysis. The individual tests are designed to (a) protect against spurious early stopping owing to the unstable nature of “early” results and (b) correct for the effect of repeated testing on type I error, which can also lead to spurious declaration of treatment effects that may not be reliable. Commonly used approaches include the Haybittle-Peto approach, for which each test through the penultimate look requires a constant highly significant result, such as p < 0.0001, in order to stop,49,50 and the O’Brien-Fleming approach and its subsequent approaches,51 in which the required significance level decreases over the looks, becoming less extreme as more information accumulates. There are many variations and extensions of the latter approach, with different properties and advantages in special circumstances. In addition to efficacy monitoring rules, phase III trials increasingly also incorporate formal futility stopping rules, although methods such as conditional power calculations have been available and used for some time.52,53 Futility stopping rules similarly involve setting a boundary such that when the test statistic falls beyond it, one considers stopping because the new treatment will not ultimately prevail. Stopping for futility is a complex decision requiring careful consideration,54 and methods continue to be studied and developed.55,56

Definitive Analysis and Secondary Analyses

When a trial reaches maturity either as planned or earlier as a result of the monitoring plan, then definitive analysis takes place. Prior to this, it not conventional or recommended to disclose any results from the trial,57 and this policy is adhered to in National Cancer Institute (NCI) cooperative group trials.

A critical aspect of clinical trial analysis is the definition of the analyzed cohort. The concept of analysis by intention to treat is often cited; however, the definition of this term can sometimes be unclear, thus it is best to explicitly describe which patients are included.58 In the strictest sense, the intention-to-treat cohort includes all patients randomized, regardless of eligibility, adherence to assigned treatment, or any other postrandomization deviations from protocol. However, it is often the case that patients found ineligible for the trial after randomization owing to being incorrectly staged or for other reasons are excluded from the primary analysis, and this practice (used with caution) is sometimes advocated, as it allows for evaluation of the therapy in the population for whom it was intended.2 A rarely acceptable practice involves exclusion of patients who did not or could not comply with assigned therapy regimens or received nonprotocol therapy or other postrandomization conditions. Such exclusions can easily lead to biased comparisons, and in general, any post-hoc analysis of treatment benefit by dose received is fraught with interpretational difficulties and should be avoided in primary analysis.59

Primary analysis methods follow naturally from a well-written protocol (see later discussion) and thus should be straightforward. Given that major journals increasingly require that study protocols be provided at the time of publication, and that regulatory agencies and public sponsors do likewise, it behooves the trialist to outline the analysis plan in the protocol and then carry it out at study conclusion. This does not suggest that additional analyses cannot be carried out but that having a framework for the planned analysis adds credibility to the findings.

Secondary prognostic factor analysis using statistical models or other techniques often follows primary analysis of phase III trials. Of particular interest is whether there are particularly responsive or nonresponsive subsets of patients in an attempt to render the findings more relevant to practice. The modeling process—which entails deciding what factors to include, determining the correct way to represent a given factor (i.e., in categories, on a continuous scale, etc.), consideration of interrelationships (e.g., interactions) among factors, and many other issues—can be complex, and it should be recognized that these analyses will be largely viewed as exploratory. Although possibly worth exploring, true differential effects of treatment by other factors usually require a large sample size, unless the effects are very large.60 A comprehensive review of current modeling methods applied to oncology data is provided by Schumacher et al.61

PRACTICAL ISSUES IN THE DESIGN, CONDUCT, AND REPORTING OF CLINICAL TRIALS

Protocol Document and Study Conduct

The goal of a clinical trial is to answer a well-formulated question that will change clinical practice. To achieve that, the investigators must know the current state of knowledge on the studied disease; clearly describe the eligibility criteria of the studied population; understand the number of patients who will be eligible in their institution/cooperative group; choose simple and achievable end points; establish statistical assumptions based on thorough review of pre-existing data; and collaborate with a biostatistician to decide on study design, sample size, and power.

The clinical trial protocol document must contain the title, investigator and sponsor names, phase (I, II, or III), protocol synopsis, background knowledge, study design and schema, objectives, methodology, subject selection criteria, registration procedures, treatment plan, dosing modifications, adverse events reporting, data and safety monitoring plan, study calendar, outcome measures, data reporting, statistical considerations, and the informed consent.62

Choosing the right study end points is crucial, because it needs to reflect the primary goal of the trial. Any number of end points may be of suitable scientific and clinical value, although if there is interest in regulatory approval, then obviously the end point must reflect the requirements of those parties involved. Overall survival largely remains the gold standard for a registration trial designed to gain marketing approval. However, survival length may be affected by effective salvage therapies or by patients “crossing over” to the other study arm. End points such as DFS and PFS have been used for either expedited drug approval or regular approval, depending on the disease site.63 If end points subject to assessment bias are to be used (e.g., PFS or tumor response), then appropriate bias reduction measures are needed. One approach to circumvent this problem is an independent review panel—such as radiologists reviewing baseline and follow-up images to quantify tumor responses and note the moment of tumor progression—consisting of experts not associated with the trial and unaware of the arm to which the patient was enrolled. One must also consider validity of modern end points even under unbiased review. For example, the phenomena of pseudoprogression and pseudoresponse have made imaging-based end points—including overall radiographic response and PFS—problematic.64

Successful completion of a clinical trial requires constant attention to its practical aspects. Sufficient personnel are necessary to ensure the smooth running of the study and safety of the participating subjects. Clinical research nurses, clinical research associates, data managers, and investigational pharmacists are crucial components of the research team.65 Of particular importance is careful and immediate recording and attribution of all adverse events. Severe adverse events have to be reported promptly to appropriate regulatory agencies (the IRB in the institution where the study is open, the U.S. Food and Drug Administration [FDA], and others) and to the study sponsor. Because most protocols have amendments added during their lifetime and new toxicities are reported from other studies, the research protocol commonly evolves over several successive versions. As a result, new versions of consent forms must be created as well, and IRB approval may again be required. It is imperative that patients enrolled to the study sign the most current version of the consent form. All prescribed follow-up tests (imaging, blood work, etc.) have to be scheduled ahead of time and must coincide with the study calendar. Departures from any the procedures are scored as protocol deviations during periodic audits and will impact adversely on the study’s validity. Additionally, designated independent medical monitors are assigned to high-risk trials (such as most single-institution investigator-initiated trials) to continuously review any reported events, which are later evaluated periodically by the institutional DSMC.

Data and Safety Monitoring Committees

The decision to alter a clinical trial in progress, including discontinuation of accrual and/or treatment, depending on its current state, and to release findings early is typically vested in an independent DSMC. In addition to evaluating according to the monitoring rules described earlier, the DSMC considers the information available from the trial as well as external information that bears on treatment for the disease under study. Specifically, it should be noted that the early stopping rules described previously are meant to serve as guidelines, and there may at any decision point be additional considerations that must be taken into account.66 The policies and procedures for NCI-sponsored cooperative group trials provide a good overview of DSMC structure and function.67

Trial Reporting

Once a study is completed and the data are fully analyzed, its results should be reported promptly. Publication of the results of a trial in a scientific journal represents culmination of investigator efforts and allows wide distribution of the findings. However, lack of precise requirements of reporting may lead to inaccurate or biased results presentation. The Consolidated Standards of Reporting Trials (CONSORT) statement is used worldwide to improve the quality of reporting of randomized controlled trials.68 It provides a 25-item checklist of all required elements and a flow diagram to ensure accounting of all enrolled patients. Many journals require authors to follow CONSORT guidelines because “diligent adherence by authors to the checklist items facilitates clarity, completeness, and transparency of reporting.”68 The International Committee of Medical Journal Editors (ICMJE) similarly publishes guidelines on uniform requirements for manuscripts submitted to biomedical journals.69 There have also been calls for improvements in reporting of phase I and II trials.70,71 In addition to quality with respect to content, a full disclosure of any financial conflict of interest by the investigators to the readers is necessary as well. Redundant publications (repeating the same results in several journals) are discouraged, and there is an obligation to publish negative studies. Study of the publication rate of cancer cooperative group trials regardless of findings shows that there is room for improvement with respect to a responsible approach to clinical trial conduct.72

In the United States, reporting requirements are trending toward an expansion to more “open access” sources, based on mandates arising from recently enacted legislation. ClinicalTrials.gov is the largest clinical trials database in the world, run by the National Library of Medicine at NIH. Initially including information only on NIH-sponsored studies, the database now contains studies sponsored by the pharmaceutical companies and demands “basic results” information not later than 1 year after the study’s primary completion date. The requirements for results reporting was prompted by removal of several drugs from the market because of earlier unrealized toxicity—knowledge about which was obscured by lack of publication or other public documentation.

RECENT APPROACHES TO CLINICAL TRIAL DESIGN

An Alternative Phase I Design for Radiation Oncology Trials

As mentioned earlier, phase I trials in radiation therapy present a unique challenge in that toxicities may occur long after treatment and need to be incorporated into dose evaluation. Traditional stepwise designs do not accommodate this; therefore, the time-to-event continual reassessment method (TITE-CRM), an extension of the continual reassessment method (CRM),73 was developed that incorporates the time-to-event (i.e., time-to-toxicity) information for each patient.26,74 In the TITE-CRM approach, a dose-response model is first posited that identifies the starting dose and range to be considered, along with a time frame for events occurring anywhere up to T time units from administration of therapy. Rather than waiting for each cohort of patients to be followed for this length of time, however, one can enter new patients at, say, half-month intervals. As in the original CRM, the first patient is assigned to a dose level on the basis of prior information or, as in the modified CRM, to the lowest candidate dose. At the time the next patients are to be enrolled, the observed toxicities and follow-up times of patients already entered are used to form an updated estimate of the b parameter that defines the dose-response curve, and the dose level for the next patients is selected according to the usual CRM or modified CRM criteria. In simulation studies, the TITE-CRM produced results comparable to its CRM counterpart while significantly reducing the average duration of the trial. However, the TITE-CRM method was associated with slightly more toxicities, particularly in situations where events tend to occur near the end of the observation period, because escalation to the next dose may have already occurred before toxicities were observed.74 Another problematic issue is rapid accrual, where premature escalation of dose may be indicated. A recent review and suggested modifications may make this approach even more suitable for radiation oncology trials.75

Alternative Phase II Trial Designs

As indicated earlier, the value of traditional single-arm phase II trials has been called into question in terms of providing a reliable basis for further pursuit of promising treatments. The currently favored design is a randomized phase II trial with a standard of care comparison group.12,16,76 This approach and the goal of accelerating development has led to consideration enhancements to the phase II design, including adaptive randomization, where one favors enrollment to the arm(s) that seem to be prevailing while reducing probability of enrollment on other arms, and even dropping some treatment arms. This is an idea with a long history77; however, recent innovations in computing and bayesian methods, as well as a newfound interest in accelerated development, have brought it to wider use in some settings.78 In some instances, it may offer advantages but must be weighed against simpler approaches with similar or even greater efficiency.79

Changes in therapeutic approaches also suggest design changes. Because primarily cytostatic agents (i.e., most biologic drugs) are not expected to necessarily result in tumor response in the traditional sense, there is a need to consider alternative phase II designs based on end points other than response rates. For trials enrolling patients who have failed prior therapy, Mick et al.80 propose a method that uses each patient as his or her own control, comparing the time to progression (possibly censored) under the new agent with the time to progression under prior therapy. Rosner et al.81 propose a randomized discontinuation design to evaluate cytostatic drugs in which all patients are initially treated with the experimental agent. After a specified interval, responders remain on drug and those who progress discontinue, whereas those patients with stable disease are randomized to either continued active treatment or placebo. This randomized comparison allows one to assess whether the drug is truly slowing the rate of growth of the tumor, as opposed to the investigators having simply selected patients with slow-growing tumors. Because patients with stable disease form a more homogeneous subgroup, this design also requires a smaller sample size than would a trial that randomized all patients at entry. It is important to note that the purpose of this design is to determine whether the drug is active in an explanatory sense. Whether the percentage of patients exhibiting stable disease is high or low has bearing on the efficiency of the approach, because in the latter case, the total sample size required may be quite large and any demonstration of activity in the randomized component would only be relevant to a small subset of the population. Korn et al.82 point out other caveats with this design. For example, patients may find it unattractive to potentially discontinue a treatment that they perceive to be helping their disease.

Using Biomarkers as Inclusion Criteria

It is increasingly understood that the response of tumors to targeted agents highly depends on their molecular subtype. Consequently, trials increasingly screen for molecular characteristics of tumors to use as eligibility criteria or, at a minimum, stratification factors. For example, recent and currently accruing RTOG brain tumor trials stratify patients according to whether or not the MGMT gene is methylated, whereas head and neck cancer trials require human papillomavirus (HPV) status to be determined at entry. When studies enroll sufficient numbers of patients, then treatment by marker synergisms, referred to statistically as interaction effects, can be investigated. Robust evaluation of true differential benefit according to markers requires that treatments be randomized; thus, randomized phase II and phase III trials are ideal settings for developing tailored treatments.

In many instances, potentially responsive subsets of patients may be small and may also be identified after trials have initiated enrollment. For instance, anaplastic lymphoma kinase (ALK) inhibitors are highly effective, although only in the 3% to 5% of lung cancers that have ALK gene rearrangements. It may simply not be feasible to perform separate phase III trials for each lung cancer subtype. One possible alternative is “adaptive randomized” trial designs in which the data gathered as the trial progresses is used to change some aspect of the trial as it progresses. Some recent examples are the Biomarker-integrated Approaches of Targeted Therapy for Lung Cancer Elimination (BATTLE) trial in lung cancer83 and the I-SPY trials in breast cancer.84 However, there are limitations and challenges to these complex trial designs. For example, to be able to acquire information rapidly enough to undertake weighted randomization favoring more promising arms or to eliminate nonresponsive arms, surrogate end points such as “disease control rate at 8 weeks” must be used. It is not clear whether such short-term end points are relevant to radiation oncology where local control is very frequently achieved.

Several recent papers in the clinical literature have provided excellent reviews of the opportunities and challenges involved in incorporating modern molecular medicine into clinical trial design.85,86

MOVING BEYOND CLINICAL TRIALS

Comparative Effectiveness Research

The randomized phase III clinical trial is considered the most robust method of comparing the efficacy of a new treatment with the standard of care. Grading systems for evaluating clinical evidence universally place randomized controlled trials above observational trials.87 More recently, this hierarchical approach to scientific evidence has been attacked;17 criticisms include the high fiscal cost of clinical trials, the length of time that it takes to obtain a conclusion (by which time the results are frequently no longer relevant because the standard of care has changed), and the large number of trials with negative results.

A specific criticism relates to clinical protocols that typically allow enrollment of only the fittest patients who lack comorbidities and have good performance status—criteria that subsequently limit the generalizability of the results. For example, studies have shown that older patients are excluded unnecessarily out of concern for potential adverse events.8890 There have indeed long been calls for simpler and more inclusive eligibility criteria.91

In some sense, the desire to reduce exclusivity and broaden trial enrollment to more closely match the population is antithetical to “personalized medicine” and more focused trials, as mentioned previously. However, there are some ways in which the two concepts can possibly work in concert. Larger trials that can robustly support subset analysis according to biologic, clinical, and health history/behavior factors can at once be more inclusive and address questions regarding particularly responsive subgroups.

Another response to the criticism of phase III trials has been a reappreciation of the importance of population-based retrospective studies as a way to measure a treatment’s effectiveness in the “real world.” Even more useful are well-designed prospective cohort studies, such as Cancer of the Prostate Strategic Urologic Research Endeavor (CaPSURE), which will provide information both on factors driving treatment choice and the effects of specific intervention strategies for which randomized trial evidence is currently lacking.92 Another key strategy is conducting trials in parallel with concurrent registries of patients treated according to physician and patient choice and/or common convention, such as the Trial Assigning Individualized Options for Treatment (Rx), or TAILORx.93 In this trial, 7,000 women with breast cancer are screened using a molecular profiling tool. For those with profiles in the range where the utility of the tool is uncertain, randomization between hormonal therapy alone and hormonal therapy plus chemotherapy is performed. Patients obtaining scores below or above this range are registered to accurately record treatment choices and ensure that good follow-up data is obtained. This trial will both serve to determine the utility of the profiling tool in the uncertain range and provide high-quality data on the validity of treatment decisions based on it.

Meta-analysis

A formal quantitative means of combining evidence from multiple clinical trials is by meta-analysis—a widely used analytic tool in many areas of social and medical science. Meta-analysis refers to a process whereby data from independent studies are combined to form a quantitative summary estimate of a given effect.

Meta-analyses are considered by some to be a level I evidence source along with large randomized clinical trials. A meta-analysis combines results of several studies, all of which ask a similar research question but may be individually too small to have enough statistical power to definitively answer the question. Performing a systematic analysis of data from all identified randomized trials can define a modest yet real advantage associated with a new therapeutic approach. To address the likelihood of publication bias (i.e., a greater representation of trials with positive results appearing in the literature), the meta-analysis should include unpublished studies as well, although their quality may be sometimes doubtful because of lack of peer review. The choice of trials to be included is critical, because analyzing trials with disparate patient populations or treatment methods may lead to erroneous conclusions. Despite some limitations, meta-analyses can be influential in guiding treatment practice. For example, a meta-analysis of sequential versus concurrent chemotherapy combined with thoracic radiotherapy in stage III non–small cell lung cancer94 confirmed the concurrent approach to be superior in overall survival and contributed to its validity as standard therapy. An excellent example of the methods and data summaries used in meta-analysis in oncology can be found in the reports of the Early Breast Cancer Trialists’ Collaborative Group.95,96

CONCLUSIONS

Advances in molecular biology over the previous two decades have led to an explosion in the number of new anticancer agents being developed, as well as a rapid increase in the number of clinical trials performed, with resulting improvements in the overall survival and quality of life of cancer patients. In parallel, however, the escalating costs of conducting trials, their increasing complexity, and the ever-expanding regulatory requirements have placed an undue burden on all involved, which has strained the available resources. Inadequate harmonization between countries preventing the straightforward conduct of international trials is another source of inefficiency, sometimes leading to unnecessary duplication of efforts between American and European cooperative groups. Furthermore, the imbalance of funding between academic-sponsored (e.g., NCI) versus pharmaceutical company–sponsored studies may lead to competition for available patients. The public’s and physicians’ awareness and understanding of the importance of clinical trial enrollment need to be improved; likewise, enhancing diverse socioeconomic and ethnic groups’ access to trials is essential.

Contemporary efforts to make electronic clinical trials management systems widespread and unified, to have tissue specimens and data banks accessible to all researchers, and the study results readily available to the public will facilitate the optimal utilization of these precious resources and, most importantly, lead to meaningful improvements in the care of cancer patients.

ACKNOWLEDGMENTS

Some content was modified from previous editions of this chapter. The current authors gratefully acknowledge author contributions from the previous edition.

REFERENCES

1. Crowley J, Ankerst D. Handbook of statistics in clinical oncology. New York: Chapman & Hall/CRC, 2006.

2. Green S, Benedetti J, Smith A. Clinical trials in oncology. Chapman & Hall/CRC, 2012.

3. Kelly K, Halabi S. Oncology clinical trials: successful design, conduct and analysis. New York: Demos Medical, 2009.

4. Unknown. Wikipedia. Available at: http://en.wikipedia.org/wiki/Clinical_trial.

5. Piantadosi S. Clinical trials: a methodologic perspective. John Wiley & Sons, 2005.

6. National Cancer Institute. NCI’s Clinical Trials Cooperative Group Program. Available at: http://www.cancer.gov/cancertopics/factsheet/NCI/clinical-trials-cooperative-group.

7. Hearn J, Sullivan R. The impact of the ‘Clinical Trials’ directive on the cost and conduct of non-commercial cancer trials in the UK. Eur J Cancer 2007;43:8–13.

8. Emanuel EJ, Wendler D, Grady C. What makes clinical research ethical? JAMA 2000;83:2701–2711.

9. Wendler D. The ethics of clinical research. In: Zalta EN, ed. The Stanford encyclopedia of philosophy, 2009. Available at: http://plato.stanford.edu/cgi-bin/encyclopedia/archinfo.cgi?entry=clinical-research.

10. Piantadosi S. Principles of clinical trial design. Semin Oncol 1988;15:423–433.

11. Ratain MJ, Sargent DJ. Optimising the design of phase II oncology trials: the importance of randomisation. Eur J Cancer 2009;45:275–280.

12. Mandrekar SJ, Sargent DJ. Randomized phase II trials: time for a new era in clinical trial design. J Thorac Oncol 2010;5:932–934.

13. Tang H, Foster NR, Grothey A, et al. Comparison of error rates in single-arm versus randomized phase II cancer clinical trials. J Clin Oncol 2010;28:1936–1941.

14. Leventhal BG. An overview of clinical trials in oncology. Semin Oncol 1988;15:414–422.

15. Simon R, Wittes RE, Ellenberg SS. Randomized phase II clinical trials. Cancer Treat Rep 1985;69:1375–1381.

16. Rubinstein LV, Korn EL, Freidlin B, et al. Design issues of randomized phase II trials and a proposal for phase II screening trials. J Clin Oncol 2005;23:7199–7206.

17. Vogelbaum MA. The future of clinical research beyond phase III trials. Clin Neurosurg 2009;56:37–39.

18. FitzGerald TJ, Urie M, Ulin K, et al. Processes for quality improvements in radiation oncology clinical trials. Int J Radiat Oncol Biol Phys 2008;71:S76.

19. Justin EB, Joachim Y. Quality of radiotherapy reporting in randomized controlled trials of Hodgkin’s lymphoma and non-Hodgkin’s lymphoma: a systematic review. Int J Radiat Oncol Biol Phys2009;73:492–498.

20. Morris SL, Beasley M, Leslie M. Chemotherapy for pancreatic cancer. N Engl J Med 2004;350:2713–2715; author reply 2713–2715.

21. Bydder S, Spry N. Chemotherapy for pancreatic cancer. N Engl J Med 2004;350:2713–2715; author reply 2713–2715.

22. Crane CH, Ben-Josef E, Small W Jr. Chemotherapy for pancreatic cancer. N Engl J Med 2004;350:2713–2715; author reply 2713–2715.

23. Rischin D, Peters L, Fisher R, et al. Tirapazamine, cisplatin, and radiation versus fluorouracil, cisplatin, and radiation in patients with locally advanced head and neck cancer: a randomized phase II trial of the Trans-Tasman Radiation Oncology Group (TROG 98.02). J Clin Oncol 2005;23:79.

24. Weiner MA, Leventhal B, Brecher ML, et al. Randomized study of intensive MOPP-ABVD with or without low-dose total-nodal radiation therapy in the treatment of stages IIB, IIIA2, IIIB, and IV Hodgkin’s disease in pediatric patients: a Pediatric Oncology Group study. J Clin Oncol 1997;15:2769–2779.

25. Abrams RA, Winter KA, Regine WF, et al. Failure to adhere to protocol specified radiation therapy guidelines was associated with decreased survival in RTOG 9704-A phase III trial of adjuvant chemotherapy and chemoradiotherapy for patients with resected adenocarcinoma of the pancreas. Int J Radiat Oncol Biol Phys 2012;82(2):809–816.

26. Normolle D, Lawrence T. Designing dose-escalation trials with late-onset toxicities using the time-to-event continual reassessment method. J Clin Oncol 2006;24:4426–4433.

27. Glass C, Den R, Dicker AP, et al. Toxicity of phase I radiation oncology trials: worldwide experience, abstract #1605. American Society for Therapeutic Radiation Oncology (ASTRO) 52nd Annual Meeting, October 31–November 4, 2010, San Diego, CA.

28. Liu PY, LeBlanc M, Desai M. False positive rates of randomized phase II designs. Control Clin Trials 1999;20:343–352.

29. Therasse P, Arbuck SG, Eisenhauer EA, et al. New guidelines to evaluate the response to treatment in solid tumors. European Organization for Research and Treatment of Cancer, National Cancer Institute of the United States, National Cancer Institute of Canada. J Natl Cancer Inst 2000;92:205–216.

30. Chen TT, Chute JP, Feigal E, et al. A model to select chemotherapy regimens for phase III trials for extensive-stage small-cell lung cancer. J Natl Cancer Inst 2000;92:1601–1607.

31. Buyse M, Thirion P, Carlson RW, et al. Relation between tumour response to first-line chemotherapy and survival in advanced colorectal cancer: a meta-analysis. Meta-Analysis Group in Cancer. Lancet2000;356:373–378.

32. Moertel CG. Improving the efficiency of clinical trials: a medical perspective. Stat Med 1984;3:455–468.

33. Gaynor JJ, Feuer EJ, Tan CC, et al. On the use of cause-specific failure and conditional failure probabilities: examples from clinical oncology data. J Am Stat Assoc 1993:400–409.

34. Dignam JJ, Kocherginsky MN. Choice and interpretation of statistical tests used when competing risks are present. J Clin Oncol 2008;26:4027–4034.

35. Hudis CA, Barlow WE, Costantino JP, et al. Proposal for standardized definitions for efficacy end points in adjuvant breast cancer trials: the STEEP system. J Clin Oncol 2007;25:2127–2132.

36. Sargent DJ, Wieand HS, Haller DG, et al. Disease-free survival versus overall survival as a primary end point for adjuvant colon cancer studies: individual patient data from 20,898 patients on 18 randomized trials. J Clin Oncol2005;23:8664–8670.

37. Prentice RL. Surrogate endpoints in clinical trials: definition and operational criteria. Stat Med 1989;8:431–440.

38. Buyyounouski MK, Hanlon AL, Horwitz EM, et al. Interval to biochemical failure highly prognostic for distant metastasis and prostate cancer-specific mortality after radiotherapy. Int J Radiat Oncol Biol Phys 2008;70:59–66.

39. Denham JW, Steigler A, Wilcox C, et al. Time to biochemical failure and prostate-specific antigen doubling time as surrogates for prostate cancer-specific mortality: evidence from the TROG 96.01 randomised controlled trial. Lancet Oncol 2008;9:1058–1068.

40. Schatzkin A, Gail M. The promise and peril of surrogate end points in cancer research. Nat Rev Cancer 2002;2:19–27.

41. Halpern SD, Karlawish JH, Berlin JA. The continuing unethical conduct of underpowered clinical trials. JAMA 2002;288:358–362.

42. Shuster J. Power and sample size for phase III clinical trials of survival. Handbook of statistics in clinical oncology. New York: Chapman & Hall/CRC, 2006: 207–226.

43. Freedman LS. Tables of the number of patients required in clinical trials using the logrank test. Stat Med 1982;1:121–129.

44. Ahnn S, Anderson SJ. Sample size determination in complex clinical trials comparing more than two groups for survival endpoints. Stat Med 1998;17:2525–2534.

45. Lachin JM, Foulkes MA. Evaluation of sample size and power for analyses of survival with allowance for nonuniform patient entry, losses to follow-up, noncompliance, and stratification. Biometrics1986;42:507–519.

46. Shih JH. Sample size calculation for complex clinical trials with survival endpoints. Control Clin Trials 1995;16:395–407.

47. Simon R. Optimal two-stage designs for phase II clinical trials. Control Clin Trials 1989;10:1–10.

48. Jennison C, Turnbull BW. Group sequential methods with applications to clinical trials. London: Chapman & Hall/CRC, 2000.

49. Haybittle JL. Repeated assessment of results in clinical trials of cancer treatment. Br J Radiol 1971;44:793–797.

50. Peto R, Pike MC, Armitage P, et al. Design and analysis of randomized clinical trials requiring prolonged observation of each patient. II. Analysis and examples. Br J Cancer 1977;35:1–39.

51. Fleming TR, Harrington DP, O’Brien PC. Designs for group sequential tests. Control Clin Trials 1984;5:348–361.

52. Halperin M, Lan KK, Ware JH, et al. An aid to data monitoring in long-term clinical trials. Control Clin Trials 1982;3:311–323.

53. Lan KK, Wittes J. The B-value: a tool for monitoring data. Biometrics 1988;44:579–585.

54. Dignam JJ, Bryant J, Wieand HS, et al. Early stopping of a clinical trial when there is evidence of no treatment benefit: protocol B-14 of the National Surgical Adjuvant Breast and Bowel Project. Control Clin Trials1998;19:575–588.

55. Freidlin B, Korn EL, Gray R. A general inefficacy interim monitoring rule for randomized clinical trials. Clin Trials 2010;7:197–208.

56. Freidlin B, Korn EL. A comment on futility monitoring. Control Clin Trials 2002;23:355–366.

57. Fleming TR, Sharples K, McCall J, et al. Maintaining confidentiality of interim data to enhance trial integrity and credibility. Clin Trials 2008;5:157–167.

58. Gail MH. Eligibility exclusions, losses to follow-up, removal of randomized patients, and uncounted events in cancer clinical trials. Cancer Treat Rep 1985;69:1107–1113.

59. Redmond C, Fisher B, Wieand HS. The methodologic dilemma in retrospectively correlating the amount of chemotherapy received in adjuvant therapy protocols with disease-free survival. Cancer Treat Rep 1983;67:519–526.

60. Schmoor C, Sauerbrei W, Schumacher M. Sample size considerations for the evaluation of prognostic factors in survival analysis. Stat Med 2000;19:441–452.

61. Schwarzer G, Schumacher M, Sauerbrei W, et al. Prognostic factor studies. In: Crowley J, Ankerst DP eds. Handbook of Statistics in Clinical Oncology, 2nd ed. New York: Chapman & Hall, 2005:289–333.

62. Grant N, Sacatos M, Kelly K. The trials and tribulations of writing an investigator initiated clinical study. In: Kelly K, Halabi S, eds. Oncology clinical trials: successful design, conduct and analysis. New York: Demos Medical, 2009:119–130.

63. FDA. Guidance for industry: clinical trial endpoints for the approval of cancer drugs and biologics. Available at: http://www.fda.gov/downloads/drugs/GuidanceComplianceRegulatoryInformation/Guidances/UCM071590.pdf. Accessed November 9, 2011.

64. Brandsma D, Stalpers L, Taal W, et al. Clinical features, mechanisms, and management of pseudoprogression in malignant gliomas. Lancet Oncol 2008;9:453–461.

65. De Pourcq F. Defining the roles and responsibilities of study personnel. In: Kelly K, Halabi S, eds. Oncology clinical trials: successful design, conduct and analysis. New York: Demos Medical, 2009:321–326.

66. Lan KK, Lachin JM, Bautista O. Over-ruling a group sequential boundary—a stopping rule versus a guideline. Stat Med 2003;22:3347–3355.

67. Smith MA, Ungerleider RS, Korn EL, et al. Role of independent data-monitoring committees in randomized clinical trials sponsored by the National Cancer Institute. J Clin Oncol 1997;15:2736–2743.

68. Schulz KF, Altman DG, Moher D. CONSORT 2010 statement: updated guidelines for reporting parallel group randomised trials. BMJ;340:c332.

69. International Committee of Medical Journal Editors (ICMJE). Uniform requirements for manuscripts submitted to biomedical journals: writing and editing for biomedical publication. Haematologica2004;89:264.

70. Mariani L, Marubini E. Content and quality of currently published phase II cancer trials. J Clin Oncol 2000;18:429–436.

71. Zohar S, Lian Q, Levy V, et al. Quality assessment of phase I dose-finding cancer trials: proposal of a checklist. Clin Trials 2008;5:478–485.

72. Krzyzanowska MK, Pintilie M, Tannock IF. Factors associated with failure to publish large randomized trials presented at an oncology meeting. JAMA 2003;290:495–501.

73. O’Quigley J, Pepe M, Fisher L. Continual reassessment method: a practical design for phase 1 clinical trials in cancer. Biometrics 1990;46:33–48.

74. Cheung YK, Chappell R. Sequential designs for phase I clinical trials with late-onset toxicities. Biometrics 2000;56:1177–1182.

75. Polley MY. Practical modifications to the time-to-event continual reassessment method for phase I cancer trials with fast patient accrual and late-onset toxicities. Stat Med 2011;30:2130–2143.

76. Cannistra SA. Phase II trials in Journal of Clinical Oncology. J Clin Oncol 2009;27:3073–3076.

77. Zelen M. Play the winner rule and the controlled clinical trial. J Am Stat Assoc 1969;64:131–146.

78. Biswas S, Liu DD, Lee JJ, et al. Bayesian clinical trials at the University of Texas M.D. Anderson Cancer Center. Clin Trials 2009;6:205–216.

79. Korn EL, Freidlin B. Outcome—adaptive randomization: is it useful? J Clin Oncol 2011;29:771–776.

80. Mick R, Crowley JJ, Carroll RJ. Phase II clinical trial design for noncytotoxic anticancer agents for which time to disease progression is the primary endpoint. Control Clin Trials 2000;21:343–359.

81. Rosner GL, Stadler W, Ratain MJ. Randomized discontinuation design: application to cytostatic antineoplastic agents. J Clin Oncol 2002;20:4478–4484.

82. Korn EL, Arbuck SG, Pluda JM, et al. Clinical trial designs for cytostatic agents: are new approaches needed? J Clin Oncol 2001;19:265–272.

83. Kim ES, Herbst RS, Wistuba II, et al. The BATTLE trial: personalizing therapy for lung cancer. Cancer Discov 2011;1:44–53.

84. Barker AD, Sigman CC, Kelloff GJ, et al. I-SPY 2: an adaptive breast cancer trial design in the setting of neoadjuvant chemotherapy. Clin Pharmacol Ther 2009;86:97–100.

85. Freidlin B, McShane LM, Korn EL. Randomized clinical trials with biomarkers: design issues. J Natl Cancer Inst 2010;102:152–160.

86. Simon R. The use of genomics in clinical trial design. Clin Cancer Res 2008;14:5984–5993.

87. Harbour R, Miller J. A new system for grading recommendations in evidence based guidelines. BMJ 2001;323:334–336.

88. Hutchins LF, Unger JM, Crowley JJ, et al. Underrepresentation of patients 65 years of age or older in cancer-treatment trials. N Engl J Med 1999;341:2061–2067.

89. Kumar A, Soares HP, Balducci L, et al. Treatment tolerance and efficacy in geriatric oncology: a systematic review of phase III randomized trials conducted by five National Cancer Institute–sponsored cooperative groups. J Clin Oncol 2007;25:1272–1276.

90. Lewis JH, Kilgore ML, Goldman DP, et al. Participation of patients 65 years of age or older in cancer clinical trials. J Clin Oncol 2003;21:1383–1389.

91. George SL. Reducing patient eligibility criteria in cancer clinical trials. J Clin Oncol 1996;14:1364–1370.

92. Lubeck DP, Litwin MS, Henning JM, et al. The CaPSURE database: a methodology for clinical practice and research in prostate cancer. CaPSURE Research Panel. Cancer of the Prostate Strategic Urologic Research Endeavor. Urology 1996;48:773–777.

93. Sparano JA. TAILORx: trial assigning individualized options for treatment (Rx). Clin Breast Cancer 2006;7:347–350.

94. Auperin A, Le Pechoux C, Rolland E, et al. Meta-analysis of concomitant versus sequential radiochemotherapy in locally advanced non-small-cell lung cancer. J Clin Oncol 2010;28:2181–2190.

95. Clarke M, Collins R, Darby S, et al. Effects of radiotherapy and of differences in the extent of surgery for early breast cancer on local recurrence and 15-year survival: an overview of the randomised trials. Lancet2005;366:2087–2106.

96. Davies C, Godwin J, Gray R, et al. Relevance of breast cancer hormone receptors and other factors to the efficacy of adjuvant tamoxifen: patient-level meta-analysis of randomised trials. Lancet2011;378:771–784.



If you find an error or have any questions, please email us at admin@doctorlib.org. Thank you!