Xiaomei Ma and Herbert Yu
INTRODUCTION
Epidemiology is the study of the distribution and determinants of health-related states or events in specified populations and the application of this study to control health problems.1 Epidemiologic principles and methods have long been applied to cancer research, with the assumptions that cancer does not occur at random and the nonrandomness of carcinogenesis can be elucidated through systematic research. An example of such applications is the lung cancer study conducted by Doll and Hill in the early 1950s, which linked tobacco smoking to an increased mortality of lung cancer in over 40,000 medical professionals in the United Kingdom.2 The observation from this study and many other studies, in conjunction with laboratory findings regarding the underlying biologic mechanisms for the effect of tobacco smoking, helped establish the role of tobacco smoking in the etiology of lung cancer. Epidemiologic methods are also used in clinical settings, where trials are conducted to evaluate the efficacy of new treatment protocols or preventive measures and where observational studies of prognostic factors are done.
Epidemiologic studies can take different forms, but generally they can be classified into two broad categories, observational studies and experimental studies (Fig. 11.1). In experimental studies, an investigator allocates different study regimens to the subjects, usually with randomization (experimental studies without randomization are sometimes referred to as “quasi-experiments”).3 Experimental studies can be individual based or community based. An experimental study most closely resembles laboratory experiments in that the investigator has control over the study condition. Experimental studies can be used to evaluate the efficacy of a treatment protocol (e.g., low-dose compared with standard-dose chemotherapy for non-Hodgkin’s lymphoma)4 or preventive measures (e.g., tamoxifen for women at an increased risk of breast cancer).5 Although experimental studies are often considered the “gold standard” because of well-controlled study situations, they are only suitable for the evaluation of effects that are beneficial or at least not harmful due to ethical concerns. Experimental studies are discussed in detail in other chapters of this book. This section will focus on observational studies.

Observational studies do not involve the artificial manipulation of study regimens. In an observational study, an investigator stands by to observe what happens or happened to the subjects, in terms of exposure and outcome. Observational studies can be further divided into descriptive and analytical studies (see Fig. 11.1). Descriptive studies focus on the distribution of diseases with respect to person, place, and time (i.e., who, where, and when), whereas analytical studies focus on the determinants of diseases. Descriptive studies are often used to generate hypotheses, whereas analytical studies are often used to test hypotheses. However, the two types of studies should not be considered mutually exclusive entities; rather, they are the opposite ends of a continuum. Descriptive studies are discussed in detail in other chapters of this book.
ANALYTICAL STUDIES
Ecologic Studies
As in experimental studies, the unit of analysis can be individuals or groups of people in observational studies. Studies that use groups of people as the unit of analysis are called ecologic studies, which are relatively easy to carry out when group level measures are available. However, a relationship observed between variables on a group level does not necessarily reflect the relationship that exists at an individual level. For example, the fraction of energy supply from animal products was found to be positively correlated with breast cancer mortality in a recent ecologic study, which used preexisting data on both dietary supply and breast cancer mortality rates from 35 countries.6 Because the data were country based, no reliable inference can be made at an individual level. Within each country, it could be that the people who had a low fraction of energy supply from animal products were actually dying from breast cancer. Results from ecologic studies are useful for inference at an individual level only when the within-group variability of the exposure is low so that a group-level measure can reasonably reflect exposure at an individual level. Alternatively, if the implications for prevention or intervention are at a group level (e.g., taxation of cigarettes to reduce smoking), results from ecologic studies are very useful.
Cross-Sectional Studies
There are three main types of analytical studies in which the unit of analysis is individuals: cross-sectional, cohort, and case-control studies. In a cross-sectional study, the information on various factors is collected from the study population at a given point in time. From a public health perspective, data collected in cross-sectional studies can be of great value in assessing the general health status of a population and allocating resources. For example, the National Health and Nutrition Examination Survey has provided valuable national estimates of health and nutritional status of the US civilian, noninstitutionalized population.7 Findings from cross-sectional studies can also help generate hypotheses that may be tested later in other types of studies. However, it should be noted that cross-sectional studies have serious methodologic limitations if the research purpose is etiologic inference. Because exposures and disease status are evaluated simultaneously, it is usually not possible to know the temporality of events unless the exposure cannot change over time (e.g., blood type, skin color, race, country of birth). If one observes that more brain cancer patients are depressed than people without brain cancer in a cross-sectional study, the correlation does not necessarily mean that depression causes brain cancer. Depression may simply have resulted from the pathogenesis and diagnosis of brain cancer, or depression may have caused brain cancer in some patients and resulted from brain cancer in other patients. Without additional information on the timing of events, no conclusions can be made. Another concern in cross-sectional studies is the enrollment of prevalent cases, who survived different lengths of time after the incidence of disease. Factors that affect survival may also influence incidence. Prevalent cases may not be representative of incident cases, which makes etiologic inferences based on cross-sectional studies suspect at best.
Cohort Studies
In a cohort study, a study population free of a specific disease (or any other health-related condition) is grouped based on their exposure status and followed up for a certain period of time. Then the exposed and unexposed subjects are compared with respect to disease status at the end of the follow-up. The objective of a cohort study is usually to evaluate whether the incidence of a disease is associated with an exposure. The cohort design is fundamental in observational epidemiology and is considered “ideal” in that, if unbiased, cohort data reflect the real-life cause/effect sequence of disease.8 Subjects in cohort studies may be a sample of the general population in a geographic area, a group of workers who are exposed to certain occupational hazards in a specific industry, or people who are considered at a high risk for a specific disease. A cohort study is considered prospective or concurrent if the investigator starts following up the cohort from the present time into the future, and retrospective or historical if the cohort is established in the past based on existing records (e.g., an occupational cohort based on employment records) and the follow-up ends before or at the time of the study. Alternatively, a cohort study can be ambidirectional in that data collection goes both directions.9 Whether a cohort study is prospective, retrospective, or ambidirectional, the key feature is that all the subjects were free of the disease at the beginning of the follow-up and the study tracks the subjects from exposure to disease. Follow-up time, ranging from days to decades, is an essential element in cohort studies.
In a cohort study, the incidence of disease in the exposed group and the unexposed group is compared. The incidence measure can be cumulative incidence or incidence density, depending on the availability of data. When comparing the incidence in the two groups, both relative differences and absolute differences can be assessed. In cohort studies, the relative risk of developing the disease is expressed as the ratio of the cumulative incidence in the exposed group to that in the unexposed group, which is also called cumulative incidence ratio or risk ratio. If we have data on the exact person-time of follow-up for every subject, we can also calculate an incidence density ratio (also called rate ratio) in a similar way. The numeric value of the risk or rate ratio reflects the magnitude of the association between an exposure and a disease. For example, a risk ratio of 2 would be interpreted as exposed individuals have a doubled risk of developing a disease than unexposed individuals, whereas a risk ratio of 5 indicates that exposed individuals have 5 times the risk of developing a disease compared with unexposed individuals. To put in another way, a factor with a risk ratio of 5 has a stronger effect than another factor with a risk ratio of 2. In addition to risk ratio and rate ratio, another relative measure called probability odds ratio can be calculated in cohort studies. The probability odds of disease is the number of subjects who developed a disease divided by the number of subjects who did not develop the disease, and the probability odds ratio is the probability odds in the exposed group divided by the probability odds in the unexposed group. Many investigators prefer risk ratio or rate ratio to probability odds ratio in cohort studies, because the ability to directly measure the risk of developing a disease is one of the most significant advantages in cohort studies. In practice, however, a probability odds ratio is often used as an approximation for risk or rate ratio, especially when multivariate logistic regression models are employed to adjust for the effect of other factors that may influence the relationship between an exposure and a disease.
As for absolute differences, a commonly used measure is called attributable risk in the exposed, which is the incidence in the exposed group minus the incidence in the unexposed group. Attributable risk reflects the disease incidence that could be attributed to the exposure in exposed individuals and the reduction in incidence that we would expect if the exposure can be removed from the exposed individuals, provided that there is a causal relationship between the exposure and the disease. Another absolute measure called population attributable risk extends this concept to the general population; it estimates the disease incidence that could be attributed to an exposure in the general population. Because both relative and absolute differences can be assessed in cohort studies, a natural question to ask is what measures to choose. In general, the relative differences are used more often if the main research objective is etiologic inference, and they can be used for the judgment of causality. Once causality is established, or at least assumed, measures of absolute differences are more important from a public health perspective. This point can be illustrated using the following hypothetical example. Assume the following: toxin X in the environment triples the risk of bladder cancer and toxin Y doubles the risk of bladder cancer, the effects of X and Y are entirely independent of each other, the prevalence of exposure to toxin Y in the general population is 20 times higher than the prevalence of exposure to toxin X, and there are only resources available to reduce the exposure to one toxin. It would be more effective to use the resources to reduce the exposure to toxin Y instead of toxin X. This is because the population attributable risk due to Y is higher than that due to X, although the risk ratio associated with toxin Y is smaller than that associated with toxin X.
Cohort studies have many advantages. A cohort design is the best way to study the natural history of a disease.9 There is usually a clear temporal relationship between an exposure and a disease because all the subjects are free of the disease at the beginning of the follow-up (it can be a problem if a subject has a subclinical disease such as undetected prostate cancer). Furthermore, multiple diseases can be studied with respect to the same exposure. On the other hand, cohort studies, especially prospective cohort studies, are costly in terms of both time and money. A cohort design requires the follow-up of a large number of study participants over a sometimes extremely lengthy period of time and usually extensive data collection through questionnaires, physical measurements, and/or biologic specimens at regular intervals. Participants may be “lost” during the follow-up because they became tired of the study, moved away from the study area, or died from some causes other than the disease under study. If the subjects who were lost during the follow-up are different from those who remained under observation with respect to exposure, disease, or other factors that may influence the relationship between the exposure and the disease, results from the study may be biased. To date, cohort studies have been used to study the etiology of a wide spectrum of diseases, including different types of cancer. If a cohort study is conducted to evaluate the etiology of cancer, usually the study sample size would need to be very large (such as the National Institutes of Health-AARP Diet and Health Study, which included more than half million subjects10) and the follow-up time would need to be long, unless the cohort selected is a high-risk population.
For simplicity, we have discussed cohort studies in which the outcome of interest is the incidence of a specific disease and there are only two exposure groups. In practice, any health-related event can be the outcome of interest, and multiple exposure groups can be compared.
Case-Control Studies
Case-control design is an alternative to cohort design for the evaluation of the relationship between an exposure and a disease (or any other health condition). A case-control approach compares the odds of past exposure between cases and noncases (controls) and uses the exposure odds ratio as an estimate for relative risk. A primary goal in a case-control study is to reach the same conclusions as what would have been obtained from a cohort study, if one had been done.11 If appropriately designed and conducted, a case-control study can optimize speed and efficiency as the need for follow-up is avoided.8 The starting point of a case-control study is a source population from which the cases arise. Instead of obtaining the denominators for the calculation of risks or rates in a cohort study, a control group is sampled from the entire source population. After selecting control subjects, who ideally would have become cases had they developed the disease, an investigator collects data on past exposures from both the cases and the controls and then calculates an odds ratio, which is the odds of exposure in the cases divided by the odds of exposure in the controls.
There are two main types of case-control studies: case-based case-control studies and case-control studies within defined cohorts.8 Some variations of the case-control design also exist. For instance, if the effect of an exposure is transient, sometimes a case can be used as his/her own control (case cross-over design). In case-based case-control studies, cases and controls are selected at a given point in time from a hypothetical cohort (e.g., at the end of follow-up). A cross-sectional ascertainment of cases will result in a case group that mostly contains prevalent cases who may have survived for different lengths of time after disease incidence. Cases who died before an investigator began subject ascertainment would not be eligible to be included in the study. As a result, the cases finally included in the study may not be representative of all the cases from the entire hypothetical cohort. Another disadvantage of enrolling prevalent cases is that cases that were diagnosed a long time ago will likely have difficulties recalling exposures that occurred before the disease incidence. In case-control studies, it is preferable to ascertain incident cases as soon as they are diagnosed and to select controls as soon as cases are identified. Case-control studies that enroll only incident cases are sometimes called prospective case-control studies because the investigators need to wait for the incident cases to develop and get diagnosed. For cancer studies, the cases can be ascertained from population-based cancer registries or hospitals. A major advantage of using a cancer registry is the completeness of case ascertainment; however, the reporting of cancer cases to registries is usually not instantaneous. There could be a lag time of several months or even over a year, and some cases could have died during the lag time. If the cancer under study has a poor survival rate and/or clinical specimens need to be obtained in a timely manner, it may be preferable to identify cases directly from hospitals using a rapid ascertainment protocol. As for the selection of controls, the key issue is that controls should be representative of the source population from which the cases arise and, theoretically, the controls would have been ascertained as cases had they developed the disease. The most common types of controls include population-based controls (often selected through random digit dialing in case-control studies of cancer etiology), hospital controls, and friend controls. The advantages and disadvantages of different types of controls have been nicely summarized by Wacholder et al.12 Because no follow-up is involved in case-based case-control studies, the incidence risk or rate cannot be calculated directly for case and control groups. The odds ratio will be a good estimate of relative risk if the disease is uncommon.
In addition to case-based case-control studies, there are also case-control studies within defined cohorts (also known as hybrid or ambidirectional designs), including case-cohort studies and nested case-control studies. In case-cohort studies, cases are identified from a well-defined cohort after some follow-up time, and controls are selected from the baseline cohort. In nested case-control studies, cases are also identified from a cohort, but controls are selected from the individuals at risk at the time each case occurs (i.e., incidence density sampling).8 In these types of designs, controls are a sample of the cohort and the controls selected can theoretically become cases at some point. The possibility of selection bias in case-control studies within defined cohorts is lower than that in case-based case-control studies because the cases and the controls are selected from the same source population. Because of an increased awareness of the methodological issues inherent in the design of case-based case-control studies and the availability of a growing number of large cohorts, case-control studies within defined cohorts have become more common in recent years. The advantage of case-control studies within cohorts over traditional cohort studies is mainly the efficiency in additional data collection. For instance, a recent nested case-control study evaluated the relationship between endogenous sex hormones and prostate cancer risk.13Instead of measuring the serum hormones levels of the entire cohort (over 12,000 subjects), investigators chose to measure 300 cases and 300 controls selected from the cohort. Doing so not only significantly reduced the cost of measurements and the time it took to address the research question, but also helped preserve valuable serum samples for possible analyses in the future. In a case-cohort design, an odds ratio estimates risk ratio; in a nested case-control design, an odds ratio estimates rate ratio. In both designs, the disease under study does not have to be rare for the odds ratio to be a good estimate of the risk ratio or rate ratio.8,14
The biggest advantage of a case-control design is the speed and efficiency of obtaining data. It is claimed that investigators implement case-control studies more frequently than any other analytical epidemiologic study.15 Because most types of cancer are uncommon and take a long time to develop, to date, most epidemiologic studies of cancer have been case-control instead of cohort in design. A case-control study can be conducted to evaluate the relationship between many different exposures and a specific disease, but the study will have limited statistical power if the exposure is rare. In general, a case-control design tends to be more susceptible to biases than a cohort design. Such biases include, but are not limited to, selection bias when choosing and enrolling subjects (especially controls) and recall bias when obtaining data from the subjects. The status of the subjects—that is, case or control—may affect how they recall and report previous exposures, some of which occurred years or even decades ago. It is important for investigators to explicitly define the diagnostic and eligibility criteria for cases, to select controls from the same population as the cases independent of the exposures of interest, to blind data collection staff to the case or control status of subjects and/or the main hypotheses of the study, to ascertain exposure in a similar manner from cases and controls, and to take into account other factors that may influence the relationship between an exposure and a disease.15
INTERPRETATION OF EPIDEMIOLOGIC FINDINGS
We have discussed measures of effects in various study designs. However, a risk ratio of 3 from a cohort study or an odds ratio of 2.5 from a case-control study does not necessarily mean that there is an association between an exposure and a disease. Several alternative explanations need to be assessed, including chance (random error), bias (systematic error), and confounding. Potential interaction also needs be evaluated.
Statistical methods are required to evaluate the role of chance. A usual way is to calculate the upper and lower limits of a 95% confidence interval around a point estimate for relative risk (risk ratio, rate ratio, or odds ratio). If the confidence interval does not include one, one would say that the observed association is statistically significant; if the confidence interval includes one, one would say that the observed relationship is not statistically significant. The width of a confidence interval is directly related to the number of participants in a study, which is called sample size. A larger sample size leads to less variability in the data, a tighter confidence interval, and a higher possibility in finding a statistically significant association if one truly exists. A 95% confidence interval means that if the data collection and analysis could be replicated many times, the confidence interval should include the correct value of the measure 95% of the time.16 It is better to consider a confidence interval to be a general guide to the amount of random error in the data but not necessarily a literal measure of statistical variability.16
Bias can be defined as any systematic error in an epidemiologic study that results in an incorrect estimate of the association between exposure and disease, and it can occur in every type of epidemiologic study design. There are two main types of bias: selection bias and information bias. Selection bias is present when individuals included in a study are systematically different from the target population. For example, a selection bias would occur if a study aimed to generate a sample representing all women in the United States, but of the women contacted, more with a family history of breast cancer agreed to participate. This sample would be at a higher risk for breast cancer than the target population. Refusal to participate poses a constant challenge in epidemiologic studies. As individuals have become more concerned about privacy issues and as studies have become more demanding of time, biologic specimens, and other impositions, participation rates have dropped substantially in recent years. If nonparticipants are different from the participants with respect to study-related characteristics, the validity of the study is threatened. Information bias occurs when the data collected from the study subjects are erroneous. Information bias is also known as misclassification if the variable is measured on a categorical scale and the error causes a subject to be placed in a wrong category. Misclassification can happen to both exposure and disease. For example, in a case-control study of previous reproductive history and ovarian cancer, a woman who had an extremely early pregnancy loss might not even realize that she was ever pregnant and would mistakenly report no pregnancy, and another woman who has only subclinical presentations of ovarian cancer might be mistakenly selected as a control. Misclassification can be differential or nondifferential. An exposure misclassification is considered differential if it is related to disease status and nondifferential if not related to disease status. Similarly, a disease misclassification is considered differential if it is related to exposure status and nondifferential if not related to exposure status. If a binary exposure variable and a binary disease variable are analyzed, a nondifferential misclassification will result in an underestimate of the true association. Differential misclassification can either exaggerate or underestimate a true effect. Usually not much can be done to control or correct bias at the data analysis stage; therefore, it is important to establish research protocols that are not prone to bias. The evaluation of potential bias is critical to the interpretation of study results. An invalid estimate is worse than no estimate.
Confounding refers to a situation in which the association between an exposure and a disease (or any health-related condition) is influenced by a third variable. This third variable is considered a confounding variable or confounder. A confounder must fulfill three criteria: (1) be associated with the exposure, (2) be associated with the disease independent of the exposure, and (3) not be an intermediate step between the exposure and the disease (i.e., not on the causal pathway). Unlike bias, which is primarily introduced by the investigator or study participants, confounding is a function of the complex interrelationship between various exposures and disease.17 In a hypothetical case-control study of the effect of alcohol drinking on lung cancer, we may observe an odds ratio of 2.5 (usually called a “crude” odds ratio in the sense that no other variables were taken into account), which indicates that alcohol drinking increases the risk of lung cancer by 1.5-fold. However, if we classify all study subjects into two strata based on a history of cigarette smoking and then calculate the odds ratio in the two strata (smokers and nonsmokers) separately, we may have two stratum-specific odds ratios both equal to one, indicating that alcohol drinking is not associated with lung cancer risk. In this example, the crude odds ratio calculated to estimate the association between alcohol drinking and lung cancer without considering smoking is simply misleading. Being associated with both the exposure (i.e., alcohol drinking) and the disease (i.e., lung cancer), smoking acted as a confounder in this example. A stratified analysis is needed to evaluate the potential confounding effect of a third variable, whether it is done with pencil and paper or statistical modeling. Usually data are stratified based on the level of a third variable. If the stratum-specific effect measures are similar to each other but different from the crude effect measure, confounding is said to be present. In this section, we have illustrated basic epidemiologic principles using an overly simplified scenario and only considered a single exposure. In practice, most if not all diseases, cancer included, have a multifactorial etiology. Consequently, it is usually necessary to assess the potential confounding effect of a group of variables simultaneously using multivariate statistical models. The effect measure derived from a multivariate model will then be called an “adjusted” one in the sense that the effect of other factors was also adjusted for. Without controlling for the potential effect of other variables, an investigator cannot really judge whether an observed association between a given exposure and a specific disease is spurious.
If the effect of an exposure on the risk of a disease is not homogeneous in strata formed by a third variable, the third variable is considered an effect modifier, and the situation is called interaction or effect modification. Put in other words, interaction exists when the stratum-specific effect measures are different from each other. In the lung cancer example given previously, if the odds ratio for alcohol drinking is 1 in smokers but 3 in nonsmokers, then there is interaction and smoking is an effect modifier. The evaluation of interaction is essentially a stratified analysis, which is similar to the evaluation of confounding. Confounding and interaction can be both present in a given study. However, when interaction occurs, the stratum-specific effect measures should be reported. It is no longer appropriate to report a summary measure in the presence of interaction. Unlike confounding, which is a nuisance that an investigator hopes to remove, interaction is a more detailed description of the true relationship between an exposure and a disease.
CANCER OUTCOMES RESEARCH
The discussion of epidemiologic methods in this section focuses primarily on etiological research, which aims at identifying the risk factors of cancer. However, similar principles and methods are applicable to cancer outcomes research, which aims at studying a variety of factors related to the early identification, treatment, prognosis, health related quality of life, and cost of care. Cancer outcomes research can be experimental or observational in nature. For example, randomized clinical trials have been conducted to assess the impact of screening on prostate cancer mortality18 and to compare the effect of radical prostatectomy versus observation in patients with localized prostate cancer.19 Observational studies of cancer outcomes, especially those that build upon preexisting resources,20,21 can be carried out in a large group of patients with relatively little cost to capture the patterns and cost of care and to address many other research questions that have important clinical implications. Although the findings of such observation studies are subject to bias and confounding inherent in an observational design, these studies are complementary to experimental studies and have their unique value. Given an increasing interest in improving the effectiveness and value of cancer care, more cancer outcomes research is to be expected in the future.
MOLECULAR EPIDEMIOLOGY
Molecular epidemiology involves multidisciplinary and transdisciplinary research that entails not only traditional epidemiology and biostatistics, but also genetics, molecular biology, biochemistry, cellular biology, analytical chemistry, toxicology, pharmacology, and laboratory medicine. Unlike traditional epidemiology research of cancer, which focuses on exposures or risk factors ascertained through questionnaire-based interviews or surveys, molecular epidemiology studies expand the assessment of exposure to a much broader scope that includes an analysis of biomarkers underlying internal exposure of exogenous and endogenous carcinogenic agents or risk factors, molecular alterations in response to exposure, and genetic susceptibility to cancer. The biomarkers often measured in molecular epidemiology research include DNA, RNA, proteins, chromosomes, compound molecules (e.g., DNA and protein adducts), and various metabolites as well as other endogenous and exogenous substances (e.g., steroids, nutrients, chemical or biologic toxins, and phytochemicals). Molecular markers can reflect different aspects of the tumorigenic process, which include biomarkers of internal exposure, biomarkers of molecular or cellular changes in response to exposure, and biomarkers of precursor lesions or early diseases.22,23 Depending on the source of molecules and location of diseases, surrogates are often used in epidemiologic studies. When using a surrogate marker or tissue, the relevance of a proxy to its underlying target needs to be established or justified.23 This justification is especially important when conducting population-based epidemiologic studies that focus on organ-specific cancers, because assessing biomarkers in target tissue is difficult for controls; molecular markers from blood samples are often used as substitutes. If a biomarker in the blood does not travel to or act on the tissue or organ of interest, an association between the circulating marker and the cancer may not be relevant. Thus, establishing a close link between a surrogate and its target is crucial in molecular epidemiology research.
Gene-environment interaction plays an essential role in cancer development.24 Common genetic variations are considered an important determinant of host susceptibility and are a major focus of molecular epidemiology research. Depending on the biologic mechanism involved, genetic variations can influence every aspect of the carcinogenic process, ranging from external and internal exposure to carcinogens or risk factors to molecular and cellular damage, alteration, and response.22,23 Currently, single nucleotide polymorphisms (SNPs) are the most studied genetic variations. It is believed that even if SNPs confer a small risk, they may still be important at the population level because these variations are common in the general population. It is also important that the impacts of SNPs on cancer are considered under the context of gene–gene and gene–environment interactions. As genotyping technology has advanced substantially with respect to its analytic quality, capacity, and cost, research of genetic polymorphisms has evolved rapidly from investigations of a single SNP to studies of haplotypes and tag SNPs, and from a pathway-based candidate gene approach to genome-wide association studies (GWAS).25 A GWAS analyzes hundreds of thousands of SNPs simultaneously for hundreds or even thousands of study subjects. When these data are further combined with questionnaire information such as environmental exposures, lifestyle factors, dietary habits, and medical history, enormous information is generated, which requires a huge sample size to allow for a reliable and complete assessment of these variables individually and jointly. A single epidemiologic study can no longer provide sufficient power for this type of investigation. Multicenter investigations or study consortia that pool study information and specimens together are developed to address the sample size issue.26 False-positive findings resulting from multiple comparisons constitute a major challenge in epidemiologic studies of genetic associations with cancer.27 A meta-analysis or pooled analysis can be used to address this problem if sufficient studies are already published and available for evaluation. To address this issue at the time of study design, one may adopt a two- or multiphase study design in which study subjects are divided into two or multiple groups for genotyping and data analysis. Selected or genomewide SNPs are first screened in one group of the study subjects (discovery phase), and then the significant findings determined by stringent statistical criteria (usually p values less than 1 × 10−5 or 1 × 10−7) are reanalyzed in one or several other groups of subjects for verification (validation phase). This study design also lowers the cost of genotyping. False-positive findings can also be addressed with various statistical methods, such as bootstrap, permutation test, estimate of false positive report probability, prediction of false discovery rate, and the use of a much more stringent p value to accommodate multiple comparisons. For epidemiologic studies that are not population based or not conducted strictly following epidemiology principles, population stratification is a potential source of bias that may distort genetic associations.28
A large number of GWAS have been completed in search for SNPs that influence host susceptibility to cancer. Considering that more than 5 million SNPs are present in the human genome, the numbers of SNPs that are found to be associated with cancer risk after rigorous validation are much fewer than what one would have anticipated. In addition, the risk associations detected are quite weak, with most of the odds ratios ranging from 1.1 to 1.5, and the functional relevance or biologic implications are unclear for most of the SNPs. Furthermore, not many SNPs associated with cancer risk are located in protein-coding regions, and even fewer are in the loci of candidate genes suspected to be involved in tumorigenesis, such as oncogenes, tumor suppressor genes, DNA repair genes, and xenobiotic metabolizing or detoxification genes. Genes where SNPs are found to be linked to cancer by GWAS include FGFR2, MAP3K1, MRPS30, LSP1, TNRC9, TOX3, STXBP1, and RAD51L1 for breast cancer29–31; JAZF1, HNF1B, MSMB, CTBP2, and KLK2/KLK3 for prostate cancer29,32; SMAD7, CRAC1, EIF3H, BMP4, CDH1, and RHPN2 for colorectal cancer29,33; CHRNA3 and CHRNA5 for lung cancer34,35; ABO for pancreatic cancer36; TACC3 and PSCA for bladder cancer37,38; and KRT5 for basal cell carcinoma.39 Among these genes identified by GWAS, two findings are considered especially interesting. One is the association of lung cancer with CHRNA3 and CHRNA5, which encode neuronal nicotinic acetylcholine receptor subunits. Different genotypes of these receptor subunits appear to influence individual’s addiction to tobacco, which further leads to different smoking exposure and lung cancer risk.40,41 Another is the link of the ABO gene to pancreatic cancer. The association between pancreatic cancer risk and ABO blood type was observed 50 years ago. The GWAS finding not only confirms the relationship, but also provides new clues for understanding the underlying biologic mechanism.
Besides intragenic SNPs, GWAS also found many intergenic SNPs in association to cancer risk, which include those in the regions of 8q24, 5p15, 1p11, 1p36, 1q42, 2p15, 2q35, 3p12, 3p24, 3q28, 6p21, 6q25, 7q21, 7q32, 9p21, 9p22, 9p24, 9q22, 10p14, 11q13, 11q23, 14q13, 18q23, and 20p12.29–31,33,39,42–48 Of these loci, SNPs in 8q24 are associated with several cancer sites, including prostate, breast, colon, and bladder.29,31–33,47–49 Further analysis of 8q24 indicates that there are nine SNPs in five regions and each region is independently related to different types of cancer, with SNPs in regions 1, 4, and 5 associated exclusively with prostate cancer, a SNP in region 2 related to breast cancer, and SNPs in region 3 linked to prostate, colon, and ovarian cancers.50 No known genes are located within the region of 8q24, but an oncogene c-MYC resides about 330 kb downstream of the region.51 An initial investigation found no evidence of the SNPs’ influence on c-MYC expression,47 but a later study suggests that the SNPs in 8q24 may be distal enhancers of c-MYC, interacting with its promoter through a chromatin loop.52 Another genomic region that is associated with the risk of multiple cancer sites is 5p15, a region involving telomerase reverse transcriptase (TERT) cleft lip and palate transmembrane protein 1–like protein (CLPTM1L). Five types of cancer are found to be linked to this region, including basal cell carcinoma, lung, bladder, prostate, and cervical cancers.53 TERT extends the length of telomere and is associated with cell proliferation and abnormal telomere maintenance.54 The risk alleles of TERT are associated with shorter telomere length among the elderly and with higher DNA adduct in the lungs.53,55
GWAS has demonstrated its value in identifying disease-related SNPs in unknown regions of the genome, which provides new clues for investigators to interrogate and understand different regions of the human genome, especially in the gene-desert areas. Despite the strength, the low yield of significant findings from the GWAS has raised concerns in several areas, including the SNP coverage in the genome (rare SNPs and SNP representativeness in unknown regions), associations with low statistical significance (p value between 0.01 and 1 × 10−5, the GWAS cutoff), other forms of genetic variations (copy number variation and other structural variations), cancer subtypes, and genetic interplay with environmental factors (gene–environment interaction).56,57 To address these issues, investigators propose to perform fine-mapping and resequencing to examine genetic regions more specifically and meticulously. Epidemiologists suggest that detailed environmental exposure and lifestyle factors should be included in the next wave of GWAS. Furthermore, to make the study more reliable and compelling, DNA specimens, instead of convenient samples, should come from well-designed and well-executed epidemiologic studies that pay close attention to the selection of study subjects and the measurement of environmental and lifestyle factors to eliminate or minimize selection bias and measurement errors.
As described earlier, analytical epidemiology has two major study designs: the case-control study and the cohort study. It is important that investigators choose an appropriate study design to investigate molecular markers in epidemiologic studies. Two types of molecular markers, genotypic and phenotypic markers, can be considered. Genotypic markers refer to nucleotide sequences of genomic DNA, and all other molecules are considered phenotypic markers, including most of the chemical modifications on DNA, such as cytosine methylation. The distinction between the two is a marker’s status in relation to an outcome variable, usually a disease. Genotypic markers generally do not change over time and are not affected by the development of a disease, whereas phenotypic markers are likely to change over time or be influenced by the presence of a disease, either itself or the treatment associated with it. If measurements of a phenotypic marker are made from the specimens that are collected after or at the time of cancer diagnosis, investigators will have difficulties determining the status of the phenotypic marker before the cancer was diagnosed. A disease condition, however, does not affect genotypic markers such as SNPs; therefore, a temporal relationship can be easily established even if the samples are collected after the disease is diagnosed. Based on this distinction, one can evaluate genotypic markers either in case-control or cohort studies, but a case-control study would be the design of choice because of efficiency and cost-effectiveness. A prospective cohort study design is ideal for phenotypic markers. Investigators, however, may use other study designs if they can demonstrate that the disease status does not influence the phenotypic markers of interest. To reduce study cost, investigators usually use nested case-control or case-cohort designs to avoid analyzing specimens from the entire cohort. The main purpose in choosing a cohort study design for a molecular epidemiology investigation is to ensure that biospecimens are collected before the development of a disease so that a temporal relationship between a marker and disease development can be established.
The differences between molecular epidemiology and genetic epidemiology are the scope of the molecular analysis and the emphasis on heredity. Sometimes molecular and genetic epidemiology both investigate genetic factors in association with cancer risk, but each has its own emphasis. The former assesses genetic involvement, but not necessarily inheritance, whereas the latter focuses mainly on heredity. Because of the difference in focus, study populations are different between the two types of investigation. Molecular epidemiology studies unrelated individuals, whereas genetic epidemiology investigates family members in the format of pedigrees, parent–child trios, or sibling pairs. Given the different research focus between genetic and molecular epidemiology, these investigations evaluate different genetic markers. Genetic epidemiology research is designed to identify genetic markers with high penetrance (strong association with an underlying disease) but low prevalence in the general populations, whereas a molecular epidemiology investigation targets low penetrance markers that are commonly present in the general population. Given the difference in study design, the analysis of genetic marker’s link to cancer is also different between the studies. Relative risks or odds ratios are calculated in molecular epidemiology studies because study participants are unrelated individuals, whereas linkage analysis is used in genetic epidemiology because individuals in the study are genetically related family members. Recently, both genetic and molecular epidemiology study designs have been considered in GWAS to improve study validity and to minimize false positive findings. Another difference between genetic and molecular epidemiology research is that molecular epidemiology also studies nongenetic molecules. Thus, the scope of molecular analysis is much broader in molecular epidemiology research than in genetic epidemiology studies.
A laboratory analysis of molecular markers is another integral part of molecular epidemiology research, which has unique features that are different from basic science research. Collecting biologic specimens is difficult and expensive in population-based epidemiologic studies. It not only increases the study cost, but also imposes constraints to multiple areas of epidemiology research. Specimen collection may adversely influence the response rate of study participants, potentially compromising study validity. For organ-specific cancer research, investigating molecular markers in target tissue is difficult. Blood is the most common and versatile specimen used in molecular epidemiology research; other specimens used include urine, stool, nail, hair, sputum, buccal cells, and saliva. Tissue samples, either fresh frozen or chemically fixed, are also used, but the availability of these samples is highly limited to patients or selected subgroups of a general study population. Comparability and generalizability are always problems in epidemiologic studies involving tissue specimens, except for those investigations that focus on cancer prognosis or treatment in which only cancer patients are involved. Attempts have been made to use special body fluids for epidemiologic research, such as nipple aspirate and breast or pulmonary lavage, but the difficulty in specimen collection and preparation makes these samples impractical in large population-based studies.
Given the research value of biologic specimens and the difficulty in collecting them for population-based studies, technical issues related to specimen collection, processing, and storage become especially important in molecular epidemiology research. These include time and conditions for specimen transportation and processing, a sample aliquot and labeling system, a sample special treatment for storage and analysis, a sample storage and tracking system, as well as backup plans and equipment for unexpected adverse events during long-term storage (e.g., power failure, earthquake, flooding). Laboratory methods used to analyze biomarkers are also important in molecular epidemiology. Because large numbers of specimens are involved, laboratory methods are required to be robust, reproducible, high throughput, low cost, and easy to use. These requirements are often met in the analysis of nucleotide sequences that serve as genotypic markers. However, for phenotypic markers, many methods do not readily meet these requirements. Moreover, many phenotypic markers, such as proteins, require both qualitative and quantitative assessments. An ideal laboratory method should be quantitative (able to measure a wide range of values), sensitive (able to detect a small amount of analyte), specific (able to detect only the molecule of interest, no other molecules), reproducible (high precision and low variation), and versatile (easy to use). In addition, investigators need to implement appropriate quality assurance procedures during sample processing and testing as well as include appropriate quality control samples in specimen analysis.
Host–environment interaction is believed to play a key role in the etiology of most types of cancer. Genetic factors, including mutations and polymorphisms, are initially considered important host factors, but recent developments in cancer research has indicated that epigenetic factors may also play a critical role in cancer as a host factor involved in host–environment interaction. Epigenetic factors, which regulate the function of human genome without altering the physical sequences of nucleotides, include pretranscription regulation through nucleotide modification (e.g., cytosine methylation at CpG sites), chromosome modification (e.g., histone acetylation), and posttranscription regulation by noncoding small RNA (e.g., microRNAs). These epigenetic factors have two unique features that have captured the attention of cancer researchers, especially cancer epidemiologists who are interested in the gene–environment interaction. It is known that epigenetic factors are heritable, but these inherited features are readily modifiable by environmental and lifestyle factors. Monozygotic twins have an identical genome as well as epigenome at birth, but the latter undergoes substantial changes over time, resulting in distinct epigenetic profiles that depend heavily on their environmental exposures.58 Animal studies also indicated that the maternal intake of dietary nutrients involving one-carbon metabolism could influence offsprings’ growth phenotypes, which are regulated by DNA methylation.59 As evidence mounts on epigenetic involvement in cancer, molecular epidemiologists will start to look for clues in human populations that can link epigenetic factors to both lifestyle factors and cancer risk. Given that epigenetic regulation is tissue specific and time dependent, investigators face challenges in accurately assessing these phenotypic markers in etiologic studies. However, progress in the analysis of circulating methylation markers and microRNAs may provide an alternative to study epigenetic regulation in human cancer. Furthermore, methods for a genomewide analysis of DNA methylation have been developed and applied in epidemiologic studies, which can substantially accelerate the search for cancer-related DNA methylation. Together with the high-throughput, high-dimensional analysis of DNA methylation, two other evolving fields that will have significant impacts on molecular epidemiology of cancer research are metagenomics and metabolomics. The former focuses on environmental genomics of the microbiome that resides in our body and influences one’s biologic functions and health status. The latter refers to the analysis of hundreds or thousands of metabolites in a biologic specimen, including tissue, blood, urine, body fluids, and fecal samples. These new analyses will add tremendous value to epidemiologic studies.
REFERENCES
1. Last J. A Dictionary of Epidemiology. 3rd ed. New York: Oxford University Press; 1995.
2. Doll R, Hill AB. Lung cancer and other causes of death in relation to smoking: a second report on the mortality of British doctors. Br Med J 1956;12:1071–1081.
3. Kleinbaum D, Kupper L, Morgenstern H. Epidemiologic Research. New York: Van Nostrand Reinhold; 1982.
4. Kaplan LD, Straus DJ, Testa MA, et al. Low-dose compared with standard-dose m-BACOD chemotherapy for non-Hodgkin’s lymphoma associated with human immunodeficiency virus infection. National Institute of Allergy and Infectious Diseases AIDS Clinical Trials Group. N Engl J Med 1997;336:1641–1648.
5. Dunn BK, Kramer BS, Ford LG. Phase III, large-scale chemoprevention trials. Approach to chemoprevention clinical trials and phase III clinical trial of tamoxifen as a chemopreventive for breast cancer—the US National Cancer Institute experience. Hematol Oncol Clin North Am 1998;12:1019–1036, vii.
6. Grant WB. An ecologic study of dietary and solar ultraviolet-B links to breast carcinoma mortality rates. Cancer 2002;94:272–281.
7. National Center for Health Statistics. Third National Health and Nutrition Examination Survey, 1988-1994, Plan and Operations Procedures Manuals (CD-ROM). Hyattsville, MD: U.S. Department of Health and Human Services (DHHS), Centers for Disease Control and Prevention; 1996.
8. Szklo M, Nieto F. Epidemiology: Beyond the Basics. Gaithersburg, MD: Aspen Publishers; 2000.
9. Grimes DA, Schulz KF. Cohort studies: marching towards outcomes. Lancet 2002;359:341–345.
10. Schatzkin A, Subar AF, Thompson FE, et al. Design and serendipity in establishing a large cohort with wide dietary intake distributions: the National Institutes of Health-American Association of Retired Persons Diet and Health Study. Am J Epidemiol 2001;154:1119–1125.
11. Mantel N, Haenszel W. Statistical aspects of the analysis of data from retrospective studies of disease. J Natl Cancer Inst 1959;22:719–748.
12. Wacholder S, Silverman DT, McLaughlin JK, et al. Selection of controls in case-control studies. II. Types of controls. Am J Epidemiol 1992;135:1029–1041.
13. Chen C, Weiss NS, Stanczyk FZ, et al. Endogenous sex hormones and prostate cancer risk: a case-control study nested within the Carotene and Retinol Efficacy Trial. Cancer Epidemiol Biomarkers Prev 2003;12:1410–1416.
14. Pearce N. What does the odds ratio estimate in a case-control study? Int J Epidemiol 1993;22:1189–1192.
15. Schulz KF, Grimes DA. Case-control studies: research in reverse. Lancet 2002;359:431–434.
16. Rothman K. Epidemiology: An Introduction. New York: Oxford University Press; 2002.
17. Hennekens C, Buring J. Epidemiology in Medicine. Boston: Little, Brown and Company; 1987.
18. Andriole GL, Crawford ED, Grubb RL 3rd, et al. Mortality results from a randomized prostate-cancer screening trial. N Engl J Med 2009;360:1310–1319.
19. Wilt TJ, Brawer MK, Jones KM, et al. Radical prostatectomy versus observation for localized prostate cancer. N Engl J Med 2012;367:203–213.
20. Yu JB, Soulos PR, Herrin J, et al. Proton versus intensity-modulated radiotherapy for prostate cancer: patterns of care and early toxicity. J Natl Cancer Inst 2013;105:25–32.
21. Ma X, Wang R, Long JB, et al. The cost implications of prostate cancer screening in the Medicare population. Cancer 2014;120(1):96–102.
22. Rundle A, Schwartz S. Issues in the epidemiological analysis and interpretation of intermediate biomarkers. Cancer Epidemiol Biomarkers Prev 2003;12:491–496.
23. Shields PG. Tobacco smoking, harm reduction, and biomarkers. J Natl Cancer Inst 2002;94:1435–1444.
24. Hunter DJ. Gene-environment interactions in human diseases. Nat Rev Genet 2005;6:287–298.
25. Hirschhorn JN, Daly MJ. Genome-wide association studies for common diseases and complex traits. Nat Rev Genet 2005;6:95–108.
26. Breast Cancer Association Consortum. Commonly studied single-nucleotide polymorphisms and breast cancer: results from the Breast Cancer Association Consortium. J Natl Cancer Inst 2006;98:1382–1396.
27. Wacholder S, Chanock S, Garcia-Closas M, et al. Assessing the probability that a positive report is false: an approach for molecular epidemiology studies. J Natl Cancer Inst 2004;96:434–442.
28. Clayton DG, Walker NM, Smyth DJ, et al. Population structure, differential bias and genomic control in a large-scale, case-control association study. Nat Genet 2005;37:1243–1246.
29. Easton DF, Eeles RA. Genome-wide association studies in cancer. Hum Mol Genet 2008;17:R109–115.
30. Ahmed S, Thomas G, Ghoussaini M, et al. Newly discovered breast cancer susceptibility loci on 3p24 and 17q23.2. Nat Genet 2009;41:585–590.
31. Thomas G, Jacobs KB, Kraft P, et al. A multistage genome-wide association study in breast cancer identifies two new risk alleles at 1p11.2 and 14q24.1 (RAD51L1). Nat Genet 2009;41:579–584.
32. Thomas G, Jacobs KB, Yeager M, et al. Multiple loci identified in a genome-wide association study of prostate cancer. Nat Genet 2008;40:310–315.
33. Le Marchand L. Genome-wide association studies and colorectal cancer. Surg Oncol Clin N Am 2009;18:663–668.
34. Hung RJ, McKay JD, Gaborieau V, et al. A susceptibility locus for lung cancer maps to nicotinic acetylcholine receptor subunit genes on 15q25. Nature 2008;452:633–637.
35. Amos CI, Wu X, Broderick P, et al. Genome-wide association scan of tag SNPs identifies a susceptibility locus for lung cancer at 15q25.1. Nat Genet 2008;40:616–622.
36. Amundadottir L, Kraft P, Stolzenberg-Solomon RZ, et al. Genome-wide association study identifies variants in the ABO locus associated with susceptibility to pancreatic cancer. Nat Genet 2009;41:986–990.
37. Kiemeney LA, Sulem P, Besenbacher S, et al. A sequence variant at 4p16.3 confers susceptibility to urinary bladder cancer. Nat Genet 2010;42(5):415–419.
38. Wu X, Ye Y, Kiemeney LA, et al. Genetic variation in the prostate stem cell antigen gene PSCA confers susceptibility to urinary bladder cancer. Nat Genet 2009;41:991–995.
39. Stacey SN, Sulem P, Masson G, et al. New common variants affecting susceptibility to basal cell carcinoma. Nat Genet 2009;41:909–914.
40. Thorgeirsson TE, Geller F, Sulem P, et al. A variant associated with nicotine dependence, lung cancer and peripheral arterial disease. Nature 2008;452:638–642.
41. Spitz MR, Amos CI, Dong Q, et al. The CHRNA5-A3 region on chromosome 15q24-25.1 is a risk factor both for nicotine dependence and for lung cancer. J Natl Cancer Inst 2008;100:1552–1556.
42. Zheng W, Long J, Gao YT, et al. Genome-wide association study identifies a new breast cancer susceptibility locus at 6q25.1. Nat Genet 2009;41:324–328.
43. Gudmundsson J, Sulem P, Gudbjartsson DF, et al. Common variants on 9q22.33 and 14q13.3 predispose to thyroid cancer in European populations. Nat Genet 2009;41:460–464.
44. Gudmundsson J, Sulem P, Gudbjartsson DF, et al. Genome-wide association and replication studies identify four variants associated with prostate cancer susceptibility. Nat Genet 2009;41:1122–1126.
45. Song H, Ramus SJ, Tyrer J, et al. A genome-wide association study identifies a new ovarian cancer susceptibility locus on 9p22.2. Nat Genet 2009;41:996–1000.
46. Stacey SN, Gudbjartsson DF, Sulem P, et al. Common variants on 1p36 and 1q42 are associated with cutaneous basal cell carcinoma but not with melanoma or pigmentation traits. Nat Genet 2008;40:1313–1318.
47. Zanke BW, Greenwood CM, Rangrej J, et al. Genome-wide association scan identifies a colorectal cancer susceptibility locus on chromosome 8q24. Nat Genet 2007;39:989–994.
48. Haiman CA, Patterson N, Freedman ML, et al. Multiple regions within 8q24 independently affect risk for prostate cancer. Nat Genet 2007;39:638–644.
49. Kiemeney LA, Thorlacius S, Sulem P, et al. Sequence variant on 8q24 confers susceptibility to urinary bladder cancer. Nat Genet 2008;40:1307–1312.
50. Ghoussaini M, Song H, Koessler T, et al. Multiple loci with different cancer specificities within the 8q24 gene desert. J Natl Cancer Inst 2008;100:962–966.
51. Harismendy O, Frazer KA. Elucidating the role of 8q24 in colorectal cancer. Nat Genet 2009;41:868–869.
52. Wright JB, Brown SJ, Cole MD. Upregulation of c-MYC in cis through a large chromatin loop linked to a cancer risk-associated single-nucleotide polymorphism in colorectal cancer cells. Mol Cell Biol 2010;30:1411–1420.
53. Rafnar T, Sulem P, Stacey SN, et al. Sequence variants at the TERT-CLPTM1L locus associate with many cancer types. Nat Genet 2009;41:221–227.
54. Fernandez-Garcia I, Ortiz-de-Solorzano C, Montuenga LM. Telomeres and telomerase in lung cancer. J Thorac Oncol 2008;3:1085–1088.
55. Zienolddiny S, Skaug V, Landvik NE, et al. The TERT-CLPTM1L lung cancer susceptibility variant associates with higher DNA adduct formation in the lung. Carcinogenesis 2009;30:1368–1371.
56. Ioannidis JP, Thomas G, Daly MJ. Validating, augmenting and refining genome-wide association signals. Nat Rev Genet 2009;10:318–329.
57. Chung CC, Magalhaes WC, Gonzalez-Bosquet J, et al. Genome-wide association studies in cancer—current and future directions. Carcinogenesis 2010;31:111–120.
58. Fraga MF, Ballestar E, Paz MF, et al. Epigenetic differences arise during the lifetime of monozygotic twins. Proc Natl Acad Sci U S A 2005;102:10604–10609.
59. Dolinoy DC, Weidman JR, Waterland RA, et al. Maternal genistein alters coat color and protects Avy mouse offspring from obesity by modifying the fetal epigenome. Environ Health Perspect 2006;114:567–572.