Understanding the organization, variation, and transmission of the human genome is central to appreciating the role of genetics in medicine, as well as the emerging principles of genomic and personalized medicine. With the availability of the sequence of the human genome and a growing awareness of the role of genome variation in disease, it is now possible to begin to exploit the impact of that variation on human health on a broad scale. The comparison of individual genomes underscores the first major take-home lesson of this book—every individual has his or her own unique constitution of gene products, produced in response to the combined inputs of the genome sequence and one's particular set of environmental exposures and experiences. As pointed out in the previous chapter, this realization reflects what Garrod termed chemical individuality over a century ago and provides a conceptual foundation for the practice of genomic and personalized medicine.
Advances in genome technology and the resulting explosion in knowledge and information stemming from the Human Genome Project are thus playing an increasingly transformational role in integrating and applying concepts and discoveries in genetics to the practice of medicine.
Chromosome and Genome Analysis in Clinical Medicine
Chromosome and genome analysis has become an important diagnostic procedure in clinical medicine. As described more fully in subsequent chapters, these applications include the following:
• Clinical diagnosis. Numerous medical conditions, including some that are common, are associated with changes in chromosome number or structure and require chromosome or genome analysis for diagnosis and genetic counseling (see Chapters 5 and 6).
• Gene identification. A major goal of medical genetics and genomics today is the identification of specific genes and elucidating their roles in health and disease. This topic is referred to repeatedly but is discussed in detail in Chapter 10.
• Cancer genomics. Genomic and chromosomal changes in somatic cells are involved in the initiation and progression of many types of cancer (see Chapter 15).
• Disease treatment. Evaluation of the integrity, composition, and differentiation state of the genome is critical for the development of patient-specific pluripotent stem cells for therapeutic use (see Chapter 13).
• Prenatal diagnosis. Chromosome and genome analysis is an essential procedure in prenatal diagnosis (see Chapter 17).
The Human Genome and the Chromosomal Basis of Heredity
Appreciation of the importance of genetics to medicine requires an understanding of the nature of the hereditary material, how it is packaged into the human genome, and how it is transmitted from cell to cell during cell division and from generation to generation during reproduction. The human genome consists of large amounts of the chemical deoxyribonucleic acid (DNA) that contains within its structure the genetic information needed to specify all aspects of embryogenesis, development, growth, metabolism, and reproduction—essentially all aspects of what makes a human being a functional organism. Every nucleated cell in the body carries its own copy of the human genome, which contains, depending on how one defines the term, approximately 20,000 to 50,000 genes (see Box later). Genes, which at this point we consider simply and most broadly as functional units of genetic information, are encoded in the DNA of the genome, organized into a number of rod-shaped organelles called chromosomes in the nucleus of each cell. The influence of genes and genetics on states of health and disease is profound, and its roots are found in the information encoded in the DNA that makes up the human genome.
Each species has a characteristic chromosome complement (karyotype) in terms of the number, morphology, and content of the chromosomes that make up its genome. The genes are in linear order along the chromosomes, each gene having a precise position or locus. A gene map is the map of the genomic location of the genes and is characteristic of each species and the individuals within a species.
The study of chromosomes, their structure, and their inheritance is called cytogenetics. The science of human cytogenetics dates from 1956, when it was first established that the normal human chromosome number is 46. Since that time, much has been learned about human chromosomes, their normal structure and composition, and the identity of the genes that they contain, as well as their numerous and varied abnormalities.
With the exception of cells that develop into gametes (the germline), all cells that contribute to one's body are called somatic cells (soma, body). The genome contained in the nucleus of human somatic cells consists of 46 chromosomes, made up of 24 different types and arranged in 23 pairs (Fig. 2-1). Of those 23 pairs, 22 are alike in males and females and are called autosomes, originally numbered in order of their apparent size from the largest to the smallest. The remaining pair comprises the two different types of sex chromosomes: an X and a Y chromosome in males and two X chromosomes in females. Central to the concept of the human genome, each chromosome carries a different subset of genes that are arranged linearly along its DNA. Members of a pair of chromosomes (referred to as homologous chromosomes or homologues) carry matching genetic information; that is, they typically have the same genes in the same order. At any specific locus, however, the homologues either may be identical or may vary slightly in sequence; these different forms of a gene are called alleles. One member of each pair of chromosomes is inherited from the father, the other from the mother. Normally, the members of a pair of autosomes are microscopically indistinguishable from each other. In females, the sex chromosomes, the two X chromosomes, are likewise largely indistinguishable. In males, however, the sex chromosomes differ. One is an X, identical to the Xs of the female, inherited by a male from his mother and transmitted to his daughters; the other, the Y chromosome, is inherited from his father and transmitted to his sons. In Chapter 6, as we explore the chromosomal and genomic basis of disease, we will look at some exceptions to the simple and almost universal rule that human females are XX and human males are XY.

FIGURE 2-1 The human genome, encoded on both nuclear and mitochondrial chromosomes. SeeSources & Acknowledgments.
In addition to the nuclear genome, a small but important part of the human genome resides in mitochondria in the cytoplasm (see Fig. 2-1). The mitochondrial chromosome, to be described later in this chapter, has a number of unusual features that distinguish it from the rest of the human genome.
Genes in the Human Genome
What is a gene? And how many genes do we have? These questions are more difficult to answer than it might seem.
The word gene, first introduced in 1908, has been used in many different contexts since the essential features of heritable “unit characters” were first outlined by Mendel over 150 years ago. To physicians (and indeed to Mendel and other early geneticists), a gene can be defined by its observable impact on an organism and on its statistically determined transmission from generation to generation. To medical geneticists, a gene is recognized clinically in the context of an observable variant that leads to a characteristic clinical disorder, and today we recognize approximately 5000 such conditions (see Chapter 7).
The Human Genome Project provided a more systematic basis for delineating human genes, relying on DNA sequence analysis rather than clinical acumen and family studies alone; indeed, this was one of the most compelling rationales for initiating the project in the late 1980s. However, even with the finished sequence product in 2003, it was apparent that our ability to recognize features of the sequence that point to the existence or identity of a gene was sorely lacking. Interpreting the human genome sequence and relating its variation to human biology in both health and disease is thus an ongoing challenge for biomedical research.
Although the ultimate catalogue of human genes remains an elusive target, we recognize two general types of gene, those whose product is a protein and those whose product is a functional RNA.
• The number of protein-coding genes—recognized by features in the genome that will be discussed in Chapter 3—is estimated to be somewhere between 20,000 and 25,000. In this book, we typically use approximately 20,000 as the number, and the reader should recognize that this is both imprecise and perhaps an underestimate.
• In addition, however, it has been clear for several decades that the ultimate product of some genes is not a protein at all but rather an RNA transcribed from the DNA sequence. There are many different types of such RNA genes (typically called noncoding genes to distinguish them from protein-coding genes), and it is currently estimated that there are at least another 20,000 to 25,000 noncoding RNA genes around the human genome.
Thus overall—and depending on what one means by the term—the total number of genes in the human genome is of the order of approximately 20,000 to 50,000. However, the reader will appreciate that this remains a moving target, subject to evolving definitions, increases in technological capabilities and analytical precision, advances in informatics and digital medicine, and more complete genome annotation.
DNA Structure: A Brief Review
Before the organization of the human genome and its chromosomes is considered in detail, it is necessary to review the nature of the DNA that makes up the genome. DNA is a polymeric nucleic acid macromolecule composed of three types of units: a five-carbon sugar, deoxyribose; a nitrogen-containing base; and a phosphate group (Fig. 2-2). The bases are of two types, purines and pyrimidines. In DNA, there are two purine bases, adenine (A) and guanine (G), and two pyrimidine bases, thymine (T) and cytosine (C). Nucleotides, each composed of a base, a phosphate, and a sugar moiety, polymerize into long polynucleotide chains held together by 5′-3′ phosphodiester bonds formed between adjacent deoxyribose units (Fig. 2-3A). In the human genome, these polynucleotide chains exist in the form of a double helix (Fig. 2-3B) that can be hundreds of millions of nucleotides long in the case of the largest human chromosomes.

FIGURE 2-2 The four bases of DNA and the general structure of a nucleotide in DNA. Each of the four bases bonds with deoxyribose (through the nitrogen shown in magenta) and a phosphate group to form the corresponding nucleotides.

FIGURE 2-3 The structure of DNA. A, A portion of a DNA polynucleotide chain, showing the 3′-5′ phosphodiester bonds that link adjacent nucleotides. B, The double-helix model of DNA, as proposed by Watson and Crick. The horizontal “rungs” represent the paired bases. The helix is said to be right-handed because the strand going from lower left to upper right crosses over the opposite strand. The detailed portion of the figure illustrates the two complementary strands of DNA, showing the AT and GC base pairs. Note that the orientation of the two strands is antiparallel. SeeSources & Acknowledgments.
The anatomical structure of DNA carries the chemical information that allows the exact transmission of genetic information from one cell to its daughter cells and from one generation to the next. At the same time, the primary structure of DNA specifies the amino acid sequences of the polypeptide chains of proteins, as described in the next chapter. DNA has elegant features that give it these properties. The native state of DNA, as elucidated by James Watson and Francis Crick in 1953, is a double helix (see Fig. 2-3B). The helical structure resembles a right-handed spiral staircase in which its two polynucleotide chains run in opposite directions, held together by hydrogen bonds between pairs of bases: T of one chain paired with A of the other, and G with C. The specific nature of the genetic information encoded in the human genome lies in the sequence of C's, A's, G's, and T's on the two strands of the double helix along each of the chromosomes, both in the nucleus and in mitochondria (see Fig. 2-1). Because of the complementary nature of the two strands of DNA, knowledge of the sequence of nucleotide bases on one strand automatically allows one to determine the sequence of bases on the other strand. The double-stranded structure of DNA molecules allows them to replicate precisely by separation of the two strands, followed by synthesis of two new complementary strands, in accordance with the sequence of the original template strands (Fig. 2-4). Similarly, when necessary, the base complementarity allows efficient and correct repair of damaged DNA molecules.

FIGURE 2-4 Replication of a DNA double helix, resulting in two identical daughter molecules, each composed of one parental strand and one newly synthesized strand.
Structure of Human Chromosomes
The composition of genes in the human genome, as well as the determinants of their expression, is specified in the DNA of the 46 human chromosomes in the nucleus plus the mitochondrial chromosome. Each human chromosome consists of a single, continuous DNA double helix; that is, each chromosome is one long, double-stranded DNA molecule, and the nuclear genome consists, therefore, of 46 linear DNA molecules, totaling more than 6 billion nucleotide pairs (see Fig. 2-1).
Chromosomes are not naked DNA double helices, however. Within each cell, the genome is packaged as chromatin, in which genomic DNA is complexed with several classes of specialized proteins. Except during cell division, chromatin is distributed throughout the nucleus and is relatively homogeneous in appearance under the microscope. When a cell divides, however, its genome condenses to appear as microscopically visible chromosomes. Chromosomes are thus visible as discrete structures only in dividing cells, although they retain their integrity between cell divisions.
The DNA molecule of a chromosome exists in chromatin as a complex with a family of basic chromosomal proteins called histones. This fundamental unit interacts with a heterogeneous group of nonhistone proteins, which are involved in establishing a proper spatial and functional environment to ensure normal chromosome behavior and appropriate gene expression.
Five major types of histones play a critical role in the proper packaging of chromatin. Two copies each of the four core histones H2A, H2B, H3, and H4 constitute an octamer, around which a segment of DNA double helix winds, like thread around a spool (Fig. 2-5). Approximately 140 base pairs (bp) of DNA are associated with each histone core, making just under two turns around the octamer. After a short (20- to 60-bp) “spacer” segment of DNA, the next core DNA complex forms, and so on, giving chromatin the appearance of beads on a string. Each complex of DNA with core histones is called a nucleosome (see Fig. 2-5), which is the basic structural unit of chromatin, and each of the 46 human chromosomes contains several hundred thousand to well over a million nucleosomes. A fifth histone, H1, appears to bind to DNA at the edge of each nucleosome, in the internucleosomal spacer region. The amount of DNA associated with a core nucleosome, together with the spacer region, is approximately 200 bp.

FIGURE 2-5 Hierarchical levels of chromatin packaging in a human chromosome.
In addition to the major histone types, a number of specialized histones can substitute for H3 or H2A and confer specific characteristics on the genomic DNA at that location. Histones can also be modified by chemical changes, and these modifications can change the properties of nucleosomes that contain them. As discussed further in Chapter 3, the pattern of major and specialized histone types and their modifications can vary from cell type to cell type and is thought to specify how DNA is packaged and how accessible it is to regulatory molecules that determine gene expression or other genome functions.
During the cell cycle, as we will see later in this chapter, chromosomes pass through orderly stages of condensation and decondensation. However, even when chromosomes are in their most decondensed state, in a stage of the cell cycle called interphase, DNA packaged in chromatin is substantially more condensed than it would be as a native, protein-free, double helix. Further, the long strings of nucleosomes are themselves compacted into a secondary helical structure, a cylindrical “solenoid” fiber (from the Greek solenoeides, pipe-shaped) that appears to be the fundamental unit of chromatin organization (see Fig. 2-5). The solenoids are themselves packed into loops or domains attached at intervals of approximately 100,000 bp (equivalent to 100 kilobase pairs [kb], because 1 kb = 1000 bp) to a protein scaffold within the nucleus. It has been speculated that these loops are the functional units of the genome and that the attachment points of each loop are specified along the chromosomal DNA. As we shall see, one level of control of gene expression depends on how DNA and genes are packaged into chromosomes and on their association with chromatin proteins in the packaging process.
The enormous amount of genomic DNA packaged into a chromosome can be appreciated when chromosomes are treated to release the DNA from the underlying protein scaffold (see Fig. 2-1). When DNA is released in this manner, long loops of DNA can be visualized, and the residual scaffolding can be seen to reproduce the outline of a typical chromosome.
The Mitochondrial Chromosome
As mentioned earlier, a small but important subset of genes encoded in the human genome resides in the cytoplasm in the mitochondria (see Fig. 2-1). Mitochondrial genes exhibit exclusively maternal inheritance (see Chapter 7). Human cells can have hundreds to thousands of mitochondria, each containing a number of copies of a small circular molecule, the mitochondrial chromosome. The mitochondrial DNA molecule is only 16 kb in length (just a tiny fraction of the length of even the smallest nuclear chromosome) and encodes only 37 genes. The products of these genes function in mitochondria, although the vast majority of proteins within the mitochondria are, in fact, the products of nuclear genes. Mutations in mitochondrial genes have been demonstrated in several maternally inherited as well as sporadic disorders (Case 33) (see Chapters 7 and 12).
The Human Genome Sequence
With a general understanding of the structure and clinical importance of chromosomes and the genes they carry, scientists turned attention to the identification of specific genes and their location in the human genome. From this broad effort emerged the Human Genome Project, an international consortium of hundreds of laboratories around the world, formed to determine and assemble the sequence of the 3.3 billion base pairs of DNA located among the 24 types of human chromosome.
Over the course of a decade and a half, powered by major developments in DNA-sequencing technology, large sequencing centers collaborated to assemble sequences of each chromosome. The genomes actually being sequenced came from several different individuals, and the consensus sequence that resulted at the conclusion of the Human Genome Project was reported in 2003 as a “reference” sequence assembly, to be used as a basis for later comparison with sequences of individual genomes. This reference sequence is maintained in publicly accessible databases to facilitate scientific discovery and its translation into useful advances for medicine. Genome sequences are typically presented in a 5′ to 3′ direction on just one of the two strands of the double helix, because—owing to the complementary nature of DNA structure described earlier—if one knows the sequence of one strand, one can infer the sequence of the other strand (Fig. 2-6).

FIGURE 2-6 A portion of the reference human genome sequence. By convention, sequences are presented from one strand of DNA only, because the sequence of the complementary strand can be inferred from the double-stranded nature of DNA (shown above the reference sequence). The sequence of DNA from a group of individuals is similar but not identical to the reference, with single nucleotide changes in some individuals and a small deletion of two bases in another.
Organization of the Human Genome
Chromosomes are not just a random collection of different types of genes and other DNA sequences. Regions of the genome with similar characteristics tend to be clustered together, and the functional organization of the genome reflects its structural organization and sequence. Some chromosome regions, or even whole chromosomes, are high in gene content (“gene rich”), whereas others are low (“gene poor”) (Fig. 2-7). The clinical consequences of abnormalities of genome structure reflect the specific nature of the genes and sequences involved. Thus abnormalities of gene-rich chromosomes or chromosomal regions tend to be much more severe clinically than similar-sized defects involving gene-poor parts of the genome.

FIGURE 2-7 Size and gene content of the 24 human chromosomes. Dotted diagonal line corresponds to the average density of genes in the genome, approximately 6.7 protein-coding genes per megabase (Mb). Chromosomes that are relatively gene rich are above the diagonal and trend to the upper left. Chromosomes that are relatively gene poor are below the diagonal and trend to the lower right. SeeSources & Acknowledgments.
As a result of knowledge gained from the Human Genome Project, it is apparent that the organization of DNA in the human genome is both more varied and more complex than was once appreciated. Of the billions of base pairs of DNA in any genome, less than 1.5% actually encodes proteins. Regulatory elements that influence or determine patterns of gene expression during development or in tissues were believed to account for only approximately 5% of additional sequence, although more recent analyses of chromatin characteristics suggest that a much higher proportion of the genome may provide signals that are relevant to genome functions. Only approximately half of the total linear length of the genome consists of so-called single-copy or unique DNA, that is, DNA whose linear order of specific nucleotides is represented only once (or at most a few times) around the entire genome. This concept may appear surprising to some, given that there are only four different nucleotides in DNA. But, consider even a tiny stretch of the genome that is only 10 bases long; with four types of bases, there are over a million possible sequences. And, although the order of bases in the genome is not entirely random, any particular 16-base sequence would be predicted by chance alone to appear only once in any given genome.
The rest of the genome consists of several classes of repetitive DNA and includes DNA whose nucleotide sequence is repeated, either perfectly or with some variation, hundreds to millions of times in the genome. Whereas most (but not all) of the estimated 20,000 protein-coding genes in the genome (see Box earlier in this chapter) are represented in single-copy DNA, sequences in the repetitive DNA fraction contribute to maintaining chromosome structure and are an important source of variation between different individuals; some of this variation can predispose to pathological events in the genome, as we will see in Chapters 5and 6.
Single-Copy DNA Sequences
Although single-copy DNA makes up at least half of the DNA in the genome, much of its function remains a mystery because, as mentioned, sequences actually encoding proteins (i.e., the coding portion of genes) constitute only a small proportion of all the single-copy DNA. Most single-copy DNA is found in short stretches (several kilobase pairs or less), interspersed with members of various repetitive DNA families. The organization of genes in single-copy DNA is addressed in depth in Chapter 3.
Repetitive DNA Sequences
Several different categories of repetitive DNA are recognized. A useful distinguishing feature is whether the repeated sequences (“repeats”) are clustered in one or a few locations or whether they are interspersed with single-copy sequences along the chromosome. Clustered repeated sequences constitute an estimated 10% to 15% of the genome and consist of arrays of various short repeats organized in tandem in a head-to-tail fashion. The different types of such tandem repeats are collectively called satellite DNAs, so named because many of the original tandem repeat families could be separated by biochemical methods from the bulk of the genome as distinct (“satellite”) fractions of DNA.
Tandem repeat families vary with regard to their location in the genome and the nature of sequences that make up the array. In general, such arrays can stretch several million base pairs or more in length and constitute up to several percent of the DNA content of an individual human chromosome. Some tandem repeat sequences are important as tools that are useful in clinical cytogenetic analysis (see Chapter 5). Long arrays of repeats based on repetitions (with some variation) of a short sequence such as a pentanucleotide are found in large genetically inert regions on chromosomes 1, 9, and 16 and make up more than half of the Y chromosome (see Chapter 6). Other tandem repeat families are based on somewhat longer basic repeats. For example, the α-satellite family of DNA is composed of tandem arrays of an approximately 171-bp unit, found at the centromere of each human chromosome, which is critical for attachment of chromosomes to microtubules of the spindle apparatus during cell division.
In addition to tandem repeat DNAs, another major class of repetitive DNA in the genome consists of related sequences that are dispersed throughout the genome rather than clustered in one or a few locations. Although many DNA families meet this general description, two in particular warrant discussion because together they make up a significant proportion of the genome and because they have been implicated in genetic diseases. Among the best-studied dispersed repetitive elements are those belonging to the so-called Alu family. The members of this family are approximately 300 bp in length and are related to each other although not identical in DNA sequence. In total, there are more than a million Alu family members in the genome, making up at least 10% of human DNA. A second major dispersed repetitive DNA family is called the long interspersed nuclear element (LINE, sometimes called L1) family. LINEs are up to 6 kb in length and are found in approximately 850,000 copies per genome, accounting for nearly 20% of the genome. Both of these families are plentiful in some regions of the genome but relatively sparse in others—regions rich in GC content tend to be enriched in Alu elements but depleted of LINE sequences, whereas the opposite is true of more AT-rich regions of the genome.
Repetitive DNA and Disease.
Both Alu and LINE sequences have been implicated as the cause of mutations in hereditary disease. At least a few copies of the LINE and Alu families generate copies of themselves that can integrate elsewhere in the genome, occasionally causing insertional inactivation of a medically important gene. The frequency of such events causing genetic disease in humans is unknown, but they may account for as many as 1 in 500 mutations. In addition, aberrant recombination events between different LINE repeats or Alu repeats can also be a cause of mutation in some genetic diseases (see Chapter 12).
An important additional type of repetitive DNA found in many different locations around the genome includes sequences that are duplicated, often with extraordinarily high sequence conservation. Duplications involving substantial segments of a chromosome, called segmental duplications, can span hundreds of kilobase pairs and account for at least 5% of the genome. When the duplicated regions contain genes, genomic rearrangements involving the duplicated sequences can result in the deletion of the region (and the genes) between the copies and thus give rise to disease (see Chapters 5 and 6).
Variation in the Human Genome
With completion of the reference human genome sequence, much attention has turned to the discovery and cataloguing of variation in sequence among different individuals (including both healthy individuals and those with various diseases) and among different populations around the globe. As we will explore in much more detail in Chapter 4, there are many tens of millions of common sequence variants that are seen at significant frequency in one or more populations; any given individual carries at least 5 million of these sequence variants. In addition, there are countless very rare variants, many of which probably exist in only a single or a few individuals. In fact, given the number of individuals in our species, essentially each and every base pair in the human genome is expected to vary in someone somewhere around the globe. It is for this reason that the original human genome sequence is considered a “reference” sequence for our species, but one that is actually identical to no individual's genome.
Early estimates were that any two randomly selected individuals would have sequences that are 99.9% identical or, put another way, that an individual genome would carry two different versions (alleles) of the human genome sequence at some 3 to 5 million positions, with different bases (e.g., a T or a G) at the maternally and paternally inherited copies of that particular sequence position (see Fig. 2-6). Although many of these allelic differences involve simply one nucleotide, much of the variation consists of insertions or deletions of (usually) short sequence stretches, variation in the number of copies of repeated elements (including genes), or inversions in the order of sequences at a particular position (locus) in the genome (see Chapter 4).
The total amount of the genome involved in such variation is now known to be substantially more than originally estimated and approaches 0.5% between any two randomly selected individuals. As will be addressed in future chapters, any and all of these types of variation can influence biological function and thus must be accounted for in any attempt to understand the contribution of genetics to human health.
Transmission of the Genome
The chromosomal basis of heredity lies in the copying of the genome and its transmission from a cell to its progeny during typical cell division and from one generation to the next during reproduction, when single copies of the genome from each parent come together in a new embryo.
To achieve these related but distinct forms of genome inheritance, there are two kinds of cell division, mitosis and meiosis. Mitosis is ordinary somatic cell division by which the body grows, differentiates, and effects tissue regeneration. Mitotic division normally results in two daughter cells, each with chromosomes and genes identical to those of the parent cell. There may be dozens or even hundreds of successive mitoses in a lineage of somatic cells. In contrast, meiosis occurs only in cells of the germline. Meiosis results in the formation of reproductive cells (gametes), each of which has only 23 chromosomes—one of each kind of autosome and either an X or a Y. Thus, whereas somatic cells have the diploid (diploos, double) or the 2n chromosome complement (i.e., 46 chromosomes), gametes have the haploid (haploos, single) or the n complement (i.e., 23 chromosomes). Abnormalities of chromosome number or structure, which are usually clinically significant, can arise either in somatic cells or in cells of the germline by errors in cell division.
The Cell Cycle
A human being begins life as a fertilized ovum (zygote), a diploid cell from which all the cells of the body (estimated to be approximately 100 trillion in number) are derived by a series of dozens or even hundreds of mitoses. Mitosis is obviously crucial for growth and differentiation, but it takes up only a small part of the life cycle of a cell. The period between two successive mitoses is called interphase, the state in which most of the life of a cell is spent.
Immediately after mitosis, the cell enters a phase, called G1, in which there is no DNA synthesis (Fig. 2-8). Some cells pass through this stage in hours; others spend a long time, days or years, in G1. In fact, some cell types, such as neurons and red blood cells, do not divide at all once they are fully differentiated; rather, they are permanently arrested in a distinct phase known as G0 (“G zero”). Other cells, such as liver cells, may enter G0 but, after organ damage, return to G1 and continue through the cell cycle.

FIGURE 2-8 A typical mitotic cell cycle, described in the text. The telomeres, the centromere, and sister chromatids are indicated.
The cell cycle is governed by a series of checkpoints that determine the timing of each step in mitosis. In addition, checkpoints monitor and control the accuracy of DNA synthesis as well as the assembly and attachment of an elaborate network of microtubules that facilitate chromosome movement. If damage to the genome is detected, these mitotic checkpoints halt cell cycle progression until repairs are made or, if the damage is excessive, until the cell is instructed to die by programmed cell death (a process called apoptosis).
During G1, each cell contains one diploid copy of the genome. As the process of cell division begins, the cell enters S phase, the stage of programmed DNA synthesis, ultimately leading to the precise replication of each chromosome's DNA. During this stage, each chromosome, which in G1 has been a single DNA molecule, is duplicated and consists of two sister chromatids (see Fig. 2-8), each of which contains an identical copy of the original linear DNA double helix. The two sister chromatids are held together physically at the centromere, a region of DNA that associates with a number of specific proteins to form the kinetochore. This complex structure serves to attach each chromosome to the microtubules of the mitotic spindle and to govern chromosome movement during mitosis. DNA synthesis during S phase is not synchronous throughout all chromosomes or even within a single chromosome; rather, along each chromosome, it begins at hundreds to thousands of sites, called origins of DNA replication. Individual chromosome segments have their own characteristic time of replication during the 6- to 8-hour S phase. The ends of each chromosome (or chromatid) are marked by telomeres, which consist of specialized repetitive DNA sequences that ensure the integrity of the chromosome during cell division. Correct maintenance of the ends of chromosomes requires a special enzyme called telomerase, which ensures that the very ends of each chromosome are replicated.
The essential nature of these structural elements of chromosomes and their role in ensuring genome integrity is illustrated by a range of clinical conditions that result from defects in elements of the telomere or kinetochore or cell cycle machinery or from inaccurate replication of even small portions of the genome (see Box). Some of these conditions will be presented in greater detail in subsequent chapters.
Clinical Consequences of Abnormalities and Variation in Chromosome Structure and Mechanics
Medically relevant conditions arising from abnormal structure or function of chromosomal elements during cell division include the following:
• A broad spectrum of congenital abnormalities in children with inherited defects in genes encoding key components of the mitotic spindle checkpoint at the kinetochore
• A range of birth defects and developmental disorders due to anomalous segregation of chromosomes with multiple or missing centromeres (see Chapter 6)
• A variety of cancers associated with overreplication (amplification) or altered timing of replication of specific regions of the genome in S phase (see Chapter 15)
• Roberts syndrome of growth retardation, limb shortening, and microcephaly in children with abnormalities of a gene required for proper sister chromatid alignment and cohesion in S phase
• Premature ovarian failure as a major cause of female infertility due to mutation in a meiosis-specific gene required for correct sister chromatid cohesion
• The so-called telomere syndromes, a number of degenerative disorders presenting from childhood to adulthood in patients with abnormal telomere shortening due to defects in components of telomerase
• And, at the other end of the spectrum, common gene variants that correlate with the number of copies of the repeats at telomeres and with life expectancy and longevity
By the end of S phase, the DNA content of the cell has doubled, and each cell now contains two copies of the diploid genome. After S phase, the cell enters a brief stage called G2. Throughout the whole cell cycle, the cell gradually enlarges, eventually doubling its total mass before the next mitosis. G2 is ended by mitosis, which begins when individual chromosomes begin to condense and become visible under the microscope as thin, extended threads, a process that is considered in greater detail in the following section.
The G1, S, and G2 phases together constitute interphase. In typical dividing human cells, the three phases take a total of 16 to 24 hours, whereas mitosis lasts only 1 to 2 hours (see Fig. 2-8). There is great variation, however, in the length of the cell cycle, which ranges from a few hours in rapidly dividing cells, such as those of the dermis of the skin or the intestinal mucosa, to months in other cell types.
Mitosis
During the mitotic phase of the cell cycle, an elaborate apparatus ensures that each of the two daughter cells receives a complete set of genetic information. This result is achieved by a mechanism that distributes one chromatid of each chromosome to each daughter cell (Fig. 2-9). The process of distributing a copy of each chromosome to each daughter cell is called chromosome segregation. The importance of this process for normal cell growth is illustrated by the observation that many tumors are invariably characterized by a state of genetic imbalance resulting from mitotic errors in the distribution of chromosomes to daughter cells.

FIGURE 2-9 Mitosis. Only two chromosome pairs are shown. For details, see text.
The process of mitosis is continuous, but five stages, illustrated in Figure 2-9, are distinguished: prophase, prometaphase, metaphase, anaphase, and telophase.
• Prophase. This stage is marked by gradual condensation of the chromosomes, formation of the mitotic spindle, and formation of a pair of centrosomes, from which microtubules radiate and eventually take up positions at the poles of the cell.
• Prometaphase. Here, the nuclear membrane dissolves, allowing the chromosomes to disperse within the cell and to attach, by their kinetochores, to microtubules of the mitotic spindle.
• Metaphase. At this stage, the chromosomes are maximally condensed and line up at the equatorial plane of the cell.
• Anaphase. The chromosomes separate at the centromere, and the sister chromatids of each chromosome now become independent daughter chromosomes, which move to opposite poles of the cell.
• Telophase. Now, the chromosomes begin to decondense from their highly contracted state, and a nuclear membrane begins to re-form around each of the two daughter nuclei, which resume their interphase appearance. To complete the process of cell division, the cytoplasm cleaves by a process known as cytokinesis.
There is an important difference between a cell entering mitosis and one that has just completed the process. A cell in G2 has a fully replicated genome (i.e., a 4n complement of DNA), and each chromosome consists of a pair of sister chromatids. In contrast, after mitosis, the chromosomes of each daughter cell have only one copy of the genome. This copy will not be duplicated until a daughter cell in its turn reaches the S phase of the next cell cycle (see Fig. 2-8). The entire process of mitosis thus ensures the orderly duplication and distribution of the genome through successive cell divisions.
The Human Karyotype
The condensed chromosomes of a dividing human cell are most readily analyzed at metaphase or prometaphase. At these stages, the chromosomes are visible under the microscope as a so-called chromosome spread; each chromosome consists of its sister chromatids, although in most chromosome preparations, the two chromatids are held together so tightly that they are rarely visible as separate entities.
As stated earlier, there are 24 different types of human chromosome, each of which can be distinguished cytologically by a combination of overall length, location of the centromere, and sequence content, the latter reflected by various staining methods. The centromere is apparent as a primary constriction, a narrowing or pinching-in of the sister chromatids due to formation of the kinetochore. This is a recognizable cytogenetic landmark, dividing the chromosome into two arms, a short arm designated p (for petit) and a long arm designated q.
Figure 2-10 shows a prometaphase cell in which the chromosomes have been stained by the Giemsa-staining (G-banding) method (also see Chapter 5). Each chromosome pair stains in a characteristic pattern of alternating light and dark bands (G bands) that correlates roughly with features of the underlying DNA sequence, such as base composition (i.e., the percentage of base pairs that are GC or AT) and the distribution of repetitive DNA elements. With such banding techniques, all of the chromosomes can be individually distinguished, and the nature of many structural or numerical abnormalities can be determined, as we examine in greater detail in Chapters 5 and 6.

FIGURE 2-10 A chromosome spread prepared from a lymphocyte culture that has been stained by the Giemsa-banding (G-banding) technique. The darkly stained nucleus adjacent to the chromosomes is from a different cell in interphase, when chromosomal material is diffuse throughout the nucleus. SeeSources & Acknowledgments.
Although experts can often analyze metaphase chromosomes directly under the microscope, a common procedure is to cut out the chromosomes from a digital image or photomicrograph and arrange them in pairs in a standard classification (Fig. 2-11). The completed picture is called a karyotype. The word karyotype is also used to refer to the standard chromosome set of an individual (“a normal male karyotype”) or of a species (“the human karyotype”) and, as a verb, to the process of preparing such a standard figure (“to karyotype”).

FIGURE 2-11 A human male karyotype with Giemsa banding (G banding). The chromosomes are at the prometaphase stage of mitosis and are arranged in a standard classification, numbered 1 to 22 in order of length, with the X and Y chromosomes shown separately. SeeSources & Acknowledgments.
Unlike the chromosomes seen in stained preparations under the microscope or in photographs, the chromosomes of living cells are fluid and dynamic structures. During mitosis, the chromatin of each interphase chromosome condenses substantially (Fig. 2-12). When maximally condensed at metaphase, DNA in chromosomes is approximately 1/10,000 of its fully extended state. When chromosomes are prepared to reveal bands (as in Figs. 2-10 and 2-11), as many as 1000 or more bands can be recognized in stained preparations of all the chromosomes. Each cytogenetic band therefore contains as many as 50 or more genes, although the density of genes in the genome, as mentioned previously, is variable.

FIGURE 2-12 Cycle of condensation and decondensation as a chromosome proceeds through the cell cycle.
Meiosis
Meiosis, the process by which diploid cells give rise to haploid gametes, involves a type of cell division that is unique to germ cells. In contrast to mitosis, meiosis consists of one round of DNA replication followed by two rounds of chromosome segregation and cell division (see meiosis I and meiosis II in Fig. 2-13). As outlined here and illustrated in Figure 2-14, the overall sequence of events in male and female meiosis is the same; however, the timing of gametogenesis is very different in the two sexes, as we will describe more fully later in this chapter.

FIGURE 2-13 A simplified representation of the essential steps in meiosis, consisting of one round of DNA replication followed by two rounds of chromosome segregation, meiosis I and meiosis II.

FIGURE 2-14 Meiosis and its consequences. A single chromosome pair and a single crossover are shown, leading to formation of four distinct gametes. The chromosomes replicate during interphase and begin to condense as the cell enters prophase of meiosis I. In meiosis I, the chromosomes synapse and recombine. A crossover is visible as the homologues align at metaphase I, with the centromeres oriented toward opposite poles. In anaphase I, the exchange of DNA between the homologues is apparent as the chromosomes are pulled to opposite poles. After completion of meiosis I and cytokinesis, meiosis II proceeds with a mitosis-like division. The sister kinetochores separate and move to opposite poles in anaphase II, yielding four haploid products.
Meiosis I is also known as the reduction division because it is the division in which the chromosome number is reduced by half through the pairing of homologues in prophase and by their segregation to different cells at anaphase of meiosis I. Meiosis I is also notable because it is the stage at which genetic recombination (also called meiotic crossing over) occurs. In this process, as shown for one pair of chromosomes in Figure 2-14, homologous segments of DNA are exchanged between nonsister chromatids of each pair of homologous chromosomes, thus ensuring that none of the gametes produced by meiosis will be identical to another. The conceptual and practical consequences of recombination for many aspects of human genetics and genomics are substantial and are outlined in the Box at the end of this section.
Prophase of meiosis I differs in a number of ways from mitotic prophase, with important genetic consequences, because homologous chromosomes need to pair and exchange genetic information. The most critical early stage is called zygotene, when homologous chromosomes begin to align along their entire length. The process of meiotic pairing—called synapsis—is normally precise, bringing corresponding DNA sequences into alignment along the length of the entire chromosome pair. The paired homologues—now called bivalents—are held together by a ribbon-like proteinaceous structure called the synaptonemal complex, which is essential to the process of recombination. After synapsis is complete, meiotic crossing over takes place during pachytene, after which the synaptonemal complex breaks down.
Metaphase I begins, as in mitosis, when the nuclear membrane disappears. A spindle forms, and the paired chromosomes align themselves on the equatorial plane with their centromeres oriented toward different poles (see Fig. 2-14).
Anaphase of meiosis I again differs substantially from the corresponding stage of mitosis. Here, it is the two members of each bivalent that move apart, not the sister chromatids (contrast Fig. 2-14 with Fig. 2-9). The homologous centromeres (with their attached sister chromatids) are drawn to opposite poles of the cell, a process termed disjunction. Thus the chromosome number is halved, and each cellular product of meiosis I has the haploid chromosome number. The 23 pairs of homologous chromosomes assort independently of one another, and as a result, the original paternal and maternal chromosome sets are sorted into random combinations. The possible number of combinations of the 23 chromosome pairs that can be present in the gametes is 223 (more than 8 million). Owing to the process of crossing over, however, the variation in the genetic material that is transmitted from parent to child is actually much greater than this. As a result, each chromatid typically contains segments derived from each member of the original parental chromosome pair, as illustrated schematically in Figure 2-14. For example, at this stage, a typical large human chromosome would be composed of three to five segments, alternately paternal and maternal in origin, as inferred from DNA sequence variants that distinguish the respective parental genomes (Fig. 2-15).

FIGURE 2-15 The effect of homologous recombination in meiosis. In this example, representing the inheritance of sequences on a typical large chromosome, an individual has distinctive homologues, one containing sequences inherited from his father (blue) and one containing homologous sequences from his mother (purple). After meiosis in spermatogenesis, he transmits a single complete copy of that chromosome to his two offspring. However, as a result of crossing over (arrows), the copy he transmits to each child consists of alternating segments of the two grandparental sequences. Child 1 inherits a copy after two crossovers, whereas child 2 inherits a copy with three crossovers.
After telophase of meiosis I, the two haploid daughter cells enter meiotic interphase. In contrast to mitosis, this interphase is brief, and meiosis II begins. The notable point that distinguishes meiotic and mitotic interphase is that there is no S phase (i.e., no DNA synthesis and duplication of the genome) between the first and second meiotic divisions.
Meiosis II is similar to an ordinary mitosis, except that the chromosome number is 23 instead of 46; the chromatids of each of the 23 chromosomes separate, and one chromatid of each chromosome passes to each daughter cell (see Fig. 2-14). However, as mentioned earlier, because of crossing over in meiosis I, the chromosomes of the resulting gametes are not identical (see Fig. 2-15).
Genetic Consequences and Medical Relevance of Homologous Recombination
The take-home lesson of this portion of the chapter is a simple one: the genetic content of each gamete is unique, because of random assortment of the parental chromosomes to shuffle the combination of sequence variants between chromosomes and because of homologous recombination to shuffle the combination of sequence variants within each and every chromosome. This has significant consequences for patterns of genomic variation among and between different populations around the globe and for diagnosis and counseling of many common conditions with complex patterns of inheritance (see Chapters 8and 10).
The amounts and patterns of meiotic recombination are determined by sequence variants in specific genes and at specific “hot spots” and differ between individuals, between the sexes, between families, and between populations (see Chapter 10).
Because recombination involves the physical intertwining of the two homologues until the appropriate point during meiosis I, it is also critical for ensuring proper chromosome segregation during meiosis. Failure to recombine properly can lead to chromosome missegregation (nondisjunction) in meiosis I and is a frequent cause of pregnancy loss and of chromosome abnormalities like Down syndrome (see Chapters 5 and 6).
Major ongoing efforts to identify genes and their variants responsible for various medical conditions rely on tracking the inheritance of millions of sequence differences within families or the sharing of variants within groups of even unrelated individuals affected with a particular condition. The utility of this approach, which has uncovered thousands of gene-disease associations to date, depends on patterns of homologous recombination in meiosis (see Chapter 10).
Although homologous recombination is normally precise, areas of repetitive DNA in the genome and genes of variable copy number in the population are prone to occasional unequal crossing over during meiosis, leading to variations in clinically relevant traits such as drug response, to common disorders such as the thalassemias or autism, or to abnormalities of sexual differentiation (see Chapters 6, 8, and 11).
Although homologous recombination is a normal and essential part of meiosis, it also occurs, albeit more rarely, in somatic cells. Anomalies in somatic recombination are one of the causes of genome instability in cancer (see Chapter 15).
Human Gametogenesis and Fertilization
The cells in the germline that undergo meiosis, primary spermatocytes or primary oocytes, are derived from the zygote by a long series of mitoses before the onset of meiosis. Male and female gametes have different histories, marked by different patterns of gene expression that reflect their developmental origin as an XY or XX embryo. The human primordial germ cells are recognizable by the fourth week of development outside the embryo proper, in the endoderm of the yolk sac. From there, they migrate during the sixth week to the genital ridges and associate with somatic cells to form the primitive gonads, which soon differentiate into testes or ovaries, depending on the cells' sex chromosome constitution (XY or XX), as we examine in greater detail in Chapter 6. Both spermatogenesis and oogenesis require meiosis but have important differences in detail and timing that may have clinical and genetic consequences for the offspring. Female meiosis is initiated once, early during fetal life, in a limited number of cells. In contrast, male meiosis is initiated continuously in many cells from a dividing cell population throughout the adult life of a male.
In the female, successive stages of meiosis take place over several decades—in the fetal ovary before the female in question is even born, in the oocyte near the time of ovulation in the sexually mature female, and after fertilization of the egg that can become that female's offspring. Although postfertilization stages can be studied in vitro, access to the earlier stages is limited. Testicular material for the study of male meiosis is less difficult to obtain, inasmuch as testicular biopsy is included in the assessment of many men attending infertility clinics. Much remains to be learned about the cytogenetic, biochemical, and molecular mechanisms involved in normal meiosis and about the causes and consequences of meiotic irregularities.
Spermatogenesis
The stages of spermatogenesis are shown in Figure 2-16. The seminiferous tubules of the testes are lined with spermatogonia, which develop from the primordial germ cells by a long series of mitoses and which are in different stages of differentiation. Sperm (spermatozoa) are formed only after sexual maturity is reached. The last cell type in the developmental sequence is the primary spermatocyte, a diploid germ cell that undergoes meiosis I to form two haploid secondary spermatocytes. Secondary spermatocytes rapidly enter meiosis II, each forming two spermatids, which differentiate without further division into sperm. In humans, the entire process takes approximately 64 days. The enormous number of sperm produced, typically approximately 200 million per ejaculate and an estimated 1012 in a lifetime, requires several hundred successive mitoses.

FIGURE 2-16 Human spermatogenesis in relation to the two meiotic divisions. The sequence of events begins at puberty and takes approximately 64 days to be completed. The chromosome number (46 or 23) and the sex chromosome constitution (X or Y) of each cell are shown. SeeSources & Acknowledgments.
As discussed earlier, normal meiosis requires pairing of homologous chromosomes followed by recombination. The autosomes and the X chromosomes in females present no unusual difficulties in this regard; but what of the X and Y chromosomes during spermatogenesis? Although the X and Y chromosomes are different and are not homologues in a strict sense, they do have relatively short identical segments at the ends of their respective short arms (Xp and Yp) and long arms (Xq and Yq) (see Chapter 6). Pairing and crossing over occurs in both regions during meiosis I. These homologous segments are called pseudoautosomal to reflect their autosome-like pairing and recombination behavior, despite being on different sex chromosomes.
Oogenesis
Whereas spermatogenesis is initiated only at the time of puberty, oogenesis begins during a female's development as a fetus (Fig. 2-17). The ova develop from oogonia, cells in the ovarian cortex that have descended from the primordial germ cells by a series of approximately 20 mitoses. Each oogonium is the central cell in a developing follicle. By approximately the third month of fetal development, the oogonia of the embryo have begun to develop into primary oocytes, most of which have already entered prophase of meiosis I. The process of oogenesis is not synchronized, and both early and late stages coexist in the fetal ovary. Although there are several million oocytes at the time of birth, most of these degenerate; the others remain arrested in prophase I (see Fig. 2-14) for decades. Only approximately 400 eventually mature and are ovulated as part of a woman's menstrual cycle.

FIGURE 2-17 Human oogenesis and fertilization in relation to the two meiotic divisions. The primary oocytes are formed prenatally and remain suspended in prophase of meiosis I for years until the onset of puberty. An oocyte completes meiosis I as its follicle matures, resulting in a secondary oocyte and the first polar body. After ovulation, each oocyte continues to metaphase of meiosis II. Meiosis II is completed only if fertilization occurs, resulting in a fertilized mature ovum and the second polar body.
After a woman reaches sexual maturity, individual follicles begin to grow and mature, and a few (on average one per month) are ovulated. Just before ovulation, the oocyte rapidly completes meiosis I, dividing in such a way that one cell becomes the secondary oocyte (an egg or ovum), containing most of the cytoplasm with its organelles; the other cell becomes the first polar body (see Fig. 2-17). Meiosis II begins promptly and proceeds to the metaphase stage during ovulation, where it halts again, only to be completed if fertilization occurs.
Fertilization
Fertilization of the egg usually takes place in the fallopian tube within a day or so of ovulation. Although many sperm may be present, the penetration of a single sperm into the ovum sets up a series of biochemical events that usually prevent the entry of other sperm.
Fertilization is followed by the completion of meiosis II, with the formation of the second polar body (see Fig. 2-17). The chromosomes of the now-fertilized egg and sperm form pronuclei, each surrounded by its own nuclear membrane. It is only upon replication of the parental genomes after fertilization that the two haploid genomes become one diploid genome within a shared nucleus. The diploid zygote divides by mitosis to form two diploid daughter cells, the first in the series of cell divisions that initiate the process of embryonic development (see Chapter 14).
Although development begins at the time of conception, with the formation of the zygote, in clinical medicine the stage and duration of pregnancy are usually measured as the “menstrual age,” dating from the beginning of the mother's last menstrual period, typically approximately 14 days before conception.
Medical Relevance of Mitosis and Meiosis
The biological significance of mitosis and meiosis lies in ensuring the constancy of chromosome number—and thus the integrity of the genome—from one cell to its progeny and from one generation to the next. The medical relevance of these processes lies in errors of one or the other mechanism of cell division, leading to the formation of an individual or of a cell lineage with an abnormal number of chromosomes and thus an abnormal dosage of genomic material.
As we see in detail in Chapter 5, meiotic nondisjunction, particularly in oogenesis, is the most common mutational mechanism in our species, responsible for chromosomally abnormal fetuses in at least several percent of all recognized pregnancies. Among pregnancies that survive to term, chromosome abnormalities are a leading cause of developmental defects, failure to thrive in the newborn period, and intellectual disability.
Mitotic nondisjunction in somatic cells also contributes to genetic disease. Nondisjunction soon after fertilization, either in the developing embryo or in extraembryonic tissues like the placenta, leads to chromosomal mosaicism that can underlie some medical conditions, such as a proportion of patients with Down syndrome. Further, abnormal chromosome segregation in rapidly dividing tissues, such as in cells of the colon, is frequently a step in the development of chromosomally abnormal tumors, and thus evaluation of chromosome and genome balance is an important diagnostic and prognostic test in many cancers.
General References
Green ED, Guyer MS, National Human Genome Research Institute. Charting a course for genomic medicine from base pairs to bedside. Nature. 2011;470:204–213.
Lander ES. Initial impact of the sequencing of the human genome. Nature. 2011;470:187–197.
Moore KL, Presaud TVN, Torchia MG. The developing human: clinically oriented embryology. ed 9. WB Saunders: Philadelphia; 2013.
References for Specific Topics
Deininger P. Alu elements: know the SINES. Genome Biol. 2011;12:236.
Frazer KA. Decoding the human genome. Genome Res. 2012;22:1599–1601.
International Human Genome Sequencing Consortium. Initial sequencing and analysis of the human genome. Nature. 2001;409:860–921.
International Human Genome Sequencing Consortium. Finishing the euchromatic sequence of the human genome. Nature. 2004;431:931–945.
Venter J, Adams M, Myers E, et al. The sequence of the human genome. Science. 2001;291:1304–1351.
Problems
1. At a certain locus, a person has two alleles, A and a.
a. What alleles will be present in this person's gametes?
b. When do A and a segregate (1) if there is no crossing over between the locus and the centromere of the chromosome? (2) if there is a single crossover between the locus and the centromere?
2. What is the main cause of numerical chromosome abnormalities in humans?
3. Disregarding crossing over, which increases the amount of genetic variability, estimate the probability that all your chromosomes have come to you from your father's mother and your mother's mother. Would you be male or female?
4. A chromosome entering meiosis is composed of two sister chromatids, each of which is a single DNA molecule.
a. In our species, at the end of meiosis I, how many chromosomes are there per cell? How many chromatids?
b. At the end of meiosis II, how many chromosomes are there per cell? How many chromatids?
c. When is the diploid chromosome number restored? When is the two-chromatid structure of a typical metaphase chromosome restored?
5. From Figure 2-7, estimate the number of genes per million base pairs on chromosomes 1, 13, 18, 19, 21, and 22. Would a chromosome abnormality of equal size on chromosome 18 or 19 be expected to have greater clinical impact? On chromosome 21 or 22?