Primer of Genetic Analysis: A Problems Approach 3rd Ed.

CHAPTER NINTEEN Overview of Molecular Biology Techniques

By the early 1970s, basic research in microbial genetics and biochemistry had provided a fundamental understanding of how genetic information was stored and transmitted, as well as how genes are expressed in response to specific physiological conditions. This knowledge was vital in harnessing microorganisms to replicate foreign DNA, generating the flood of technical applications that are often grouped under the banner of “recombinant DNA technology.” This overview represents a primer in the basic tools used by biologists to introduce foreign DNA into cells. These protocols make extensive use of enzymes produced by microorganisms and used in replicating, repairing, or protecting their own genetic material. It also describes some of the applications of these methodologies in basic research, medicine, and industry.

I. Restriction endonucleases. Restriction enzymes are site-specific nucleases that produce internal breaks at distinct nucleotide sequences in double-stranded DNA.

A. In bacteria, they are part of a protection system against viral infection that allows the host to recognize invading DNA and degrade it. Each restriction enzyme will bind to a particular base sequence and cleave DNA at that base sequence. The DNA of the host is modified by methylases, enzymes that add methyl groups to certain bases within the endonuclease recognition pattern. This protects the host genetic material from digestion. Hundreds of different restriction endonuclease/ methylase modification systems have been identified from a large variety of different microorganisms. For example:

EcoRI (from certain E. coli. strains)

Recognition sequence GAATTCCTTAAG

HindIII (from certain Hemophilus strain)

Recognition sequence AAGCTTTTCGAA

B. Note that the recognition sites for these enzymes are often palindromes, reading the same way in the 5′ U+2192 3′ direction on both strands (some call these the Watson and Crick strands). Note also that some restriction enzymes cleave each individual DNA strand off-center, producing a “staggered” cut that leaves a single-stranded tail:

As discussed later, these overlapping, or “sticky,” ends are useful in joining DNA molecules that have been cut with the same enzyme.

II. Use of vectors in the propagation of foreign DNA in bacteria

A. Plasmid vectors. By the use of specific nucleases and DNA ligase, it is possible to introduce foreign DNA molecules (i.e., genes) into plasmids, transform this re-combinant plasmid DNA into bacteria, and use the bacterial cell to produce large amounts of the foreign DNA (DNA cloning, described later). Plasmids have several features that make them highly suitable for the replication of foreign DNA in bacterial cells (i.e., as cloning vectors).

1. They are capable of autonomously replicating, since they possess their own origin of replication. Most plasmids used for DNA cloning replicate many times per bacterial chromosome replication cycle and are present in multiple copies per cell.

2. They replicate as small, supercoiled, covalently closed circles. This topological property allows them to be easily purified away from the large bacterial chromosomal DNA.

3. They often carry genes that can serve as selectable markers, such as antibiotic resistance. Thus, the bacterial cells that carry them can be isolated by growing on the appropriate culture media. Bacteria thus become factories for the production of foreign DNA, or in specialized cases, for proteins that can be expressed from the introduced genetic material.

B. Bacteriophage vectors: Bacteriophages (bacterial viruses) are capable of infecting a host bacterial cell, entering into a lytic cycle, and producing large numbers of progeny phage from a single infection. They can also serve as cloning vectors for propagating foreign DNA. The foreign DNA can be biochemically inserted to replace regions of the phage genome normally used in the lysogenic pathway. After infection, the recombinant bacteriophage will replicate in the bacterial cell, produce numerous progeny phage (containing the foreign DNA of interest), and go on to infect other cells. This will produce lysed regions upon a bacterial lawn that contain large amounts of the same recombinant phage viral particles.

C. Specialized vectors: Specialized plasmid and viral vectors allow propagation and/or expression of cloned genes in a large variety of different prokaryotic and eukaryotic model systems. Many specialized vectors incorporate features that allow for increased ease of cloning, maximal recovery of large contiguous pieces of nucleic acid, and/or targeted gene expression under specific environmental or developmental conditions (see applications, below). They are often composites of different genetic elements important for performing a specific function. Yeast Artificial Chromosomes (YACs), for example, allow for the propagation of very large pieces of foreign DNA in yeast, since they are capped by telomeres and contain an origin of replication and centromere. Similarly, the transposable element sequences engineered into Drosophila P-element vectors promote the integration of a cloned gene of interest directly into the fly genome. Many vectors contain a promoter element that allows a particular DNA sequence to be expressed only in response to a precise environmental stimulus (e.g., induction of the lac promoter by a lactose analog).

III. Construction of a recombinant genomic DNA library in a bacterial vector and introduction into bacteria

A. Vector DNA and the foreign DNA of interest (e.g., genomic DNA from a eukaryotic organism) are purified biochemically and then cut with a restriction endonuclease that produces complementary staggered ends in both molecules. An enzyme is chosen that will produce only one distinct cut within the bacterial vector, in a region of the molecule that does not contain any gene essential for its replication or selection. The eukaryotic DNA, which has a much greater sequence complexity, will be cut hundreds or thousands of times, that is, whenever the recognition site for the enzyme appears in the genome. The two DNA populations are mixed under conditions in which the staggered ends of vector and genomic DNA can base pair, and DNA ligase is added to the reaction. The ratio of the insert and vector DNA is also adjusted, so that most of the time, one vector molecule will base pair to one molecule of foreign DNA. If bacteriophage vectors are used, the ligation products may be combined with a bacterial cell extract containing viral coat proteins, and the recombinant DNA is “packaged” – allowed to self-assemble into infective phage particles.

One trick used to assure that only recombinant phage are packaged is to engineer the phage vector DNA to be too small to be efficiently packaged by itself: only phage containing a “stuffer” fragment of foreign DNA will be packaged into an infective phage particle. If plasmids are used, some of the plasmid molecules will also recircularize during the ligase reaction, either by themselves or after a piece of eukaryotic DNA has been joined to them. There are also a number of tricks to limit the number of plasmid molecules that do not pick up a foreign DNA insert; one is to treat the vector DNA with the enzyme alkaline phosphatase, which removes 5′ phosphates and prevents the vector molecule from recircularizing without an insert.

B. The recombinant molecules produced in vitro can now be used to infect or transform E. coli cells. For viral vectors, individual phage infections upon a bacterial lawn in a petri dish will give rise to clones of viruses clustered within zones of lysed cells (plaques). For plasmids, a recipient strain of E. coli, sensitive to a particular antibiotic, is chemically treated to increase the probability of DNA uptake and then exposed to recombinant plasmids. As discussed previously, plasmid vectors contain genes that can confer selective advantage, such as a gene that encodes an enzyme that degrades antibiotics. Successful tranformants are selected by plating on the appropriate medium (e.g., agar plates containing an antibiotic) such that only cells containing the plasmid will survive.

In library construction, each bacterial colony growing on the selective media will have resulted from a single plasmid transformation. Each individual colony will be replicating the same foreign DNA. Each initial transformation, however, will be unique to a particular DNA molecule, and the vast majority of the transformed cell population will therefore be replicating different pieces of foreign DNA. The same situation holds for viral vectors. A single recombinant phage will infect a cell, be replicated, lyse that cell, and go on to infect neighboring cells in a chain reaction that propagates a single phage’s recombinant DNA insert within a single plaque. The initial infecting bacteriophage, however, will be derived from an independent packaging event and each plaque will contain phage harboring different pieces of foreign DNA. It is thus possible to construct a “library” of individual clones, each of which contains a different part of the genome of the eukaryotic organism. Virtually all of a particular eukaryotic genome may therefore be represented in a recombinant DNA library consisting of hundreds of thousands of individual recombinant DNA clones.

C. cDNA libraries: In addition to producing recombinant DNA libraries that are representative of an organism’s entire genome, it is possible to construct libraries that are representative of only those genes that are expressed in a particular tissue or during a specific stage in an organism’s developmental history. This is accomplished by first isolating messenger RNA from a tissue of interest. If, for example, we were interested in isolating a DNA sequence that encoded the hormone insulin, pancreatic tissue would be the logical source of mRNA. Once biochemically purified, the mRNA can serve as a template for producing a “copy DNA” or cDNA. The process of efficiently converting mRNA into a double-stranded DNA molecule that can be inserted into a bacterial vector is accomplished in vitro, using many of the bacterial enzymes described. The critical first step in cDNA synthesis, however, requires reverse transcriptase, an enzyme purified from the RNA viruses of vertebrates. This enzyme is an RNA-dependent DNA polymerase; that is, it can copy an RNA sequence into DNA, an activity critical in the viral reproductive cycle.

To construct cDNA libraries from purified mRNAs, an oligo-dT primer is often used. This primer will bind to the poly(A) tail found at the 3′ end of most eukaryotic mRNAs, thus serving as a “universal” primer for all mRNAs in the population. Another approach to obtaining cDNAs representative of the entire mRNA population is to use a mixture of short (hexamer) random oligonucleotides, which will prime wherever they find a complementary sequence on the mRNA. Sequence analysis of these cDNA libraries provides the researcher with data on what genes are activated in a particular tissue (termed expressed sequence tags, or ESTs).

IV. Identifying a specific DNA clone within a recombinant DNA library

A. Given that a particular DNA sequence may be represented once among thousands of individual clones, how does one go about isolating from a recombinant DNA library a particular clone carrying a specific gene sequence? A number of methods are used, but the most common approach is to produce a single-stranded nucleic acid “probe” that is capable of base-pairing with complementary sequences present in the cloned DNA. The “probe” nucleic acid sequence is tagged with a covalently attached radioactive or fluorescent molecule, a molecular label that allows small amounts of the probe to be detected. There are several ways to produce labeled single-stranded nucleic acids complementary to a particular gene sequence. For example, if a small portion of the protein sequence of a particular gene product has been determined, it is possible to synthesize a small, labeled oligonucleotide (18–30 oligomers) based on the determined peptide sequence

Note that since the code is degenerate, several codon possibilities may exist for any given peptide sequence. This is not normally a problem since a number of different oligonucleotides can be synthesized and used in combination as a probe.

B. The bacterial colonies or plaques representing the library can be replica-plated onto a piece of filter paper, such that their orientation in relation to the agar plate can be identified. The bacterial cells or phage transferred to such filters can now be lysed and the DNA in them denatured to form single-stranded molecules. By incubating the filters with the probe under controlled salt and temperature conditions, the probe will only base pair with complementary sequences in the genomic DNA library. This technique is termed nucleic acid hybridization. If, for example, the probe molecule were radioactively labeled, the plaques or colonies that hybridized to the probe could be detected by covering the filters with photographic film and allowing the radioactivity to expose the film (autoradiography). Developing the film will reveal the positions of the colonies or phage containing the gene of interest. Returning to the original plate, the individual colonies or phage containing this foreign gene can now be grown, the recombinant DNA isolated biochemically, and the foreign DNA molecule studied.

V. Polymerase chain reaction

A. In 1985, a technique called the polymerase chain reaction (PCR) was introduced. It significantly reduced the time and effort necessary to generate large amounts of a specific DNA sequence. A short pair of oligonucleotides (primers) complementary to opposite strands of a given DNA sequence could be synthesized, and the region of a source DNA flanked by these oligonucleotides amplified a million-fold in an in vitro reaction. The source DNA sample could be extremely heterogeneous (e.g., genomic DNA) and the concentration of the sequence of interest within the source material exceedingly small (theoretically, a single molecule).

B. The protocol requires two oligonucleotide primers that are capable of base-pairing with the nucleotide sequences at the two 3′ ends of the DNA sequence of interest. These primers, along with deoxynucleotide triphosphate (dNTP) precursors, are introduced into the reaction in great molar excess relative to the source DNA molecule. The reaction is heated to denature the source DNA, then cooled to permit the single-strand primers to anneal to complementary regions of the sample. DNA polymerase is now added, and the oligonucleotides, which serve as primers, are elongated in the 5′ to 3′ direction. This round of denaturation, annealing and synthesis leads to a doubling of the DNA in the region flanked by the primers. Indeed, successive cycles of synthesis will lead to an exponential increase of product: 1, 2, 4, 8, 16, and so on. As a result of asynchrony in synthesis, molecules that are longer than the region flanked by primers will be synthesized in each cycle, but they will accumulate linearly and be a minor component of the final product.

C. When first introduced, this protocol was performed with E. coli DNA polymerase I. Since this enzyme is heat labile, it was destroyed after every denaturation, requiring that fresh enzyme be added during each cycle of synthesis. This inconvenience was eliminated with the isolation (and gene cloning) of DNA polymerases isolated from thermophilic bacteria, such as those inhabiting hot springs. These enzymes can be added just once to the reaction tube and the entire process automated on thermal cyclers, which are reaction chambers that can be programmed to switch progressively between the different temperature conditions required for denaturation, annealing, and DNA synthesis.

D. One variation of the PCR technique is to use a reverse transcriptase reaction to first copy mRNA molecules into a single stand of DNA, then use PCR to amplify that DNA using gene-specific primers, a process termed RT-PCR (reverse-transcriptase/polymerase chain reaction). As in genomic DNA amplification, a particular sequence found in low amounts can be amplified from a complex mixture of templates, using the appropriate primers. This technique can therefore recover the sequence of an mRNA of interest, even if the gene encoding that mRNA is expressed at very low levels in a particular cell type. If a small amount of sequence information is available to design specific primers, this technology circumvents the need for building and screening a cDNA library. To recover a clone of a certain gene sequence, one can amplify it directly from the appropriate mRNA population.

VI. Applications

A. Restriction enzyme mapping: If large amounts of a particular DNA sequence are available, one obvious benefit is the ability to subject that DNA to detailed biochemical analysis. Often the first step in characterizing cloned DNA is the construction of a restriction map. The DNA in question is usually digested with a variety of different enzymes, both singly and in various enzyme combinations. The fragments produced from the digestions are then separated by gel electrophoresis. In this technique, the negatively charged DNA molecules move to the positive pole and are separated on the basis of size, with the smaller molecules migrating fastest through the gel matrix. By comparing the sizes of the digestion products from different combinations of enzymes, the restriction sites in the DNA can be positioned relative to one another. The order of specific enzyme sites constitutes a “map,” which is a characteristic of any given DNA molecule.

B. DNA sequencing: If the linear array of nucleotides in a DNA molecule could be determined, then one could potentially examine what proteins it encodes or what structural features may be important for gene expression. In 1977, two techniques were introduced that permitted researchers to determine the linear order of nucleotides present in a DNA molecule. The technique that is now most widely applied, termed the chain termination method, is again based on an enzymatic reaction using DNA polymerase. As in PCR, the source DNA is denatured by heating and a synthetic oligonucleotide primer complementary to a small region of the DNA is allowed to base pair to one of the denatured strands. A DNA polymerase reaction can be initiated off the primer oligonucleotide; the enzyme will then synthesize a complementary copy of the DNA of interest. In addition to the four dNTPs needed for synthesis (which can be radioactively labeled to detect the DNA being synthesized) a nucleotide analogue, missing a 3′ hydroxyl group (a dideoxynucleotide, or ddNTP) is also included in the enzymatic reaction. Four separate enzymatic reactions are therefore set up, all of which include the four dNTPs, a primer, and the source DNA of interest, but which differ in the type of dideoxy-analogue added: ddATP added in reaction 1, ddGTP added in reaction 2, ddCTP added in reaction 3, and ddTTP added in reaction 4. The purpose of this is to produce a controlled interruption of enzymatic replication at particular sites; in each of the four reactions, the synthesis will be randomly terminated at A, G, C, and T, respectively. If a dideoxynucleotide is inserted instead of the normal nucleotide (e.g., ddCTP instead of dCTP) the chain cannot be continued, since the analogue lacks the 3′ hydroxyl terminus necessary for the next phosphodiester bond. Remember that, in each reaction, termination will occur randomly at only one type of nucleotide, producing a discrete fragment. Each of the fragments in the four reaction tubes will differ in length from the next largest fragment by a single nucleotide. By sizing the length of the fragments produced in each reaction on an electrophoretic gel, the sequence of the DNA can be read 5′ to 3′, starting from the bottom of the gel (the closest to the priming site) and moving upward:

5′ CGGGAGCGCTCGCTCGAGCTTTCAGA-3′

Remember that the chain termination technique requires an oligonucleotide primer of known sequence to initiate synthesis (although this is not required for the other method of DNA sequencing mentioned). For many vectors, the DNA sequences flanking the restriction enzyme sites used for DNA cloning have been determined. Short oligonucleotides complementary to these sequences are available commercially and are essentially universal primers that can be used to sequence any DNA introduced into the restriction enzyme insertion site.

With the advent of the human genome project, automated technologies to prepare and sequence DNA were perfected. All large-scale sequencing projects are now performed using robotic methodologies based on the chain termination method. The major modification in the automation of this protocol is the use of different fluorescent dyes to distinguish the termination products for each of the four bases. The termination products from each base are marked by a specific dye, each with its characteristic emission color. Since the base at the end of each termination product can be discriminated by its color, they can be run together on a single lane of an electrophoretic gel. A laser built into the electrophoresis apparatus excites each fragment as it passes down the gel and its color is detected; and the base assignment for each fragment, in succession as it passes by the laser, is recorded directly on a computer.

C. Probes for gene expression: Cloned DNA sequences are powerful tools that can be used to monitor gene expression in a particular tissue. It is possible to isolate mRNA biochemically from any given tissue type, then use nucleic acid hybridization techniques (Northern blots) to quantify the amount of that RNA that is complementary to a cloned gene probe. A variant of this technique, termed in situ hybridization, can also be applied directly to tissue sections, which can then be examined under a microscope, allowing researchers to identify individual cells within the tissue sample expressing the gene of interest. The use of nucleic acid hybridization in the localization of specific gene sequences also will be discussed in the next chapter.

The RT-PCR technique also has applications in the rapid detection and quantification of specific mRNA molecules (Q-PCR, or quantitative PCR). Using fluorescent probes that can be used to monitor experimental and control reactions simultaneously, thermal cyclers with integrated detectors have been designed to measure and quantify the kinetics of the amplification reactions as they occur in real time. This technology allows for accurate estimations of the specific starting mRNA concentrations in the sample and facilitates rapid analysis. RT- and Q-PCR are seeing broad applications in gene expression studies (e.g., monitoring changes accompanying drug treatments, oncogenic transformation, and others) and in the detection of viral and bacterial infection/contamination in clinical or commercial environments.

D. Array analysis: The computational analysis of genome sequence information, coupled with the sequencing of expressed cDNAs (ESTs) from various tissues, has led to the ability to use parallel hybridization technologies to compare the expression patterns of thousands of genes in the same experiment. DNA molecules representing specific genes are immobilized at a precise spot on a solid matrix, such as a glass slide. Depending on the type of array, these DNA molecules vary in their size and sequence complexity; they may represent an entire cDNA clone, a PCR product derived from an EST clone, or a chemically synthesized oligonucleotide based on a particular gene sequence. These DNA arrays can contain a very large amount of genetic information in a small area (microarrays), particularly if oligonucleotides are used to represent the gene sequences (more than 50,000 genes/slide). Generally, the shorter the DNA molecule used, the greater the density of the array.

RNA or cDNA isolated from a particular tissue is labeled, usually using a fluorescent dye, and hybridized to the DNA present in the array. Under the appropriate hybridization conditions, the levels of fluorescent dye bound to the array are proportional to the expression level for a given gene. If fluorescent dyes of different colors are used to label two separate mRNA populations, which are then hybridized to an array together, the relative differences in dye intensities bound to the array serve as an index of differences in expression for any given gene. In this manner, we can simultaneously compare differences in gene expression between normal and abnormal tissue, cells at different states of development, responses to changing environmental conditions, and so forth.

E. Production of recombinant proteins in microorganisms: It is possible to link the coding sequences for eukaryotic proteins to transcriptional control sequences and produce mRNAs that will be translated into proteins in microorganisms, such as bacteria or yeast. These fusion genes can be placed under the control of inducible “strong” promoters (such as the control sequences for the lactose operon) so that large amounts of protein can be synthesized and purified from recombinant cell extracts. This process has had enormous impact in the production of pharmaceuticals. Compounds of therapeutic value that are made in minuscule amounts in their host organism (insulin, growth hormones, cytokines, clotting factors) can now be synthesized in large amounts for use in the treatment of disease. In some respects, using proteins produced by recombinant DNA techniques may be safer than using the same molecules isolated from the host organism; the spread of human immunodeficiency virus (HIV) infection to hemophilia sufferers treated with pooled human plasma is one example.

F. Transgenic organisms: Standard genetic crosses generate recombinant DNA via meiosis, and this process has long been used by humans to select for desirable phenotypic characteristics. Techniques are currently being developed, however, that permit the introduction of cloned genes into the genome of a target eukaryotic organism, with the goal of influencing host phenotype. As in bacteria, introduction of the DNA is often vector mediated, although the vectors usually are capable of integrating into the host genome. The transformation events may involve introducing genes normally associated with the host genome (e.g., genes involved in fruit ripening) or genes that were never part of the genetic repertoire of the host species (e.g., genes involved in pesticide resistance, or nitrogen fixation). Creation of transgenic plants and animals has broad agricultural and industrial applications, and many examples have reached the marketplace. Like many technological breakthroughs, the ability to generate transgenic organisms has also stimulated public controversy over perceived misapplications of the technology or potential unforeseen environmental consequences.

G. Diagnosis of genetic disease and somatic gene therapy: Over the past decade, nucleic acid hybridization techniques using recombinant DNA probes have been increasingly used to identify individuals at risk for particular genetic diseases. If a mutation at a specific gene locus can be defined, differences between normal and abnormal alleles in their nucleotide sequence can be used to identify genotypes. In many cases, however, the transmission of an abnormal allele may be traced from parents to offspring by detecting differences not in the defective gene itself, but in DNA sequences closely linked to the disease gene. Variations in DNA sequence are often detected by population differences in restriction enzyme sites within a specific chromosomal region. This methodology, restriction fragment length polymorphism(RFLP) analysis, is a general tool for assessing variation in populations and will be discussed in more detail in the next chapter. If individuals at risk for genetic disease can be identified, however, can transgenic techniques be applied to rectify the defect via introduction of a functional gene? Treatment regimens and clinical trials attempting such intervention are still in their infancy but have been initiated for a number of single-gene disorders (e.g., cystic fibrosis).



If you find an error or have any questions, please email us at admin@doctorlib.org. Thank you!