The Family Tree Guide to DNA Testing and Genetic Genealogy

12

The Future of Genetic Genealogy

As with computers and the Internet, DNA has become an important component of modern genealogical research. Although this tool only became available a few short years ago, it is now a source of evidence for thousands of genealogists and of fascination and excitement for millions of test-takers. Over the course of the next ten to twenty years, new advancements in DNA technology and techniques will change and expand how genealogists obtain DNA test results and how those test results are applied to genealogical questions. In this chapter, we will peer into our crystal ball to look at current trends in DNA technology and see how they will affect genetic genealogy.

The Future of Y-DNA Testing

One area of genetic genealogy that is likely to change considerably over the next decade is Y-chromosomal (Y-DNA) testing. Currently, Y-DNA testing comprises either a handful of Y-STRs (short tandem repeats on the Y chromosome), usually between 37 and 111, or a few thousand Y-SNPs (single nucleotide polymorphisms on the Y chromosome). Some Y-DNA tests, such as the Big Y test from Family Tree DNA <www.familytreedna.com>, examines approximately 15 to 25 million base pairs of the Y chromosome, which is just 25 to 45 percent of the entire Y chromosome, much of which potentially contains information about ancestry, although these tests are only just beginning to yielding information that will be helpful to genealogists.

While future researchers would ideally sequence the entire Y chromosome and analyze it for ancestral information (including identifying and characterizing new STRs and SNPs), the Y chromosome presents some unique challenges that prevent affordable or accurate whole Y-DNA sequencing with current technology. For example, much of the Y-chromosome is either highly repetitive, palindromic, or nearly identical to the X chromosome. Since current DNA sequencing technology sequences many short overlapping fragments (called “reads”) of a chromosome and then pieces them back together by mapping them to a human genome reference, repetitive or palindromic sequences can make this difficult (if not impossible).

New sequencing technology, however, might provide testing companies like Family Tree DNA with new opportunities. For example, DNA sequencers that obtain very long high-quality reads will be much better equipped to piece those long reads together. Ideally, sequencing technology in the future might start at one end of a chromosome and, with one single read, sequence all the way to the other end of the chromosome.

Once raw data is available, analyzing the data to extract the STR and SNP information is an easy step in the process. Just as the influx of Big Y results from Family Tree DNA in 2015 resulted in the so-called “SNP tsunami,” the data flowing in from future sequencing technology will result in a massive amount of information that will have to be analyzed and categorized in the context of the human Y-DNA family tree. There will likely be numerous “family-specific SNPs” within these results—Y-SNP variations that are found among men who have a recent shared paternal ancestor (within about the past 100 to 250 years).

There is a strong economic pressure to create improved DNA sequencing technology, including the desire to use low-cost DNA testing for health assessment and treatment. As a result (despite some technical setbacks), Y-DNA sequencing developments will likely occur within the next five to ten years.

The Future of mtDNA Testing

Since the entire mtDNA molecule is already sequenced by current testing, there will likely be only a few developments in mtDNA testing in the future. It is not possible to extract any additional information from the sequence of A, T, C, and G’s in the mitochondrial genome.

Rather, the biggest development in mtDNA testing will likely come from a greater number of people testing, meaning that the likelihood of finding a meaningful match will increase. Although many people find matches when they test and for the reasons we learned in the mtDNA chapter (namely that mtDNA mutates very slowly and thus many, many people have the same mtDNA), it is very rare that a test-taker finds a meaningful match that helps them with their genealogical research. Additionally, mtDNA is often associated with a shorter pedigree because of the challenges of researching a maternal line with a name change every generation. As the database gets larger however, the likelihood of finding an important close match that fits within your pedigree increases significantly.

Another important development in mtDNA testing might be epigenetic testing. Like atDNA, mtDNA is packaged into an organized structure with proteins and chemical groups that associate with it. If this epigenetic structure is heritable, as recent studies suggest it is, then it could be analyzed and exploited for genealogical purposes. People who are more closely related would be expected to have more similar epigenetic structure, and thus the epigenetic structure of mtDNA might single out close relatives among a list of mtDNA matches. Epigenetic testing will be explained in much greater detail later in this chapter.

The Future of atDNA Testing

The biggest changes in genetic genealogy are expected to occur in the area of autosomal-DNA (atDNA) testing, for some of the same reasons as Y-DNA testing. Namely, new developments in DNA sequencing will both improve sequencing and lower the cost of testing.

When whole-genome sequencing reaches a price point similar to current atDNA testing, genealogists will be a driving force behind the use of whole-genome results for genealogy, including improved cousin identification and relationship estimation. Whole-genome testing will provide some benefits to cousin identification and relationship estimation, although it likely won’t identify new close cousins (closer than fourth cousin, for example) who couldn’t be identified with current testing. In a similar way, an affordable whole-genome test for genealogists likely won’t provide relationship predictions that are stunningly accurate, instead providing new distant cousins and improving the confidence in relationship estimates.

In addition to whole-genome sequencing, new atDNA methodologies will be created in the next decade. These methodologies will not only be the result of research and experimentation by scientists and genetic genealogists, but will also be possible due to the sheer sizes of atDNA databases. Once the databases comprise many millions of people, new tools can be developed that were either not apparent when the database was smaller or was not possible using a smaller database.

Genetic Reconstruction: Piecing Together Genomes of the Dead

Piecing together a portion or all of the genomes of our ancestors will enable genealogists to learn things about them that we might otherwise not be able to, such as their ethnicity, health, and recent genealogical relationships. Genealogists might also learn about some of their physical characteristics like eye color and hair color, although these are not always a perfect estimate based on DNA alone. In this section, we’ll discuss what genetic reconstruction is, how it works, and why it may be of interest to genealogists.

Genetic reconstruction is made possible by testing many descendants of an ancestor or ancestral couple. For example, imagine John and Jane Smith, living in New England in the mid-1700s. They had twelve children, ten of whom lived to adulthood, and thus they now have many, many thousands of descendants living today. A handful of these descendants possess random segments of DNA handed down from John and Jane; the more children and descendants the ancestral couple had, the more DNA from that couple that is likely to exist in test-takers today.

Segments of DNA that potentially came from this ancestor or ancestral couple can be identified using well-researched family trees, then woven together to create as much of the couple’s genome as possible. Some segments will be lost forever, although often these missing segments can be derived or estimated.

In image A, segments of DNA from a couple in the mid-1700s are found in living descendants. Some descendants, such as #4, either did not inherit any DNA from the couple or does not share any of the inherited DNA with another relative, and thus will likely not contribute to the genetic reconstruction. Others descendants—or relatives—such as descendant #6 may have no documentation that they are related to the couple. Only segments that can be reliably assigned to the ancestor or ancestral couple will be mapped. Usually, this will involve identifying segments of DNA shared by two or more descendants of the ancestor or ancestral couple.

By looking several of a couple’s descendants, you can cobble together an approximation of their DNA makeup. However, only some of these descendants will contain the original couple’s DNA; here, descendant #4 doesn’t have this ancestral DNA and so will not be helpful for these purposes.

Slowly, the genomes of hundreds and possibly even thousands of early ancestors may be generated as millions of DNA samples and family trees are entered into massive databases. There will undoubtedly be numerous errors introduced by both poor quality trees—or trees that are incorrect due to misattributed parentage events—and from improperly assigning shared segments to one ancestor versus another. However, most of these errors will be resolved over time as more samples and trees are entered into the system and processed and as genealogists and citizen scientists pour over the trees.

Interestingly, these recreated genomes sometimes will belong to unknown or unidentified ancestors (DNA-Only Ancestors), as discussed later. For example, shared segments of DNA may appear to be from a couple who probably lived in the area of Boston and had children in the early 1700s, but no known couple can be found in existing records.

To be successful, this future methodology will of course require numerous, wide-ranging, and extremely well-researched family trees, as well as DNA samples from several millions of individuals.

With reconstructed genomes, it might also be possible to estimate what our ancestors looked like, even if no picture of that ancestor was ever taken or has survived. Everyone knows the age-old game of guessing which parent or sibling a new child looks like, or guessing which identical twin is which. These scenarios demonstrate the existence of a relationship between DNA and appearance. Accordingly, by examining and understanding this relationship, we can theoretically predict appearance based on DNA alone.

For example, in 2014, scientists published a study that identified twenty-four gene variants across twenty genes that affect facial structure <journals.plos.org/plosgenetics/article?id=10.1371/journal.pgen.1004224>. The researchers then used DNA profiles from volunteers to create approximations of the volunteer’s facial structure. In addition to having numerous potential forensic and law-enforcement applications, this technology could assist genealogists who are recreating the appearance of long-dead ancestors. Facial structure estimates could be combined with other physical information mined from the DNA sequence, including eye color, hair color, height, skin color, and other physical characteristics, to create a composite image of the ancestor. The DNA information could also be supplemented with cultural and socioeconomic information to predict hair styles and other features.

In another example, AncestryDNA <dna.ancestry.com> announced in 2014 that it had successfully recreated significant fragments of the genome of David Speegle (1806–1890) and his two wives, Winifred Cranford and Nancy Garren. This was accomplished by analyzing the DNA of hundreds of Speegle’s descendants through his twenty-six children and piecing together shared segments of DNA using two different methods. With many children between the two marriages during his lifetime, David and his spouses were excellent candidates for reconstruction given the number of living descendants who all potentially carry a piece of their DNA. Indeed, according to Speegle’s obituary in 1890, he had at least three hundred descendants at the time of his death, suggesting why DNA from David Speegle and his wives was so prevalent in the AncestryDNA database.

Using these recreated partial genomes, AncestryDNA analysts learned that David or one of his wives had a gene variant that increases the likelihood of male pattern baldness, and that David had at least one copy of the gene variant for blue eyes. See <blogs.ancestry.com/techroots/ancestrydna-achieves-scientific-advancement-in-human-genome-reconstruction/> for more information.

Generating Family Trees

So what could atDNA advancements do for your documented family research? In theory, once the genomes of hundreds or thousands of seventeenth-, eighteenth-, or nineteenth-century ancestors are created and collated into a massive family tree, they can be used to recreate portions of the family tree of modern-DNA test-takers using just the results of a DNA test. This is done by first identifying potential ancestors based on the results of an atDNA test, then by fitting those identified ancestors into a family tree for the test-taker. This might be augmented, for example, by any known genealogy for the test-taker.

Family tree prediction or reconstruction is possible because identified ancestors will only fit together into a family tree in a limited number of ways. For example, let’s assume you’ve taken an atDNA test and the testing company has identified twenty ancestors using only its database of reconstructed ancestors and your DNA test results. Statistically speaking, there are a limited number of ways to collate those twenty ancestors into a single family tree; only a limited number of lines of descent lead from all of these twenty different ancestors to you. A future service could suggest having a certain relative tested (e.g., “We suggest having a descendant of your great-grandfather—your second cousin—tested to further refine your reconstructed family tree”), or ask you a series of questions to more accurately resolve the conflicts in the tree (“What was your maternal grandmother’s name?” “What was your great-grandmother’s name and date of birth?,” and so on). By asking for user feedback, the program could select the most likely path of descent from your twenty ancestors to you and construct a likely family tree based on that information.

In the example shown in image B, Client B6429 possesses segments of DNA from four different reconstructed genomes. This information is used to create a reverse or reconstructed family tree with the identified ancestors mapped to it in the most likely configuration based on the size of the segments, established genealogies, and several other factors.

Future test-takers, like Client B6429 could theoretically construct family trees based on the segments of DNA they hold, inheritance patterns, and other factors.

This reconstructed family tree process could also be used alongside traditional genealogical research, and in fact is used for identifying the families of adoptees. For example, a genealogist who knows a client is descended from ten people could easily recreate the probable family tree of that individual. Rather than creating a family tree entirely from scratch, the program pieces together portions of existing family trees in the database to generate possibilities for the client. While the most recent three to five generations might have to be filled in by the client (since these generations are the least likely to be included in the company’s database), much of the tree could be completed based only on the DNA results.

There are, of course, many caveats with this method, and any computer-generated family tree should be confirmed by traditional research. Poor-quality trees, for example, will present a major challenge to this process, although they will not provide a complete stumbling block. Indeed, DNA evidence will likely ameliorate poor-quality trees to a significant extent. In addition to poor-quality trees, well-researched and well-documented family trees can be incorrect due to otherwise undetectable misattributed parentage events such as adoption, name change, or infidelity. These trees can also be detected and analyzed using the methods described above.

In addition, possessing DNA from a reconstructed genome does not automatically mean that the test-taker is descended from the person who possessed that genome. Instead, the test-taker may only be related to that person. For example, the test-taker may descend from the relatively unknown brother of John Smith who has very few living descendants, instead of being a descendant of John Smith himself. The test-taker could still possess any of the DNA hypothesized to be from John Smith. While the methods described above will likely ultimately focus on characterizing the genomes of “branch point ancestors” (such as immigrants, founders of unique haplotypes, and others) to avoid this problem, the existence of unknown or lesser known relatives of the branch point ancestor will temporarily throw a monkey wrench in the process.

Despite these caveats, ancestor and family tree reconstruction is likely to have an enormous impact on genealogical research in the next few decades, providing valuable ancestral information to users (especially adoptees).

Creating DNA-Only Ancestors

In the very near future, DNA from genetic cousins will be used to recreate the genomes of unknown ancestors who reside completely behind brick walls. While traditional research will often be able to provide a potential identity for the recreated genome, the individual will sometimes forever be known only by his reconstructed DNA. The DNA of these “DNA-Only Ancestors” is dispersed among living descendants, and some of it is already found within the testing companies’ databases.

As we saw with the Speegles, it is possible to recreate at least a portion of an ancestor’s genome if there are enough descendants. This process is greatly simplified if the family trees of those descendants are known and well researched. However, it will still be possible to use the DNA of descendants of an ancestor who is not known to recreate portions of that unknown ancestor’s genome.

Let’s assume, for example, that a group of individuals trace their particular family tree to Akron, a small town in upstate New York in the early 1800s. The lines all end there with no known shared ancestor, and there are no traditional paper clues or shared surnames. However, all of these families are genetically related to one another, and based on their extensive research, they don’t appear to share any other lines. If traditional research is exhausted, how can these relatives learn about their common ancestor?

Using the most advanced atDNA techniques, the DNA shared among the descendants could be assigned to an ancestor or ancestral couple (image C). The recreated partial genome will then provide other information about the DNA-only ancestor, such as predicted eye color, hair color, medical conditions, and traits. It could also be used to find other descendants or relatives.

Shared DNA from suspected descendants could one day be combined to piece together information about an ancestral couple.

Indeed, once a potential DNA-only ancestor is identified, some clues could help identify the name of the ancestor, such as inherent phenotypic information (a medical problem, for example) or other relatives who now show as a match because of the recreated genome (a Johnson with a strong oral record for a family in the same town, for example). Seeing a scholarly article in a national genealogy journal with the title “Pinpointing the Likely Identity of a DNA-Reconstructed Ancestor in Akron, New York,” is not as far away as you might think.

While it will be ideal to identify the name and family of the DNA-only ancestor, for many it will be impossible. This is especially true for regions and time periods where records are too sparse for such identification, such as eighteenth- and nineteenth-centuries in Ireland, African-American ancestry, Native American ancestry, and so on. For each of these regions, there will be many different DNA-only ancestors. While we may not know their names, we can fill in those gaps with any pieces of information we do have or complement the DNA-only identification with historical events or the life of a common individual in that time frame. For example, the profile for a DNA-only ancestor might look something like this:

· AkronNY-1800s-Male-1: Likely lived in Akron, South County, New York, between approximately 1800 and 1820. Had at least three children, probably daughters. Akron was first settled in 1797, so AkronNY-1800s-Male-1 was likely an early settler of the town, which prospered during its first two decades. Earliest known descendants are grandchildren Susannah (Unknown) Smith, Rebekah (Unknown) Mullen, and Sarah (Unknown) Johnson. AkronNY-1800s-Male-1 was of Irish descent and had blue eyes.

Although this example focuses on someone who was suspected of existing in a particular time and place, it will also be possible to recreate the genomes of individuals who were previously completely unknown to history and for which there are no paper or oral records of any kind.

Epigenetic Testing

All current DNA testing for genealogy looks at the sequence of A (adenines), T (thymine), C (cytosines), and G’s (guanines) along a chromosome or the mitochondrial genome. However, DNA comprises significant amounts of information beyond the order of A, T, C, and G’s. For example, in order to be a manageable size within the nucleus of the cell, DNA is packaged into a tight structure called chromatin, a complex bundle of DNA and packaging proteins. Some of the chromatin—called heterochromatin—is highly packaged and not actively used by the cell. Other portions of the chromatin—called euchromatin—is less tightly packaged and can be used by the cell. The packaging proteins themselves can be modified to affect how tightly packaged, and thus how active, portions of the genome are. In addition, the DNA itself can be tagged with chemical groups such as “methyl groups” that affect the accessibility and/or activity of that DNA (image D). Together, these epigenetic mechanisms have a direct and important impact on DNA activity.

Epigenetic mechanisms, which control how genetic information is bound together in methyl groups that are arranged into chromatin, are beyond the scope of this book. However, future advancements in technology could uncover if and how this information can be useful to genealogists. Courtesy the National Institutes of Health.

Recent research has shown that at least some of the epigenetic structure of the DNA is likely inherited from one generation to the next <www.discovermagazine.com/2013/may/13-grandmas-experiences-leave-epigenetic-mark-on-your-genes>. For example, preliminary studies have suggested that people who experience trauma as a child—or had a parent or grandparent who experienced trauma—have different epigenetic profiles than those who did not experience similar trauma.

If epigenetic structure of DNA is indeed inherited, it can be utilized for genetic genealogical analysis. Although it is currently unclear whether epigenetic structure is stable and inherited for more than a handful of generations, it appears from current research that this epigenetic structure will at least be useful for examining recent and close relationships. For example, when determining whether someone is an aunt or half-sibling—or a third cousin versus a second cousin twice removed—epigenetic information might provide enough additional information to differentiate between possible relationships.

Additionally, epigenetic information might be useful to learn about the life experiences of our ancestors. There may be epigenetic markers of trauma (or of contentment) that have been inherited through generations. Once we assign a segment of DNA to an ancestor, for example, we might be able to characterize the epigenetics of that segment of DNA to learn about that ancestor or ancestral line.

Whether it is used only to differentiate between close cousin relationships or to learn about the life experiences of our ancestors, epigenetic analysis will almost certainly form an important part of genetic genealogy testing in the near future.

Other Advances in Genetic Genealogy

The developments discussed above are just some of the potential new tests or techniques that will benefit genealogy over the next few years. Of course, other improvements that will occur are impossible to predict.

Another, less predictable potential development in the field of genetic genealogy is improved cousin identification and characterization using both new testing technology such as whole-genome sequencing and massive family trees linked to DNA. For example, in the near future, cousins will likely be identified not only as a potential cousin, but as a statistically likely genealogical relationship, complete with an identified common ancestor. AncestryDNA’s New Ancestor Discoveries attempts to do this, although it is a very early version.

Another area of development will be the increased use of DNA by lineage societies. As DNA becomes an increasingly important piece of genealogical evidence, it will continue to be adopted as evidence by lineage societies. The Daughters of the American Revolution, for example, announced a re-vamped DNA policy in 2013 that allows for limited uses of Y-DNA www.dar.org/national-society/genealogy/dna-and-dar-applications>. Other organizations are also allowing Y-DNA evidence to prove or support membership claims. In the future, however, results from other types of DNA tests might be available to create potential lineage lines for society members, particularly as genetic genealogists continue to demonstrate the power and efficacy of DNA testing for genealogy. Additionally, at some point, DNA evidence alone might be sufficient for membership, either because the DNA adequately establishes descent from an ancestor who qualifies for membership or because the lineage society itself is based on the DNA of particular ancestors.

CORE CONCEPTS: THE FUTURE OF GENETIC GENEALOGY

Genetic genealogy is still a new field of scientific research. Future developments in both DNA testing technology and DNA analysis methodologies promise to add important new information to genealogical research.

Improvements in Y-DNA sequencing will enable the discovery of new genealogically relevant Y-STRs and Y-SNPs to test and could enable even more refined estimates of paternal relationships.

Affordable whole-genome sequencing of atDNA will allow for refined relationship predictions.

Genealogists will recreate significant portions of their ancestors’ genomes, revealing information about their lives, health, and physical appearance. Eventually, genealogists will be able to estimate the facial structure of ancestors.

Some of these recreated genomes will belong to DNA-only ancestors who do not have a name or identity associated with them due to a lack of traditional genealogical records.

Genealogists will be able to reconstruct portions of family trees from just the results of a DNA test, in conjunction with vast databases that combine family trees and DNA.

Genealogists will use epigenetic testing to examine recent genealogical relationships and possibly to learn about the lives of our ancestors.

Lineage societies will increase their acceptance of DNA evidence, and some may even rely entirely on DNA evidence.



If you find an error or have any questions, please email us at admin@doctorlib.org. Thank you!