Showing posts with label linkage. Show all posts
Showing posts with label linkage. Show all posts

Friday, October 17, 2014

BigData: scaling up, but no Big message

As technology has advanced dramatically during the past few decades, we have been able to look at the relationship between genotypes and phenotypes in ever more detail, and to address phenogenetic questions, that is to search for putative genomic causes of traits of interest on an ever and dramatically increasing scale. At each point, there has been excitement and a hope (and widespread promises) that as we overcome barriers of resolution, the elusive truth will be found. The claim has been that prior methods could not identify our quarry, but the new approach can finally do that. The most recent iteration of this is the rush to very Big Data approaches to genetics and disease studies.

Presumably there is a truth, so it's interesting to review the history of the search for it.

The specifics in this instance
In trying to understand the genetic basis of disease, modern scientific approaches began shortly after Mendel's principles were recognized, around 1900. At that time, the transmission of human traits, like important non-infectious diseases, was documented by the same sorts of methods, called 'segregation analysis', that Mendel used on his peas. Specific traits, including some diseases in humans or pea-color in plants, could be shown to be inherited as if they were the result of single genetic causal variants.

Segregation analysis had its limits, because while a convincing disorder might be usefully predicted, as for couples wanting to know if they are likely to have affected children, there was rarely any way to identify the gene itself. However, as methods for identifying chromosomal or metabolic anomalies grew, the responsible gene for some such disorders, usually serious pediatric traits present at birth, could be identified. For decades in the 20th century this was a matter of luck, but by the '80s methods were developed to search the genome systematically for causal gene locations in the cases of clearly segregating traits. The idea was to trace known genetically varying parts of the genome (called genetic 'markers', typed because they were known to vary not because of anything they might actually cause) and search for co-occurrence among family members between a marker and the trait. This was called 'linkage mapping' because it tracked markers and causal sites that were together--were 'linked'--on the same chromosome.

Linkage analysis was all that was possible for many years, for a variety of reasons. Meanwhile, advances in protein and enzyme biology identified many genes (or their coded proteins) that were involved in particular physiology and whose variants could be associated causally with traits like disease. A classic example was the group of hemoglobin protein variants associated with malaria or anemias. The field burgeoned, and identified many genes that were likely 'candidates' for involvement in diseases of the physiology the gene was involved in. So candidate genes were studied to try to find association between their variation, and disease presence or traits.

Segregation and linkage analysis helped find many genes involved in serious early onset disease, and a few late-onset ones (like a subset of breast cancers due to BRCA genetic variants). But too many common traits, like stature or diabetes, were not mappable in this way. The failure of diseases of interest to have clear Mendelian patterns (to ‘segregate’) in families showed that simple high-effect variants were not at work or, more likely, that multiple variants with small effects were. Candidate genes accounted for some cases of a trait, but they also failed to account for the bulk. Again, this could be because responsible variants had small effects and/or were just not common enough in samples to be detected. Finding small effects requires large samples.

So things went for some time. But a practical barrier was removed and a different approach became possible.

Association mapping is linkage mapping and it is candidate gene testing
As it became possible to type many thousands, and now millions, of markers across the genome, refined locations of causal effects became possible, but such high-density mapping required large samples to resolve linkage associations. Family data are hard and costly to collect, especially families with many members affected a given disease.

However, linkage does not just occur in close relatives, because linkage relationships between close sites on a chromosome last for a great many generations, so a new approach became possible. Variants arise and are transmitted (if not lost from the population) generation upon generation in expanding trees of descend. So if we just compare cases and controls, we can statistically associate map locations with causal variants: we know the marker location, and the association points to a chromosomally nearby causal site.

It is not widely appreciated perhaps, especially by those with only a casual background in evolutionary genetics, but the reason markers can be associated with a trait in case-control and similar comparisons is that such 'association' studies are linkage studies because we assume the sets of individuals sharing a given marker do so because of some, implicit--unknown but assumed--family connection, perhaps many generations deep. The attraction of association studies is that you don't have to ascertain all family members that connect these individuals, though of course such direct-transmission data provide much statistical power to detect true effects. Still, if causation is tractably simple, GWAS, a statistically less powerful form of linkage analysis, should work with the huge samples available.

And GWA studies are essentially also a form of indirect candidate gene studies. That's because they identify chromosomal locations that then are searched for plausible genetic candidates for affecting the trait in question. GWAS are just another way of identifying what is a functional candidate. If the candidate's causal role is tractably simple, candidate gene studies should work with the huge samples available.

But by now we all know what’s been found so far—and it’s been just what good science should be: consistent and clear. The traits are not simple, no matter how much we might wish that. All the methods, from their cruder predecessors to their expansive versions today, have yielded consistent results.

Where the advocacy logic becomes flawed 
This history shows the false logic in the claim typically raised to advocate even larger and more extensive studies. The claim is that, since the earlier methods have not answered our questions, therefore the newly proposed method will do so. But it's simply untrue that because one approach failed, some specific other approach must thus succeed. The new method may work, but there is no logical footing for asserting certainty. And there is a logical reason for doubting it.

The argument has been that prior methods failed because of inadequate sample size and hence poor resolution of causal connections. And since GWAS is in essence just scaled up candidate-gene and linkage analysis, it should ‘work’ if the problem in the first place was just one of study size. Yet, clearly, the new methods are not really working, as so many studies have shown (e.g., the recent report on stature). But there's an important twist. It isn't true that the older methods didn't work!

In fact, the prior methods have worked: They first of all stimulated technology development that enabled us to study greatly expanded sample sizes. By now, those newer, dense mapping methods have been given more than a decade of intense trial. What we now know is that the earlier and the newer methods both yield basically the same story. That is a story that we have good reason to accept, even if it is not the story we wished for, that causation will turn out to be tractably simple. From the beginning the consistent message has been one of complexity of weak, variable causation.
Indeed, all the methods old and new also work in another, important sense: when there's a big strong genetic signal, they all find it! So the absence of strong signals means just what it says: they're not there.

By now, asserting 'will work' or even 'might work' should be changed to 'are unlikely to work'. Science should learn from experience, and react accordingly.

Friday, August 29, 2014

Genomic cold fusion? Part II. Realities of mapping

Mapping to find genomic causes of a trait of interest, like a disease, is done when the basic physiology is not known—maybe we have zero ideas, or the physiology we think is involved doesn’t show obvious differences between cases and controls.  If you know the biology, you won't have to use mapping methods, because you can explore the relevant genes directly.  Otherwise, and today often, we have to go fishing, in the genome, to find places that may vary in association—statistical regularity—with the trait.

The classical way to do this is called linkage analysis.  That term generally refers to tracing cases and marker variants in known families.  If parents transmit a causal allele (variant at some place in the genome) to their children, then we can find clusters of cases in those families, but no cases in other families (assuming one cause only).  We have Mendel’s classical rules for the transmission pattern and can attempt to fit that pattern to the data—for example, to exclude some non-genetic trait sharing.  After all, family members might share many things just because they have similar interests or habits.  Even disease can be due to shared environmental exposures. Mendelian principles allow us, with enough data, to discriminate.

“Enough data” is the catch.  Linkage analysis works well if there is a strong genetic signal.  If there is only one cause, we can collect multiple families and analyze their transmission patterns jointly.  Or, in some circumstances, we can collect very large, multi-generational families (often called pedigrees) and try to track a marker allele with the trait across the generations.  This has worked very well for some very strong-effect variants conferring very high risk for very specific, even quite rate, disorders.  That is because the linkage disequilibrium—the association between a marker allele and a causal variant due to their shared evolutionary history (as described in Part I) ties the two together in these families.

But it is often very costly or impractical to collect actual large pedigrees that include many children each generation, and multiple generations.  Family members who have died cannot be studied and medical records may be untrustworthy, or family members may have moved, refuse to participate in a study, or be inaccessible for many reasons.  So a generation or so ago the idea arose that if we collect cases from a population we may also collect copies of nearby marker alleles in linkage disequilibrium—shared evolutionary history in the population—so that, as described in Part I, a marker allele has been transmitted through many generations of unknown but assumed pedigree, so that the marker will have been transmitted in the pedigree along with the causal variant.  This is implicit linkage analysis, called genomewide association analysis (GWAS), about which we’ve commented many times in the past.  GWAS look for association between marker and causal site in implicit but assumed pedigrees, and is another form of linkage analysis.

When genetic causation is simple enough, this will work.  Indeed, it is far easier and less costly to collect many cases and controls than many deep pedigrees, so that a carefully designed GWAS can identify causes that are reasonably strong.  But this may not always work, when a trait is ‘complex’, and has many different genetic and/or environmental contributing causes.

If causation is complex, families provide a more powerful kind of sample to use in searching for genetic factors.  The reason is simple: in general a single family will be transmitting fewer causal variants than a collection of separate families.  Related to this is the reason that isolate populations, like Finland or Iceland, can in principle be good places to search, because they represent very large, even if implicit, pedigrees.  Sometimes the pedigree can actually be documented in such populations.

If causation is complex, then linkage analysis in families will hopefully be better than big population samples for finding causal contributors, simply because a family will be segregating (transmitting) fewer different causal variants than a big population.  We might find the variant in linkage analysis in a big family, or an isolate population, but of course if there are many different variants, a given family may point us only to one or two of them.  For this reason, many argue that family analysis is useless for complex traits—one commenter on a previous Tweet we made from our course, likened linkage analysis for complex traits to ‘cold fusion’.  In fact, this was a mistake and is incorrect. 

Association analysis, the main alternative to linkage analysis, is just a combining of many different implicit families, for the population-history reason we’ve described here and in Part I.  The more families you combine, whether they are explicit or implicit, the more variation, including statistical ‘noise’, you incorporate.  The rather paltry findings of many GWAS are a testament to this fact, explaining as they have only a small fraction of most traits to which that method has been applied.  Worse, the greater the sample of this type, like cases vs controls, the more environmental variation you may be grouping together, again greatly watering down even the weak signal of many or, probably, by far most genetic causal factors.

In fact, if you are forced to go fishing for genetic cause, you may well be fishing in dreamland because you may simply be in denial of the implications of causal complexity.  In fact, all mapping is a form of linkage analysis.  Instead, one should tailor one’s approach to the realities of data and trait.  Some complex trait genes have been found by linkage analysis (e.g., the BRCA breast-cancer associated genes), though of course here we might quibble about the definition of 'complexity'. 

Sneering at linkage analysis because it is difficult to get big families, or  because even single deep families may themselves be transmitting multiple causes (as is often found in isolate studies, in fact), is often simply a circle-the-wagon defense of Big Data studies, that capture huge amounts of funding with relatively little payoff to date.

A biological approach?
Many linkage and association analyses are done because we don’t understand the basic biology of a trait well enough to go straight to ‘candidate’ genes to detect, prevent, or develop treatment for a trait.  Today, even though this approach has been the rule for nearly 20 years now, with little payoff, the defense is often still that more, more and even more data will solve the problem.  But if causation is too complex this can also be a costly, self-interested, weak defense.

If we have whole genome sequence on huge numbers of people, or even everyone in a population, or in many populations so we can pool data, that we will find the pot of gold (or is it cold fusion?) at the end of the rainbow.

One argument for this is to search population-wide genome sequenced biomedical data bases for variants that may be transmitted from parents to offspring, but that are so rare that they cannot generate a useful signal in huge, pooled GWAS studies.  This usually will still be in the form of linkage analysis if a marker in a given causal gene is transmitted with the trait in occasional families but the same gene is identified, even if via different families.  That is, if variation in the same gene is found to be involved in different individuals, but with different specific alleles, then one can take that gene seriously as a causal candidate.

This sometimes works, but usually only when the gene’s biology is known enough to have a reason to suspect it.  Otherwise, the problem is that so much is shared between close family members (whether implicitly or explicitly in known pedigrees) that if you don’t know the biology there will be too much to search through, too much co-transmitted variation.  Causal variation need not be in regular ‘genes’, but can be, and for complex traits seems typically to be, in regulatory or other regions of the genome, whose functional sites may not be known.  Also, we all harbor variation in genes that is not harmful, and we all carry ‘dead’ genes without problems, as many studies have now shown.

If one knows enough biology to suspect a set of genes, and finds variants of known effect (such as truncating a gene’s coding region so a normal protein isn’t made) in different affected individuals, then one has strong evidence s/he has found a target gene.  There are many examples of this for single-gene traits.  But for complex traits, even most genes that have been identified have only weak effects—the same variant most of the time is also found in healthy, unaffected individuals.  In this case, which seems often to be the biological truth, there is no big-cause gene to be found, or a gene has a big-cause only in some unusual genotypes in the rest of the genome.

Even knowing the biology doesn't say whether a given gene's protein code is involved rather than its regulation or other related factors (like making the chromosomal region available in the right cells, downregulating its messenger RNA, and other genome functions).  Even in multiple instances of a gene region, there may be many nucleotide variants observed among cases and controls.  The hunt is usually not easy even knowing the biology--and this is, of course, especially true if the trait isn't well-defined, as is often the case, or if it is complex or has many different contributors.

Big Data, like any other method, works when it works.  The question is when and whether it is worth its cost, regardless of how advantageous for investigators who like playing with (or having and managing) huge resources.  Whether or not it is any less ‘cold fusion’ than classical linkage analysis in big families, is debatable.  

Again, most searches for causal variation in the genome rest on statistical linkage between marker sites and causal sites due to shared evolutionary history.  Good study design is always important.  Dismissal of one method over another is too often little more than advocacy of a scientist’s personal intellectual or vested interests.

The problem is that complex traits are properly named:  they are complex. Better ideas are needed than what are being proposed these days.  We know that Big Data is ‘in’ and the money will pour in that direction.  From such data bases all sorts of samples, family or otherwise, can be drawn.  Simulation of strategies (such as with programs like our ForSim that we discussed in our recent Logical Reasoning course in Finland) can be done to try to optimize studies. 

In the end, however, fishing in a pond of minnows, no matter how it’s done, will only find minnows. But these days they are very expensive minnows.