Showing posts with label rare variants. Show all posts
Showing posts with label rare variants. Show all posts

Tuesday, November 11, 2014

Genomics: finding a lid to fit one's kettle?

I recently read Laurence Sterne's Tristram Shandy, and this led me to begin re-reading one of the books that was a precursor to the jumbled, chaotic, but often hilarious adventures of Tristram, namely, the 16th century Gargantua and Pantagruel by Francois Rabelais.  In the Preface to Book I, I noticed that Rabelais spoke of one Friar Lubin, who went to great lengths "to find a lid to fit his kettle."



The context was Rabelais' argument that a lot of retrospective meanings were often assigned, by presumed sages, to the works of the classics such as Homer and Ovid.  Rabelais' idea was that these authors wrote wonderful stuff but afterwards scholars combed through it to find subtle meanings that were never really there.  The relevance to science, as I interpret this, is to the widespread, perhaps quite natural tendency for investigators convinced something is true to force results into interpretation consistent with that conviction.  Could this lead geneticists to make more of a specific mapping-based genome location of functions that are not really there, and thus to be distracted from functions that might be more important?

A century of work has found many normal and disease traits that are tractably genetic in the sense that one or at most a few, or a choice of one among a few, identifiable genetic loci are responsible.  But simple genetic causation is far from what is being routinely promised as the more general case, especially for the common, important complex traits that are the major public health problems, and a main target of genomics today.  For those traits, rather than single genes, tens to even thousands of different genome regions are being found to have statistically detectable association with the traits, usually detectable only in huge samples and/or with individually very small effects.  Yet such findings are commonly claimed as triumphs, and the investigators go to great lengths to find in them a lid to fit their kettle.

One possible current example of this kind of Procrustean approach is a paper in the current issue of Cell ("Lessons from a Failed γ-Secretase Alzheimer Trial," De Strooper, Cell, Nov 6,  2014).  Protein complexes called γ-secretases have been thought by various criteria related to the amyloid plaques associated with Alzheimer disease, to be likely candidates for inhibitors to have therapeutic effects, but a directly relevant drug trial study found some negative consequences but failed to find the positive effect.  The authors of the Cell paper argue that this 'No' actually means 'Yes' if you just let us continue the research:  "This pessimism is unwarranted: analysis of available information presented here demonstrates significant confounds for interpreting the outcome of the trial and argues that the major lessons pertain to broad knowledge gaps that are imperative to fill."

Strooper presents a vigorous and technically specific set of arguments, and he may be right, of course, but even if so in this case, No-means-Yes arguments are seen rather more often than any actual beef.  It's easy to make fun of, and indeed if there is good plausibility evidence, a single study, especially with statistically-based inference, may not be a definitive refutation of an idea.  If one has what seems like a good idea it is natural and right not to give up on it too easily.  But the frequency of this persistence and the rather typical lack of strong follow-up confirmation at least raises serious questions about our criteria for inference and for giving up on an idea that isn't panning out.

If we choose in advance some significance level, say p = 0.05, as a cutoff for finding a signal, and we design a sample that according to our model should be able to detect an effect of the size we expect, but the study arrives at a p-value of, say, 0.06, in technical terms we should abandon our hypothesis, but of course we usually don't.  We call 0.07 'suggestive' and press ahead with our hypothesis.  This seems like cheating, and in a sense it is.  But in a deeper sense, if we realize the arbitrariness of all our inferential criteria (parsimony, falsifiability, significance....) then we realize that inference is a subjective kind of collective sense of acceptance (or not) of hypotheses.  In that light, the γ-secretase Carry On Regardless attitude may not be so wrong--even if it shows that belief, not just objectivity, is important in sciences like genomics that would fancy themselves rigorously objective.

Another example is the search that some investigators are making to find rare rather than common variants causing disease.  In principle this makes sense since most variants in the human genome are rare and this may be especially true of harmful variants because evolution (natural selection) will on average work against them.  There are various techniques for a rare-variant approach, such as finding a given gene in which different sequence variants are seen in different cases of the disease.  This is persuasive, not in the sense that the nature of the specific variants themselves shows why they are pathogenic, but because multiple observations of the same gene at least suggests it might be causal.  Historically, once relatively common variants were used to map causation of some pediatric traits, like PKU or Cystic Fibrosis, subsequent sequencing of the gene in patients has found a large variety of different variants--typically hundreds!--that are themselves too rare to generate statistical association on their own.  If we can now assume that the gene is the cause, then we can infer that the newly found mutations are causal.  That is an assumption that can be questioned, because under the assumption many strange seemingly innocuous variants (e.g., in noncoding or intronic or synonymous sites) are blamed as being causal.

Another tactic for attributing cause to rare variants is to find the same variant in affected relatives, especially a parent and offspring.  This plausibly appears as high-penetrance (Mendelian dominant) inheritance.  Of course, roughly half of all sequence variants found in any parent will be found in any given offspring, but if there is functional, experimental, or other substantial reason to suspect a particular gene, or the inherited rare variant seems culpable (e.g., a premature stop codon), then such transmission would seem to be at least plausibly convincing.  This seems to be widely accepted logic, but is it right?

Fitting data to prior ideas: Procrustean beds or lidless kettles?
The answer is, undoubtedly sometimes, but probably in most cases not really.  How can that be?  The reason, if not the trait, is simple: genetic variants have their effects only in their environmental and genomic context.  If a given variant is not always seen in association with a trait, or there is no particular known functional reason to 'blame' a given genome location for the trait, then one has to ask why one can make a causal assumption.  We know from many mapping studies by now that variant-specific risks are usually very small, often detectable only in huge samples.  In other words, by far most people carrying the variant don't get the disease, so it's a tad strange to think of it as a 'causal' finding. 

Even genes in which known variants with clearly very strong unquestioned effect are widely accepted (e.g., major mutations in the BRCA genes in relation to breast cancer).  But even then the risk estimated from samples is neither 100%, nor similar across cohorts.  Something differs among affected carriers of these variants, and that is context.  The context  is either environmental or genomic. In the case of BRCA, the genes are thought to function to detect mutations in the cell and stimulate their correction or to kill the cell.  Their role in cancer is that in a meaningful sense the tumor is caused by other variants in the genome, not BRCA itself.  But lifestyle factors somehow seem clearly also to be involved--how would that be if the BRCA is a mutation-repair related gene?

Parent-offspring transmission of rare variants certainly may indicate that they play some role in the outcome, but it's possibly (perhaps likely?) because of other genetic (or environmental) co-conditions in the individuals.  Offspring inherit much besides a single variant from their parents, after all. 

Various studies, based on DNA sequence analysis, have by now shown that we each typically carry around tens or more defunct or seriously damaged genes.  The variants may be pathogenic in some individuals but not in others.  The reason again must be context, that is, something other than the gene itself.  If not, it is some probabilistic aspect of causation about which we can usually only speculate (or assume without even a guess about mechanism)--or simply use 'probability' as a fudge factor to make our story seem scientifically convincing.

Weak signals may not be low fruit, but pointers to elsewhere
Ironically and oddly, finding rare variants in various individuals or finding variants common enough to generate a statistically significant association test but with only low relative risk may mainly mean that the bearers also carry other risk factor(s) that made the target variant 'causal' in the few observed cases.  Most of the time--in most contexts--there seems to be no excess risk, or else one would expect the gene to be easily identified even in modest samples, as CF and PKU and many other traits were.  Small effect is what small relative risks mean, and small relative risks are by far the rule in mapping studies.

Indeed, claiming success by forcing the conclusion that the identified gene is 'the' cause in these individuals, even parent-offspring pairs, may be another way of finding lids to fit investigators' kettles.  Again, if there is a conclusion, it might better be that when small-effect variants are found, it is the context of the rest of the genome (plus life experience of the cases) that are as key to understanding the trait as the target 'hit' site itself.  The discovered hit may be involved, but mainly acting as a pointer to some other factor(s) that really account for the effect.  If the identified gene itself is so important, why do we only identify a few rare variants in that gene associated with risk, even if transmitted in families?  That is, why don't we see some higher-frequency mutations in the same gene as we do with many of the other largely single-allele traits?


These questions apply even to those who argue that finding these cases, the 'low hanging fruit' as such things are often called, is a worthy objective that we can attain, even in the face of complexity.   Of course, there are population genetic (evolutionary history) reasons why this may be so, since variant frequencies are affected by chance among other things.  And finding a cherry is not evidence against it's involvement.  When an inactivating variant is found to be transmitted, this is certainly plausibility evidence worth following and, after all, many single-gene disorders have been identified once there is a clear-enough trail to follow.  Still, even knockout mouse confirmations are not always definitive support by any means, and as we noted above, healthy people may harbor as many 'bad' genetic variants as those affected. But if the finding is confirmed, then of course therapeutic approaches can be contemplated.

However, the great lack of clear therapeutic consequents of the vast majority of GWAS-like findings is consistent with the idea that the target site is in truth mainly pointing us to other things that are what we need to know about. That is, thinking of the picked cherry as really causal may be a mistaken way to interpret genomic data, even if the cherry is a small part of the story.  If this is being too critical then it is only to match the predominant view which is being too promotional.

An upside of these ideas could be to lead investigators to take the context-dependent aspect of such findings more seriously and see what else may be accompanying the rare variant in question, or what it may interact with.  There are, of course, efforts to do this, but it is not an easy problem, because such follow-ups lead back into the web of complexity; but perhaps using these situations as entry points we can find some order there.

To make more of it than that may suggest that oftentimes we've decided ahead of time what sort of kettle we have, we will fit what lids we find to it.

Tuesday, August 21, 2012

The rare variant safety valve

We are desperate to find a genetic cause, or at least a tractably small number of genetic causes of every trait we want to study.  It's understandable that we want this kind of answer.  Unfortunately, GWAS has not accounted for much of the heritability (estimated genetic fraction of causation) of most diseases and other similarly complex traits that have been studied, as we've often pointed out.

Rather than abandon the game, and especially since the dream of common variants having major effects on common disease (so they would be a usefully large market for Pharma to invest in targeting) is diminishing, we have had to become contortionists to try to find how or why genomics is still the way to approach such traits.

Our approach to this question is statistical and hence is based on repeated observation of enough observed instances of a variant for it to achieve statistical 'significance' in our data, whether or not that makes its effect of enough importance'  for counter measures to be develop that are targeted against it.  But this means that in any practical sense, we can't get large enough samples to detect the effects of very rare variants.  We need other approaches.

One is to track variants in key gene regions among family members, looking for correlations between the presence of the variant and that of disease.  How effective this will be depends on whether we know enough about the trait or about genes to find those variants that have such a track.  If we find enough different variants in the same gene doing this in different families, that's strong evidence.

There are two ways, however, for rare variants to work.  One, as just described, is for the variant to have a major effect all on its own.  That could be detectable.  But if combinations of many different rare variants are required, the variants coming from a number of genes and many different combinations having similar effects, this method may not work well.  Unfortunately, there are theoretical reasons to think this will likely be the case.

Recently there has been a story in Science News about the number of rare variants that we each carry around.  The story summarizes various recent papers, citing the authors.  The following graph shows the results of various studies of genome sequencing of different individuals:

This shows that each of us carries quite a few variants.  The estimate is that there is about one variant per 1000 sites,  if you compare both instances of the genome in a given person (or any two randomly chosen copies), which means about 3.1 million variants per pair compared.  But the more people you look at the larger the number of sites that you'll see varying (even if the frequency of the rarer variant at the site is increasingly lower).

Obviously, in a huge population, almost any site will vary, and if lots of sites can potentially contribute to a disease, there will be lots of instances of 'causal' variants per gene, but these won't be detectable by group studies where one needs statistical association between the variant and the outcome (again, because statistical significance can't be achieved with very rare observations in this kind of study design).

The story was about disease hunting, which seems to be the international obsession (or, more accurately perhaps, rationale for funding the work).  However, sites contributing to a trait's variation today will also potentially contribute to its evolution.  Thus, in a subtle way too complex to go into here, hunting for the genes or variants that are responsible for the evolution of a trait is going to be very challenging, to say the least.  It's hard enough to explain the selective (or chance) reasons for the trait's presence, much less what genes were responsible for its evolution.

The story also cites work by Andy Clark and Alon Keinan who have pointed out that the very rapid expansion of the human species in the 10,000 years since agriculture has generated a massive number of rare to very to very very rare variants.  In a statistical sense, each gene lineage present at the beginning of that time has a million descendants today.  This is not new or speculative theory but simply the consequence of the sequence-nature of genes: long strings of nucleotides mean many places where a nucleotide can change.  Even if any given change is very rare, the genome and number of people born each generation are large.  The variation is being found now that we can sequence at a high scale, as reports such as those mentioned here clearly show.

One sobering implication is that if we are concerned with the ways one can get a common disease or trait (behavior, morphology, or whatever, normal or not), then we face trying to work out this sea of nearly unique variation.  It could be a hopeless task!

However, comparing close species with and without the trait in a sense aggregates the results of countless variants and genes and individuals over countless generations.  In a subtle and statistically detectable way this could point to responsible genes , because the sample size that generated the result over time might leave enough evidence.  Some methods to find such evidence are available (one test is called the Macdonald-Kreitman test, after its developers), though so far they are statistically rather weak for close species.  Perhaps creative thinking will lead to new and better ideas.

Whether approaching disease in this way is the best thing to do is a separate question from how it can be done in practice if that is what, as currently, people are deciding they need to do.

Saturday, May 23, 2009

Genetic leaf-litter

There are many ways in which everyone is a conceptual prisoner, encaged in culturally based limits. We are born to, and trained in and entrained by our circumstances, and these in turn are a legacy of history. We can try to escape from this but probably the most we can hope for is to keep subtle assumptions and constraints at bay. In genetics, there is a pervasive concept of the 'wild type', a concept that goes back into the history of genetic research, referring to the natural allele at a gene, that was favored by a history of selection, relative to which other alleles (mutational variants) were viewed as generally rare and harmful (waiting to be shortly removed by natural selection).

There is a tacit extension of this gene-specific concept to the whole genome (or even organism) as when 'normal' inbred laboratory mice are referred to as the 'wild type' relative to an experimental modification such as a transgenic gene knockout mouse of the same strain.
Sometimes this is clear shorthand, but beware of conceptual shorthand! An implication of this kind of genetic thinking is that in regard to human traits, including especially disease, there is the normal human genome as represented by 'the' human genome sequence available in genome data bases, and the disease-causing mutants. But in fact genomes are very large sequences of DNA that serve as targets for mutation in every cell, every individual, every generation.

We know that biological traits are the result of developmental processes that include countless genes (of the classical protein-coding type as well as many other functional DNA sequence elements). Species contain large numbers of members--there are about 7
billion of us humans stalking the Earth. What this means is that there is a potentially huge amount of variation at most if not all viable spots in our genome. After a mutation occurs, it may proliferate if its bearer successfully reproduces. Over time, some of these alleles grow in frequency to become quite common.

When genomic DNA is sequenced in a number of individuals, this variation is easily detected. But whether affected by natural selection or just by the chance aspects of reproductive success or failure, most allelic variation that is present in genomes at any given time is rare. Relative to the more common variants, this genetic variation is a kind of leaf-litter of variation. Even with hundreds of thousands or, indeed, hundreds of millions of very rare variants present in our species, any small sample will pick up some of them by chance.

In a small sample, those will seem
to be more common than they are; so if we sequenced 5 people (10 copies of the genome) the lucky variants whose true population frequency is only a few in a billion that by chance are in the 5 people we sample, will seem to have a frequency of at least 10% (one copy of the 10 we sampled being the variant). The tip-off that this genomic leaf-litter exists is that most of the variants are not seen in other samples, or if common enough to be sampled more than once, usually only seen in samples from the same geographic region (because that's where they arose as new mutations, and were transmitted to descendants who remained living in the same continent). In developed countries, variants that cause disease will show up in specialty clinics at major medical centers.

In trying to find variants by mapping, as in genomewide association studies (GWAS) that compare sequences between cases and controls, we may feel that we have so far detected the common, but not all the rare causal variants that exist. But we may also feel that if we can just enlarge our samples, we'll get a much better handle on the nature of the effects of these variants, or we'll detect the remaining variants that haven't yet been detected.


This is likely to be an illusion, as the growing number of those of us who argue that very large GWAS will not bring a big payoff of the kind envisioned and promised by those who argue for this kind of project. There are several reasons for this skepticism.
First, it is hard to detect rare things with statistical significance, much less to get a good idea of their effects and action. One needs huge samples to get enough instances to show that the variant is meaningfully more common in cases than controls.

But second, the leaf-litter phenomenon means that as sample sizes increase, more and more rarer and rarer variants will be picked up. It will be difficult to show clearly that they are causally involved with our trait, but even if they are they will have less and less effect on public health. They will vary from population to population, and sample to sample from the same population. Environments may affect whether carriers of the variant manifest the disease, and most such variation will at most have minor effect on risk of disease (if the effect were stronger, the allele would have been removed by selection, or we would have been able to detect it in family studies).

And if it requires more than one such variant, or even many of them, to combine to produce disease, the detection and evaluation situation will be that much more challenging, if not pointless.
There will always be exceptions, as is true about the nature of life. But the leaf-litter phenomenon is real and there is plenty of evidence for it. It is predicted by population genetics theory. And it is consistent with results of mapping studies that have been done to date. Ironically, perhaps, while the individual rare alleles have little detectable effects, their aggregate effects in the population may account for the observed heritability (familial aggregation of risk, or similarity of trait values) of most traits, including disease. That heritability, which is clearly there, is what has been considered mysterious given the failure of linkage or GWAS studies to find the genes that are responsible.

We are presented with a kind of epistemological paradox: the genetic variation exists, but we may have insurmountable challenges to find most of it. Indeed, it is somewhat mystical even to argue that it exists as individual effects, if they cannot be found or replicated by current statistical genetic methods.
Evolution 'cares' about reproductive success, not about simplicity in genetic causation. From a population perspective, evolution occurs because mutations occur generating variation that selection and chance can effect from one generation to the next.

Genetic leaf-litter is thus the fuel for evolution. We may care to know the cause of each instance of a trait or disease, but Nature has only cared about viability and success, and tolerates the leaf-litter. As massive amounts of human DNA sequence are produced, we will see this. It will be an incredible playground for population and evolutionary geneticists. But what we do with it, in terms of identifying disease causation, is not clear.