Showing posts with label polygenes. Show all posts
Showing posts with label polygenes. Show all posts

Thursday, April 30, 2015

The tail that doesn't wag: Why?

We know of countless traits that are due to mutations in single genes.  The same allele may not always confer exactly the same disease or other trait, because other genes or environmental factors may contribute, but for most purposes this additional variation is unimportant or at least tractable.  These traits appear in families in roughly Mendelian proportions, as has long been known.

However, most traits including most of the common diseases, do not segregate in families. Instead, it is clear that many different genetic factors are contributing, and most of these effects are individually small.  Further, lifestyle factors are usually even more important, in aggregate.

GWAS and other gene-searching methods have shown that this is likely to be the case for many important diseases, as we’ve written about here many times (and many others have written about elsewhere).  The state of such traits can be depicted schematically as in this figure, where just 3 different genes, each with two states, are shown.  The different combinations of variants are shown with their average associated trait value.  Here, the capital letter allele at each gene confers greater trait value.  This is highly simplified but represents our usual reconciliation of Mendelian genetic effects and the complex traits we observe.  The simplification, as far as it goes, doesn’t affect the main point.  In essence the causal model is that a huge (in many applications, essentially infinite) number of contributing genes is involved, each making a minor contribution.

Schematic distribution of stature and contributing genotypes

The causal complexity so often observed is a problem for the understandable yearning for simple causation—for the promise of ‘precision’ based or highly predictive medicine based on genotype.  We won’t here yet again belabor the reasons this is a culpable fantasy being perpetrated on the paying public, because we’re after a different point. 

If most individuals are in the middle of the population’s trait distribution (e.g., are of roughly average height in the figure), you can see that there are many different genotypes that confer the same trait value.  The sample size is large for the middle group, the bulk of the population, but no single genotype stands out as ‘causing’ average height.  But this is a general model of equal effects, and perhaps there are ways to go beyond such population averages and mine the data for those individual variants that do have identifiably strong effects.

In this figure, as in most textbook illustrations of the point, only a few genes (and those with only two variants each) are shown. But really we know there are tens, hundreds, or even thousands of genome regions that may contribute. This might suggest some strategies. Perhaps the extremes—the tails—can be used to inform us about what is going on in the whole distribution. We can let the tail wag the genetic causal dog in a few ways, perhaps.

The idea of tractability
The search for causal relationships is necessarily reductionist and naturally leads to the search for study design or analysis 'tricks' to turn complexity into simplicity, or at least tractability by some meaningful standard.  Can things be found by some approaches not to be so complex, with at least for some segments of the population, simpler genomic causation?  The following are instances in which this is thought possibly to be so.

1.  Rare variants in families.  Rare variants with strong effect can sometimes be found in close relatives with the same trait.  This may mean a clear-cut and hence typically also rare trait--something far from the norm.  Buried in a population sample, they could simply not be frequent enough to generate a statistical signal.  But, close relatives share big chunks of their genome as well as environments, so one must have some criterion for assigning causation to a genetic variant. A variant that creates a stop codon in a physiological relevant ('candidate'?) gene would be one such.  Even in the huge general population samples that are being collected, families will be identifiable, so a once-old-now-new family-based approach may be able to find some important variants.

2.  Multiple rare variants in the same gene, but different variants in different affected persons, especially when found in a gene known by more common variants or for some other physiological or functional reason to be a plausible causal factor.  The figure only shows two variants per each (A,B,C) gene. But different people may have different variants in the same gene.  They aren't likely to be found in simple whole-population studies, at least not initially.   But if in some way a strong variant identifies the gene as a possible candidate, and then an examination of population data shows other people with similar traits having other variants in the same gene, the gene gains causal plausibility, even if these variants don't seem to be sufficient on their own to be detected in association studies..

3.  Tail-wagging.  If concentrations of individually rare variants are found in individuals with phenotypes at the extremes of a trait distribution, they may be suspected as being causal.  People in the tail might share similar multi-site genotypes, even if the individual variants are generally rare. If the variants consistently contribute in non-trivial ways to the trait, then maybe if we look at those individuals who clearly would have collections of such variants, we might find them.  The tail of the distribution will show us the genes and then we can search for their effects in individuals with less extreme trait values.

In the figure this is clear.  All those individuals with very low or with very high trait values have the same or nearly the same genotype (e.g., AaBBCC, AABbCC, AABBCc, and AABBCC for the larger trait values).  The same thing goes for the lower tail (arrows in the figure).  The role of the 'capital letter' variants in these genes would be clear.  Of course, this classroom figure only shows 3 different genes but one can easily imagine the same sort of thing if there were 10 or hundreds of such contributors.  Environmental effects will of course make this less clear, but the tail might still wag the causal dog for us.  So what has been found?

Checking for the wag
Unfortunately, studies of extreme phenotypes have not yielded much, except for those already long-known because they are basically single-gene traits (CF, PKU, Tay-Sachs, MS,....), which mainly didn't require GWAS etc. to find.  Once they have been found, other variants of lesser effects have indeed been found, and though the story is more complex than that, for our purposes the single-gene effects did their job.

More importantly, for common traits one might hope to find clearer causation in very-high phenotype individuals.  However, where this has been looked for, such as in studies to map the genetic effects on traits like intelligence and stature, investigators have not found much tail-vs-middle difference, as far as we are aware.  A new study of supercentenarians, people in the extreme of the longevity distribution, did not find anything that explained their long survivorship.  The tail is not wagging the dog!

How can this be?  Is environment obscuring things even in the extremes?  Is it the reason those few people are in the extremes?  Or are we making some other mistaken assumption?  How can the extremes not be causally simpler than in the bulk of the population?  This seems a conundrum.

If we believe the evidence, there seem to be as many ways to be in the tails as to be in the middle of the trait distribution, with each person being genotypically unique.  The tails are not wagging the causal dog.  But why not?

What might this mean?  
This is curious because if there are finite numbers of contributing genes, with a distribution of allelic effects, the normal (unimodal, or central tendency) trait distribution, with most people in the middle, would suggest that there are more ways to get there than there are to be in the extremes.  This should also be true of mixtures of rare variants, shouldn't it?  Maybe not!

The figure is a simplified representation of the classical model for polygenic traits due to RA Fisher that basically was essentially of an infinite number of sites each contributing infinitesimal amounts.  In the limit, there are infinitely many ways to be in any part of the distribution. The model is powerful in its applications to various areas of genetics and seems to be basically right, but perhaps there are some key problems with the infinities and infinitesimals underlying the model.

Historically these 'infinities' play a major role in reconciling discrete Mendelian inheritance with quantitative traits and their inheritance, which was an important factor in the 'modern evolutionary synthesis' in the 1930s.  The general theory has served evolutionary and experimental breeders very well for nearly a century.  But is it correct or have we now found a problem area?

I raise this question because in the limit we must be reaching different levels of 'infinity' if our notion is that the reason for the central tendency is that there are more ways to get there than there are to be in the tails--just the assumption we are testing.  But in the limit, to get a smooth distribution and its properties, we essentially assume a greater infinity of ways to be modal than the infinity of ways to be in the tails, and this may be an approximation that makes little sense in the genomic enumeration era--or else that tells us something we need to know.  Infinities are approximations, but maybe the idea of very many contributors runs into practical issues in the kinds of data that GWAS and other studies are using, even their enormous samples.

Could the lag of 'wag' be that we are not dealing with what mathematicians or physicists would call 'well-posed' questions? Maybe stature, obesity, diabetes, or heart disease are not biologically unitary traits.  Then, if they are instead complexes of multiple partly independent (and separately evolved) traits, maybe being simple in the tail is not what we should expect.  Maybe what we are calling a trait is not what evolution 'called' any such thing.

Or it could be that 'infinity' here just means a great many, so that the gist of things is that even in the tails, there really is an essentially unlimited number of ways to inherit few or many small (left tail) or large-effect (right tail) alleles.  And since in any case the number of different combinations is large, and the presence of specific variants small, statistical methods can't enumerate them very well.  One may get into the tails because his/her huge collection of rare or even unique variants have in aggregate more 'large' effect than the collections of those in the middle of the distribution, but each person is unique and the extra individual effects are trivially small.  Our methods are not suited to detect this.  Or maybe the same variants are found across the distribution, but they are slightly more individually common in those with traits in the tail.

Maybe these differences, and/or even the specific variants involved, are so small or the variants uniquely rare, that aggregate, statistical, probabilistic or distribution-based kinds of approaches just won't find them, or just can't find them.  If that is the case, then neither our questions nor our methods are well-posed.

It is perhaps relevant that for many traits we really do have a simplified tail in the population: the very rare, very pathological, usually early onset and severe instances of traits that do turn out to be largely single-gene in their causation.  But they are not a sufficient part of the samples being studied by most mapping efforts.  At the same time, it is all too easy to forget or conveniently ignore, the massive effects lifestyle factors can make in achieved traits be they physical or behavioral.  No wonder, even in the tails of the distribution, we don't get a clear genetic wag!

Here, at least, there are things to think about.  These are real questions.  Technical statisticians (which we're not!) may have explanations--but if so, they haven't led to much in the way of clear causal tractability or these general issues about mapping would not be of such widespread concern. How can the tail not be notably simpler in its causation?  Has our explanation here missed something important, or are geneticists as a community missing something?

The questions have both empirical and theoretical meaning.  Whatever one's view about GWAS and massive whole genome sequencing with the goal of predictive medicine, at least the issues raised are perhaps things we all could agree about.

Wednesday, December 19, 2012

Is it 'progress' to identify 100s of genes for a trait? If not...what is it?

What is known
Crohn's disease (CD), an inflammatory bowel disease, has a large genetic component, but specific genes, as for most such complex diseases, have been elusive.  A paper in this month's American Journal of Human Genetics, "Refinement in Localization and Identification of Gene Regions Associated with Crohn Disease," Elding et al, reports that they believe they are zeroing in on the answer.

The gene most closely associated with Crohn's disease is one that plays a role in the immune response, NOD2 which codes for a protein that recognizes peptidoglycans, or bacterial molecules, and stimulates the immune system to respond.  It makes sense that genes involved in immunity would be involved, as the disease is inflammatory in nature, perhaps due to an impaired innate immune system which leads to chronic inflammatory response by the adaptive immune system to microbes in the gut.

A number of genomewide association studies (GWAS) of Crohn's disease have been done, but none has identified genes with large explanatory power.  A recent meta-analysis of six studies (reported here) identified 32 new loci associated with Crohn's, which, added to the 31 that had been identified in 2010, brought the total to 71.  This doesn't mean 71 single genes but instead stretches of chromosomes that generally contain multiple genes -- sometimes hundreds -- and 'gene' may mean other kinds of function than protein coding, such as regulatory, directly functional RNA, and so on.  The 2010 report explained 20% of the variation in the disease, and the additional 32 brought that total to 23.2%, which indicated that most of these loci represented genes with very small effect, and that many more loci were left to be found. And what about the 77% that is still unexplained?

The Elding et al. paper reports use of a "mapping approach that localizes causal variants based on genetic maps in linkage disequilibrium units (LDU maps)." That means chromosome locations, but not specific to any nucleotide or functional element.  The authors confirm 66 of the previously reported 71 loci, and narrow in on "more precise location estimates" in those intervals (that is, they come closer to identifying candidate genes rather than just chromosomal intervals). They identified 78 additional regions that were statistically significant, and which provide "strong evidence for 144 genes." They also found 56 "nominally significant signals, but with more stringent and precise colocalization." So, this paper reports 200 gene regions in total associated with Crohn's disease, most of which, the authors say, unambiguously implicate single genes. The reason for that inference isn't clear, since clusters of DNA units can function together.  Again, many of these loci contain genes involved in the immune system. The authors suggest that "The precise locations and the evidence that some genes reflect phenotypic subgroups will help identify functional variants and will lead to greater insight of CD etiology."

The immune system is complex and involves many components so that mapping that 'hits' in a region that has some immune system elements might happen by chance if you have 200 hits.  Also, the immune system is involved in response to external threats (like viruses and bacteria) as well as to internal problems (recognizing and repairing damaged tissues), so the reason for 'immune' involvement is unclear -- and a challenge to determining what is responding to what.

The authors have previously demonstrated genetic heterogeneity within the NOD2 gene region -- that is, that different genes explain risk in different people. They also found independent involvement of a nearby gene, CYLD. They further demonstrated the importance of precise definition of the phenotype to identify loci that might explain risk in multiple cases.  In this new paper they use a high-resolution linkage disequilibrium map, basically meaning that their test markers are closely spaced so that implicated regions are fairly short, which helps to identify genes in the implicated region of the chromosome, and fine-tune phenotype definition as well.  They were able to replicate 66 of the previous 71 gene locations, and identified an additional 134 signals, many of which contain genes.  One might always quibble with this or the other statistical issue, but the overall conclusion is unlikely to change.  In the authors words:
This is a major step forward in identifying the relevant genes and functional variants and thus elucidating the genetics of CD etiology. The very large numbers of genes [we identify] confirm that CD is truly polygenic and complex in nature. Many genes show functions that are compatible with involvement in immune and/or inflammatory processes as well as integrity of the intestinal epithelium and differentiation.  
But
In fact, do we know anything more than we did before this study?  It has long been clear that CD is polygenic and complex, with immune and/or inflammatory gene involvement.  Whether 71 genes are involved or 200, it means that this disease is another instance of many pathways leading to a complex trait.

Much money and effort by the highest quality investigators has been expended on this disease.  This means that by now we probably can't claim the complexity is an artifact of imprecise methods.  The complexity seems to be real.  Each CD individual is affected by a different combination of variants at their two copies of these 200 genes -- and, if this is 'simple' complexity, another 600 unidentified genes to top up the current 23% explained causation.

This also doesn't include any serious identification of environmental contributions.  That is, environments like, say, diet at some point in  life, may have different effects on different genotypes, so that a given genotype may have no harmful consequences in one lifestyle and be harmful in another.  And if there are complex interactions among the different contributing genomic regions, then enumerating the variants at 200 (much less 800) genome regions will not make prediction or perhaps not even treatment very genome-dependent.

It is likely that after all of this, some genetic variants will have predictive power or medical treatment relevance, and that of course will be genuine progress.  But it leaves the question that so many traits leave: how do we actually deal usefully with this sort of routine-level complexity?

Thursday, June 9, 2011

The Mayor of Casterbridge's daughter: family resemblance close and far

Thomas Hardy's 1886 novel The Mayor of Casterbridge has a complex plot.

Casterbridge, imaginary Wessex village
Hardy
                                      
The main protagonist, Michael Henchard, opens the story in a drunken state, in which he sells his wife and infant daughter for 5 pounds, to a mariner named Richard Newson, who had dropped casually by for a quaff.  Eighteen years later, having become the Mayor of Casterbridge, Henchard's former wife appears in town with their grown-up daughter, Elizabeth-Jane....or is she?  Henchard adopts her and restores his honor by re-marrying his wife and adopting E-J.  Yet, hearing that she may have died and been replaced by a daughter sired by the mariner, he observes her sleeping.
In sleep there come to the surface buried genealogical facts, ancestral curves, dead men's traits, which the mobility of daytime animation screens and overwhelms.  In the present statuesque repose of the young girl's countenance Richard Newson's was unmistakably reflected.  He could not endure the sight of her, and hastened away.
This is an interesting reflection of common views at the time.  Identical twins show that genes can have remarkable ability to generate even subtle aspects of a person's (or other species') traits.  This is often the subtle undercurrent behind genetic determinism.  Darwin knew that environmental effects could be important, but perceptively opined that a rare trait found in close relatives was probably a genetic (inherited) trait.  This was consistent with his theory of inheritance.  But what about atavistic traits, or 'ancestral curves' that seem to come back generations later?

Mendel gave us one explanation, recessiveness.  Truly atavistic traits like short tails in humans or toes in horses, are explained as developmental anomalies, but ordinary recessive traits are around and noted, but skip generations.  Darwin certainly knew of these, and in The Variation of Animals and Plants Under Domestication he tried valiantly but rather forcedly to explain them within his pangenesis theory of inheritance.

But what about things non-Mendelian that, as in Hardy's tale, seem to come back as vague shadows of the past?  These don't segregate in the fashion of classical dominant or recessive traits, though we are still taught that such traits are fundamental in genetics.  Instead, we see quantitative resemblances among relatives, that can be uncannily vague but clearly inherited.  Such traits, not appearing faithfully generation upon generation, clearly happen.

They are likely to be 'polygenic'--involve many different contributing genes.  Biology is rife with such traits, such as glucose levels, blood pressure, or stature in humans.  Relatives resemble each other due to the roughly additive contributions of countless loci.  You're expected to have a trait mid-way between that of your parents, unless environments play a major role.  And it has become the game (or, sometimes, shame) of GWAS studies to chase such traits and promise that the genes will be found and each person's future will be predictable from them.

But the vague family resemblances of complex, multivariable things like facial characteristics (or, for anthropologists, cranial size and shapes) have a rather unclear place in modern genetics.  We can see the resemblance but it is often difficult even to define just what it is.  It is not traits that segregate, even though people search for the genes 'for' them.  Yet they are not simple regressions of relatives' measurements, the way milk yield in cattle or stature in humans are.

The way such traits 'skip' generations presumably has to do with the effects of combinations of contributing alleles that, depending on what the other parent contributes in a given generation, are more or less easily recognized by our amazing image-processing software (our brains).  The more generations of removal, the fewer the contributing alleles that are still present in a parent.  The traits clearly are genetic, can occasionally seem to be very primitive ('atavistic') from eons past, yet they are not Mendelian the way green or yellow peas were for Mendel.  Yet these subtle aspects of relationships can be the stuff of evolution--the traits they affect change due to chance and natural selection, gradually molding the variation in the contributing genes.

The old quip is that if a child looks like the man of the house, it's genetic, but if it resembles the postman, it's environment.  To poor Henchard, one glimpse of Elizabeth-Jane's visage was devastating.  Had she resembled himself, he would have known it was genetics in action, and she was his own.  But he wasn't a modern man, so when he saw the resemblance to Newson, he wasn't coy enough to blame it on the environment!

Friday, February 25, 2011

The complexity of simple genetic disease

Cystic fibrosis is an ion channel disease that interrupts the flow of salts and fluids into and out of cells, and this affects multiple organs. The most serious consequence of the disease is the production of thick mucus in the intestines and lungs, which leads to respiratory complications, the leading cause of death among people with CF.

Cystic fibrosis is an inherited disease.  The causative gene, CFTR, was identified 22 years ago.  Over 1000 mutations associated with CF have been identified since then, many seen in only one patient or a single family.  In the US the most common mutation is F508del; this designation means that the amino acid that is normally the 508th amino acid along the chain that makes the CFTR protein has been deleted. Another mutation, G551D, is found in about 4% of patients in the US -- this mutation replaces one amino acid with another at the 551th position in the protein chain.  

Identifying the CFTR gene created quite a lot of excitement about the potential for gene therapy, but the initial enthusiasm was pretty quickly dampened by the difficulty in transporting a normal copy of the gene to the required sites in the body. A different therapeutic approach was described in a paper in PNAS in 2009.
Most CF mutations either reduce the number of CFTR channels at the cell surface (e.g., synthesis or processing mutations) or impair channel function (e.g., gating or conductance mutations) or both. There are currently no approved therapies that target CFTR. Here we describe the in vitro pharmacology of VX-770, an orally bioavailable CFTR potentiator in clinical development for the treatment of CF. In recombinant cells VX-770 increased CFTR channel open probability (Po) in both the F508del processing mutation and the G551D gating mutation. VX-770 also increased Cl secretion in cultured human CF bronchial epithelia (HBE) carrying the G551D gating mutation on one allele and the F508del processing mutation on the other allele by ≈10-fold, to ≈50% of that observed in HBE isolated from individuals without CF. Furthermore, VX-770 reduced excessive Na+ and fluid absorption to prevent dehydration of the apical surface and increased cilia beating in these epithelial cultures. These results support the hypothesis that pharmacological agents that restore or increase CFTR function can rescue epithelial cell function in human CF airway. 
The pharmaceutical company that makes VX-770 has just announced the successful completion of a 48 week clinical trial of the drug. The results are impressive. Lung function was significantly improved, and
[h]ighly statistically significant improvements in key secondary endpoints in this study were also reported through week 48. Compared to those treated with placebo, people who received VX-770 were 55 percent less likely to experience a pulmonary exacerbation (periods of worsening in signs and symptoms of the disease requiring treatment with antibiotics) and, on average, gained nearly seven pounds (3.1 kilograms) through 48 weeks. There was a significant reduction in the amount of salt in the sweat (sweat chloride) among people treated with VX-770 in this study. Increased sweat chloride is a diagnostic hallmark of CF. Sweat chloride is a marker of CFTR protein dysfunction, which is the underlying molecular mechanism responsible for CF. People who received VX-770 also reported having fewer respiratory symptoms.     
This is exciting news for the CF community, even for those who don't have the G551D mutation, because the same company is currently testing a drug to correct for the effects of the F508del mutation.
In people with the G551D mutation, CFTR proteins are present on the cell surface but do not function normally. VX-770, known as a potentiator, aims to increase the function of defective CFTR proteins by increasing the gating activity, or ability to transport ions across the cell membrane, of CFTR once it reaches the cell surface. In people with the F508del mutation, CFTR proteins do not reach the cell surface in normal amounts. VX-809, known as a CFTR corrector, aims to increase CFTR function by increasing the amount of CFTR at the cell surface. 
This all has the potential to change the future for people with CF.  And it also means that if the function of a gene and mutations in that gene are understood, the parameters are there for potentially developing therapies.  We seem to understand a lot about this ion channel.  In fact, if these results are real--general, long-lasting, and clinically or lifestyle-important as well as statistically significant -- they probably will apply to many other CF patients with other mutations that are individually rarer but have similar effects on the CFTR protein. 

But, the gene for CF has been known for 2 decades, and a treatment for just  4% of people with the disease is only now beginning to look promising.  The difficulty of getting to just this point is a sobering reminder that 'personalized medicine' is going to be an order of magnitude harder for polygenic diseases.  And hopefully, we won't have to take back these positive feelings about these potentially life changing results, and this 48 week trial is not being reported prematurely to boost stock prices or anything cynical like that.

If it works as the current story suggests, these results exemplify what we personally have repeatedly said about medical genetics.  There really are good ways to spend genetics research effort, not on mindless GWAS mapping, but on traits that are tractably simple and that really are genetic in a meaningful sense.  This seems like a very good example of that principle even if, as is the case with other instances, only a fraction of all CF patients will benefit directly.

Wednesday, December 22, 2010

A new broom sweeps gene?

Today we write about a story that isn't hot off the press, but was published in Nature a few months ago, and  happens to be one of those findings that we think more people should know about:.  The issue is the genetic signature of selection, something that has become the focus of much anthropological and population genetic research with the advent of whole genome sequencing data.

What the authors did was to follow populations of fruit flies over 600 generations and comb the entire genome of 260 of them for variation after applying intense artificial selection on the measurable, and malleable, traits of accelerated development and early fertility.  They bred a population in which development was about 20% faster than in unselected populations.  The question was whether a single gene or multiple genes would be responsible for the change.

They compared the genomes of the selected and control populations with the reference fruit fly genome, and found hundreds of thousands of SNPs, or single nucleotide polymorphisms, differences between the populations.  Of these, they found tens of thousands of amino acid-changing SNPs, about 200 segregating stop codons and 118 segregating splice variants -- that is, variants that could be responsible for the phenotypic changes they had selected for.  They further narrowed down these candidate loci to 662 SNPs in 506 genes that they considered to be potential candidates "for encoding the causative differences between the ACO and CO populations, to the extent that those differences are due to structural as opposed to regulatory variants."
For the biological processes, there is an apparent excess of genes important in development; for example, the top ten categories are imaginal disc development, smoothened signalling pathway, larval development, wing disc development, larval development (sensu Amphibia), metamorphosis, organ morphogenesis, imaginal disc morphogenesis, organ development and regionalization. This is not an unexpected result, given the ACO [accelerated development population] selection treatment for short development time, but it indicates an important role for amino-acid polymorphisms in short-term phenotypic evolution.
Actually the idea that adaptive change was brought about by gene-impeding mutations (premature stop codons and splice variants, for example) is interesting.  It means that adaptive change under selection doesn't just improve function, but it may also destroy function--to pave the way to the change, one might surmise.

They went on to do a 'sliding window' comparison of regions of the genome that diverged significantly between the selected and control populations, and identified 'a large number'.
...it is apparent that allele frequencies in a large portion of the genome have been affected following selection on development time, suggesting a highly multigenic adaptive response.
The authors interpreted this work in terms of the 'soft' or 'hard' sweep idea that is often used to explain reduced gene frequencies ( a 'hard sweep' being when a single mutation quickly becomes fixed in a population, and a 'soft sweep' being when multiple genes influence a trait).  They suggest two explanations for their 'failure to observe the signature of a classic sweep in these populations, despite strong selection' (not enough time for the causative gene to reach fixation in the population, or that selection acts on standing, not new mutations).

RA Fisher
But, it wouldn't be a surprise to RA Fisher (this is a link to his Facebook page, by the way -- go friend him, he only has 6!) that the observed changes are due to polygenes.  But it is nice to see an experimental confirmation, and to note the implications it has for understanding complex traits.

Many biologists have been lured into single-gene thinking by the research paradigm and model set up initially by Mendel.  For decades single-gene traits formed the core of what we would call the evolving molecular genetics including Morgan's work on chromosomal arrangement of genes, many human geneticists' work on 'Mendelian' disease, the work leading to the idea that genes code for proteins, and much else.

Besides these examples of causal genetics, we had selection examples such as sickle cell anemia, that seemed to reflect evolutionary genetics and were due to single protein changes.  But we always knew (or those who cared to understand genetics should have known and could have) that traits were more complex than that as a rule.  Sewall Wright and others knew this clearly in the early 20th century.  'Quantitative genetics' going back basically to Darwin (or at least his 2d cousin Francis Galton) recognized the idea of quantitative inheritance and Fisher's influential but largely impenetrable (to mere mortals) 1918 paper was a flagship that reflected formally the growing recognition that complex traits could be reconciled with Mendelian genetics if many genes contributed to complex traits.

The idea that strong directional or 'positive' selection favored a single gene grew out of the Mendelian thread, but nobody in quantitative genetics (such as agricultural breeders or many working in population genetics theory) and those who understood gene networks, should have known that most of the time, especially given the typical weakness of selection, selection would not just find and fix a single allele in a single gene.

We had reason to know, and certainly know now that when a trait's effects are spread across many variable and contributing genes, the net selective difference on most if not all of them will be very small. The response to selection will be just what the fly experiments, and many others likewise, found.

In the television attention-seeking era we need melodramatic terms, and that is just what ideas like selective 'sweeps' are.  The circumstances under which a single allele will 'sweep' (watch out, here comes that broom sweeping clean!) would occur across an entire species' habitat replacing all other alleles that affect a trait are likely to be very restrictive.  We don't need terms like hard and soft sweeps, and should not be over-dramatizing what we find.  Even a hard 'sweep' at the phenotype level is typically 'soft' at the specific gene level, and usually also softly leaves phenotypic variation in the population after it's over.

At the same time, these experiments are giving us great detailed knowledge about how evolution works, when there is, and when there is not strong selection.  This supports long-standing theory and is no kind of 'paradigm shif', it's true, but it is new understanding of the details and genetic mechanisms by which Nature gets from here to there--whether it does that in a hurry or not.

Thursday, August 6, 2009

GWAS: Carry on regardless....(?)

From 1958-78,29 very funny British 'Carry on' movies, with titles like 'Carry on nurse,' were produced. It's a British phrase that a boss would say to an employee, but in this case, no matter how bollixed up the situation was, the idea was that people just 'carry on' regardless. The movies parodied British institutions and their behavior.

Well, though it's neither satirical nor funny, the same is often the case in science, where the thing-to-do is what's done regardless of whether it's the best thing or is working very well. We cling to our flotsam if it's all we know or seems the safest.

On July 14, we wrote about a Nature paper describing the results of a genome-wide association study (GWAS) of the genetics of schizophrenia, published online. This and two accompanying papers appear in the journal this week, reinforcing the story of schizophrenia as a polygenic trait, with many genes involved, each with very small effect. The papers suggest that the immune system may somehow be involved, as genes in the HLA system are found to have a significant, if limited effect, as well as some aspects of brain development, cognition and memory.

This is potentially very interesting because if GWAS are finding anything it may be that immune or inflammatory system genes are involved in a wide array of traits, perhaps not always previously suspected as such. Could infectious or autoimmune causes be more widespread than we have suspected? If so, it may say a lot. Partly it could be the things that go wrong over a life that's decades long, in terms of exposures and/or mutations that attack self.

But this and a host of other GWA studies have had minimal findings--hyped to death, perhaps, but usually accounting for only a minor fraction of all causation, even the known genetic component as revealed by the family cluster of disease. Each study finds one or a few genes that contribute detectable amounts to the trait. Some of these have been replicated and begin to be believable for that reason (though many if not most have no a priori plausibility as causes of the mapped disease).

This should be providing geneticists with plenty of targets for real genetics, real in the sense of figuring out what the genes do and how to attack them therapeutically. Modest they may be, but they're the strongest candidates we have.

So why, then are new GWAS still being done all over the place on the same diseases? Even strong proponents of this method have acknowledged that they are finding many genes with small effect. The same traits are being studied over and over again, often finding different genes, but which almost invariably explain very little risk.

It's time to stop paying for ever more of this, and to demand proof of principle. The principle is that (1) we can show how, why, and when these candidates (and their many mutations and regulatory sequence variants) are involved in disease, and then (2) that we can do something about them. Yet, because many investigators are set up for mapping, which is after all a rather mechanical button-pushing kind of enterprise (and very grant-able), one often hears investigators saying that 'mapping is what I do' .

That's a very poor excuse, even if in a careerist society that depends on the grant system and on intellectual inertia. What we need now is some accountability: to show that the fruits of GWAS labor to date are worth it and do, after all, have important biomedical use.

That being done, we'll know better whether we should continue down the reductionist mapping road, or whether better, more effective approaches to causation, even genetic causation, are the proper course.

In the 'Carry on' movies, things muddled along despite all the confusion, and in science we'll muddle along, too. Nobody can say that persistence with GWAS or other tactics is useless, even if it's inefficient and we know better what would be better. But biomedical science can do better, and we think that it should.