Showing posts with label genetic variation. Show all posts
Showing posts with label genetic variation. Show all posts

Thursday, February 26, 2015

Digesting yeast's message

A new paper in Nature by Levy et al. reports on the genomic consequences of large-scale selection experiments in yeast.  Yeast reproduce asexually and clones can be labeled with DNA 'barcode' tags and followed in terms of their relative frequency in a colony over time.  This study was able to deal with very large numbers of yeast cells and because they used barcodes the investigators could practicably follow individual clones without needing to do large-scale genome sequencing.  Prior to this, this sort of experiment was prohibitively costly and laborious.  So the authors add to findings in selection experiments using bacteria or flies and so on, where mostly aggregate responses could be identified.

In this case, nutrient stress was imposed, and as beneficial mutations occurred and gave their descendant cells (identified by their barcode) an advantage, the dynamics of adaptation could be followed.  The authors showed, in essence, that at the beginning the fitness of the overall colony increased as some clones, bearing advantageous mutations, rose rapidly in relative frequency.  Then, the overall colony fitness stabilized and subsequent advantageous mutations were largely kept at low frequency (most eventually went extinct).  But overall, the authors found thousands of colonies with different advantageous variants; most fitness effects were of only a small (or, for the majority, very small) percent.  Once a set of large numbers of 'fit' variants had become established, new ones had a difficult time making any difference, and hence staying around very long.

This study will be of value to those interested in evolutionary dynamics, though I think the interpretation may be rather more limited than it should, for reasons I'll suggest below.  But I would like to comment on the implications beyond this study itself.

Who cares about yeast (except bakers, brewers, and a few labs)?  You should!
This is interesting (or not) you might say, depending on whether you're running a yeast lab, or in the microbrew or bakery business. But there are important lessons for other areas of science, especially genomics and the promises being made these days.  Of course, the lesson isn't a pleasant one (which, you might correctly assume, is why we're writing about it!).

This study has important implications for basic evolutionary theory perhaps, but also for much that is going on these days in human biomedical (and also evolutionary) genetics, where causal connections between genomic genotypes and phenotypes are the interest.  In evolution, selection only works on what is inherited, mainly genotypes, but if causation is too complex, the individual genotype components have little net causal effect and as a result are hardly 'seen' by selection, and evolve largely by chance.  That's important because it's very different from Darwin's notions and the widespread idea that evolution is causally rather simple or even deterministic at the gene level.

Put another way, genomic causation evolved via the evolutionary process.  If natural selection didn't or couldn't refine causation to a few strong-effect genes, that is, to make it highly deterministic at the individual gene level, then biomedical prediction from genome sequences won't work very effectively.  This is especially true for traits, disease or otherwise, that are heavily affected by the environment (as most are) or for late-onset traits that were hardly present in the past or arose post-reproductively and hence didn't affect reproductive fitness and are not really 'specified' by genes.

There was considerable genomic variation between the authors' two replicate yeast experiments.  As one might say, meta-analysis would have some troubles here.  Likewise, from cell lineage to cell lineage, different sets of mutations were responsible for the fitness of the lineage in this controlled, fixed environment. This means that even in this very simplified set-up, genomic causation was very complex.  No 'precise' yeastomic prognostication!

In real biological history, even for yeast and much more so for sexually reproducing species in variable environments, selection has never been unitary or fixed, and genomes much more complex. Human populations have been until very recently very much smaller than 10^8 in the yeast experiments, and recent population expansion will make the number of low-frequency variants much greater, and with recombination, vastly more genomically unique.

The bottom line here is that our traits should be much less predictable from genotypes than traits in yeast. We have not reached, nor did our ancestors ever reach, the kind of fitness equilibrium reached in the yeast study under controlled selection, and fixed environments.

Somatic mutation
The authors also compare the large numbers of cells whose evolution they were able to follow with their barcode-tagging method, to the evolution of genetic variation in cancer and microbial infections, where there are even larger numbers of cells in an affected person and, importantly, clones expanding because of advantageous mutations. From the yeast results, these clonal advantages may not generally be due to one or two specific mutations (with perhaps, hopefully, exceptions when chemotherapy or antibiotics exert far stronger selection than was imposed in the yeast experiment). But the general complexity of such clonal expansions present major challenges, because they may end up with descendant branches distributed throughout the body where even in principle the responsible variation can't be directly assessed.

But the implications go far beyond cancer.  As we've recently posted, cancer is a clear but perhaps only a single manifestation of a more general phenotypic relevance of the accumulation of somatic mutations, that occur in body cells during life and can in aggregate have systemic or organismal-level implications.  The older we get the more likely we are to generate such clones, all over the body, and it seems likely that they can become manifest not just as individually ill-behaving cells, but as disease for the whole person.

But it's not just late onset implications that the yeast work may forebode.  There are already huge numbers of cells in the early embryo and fetus whose even huger descendant clades of cells during life grow many, many fold by adulthood.  There is no reason not to expect that each of us will carry clades that include differently-than-normal functioning cells in our tissues.  Let age, environmental exposure, and further mutations add to this and disease or age-related degeneration can result.  Yet none of this can be detected in the usual individual's 'genome' as currently viewed.  This is a potentially important fact that, for practical reasons or what one might call reasons of convenience, is ignored in the wealth of mega-sequencing projects being lobbied for based on genome sequencing (precision prediction being the most egregious claim).

So a bit of brewer's yeast may be telling us a lot--including a lot that we don't want to hear. Inconvenient facts can be dismissed.  Oh, well, that's just yeast!  They evolve differently!  That was just a lab experiment!  Brewers and bakers won't even care!

So let's just ignore it, as if it only applies to those rarefied yeast biologists.  Eat, drink, and be merry!

Friday, May 17, 2013

Of mice and men: Genes, environment, and whatever

Is the nature/nurture question finally solved? Are we who we are as a consequence of luck?  A paper in Science, Freund et al., "Emergence of Individuality in Genetically Engineered Mice," assesses the effects of the nonshared environment on neural and behavioral development in 40 inbred female mice living in an enriched environment compared with genetically identical mice in a non-enriched environment.  Does the plasticity of the brain in response to environmental stimuli go a long way toward explaining who we are? 

Freund et al. write:
...the emergence of experience-based individual differences within groups of genetically identical animals exposed to the same enriched environment has rarely been addressed. We used a large group of animals and a particularly complex environment to capture the emergence of individual differences in brain and behavior over time. We used exploration as a marker of behavioral development, and adult neurogenesis in the hippocampus as a marker for continued brain development.
  This figure from the paper gives a schematic of the experimental set-up.

Experimental setup and effects on body and brain weight.
(A) Schematic illustration of the large enrichment enclosure housing 40 mice including RFID antenna positions (shown as red rings). Positions of levels, water sources, nesting boxes, and connecting tubes are drawn to scale. (Inset) Schematic illustration of animal tracking; an RFID passive integrated transponder (PIT) is implanted in mouse’s neck. The electromagnetic field issued by the antenna induces the PIT to emit the number identifying the animal. This information is then picked up by the antenna and stored into a database together with spatial and temporal annotations. (B) Experimental time line. (C) Body weight development: weights (in grams) of CTR (blue) and ENR (red) mice at the beginning and end of the experiment. (D) Brain weights at perfusion (in grams). The difference in variance between CTR and ENR missed conventional statistical significance at P = 0.057. Source: Freund et al., 2013, Science.
After 3 months, mice in the enriched environment were heavier, with more variability in brain and body size than the control mice (though, there were only 8 control mice) and more neuronal connections were made in the hippocampus of the enriched adult mice than in the brains of the controls. "This finding supports the idea that the key function of adult neurogenesis is to shape hippocampal connectivity according to individual needs and thereby to improve adaptability over the life course and to provide evolutionary advantage."

While the authors point out a number of ways in which these mice may not in fact be genetically identical, primarily, they suggest, because of epigenetic changes due to such variables as position in the uterus, maternal disease, nutrition and interactions with the mother, maternal imprinting and so on (the 40 mice were randomly picked from different litters of the same strain of inbred mice to minimize these kinds of effects), they still consider them to be "identical."

But...wait a second!  Are they identical?
Much as we like the authors' general conclusion, because it moves away from the excessive level of imputed genetic determinism of our traits, we must add a word of caution.  These mice were actually not even genetically identical. Every time a cell divides, mutation is likely to happen. Based on some estimates from various kinds of data, that can be about 150 changes per cell division.

Such mutational variation accumulates from conception to an animal's sperm or egg production, and is thus transmitted across generations.  This is true, of course, even in inbred laboratory animals.  So even in a litter of 'identical' pup embryos, there is genetic variation.  But the picture is even more complex.

Every cell division during a mouse's (or your) lifetime, mutations occur.  If we assume that the rate is roughly as above, and even a mouse has millions if not billions (and you have billions if not a trillion or so) of cells, there is a lot of genetic variation within an organism.  Once such a somatic (body cell rather than germ cell) mutation occurs, when that cell divides its daughter cells, throughout future  life of the organism, inherit the change.  Thus, the earlier during embryological development that a somatic mutation occurs, the larger the tree of cellular descent--the more organs or larger the part of a developing organ--that will inherit the change.

Now most of these by far will occur in unimportant areas of the genome--not in any actual genes at all, or in genes that aren't used in the tissues in which the mutation has occurred.  But this cannot be assumed as a general fact and one has to consider the nearly inevitable likelihood that whatever trait is being studied, the animals, even inbred lab animals, are not genetically identical in respect to it.

The challenge here is to identify the variation and figure out if it matters to the trait.  To do that one needs to do tissue-specific, if not detailed individual cell-specific genome sequencing, and even then one needs to identify gene expression at the cell level--and maybe (probably) at different times during development or environmental exposure changes--to attempt to identify those genomic elements whose variation might be involved.   This is essentially impossible as a general rule, and certainly only under some unusual circumstances would it even be worth undertaking.  And, of course, you'd kill the animal in the process, so its behavior would be (only) in the mind of the investigator!  Even to justify the cost and effort, without this minor mortal stumbling block, one would once again have to believe that slicing and dicing the genome, cell by cell, minute by minute, will explain complex traits.

Of mice and men....
Is there really such variation in inbred animals?  Well, we used ForSim, a forward evolutionary simulation program developed in my lab by me and Brian Lambert, to simulate a mouse experiment.  Two independent lines of mice were simulated, using many genes and a lot of DNA to get enough statistical stability in the results.  After generating a normal level of variation, as seen in wild animals and people, we simulated selection of some trait in the opposite direction (small trait values in one, large in the other strain), and then inbreeding for a large number of generations (around 200) with the population kept small, roughly as inbred lines are developed.  Then, we examined these simulated animals for sequence variation and we found a substantial amount of it: roughly as many different sites were varying among the animals as had been fixed by selection and inbreeding.  Yet in the usual mapping and experimental approaches such variation is assumed not to exist.  But it does.

Hey, are we against genes or for genes, after all??
Given our predilection for criticizing what we believe are excessive claims of genetic causation, one has to think carefully and avoid oversimplifying.  The argument cuts all ways.  We must therefore also raise similar cautionary questions about excessive dismissal of genomic effects!

Today, our point is that even here, where traits vary to a surprising extent even in putatively identical animals, it cannot really all be attributed to 'chance' (unless that includes mutations), nor to learning, nor environment.  The causal mix is inextricably complex under widespread if not most conditions.

It is for reasons such as revealed by this study in what otherwise is a clear demonstration, that we write so often to try to temper the enthusiasm for genetically deterministic thinking, much less such gene-based predictions of individuals' futures.  But genes do vary, and they vary subtly.  There is no one crystal ball, not even for mice!

Friday, November 30, 2012

Can we or can't we explain common disease?

Rare variants don't explain disease risk
We're still catching up on readings after a long Thanksgiving weekend, so are just getting to last week's Science.  Here's a piece that's of interest -- 'Genetic Influences on Disease Remain Hidden,' Jocelyn Kaiser -- in part because it touches on a subject we often write about here, and in part because it seems to contradict a story getting big press this week, published in this week's Nature.

Kaiser reports from the Human Genetics meetings in San Francisco that finding genes for common disease is proving to be difficult.  GWAS, it turns out, are finding lots of genes with little effect on disease.  This is of course not news, though the Common Variant/Common Disease hypothesis -- the idea that there would be many common alleles that explain common diseases like heart disease and type 2 diabetes -- died far too slowly given what was obvious from the beginning (never with any serious rationale, as some of us had said clearly at the time, we may not-so-humbly add), and the rare variants hypothesis that replaced it is rather inexplicably still gasping.  Or, as Kaiser writes, "...a popular hypothesis in the field—that the general population carries somewhat rare variants that greatly increase or decrease a person's disease risk—is not yet panning out."

Apparently the idea, then, is that there's still hope. Indeed, many geneticists believe that larger samples are the answer.  That is, studies that include tens or hundreds of thousands of individuals, because these will be powerful enough to detect any strong effect rare variants may have on disease, in theory explaining the risk in the center of the graph from the paper, which we reproduce here.  Kaiser cites geneticist Mark McCarthy of the University of Oxford in the United Kingdom: “We're still in the foothills, really. We need larger sample sizes."  Further, he says, "The view that there would be lots of low frequency variants with really big effects does not look to be well supported at the moment." 

Fig from Kaiser. New studies failing to explain the genetics of common disease.  


Even with larger sample sizes, it turns out that some variants are so rare that they're only seen once.  And probably explain only a small proportion of risk anyway, even in that single individual. And certainly can't be used to predict disease. But this doesn't stop geneticists from wanting to increase sample sizes, at this point usually by doing exome sequencing (sequencing all the exons, or protein coding regions) of tens of thousands of people and looking for rare variants with large effects.  Ever hopeful.  McCarthy, a seriously non-disinterested party to any such discussion, is not likely to give up on ever-larger scale operations; that would be research-budget suicide, regardless of the plausibility of the rationales.

Rare variants do explain disease risk
Which brings us to the big news story of the week, a paper in Nature by geneticist Josh Akey et al., described in a News piece by Nidhi Subbaraman in the same journal, 'Past 5,000 years prolific for changes to human genome.'  The idea is that the rapid population growth of the last 5,000 years has resulted in many rare genetic variants, because every generation brings new mutations, and that these are the variants that are most likely to be responsible for disease because they haven't yet been weeded out of the population for being deleterious.

The research group sequenced 15,336 genes from 6,515 European Americans and African Americans and determined the age of the 1,146,401 variants they found.  "The average age across all SNVs was 34,200 ± 900 years (± s.d.) in European Americans and 47,600 ± 1,500 years in African Americans..."  They estimated that the large majority of the protein-coding, or exonic single nucleotide variants (SNVs) "predicted to be deleterious arose in the past 5,000-10,000 years."  Genes known to be associated with disease had more recent variants than did non-disease genes, and European Americans "had an excess of deleterious variants in essential and Mendelian disease genes compared to African Americans..."

They conclude that their "results better delimit the historical details of human protein-coding variation, show the profound effect of recent human history on the burden of deleterious SNVs segregating in contemporary populations, and provide important practical information that can be used to prioritize variants in disease-gene discovery."  Indeed, the proportion of SNVs in genes associated with Mendelian disorders, complex diseases and "essential genes" (those for which mouse knockouts are associated with sterility or death) that were 50,000 to 100,000 years old was higher in European Americans than in African Americans.  The authors propose that this is because these variants are associated with the Out-of-Africa bottleneck as humans migrated into the Middle East and Europe, which "led to less efficient purging of weakly deleterious alleles."

The researchers conclude:
In summary, the spectrum of protein-coding variation is considerably different today compared to what existed as recently as 200 to 400 generations ago. Of the putatively deleterious protein-coding SNVs, 86.4% arose in the last 5,000 to 10,000 years, and they are enriched for mutations of large effect as selection has not had sufficient time to purge them from the population. Thus, it seems likely that rare variants have an important role in heritable phenotypic variation, disease susceptibility and adverse drug responses. In principle, our results provide a framework for developing new methods to prioritize potential disease-causing variants in gene-mapping studies.  More generally, the recent dramatic increase in human population size, resulting in a deluge of rare functionally important variation, has important implications for understanding and predicting current and future patterns of human disease and evolution. For example, the increased mutational capacity of recent human populations has led to a larger burden of Mendelian disorders, increased the allelic and genetic heterogeneity of traits, and may have created a new repository of recently arisen advantageous alleles that adaptive evolution will act upon in subsequent generations.
This does seem to contradict the Kaiser piece we mention above, which concludes that rare variants with large effect will not turn out to explain much common disease.  This paper suggests they will -- which we don't think is right, for reasons we write about all the time.  But it does lend support to the idea that the Common Variant/Common Disease hypothesis is dead and buried. 

Serious questions
It is curious, and serious if true, that Africans harbor fewer rare variants than Eurasians.  African populations expanded rapidly since agriculture, just as Eurasians did.  It could be, but seems like rather post-hoc rationalizing, that Africa is more dangerous to live in, even for only mildly harmful variants.  Rapid expansion--the human gene lineages have expanded a million-fold in the last 10,000 years, will lead to many slightly harmful variants being around at low frequency, because slight effects aren't purged by selection as fast as they are generated in an expanding population.

In a sense the deluge has not been of functionally important but rather functionally minimal variants.  Maybe there is something about the raised probability that a person will have a combination of such variants, and the variants could be found by massive samples.  But then their individual effect probably isn't worth the cost of finding them, as a rule.

But where's the nod to complexity?
But, environments change, and genes now considered to be deleterious may not have been so in previous environments, or may even have been beneficial.  And African Americans don't represent a random sample from the entire African continent, as their ancestry is predominantly West African, and SNV patterns are likely to be different in different parts of Africa.  And, numerous studies have found that healthy people carry multiple 'deleterious' alleles, so the idea that 84% of SNVs will lead to disease is probably greatly exaggerated. Geneticists just can't bring themselves to acknowledge that complexity trivializes most individual genetic effects.

The more likely explanation for complex disease continues to be, "It's complex."

Thursday, June 7, 2012

The Vampire Monologs: Ancient DNA and the un-dead?


Tuesday's news stories reveal that Bulgaria rather than Romania (Transylvania) has provided us with some skeletal remains of what are being claimed to have been vampires.  According to a story on the BBC, "Archaeologists in Bulgaria have found two medieval skeletons pierced through the chest with iron rods to supposedly stop them from turning into vampires."  Fortunately, once living, and then not-so living, they were thus driven, so to speak, from the un-dead to the really, truly dead.   We all (or at least young, nubile ladies with accessible necks) should be relieved at this news.  But it also portends  important science, if done carefully.

Vampire skeleton. Source: BBC
We're told by Bram Stoker in the classic 1897 Dracula that vampires are recruited from the living by other vampires. I don't recall if it's explained how the first vampire came to not-be.  Nor how the legendary vampires are mainly males.  I'll leave that to others to study.

What's actually relevant for MT readers is the genetic questions posed by these discoveries.  The skeletons are recent enough that the bones will contain DNA that is in good enough condition to be sequenced.  Normally this would be called 'ancient DNA' (denoted aDNA), but a more relevant term for these rather recently un-deceased would be vDNA.  What will it show?  Will we be able to find the gene 'for' vamping?  How will the scientists do this?

Source: Flickr CC photo by Mugley
Vampires pose serious problems for evolutionary genetics, that are too intricate to go into in a mere blog post, and I plan to write about this in the near future, and report the findings to the NY Times rather than this puny blog outlet.  But there are some less grandiose issues posed by vDNA.  First, what would one expect to see?  

The sequence should show clear similarity to modern eastern Europeans.  Overall, there should be no particular trait that would reveal the vampire status of the individual.  Indeed, if vampires are randomly recruited, what we need to know is whether some people (like Mina Harker) were genetically susceptible to being vamped, or not.  If not, of course we have the conundrum that, since we're routinely informed (by the NY Times and the major journals) that everything human must have a genetic cause.  So we must assume some genetic difference.

This is a challenge, because we now know that any two copies of the human genome--even the two that you carry--differ by millions of nucleotides.  Therefore, to find the vampire susceptibility variation we might be looking for a family of genes, call them Vam1, Vam2,..., and so on that are responsible.  Since all genes are already known from human DNA sequence, we must simply have mislabeled these genes.  Many genes' functions are not known, and the Vam genes must be among them.  However, why do they even exist if only some women ever become vamps (not to mention male vampires)?  This poses one of the key evolutionary questions raised by vampirehood, since our view of evolution is that it has no foresight, so Vam genes can't have evolved for their future adaptive value in the Caucasus.

This suggests that there aren't really any Vam genes after all.  Instead we must search for variation in known genes that yield susceptibility to being vamped.  It's easy to imagine how that could be. For example, genes conferring long, luscious necks on women, or that make a woman want to wear low-necked blouses, could easily have the allure that is needed for them to be among the Chosen.
Vampire, Edvard Munch

But we don't know the genes 'for' necks (or low-neck shirt wearing), and what GWAS have clearly shown without doubt is that such traits are complex with many contributing genes.  As a result, we need to identify many places in the genome, where variants will generally only contribute a minor amount.  As we know with other disorders like diabetes, heart disease, and the genetically based Gullibility Predilection to believing that Everything is Genetic (the high frequency GPEG allele), we need large samples to find the critical variants in the sea of millions of rare but useless variants each of us carries.  Rather than Vam genes, what we seek are, shall we call them, genetic V-ariants.

That means, of course, that we must first of all do whole-genome sequencing and collect very large case-control samples to ferret out the V-ariant elements, and therein lie two V-ery serious challenges!

First, how on earth will we find enough cases?  We need to find the skeletons (or undead cadavers in current dungeon coffin residences) of a huge number of vampires--these days, the state of the art requires that we ascertain hundreds of thousands!  But how on (or under the) earth, this side of the Styx, could we find such a horde?  We need to compare their un-dead vDNA with that of the not-yet-dead DNA of living people.  We can't just dig up graveyards, or rummage around everybody's  basement, because how would we know which corpses were, or might have been, vampires?  This exemplifies the second challenge, which is how can we even obtain adequate controls?

Reading vDNA: a real problem in genetic cryptography!
You might be aware (most geneticists don't seem to be) that while controls are defined as being unaffected by the trait in question, many of them will  become future cases--that is their DNA is susceptible even if classified as 'unaffected' or 'normal'.  That means that until we know who will become un-dead in the future, we don't know whose DNA doesn't contain V-ariants!

A typical GWAS kind of approach, to salvage this situation, might be to select as controls only women who typically wear turtlenecks.  They might be the least likely to carry the susceptibility variants (if indeed susceptibility is linked to making ones neck alluringly available).  Otherwise, how do we match our cases with adequate controls?

If it turns out, as it surely will, that hundreds or thousands of genes in vDNA as in living DNA contain V-ariants, then we will face the horrible, or horrifying, problem that most of us carry some of the V-ariants, but we can't really know who.  A geneticist, even the usual type seeking attention, may be reluctant to stick her neck out with only weak evidence.  So as we walk and work among the living, we have no way to know if, at the end of the day as the phrase goes, we'll later walk among the un-dead...

Wednesday, February 22, 2012

How many genes can we live without?

It's well-known that 'the' human genome doesn't exist -- despite the hype about the complete sequencing of this thing (and despite the incompleteness of its sequencing).  In fact, each of us has a collection of DNA variants that, added together, means that our genome has never been seen before in the 3.8 billion year history of genomes, and will never be seen again.  And it's no mean collection; we all differ from each other at something like 3 million loci.

We need to understand first of all that 'the' human genome is not from one person, and even so, what it is, is a reference sequence, useful for comparing other data, but not definitive of our species.  The donors of the DNA were healthy at the time of donation, but that's about it.  They were not particularly special in any way.

Human genome by functions; Wikimedia Commons
In fact, as it turns out, not only do we differ at single nucleotides -- you have a T where I have an A, I have a G where you have a C, and so on, times 3 million -- but we are each carrying around a not insignificant number of variants that result in loss of function (LoF) of some of our protein coding genes.  And many of these are genes we think of as essential.

A paper published in last week's Science, "A Systematic Survey of Loss-of-Function Variants in Human Protein-Coding Genes," MacArthur et al., estimates that each of us has around 100 of these LoF variants, 20 of which result in complete loss of a gene. The authors looked at three pilot data sets from the 1000 Genomes Project, (58 Yoruba, from Nigeria, 60 European American from Utah, 30 Chinese individuals from Beijing and 30 Japanese from Tokyo) and a European genome, and, after filtering a larger initial set of candidate LoF variants, finally analyzed 1285 variants that they found to be likely to cause protein-coding genes to lose function.   

As MacArthur et al. point out, other recent studies of the complete DNA sequences of healthy individuals have also found many LoF variants -- from 200 to 800.  People walking around perfectly normal so far in their lives, but without the use of substantial numbers of their genes.  The specifics vary, and what you can do without, your genes that aren't working, likely depends on what is working.

Before the advent of complete genome sequencing, LoF variants were thought to be rare, largely associated with severe Mendelian disorders such as cystic fibrosis or Duchenne muscular dystrophy.  The finding that they aren't so rare after all suggests to MacArthur et al. "a previously unappreciated robustness of the human genome to gene-disrupting mutations and [this has] important implications for the clinical interpretation of human genome–sequencing data."
LoF variants found in healthy individuals will fall into several overlapping categories: severe recessive disease alleles in the heterozygous state; alleles that are less deleterious but nonetheless have an impact on phenotype and disease risk; benign LoF variation in redundant genes; genuine variants that do not seriously disrupt gene function; and, finally, a wide variety of sequencing and annotation artifacts. Distinguishing between these categories will be crucial for the complete functional interpretation of human genome sequences.
After they weeded out false positives (which were due to sequencing errors; of course, false positives are a problem in their own right in a clinical setting), the variants included indels (insertions or deletions of 1 or more nucleotides) that changed the splicing of the gene (and thus changed the amino acids that got strung together in the resulting protein), single nucleotide variants that introduced a stop codon into the gene sequence (that is, that caused transcription of the gene to halt prematurely), and large deletions that removed some of the coding sequence.  Some of the variants were found to affect all known protein-coding transcripts of the affected gene, and some affected only some of the coding transcripts.  That is, some transcripts were normal. 

The authors find that the common gene-disrupting variants described in this study are not a significant cause of complex disease.  Most of the LoF variants identified that are associated with complex diseases, all but one of which were heterozygous in these subjects (they had one functional copy, one not), are at low frequency, presumably due to purifying selection -- that is, selection against alleles that are severely harmful, thus preventing them from reaching high frequencies in any population. 

Individuals in this study have about 120 LoF variants, about 100 of these heterozygous and 20 of them homozygous -- both the person's copies are non-functional.  It's the homozygous LoF's that seem to have no effect that are of most interest.  The individuals in this study are healthy -- so either their LoF's truly have no effect, or they haven't yet had an effect.  If they are truly what MacArthur et al. are calling LoF-tolerant genes, and don't lead to disease, the authors suggest that they can be used to "define the functional and evolutionary characteristics that distinguish these genes from severe recessive disease genes."
We examined the 253 genes containing validated LoF variants that were found to be homozygous in at least one individual. These LoF-tolerant genes are significantly less conserved and have fewer protein-protein interactions than the genome average. They are also enriched for functional categories related to chemosensation, largely explained by the enrichment of olfactory receptor genes in this class (13.0% versus 1.4% genome-wide), and depleted for genes involved in embryonic development and cellular metabolism.
The finding that so many of the genes that are 'allowed' to vary are olfactory receptor (OR) genes isn't much of a surprise, as we all carry many OR genes that are pseudogenes, OR's that no longer function.  So, the researchers eliminated these from the set of genes that could lead to Mendelian disease, and then compared the remaining 213 LoF-tolerant genes with 858 known recessive disease genes.  They found these 2 categories were very different, and suggest that the characteristics of the recessive disease genes that are not shared with the LoF-tolerant genes could be used to prioritize candidate disease genes.

This can be important because we all have so many variants, and identifying which of these is or are contributing to a disease we may have is often impossible without data from other affected family members.  If likely candidates can indeed be prioritized based on the results of this study, this could be very helpful.  But, many genes are only deleterious in a given environment, or after years of exposure to environmental factors, and one person's LoF-tolerant gene might be another's disease gene.

To us, the fact that so many genes can apparently be disrupted with no discernible ill effect is further evidence of the adaptability that evolution has built in -- DNA replicating errors are common, the timing and locale of gene expression are often imprecise, and environments frequently change.  We're redundant, buffered, pretty hardy creatures!  The fact that all this can happen and we can live on with no ill effect is a beautiful fact of life.

Tuesday, February 14, 2012

Ptolemaic genetics: epicycles of lobbying

That was then...
Ibn al-Shatir's model for the
appearances of Mercury,
showing the multiplication of
epicycles in a Ptolemaic
enterprise. 14th century CE
(Wikimedia Commons).
Way back then, in the dark ol' days of science, the Roman astronomer Claudius Ptolemy (90-168AD) tried to explain the position of the planets in terms of divinely perfect circles of orbit around God's home (the Earth).  The idea that we were at the center of perfect celestial spheres was a standard 'scientific' explanation of the cosmos and our place in it.

But the cantankerous planets refused to play by the rules, and their paths deviated from perfect circles.  Indeed, occasionally the seemed to move backward through the skies!  Still, perfect circular orbits around Earth simply had to be true based on the fundamental belief system of the time, so astronomers invented numerous little deviations, called epicycles, to make the (we now know) elliptical orbital pegs fit the round holes of theory.

And then along came Nicolaus Copernicus (1473-1543 AD).  And the cosmos was turned inside out: the earth was not the center of things after all!

Thomas Kuhn famously described in The Structure of Scientific Revolutions how the best and the brightest scientists struggle valiantly to fit pegs into holes they don't really fit, until some bright person ccomes along and shows the benighted herd a better way to account for the same things.  Copernicus, Galileo, Newton, Einstein, and others were the knights in shining armor who inaugurated some of the most noteworthy of these occasional 'scientific revolutions.'  Darwin's evolutionary ideas are also a classic example.

The same kind of struggle is just what is happening now in genetics and evolutionary biology--indeed in many other fields in which statistical evidence runs headlong into causal complexity.  Whether, when, or what knightly change will occur is anyone's guess.

And this is now
Everyone remembers the hoopla the sequencing of the human genome was met with when it was announced (or rather, each time it was announced) -- we were promised that we would by now not only know why people were sick, but we'd be able to predict what we'd get sick with in future.  It was promised that this would be a silver-bullet reality by the early 21st century by no other than Francis Collins.  Others were promising lifespans in the centuries: all of us would be Methuselahs!

So, all those illnesses would now be treatable or preventable in the first place. How?  Well, the genome would allow us to identify druggable pathways, and common diseases must be due to common genetic variants (an idea that came to be known as common disease common variant, or CDCV), and if we could just identify them, we'd be in business.  After all, didn't Darwin show us that everything about everything alive was due to genetic causation and natural selection?  If that's the case, we should be able to find it, and our wizardry at engineering would take the ball and run with it.  Big Pharma jumped on the 'druggable' genome bandwagon and people running big sequencing labs jumped on the CDCV idea, and genomewide association studies (GWAS) were born.  And then the 1000 Genomes project, and all the -omics projects....  Big is better, of course!  Not that these efforts weren't questioned at the time, based on what everyone should have known about evolution and population genetics, but the powers-that-be plowed ahead anyway.

Well, we're no longer in a minority of naysayers.  It's widely recognized that GWAS haven't been very successful, relative to the loud promises being trumpeted only a few years ago.  And even the successes they have had -- and numerous genes associated with traits have been identified, it must be said -- typically explain only a small amount of the variation in disease, or any trait, in fact.  So now researchers are working on automating the prediction of disease from gene variants based on protein structure and other DNA-based clues.  But the assumption--the belief system, really--is still that the answer is in the DNA, and disease prediction is still going to be possible.

A piece in Feb 9 Nature describes a number of state-of-the-art approaches to predicting the effects of DNA variants, in part based on what amino acid changes do to proteins.  The idea now is that diseases are going to be found to be due to rare variants, and the challenge is to figure out what these variants do.  In part, evolution will help us to do this.
"Sequencing data from an increasing number of species and larger human populations are revealing which variants can be tolerated by evolution and exist in healthy individuals."
But, are we trying to explain a current disease, or predict the diseases someone will eventually get? These are different endeavors, though it may often be inconvenient to acknowledge that.  Rare pediatric diseases that are due to single genetic mutations, or genetic diseases that cluster in families (and, again, usually with young onset age and rare) are easier to parse than the complex chronic diseases that most of us will eventually get.  But, based on the comparison of the genomes that have already been sequenced, we now know that we all seem to differ from each other at something like 3 million bases.  That is, we all have a genome that has never existed before and never will again. Assigning function to all that variation is from daunting to impossible -- not least because a lot of it might not even have a function.  And the idea that we'll eventually be able to make predictions from those variants is based on questionable assumptions.

It's true in one sense that every disease we get is genetic -- everything that happens in our body is affected by genes -- but in another sense, much of what happens is a response to the environment, and so is environmentally determined--that is, is not due to genetic variation in susceptibility.  Predicting a disease from genes when it's due to combined action of genes and environment, therefore, is a very challenging problem.

Here is just one example of why: Native Americans throughout the Americas are about 65 years into a widespread epidemic of obesity, type 2 diabetes and gallbladder disease, diseases that were quite rare in these people before World War II.  There are a number of reasons to suspect that their high prevalence is due to a fairly simple genetic susceptibility.  But, if gene variants (still not identified) are responsible, they have been at high frequency in the descendants of those who crossed the Bering Straits from Siberia for at least 10,000 years -- which means that variants that are now detrimental were "tolerated by evolution and exist[ed] in healthy individuals" for a very long time.

If geneticists had wanted to predict 70 years ago what diseases Native Americans were susceptible to, these variants would have been completely overlooked, because they weren't yet causing disease.  And indeed these 'risk' genes, whatever they be, were benign -- until the environment changed.  We're all walking around with variants that would kill us in some environment or other, and since we can't predict the environments we'll be living in even 20 years from now, never mind 50 or 100, the idea that we'll be able to predict which of our variants will be detrimental when we're old is just wrong. In fact, we're each walking around with substantial numbers of mutant or even 'dead' genes, with apparently no ill effect at all -- but who knows what the effect might be in a different environment.

But, ok, some of us do have single gene variants that make us sick now.  Many of these have been identified, most readily when a family of affected individuals is examined (though the benefit of knowing the gene is rarely of use therapeutically), but many more remain to be.  The current idea is that this can be done by looking for mutations in chromosome regions that are conserved among species, and figuring out which of these change amino acids (and thus the protein coded for by the gene).  The idea is that unvarying regions are unvarying because natural selection has tested the variants that arose and found them wanting, thus eliminating them from the population.  They must, therefore, be functionally important!
A host of increasingly sophisticated algorithms predict whether a mutation is likely to change the function of a protein, or alter its expression. Sequencing data from an increasing number of species and larger human populations are revealing which variants can be tolerated by evolution and exist in healthy individuals. Huge research projects are assigning putative functions to sequences throughout the genome and allowing researchers to improve their hypotheses about variants. And for regions with known function, new techniques can use yeast and bacteria to assess the effects of hundreds of potential mammalian variants in a single experiment.
This is potentially useful, because for those with single gene mutations that cause disease -- 1 variant among 3 million other ways in which each person differs from everyone else -- homing in on the causative mutation is, again, difficult to impossible if you don't have a large family with similarly affected individuals in which to confirm the association of mutation and disease.

Well, if we can do with or without a protein (or other functional DNA element), depending on the variation we have across the genome, then even when the element is important its variation in a given individual may not be causal: there are many examples where that is clearly true.  Further, the same kind of evolutionary reasoning would say that centrally important -- and hence highly conserved -- parts of the genome probably cannot vary much without being lethal, largely to the embryo.  So, from that equally sound Darwinian reasoning, we would expect that disease-associated variation will be in the minor genes with only little effect!  So the 'evolutionary conservation' argument cuts both ways, and it's not at all clear which way its cut is sharpest.  It's a great idea, but in some ways the hope that searching for conservation will bail us out, is just more wishful thinking to save business as usual.

Methuselah (Della Francesca ca. 1550) 
To complicate things even more, not all amino acid changes cause disease, or even do much of anything.  And again, sometimes they will only be harmful in a given environment.  And, of course, not all diseases are caused by protein changing mutations -- sometimes they are caused by disturbances to gene regulation.

In fairness, the multitude of researchers trying to make sense of the limitless genetic variation that is pouring out of DNA sequencers recognize that it's complicated.  But then, why are they still saying things like this, as quoted in the Nature piece: “The marriage of human genetics and functional genomics can deliver what the original plan of the human genome promised to medicine.”

What's to the rescue?  Do we need another 'scientific revolution'?
We have no idea when or if our current model of living Nature will be shown to be naive, or whether our understanding is OK but we haven't cottoned on to a seriously better way to think about the problems, or indeed whether the hubris of computer and molecular scientists' love of technology will, in fact, be victorious.  If it comes, it could be.  But we are certainly in the midst of a struggle to fit the square truths about genetics and evolution into the round holes of Mendelian and Darwinian orthodoxy.

Perhaps the problem to be solved is how to back away from enumerative, probabilistic, reductionistic treatment of complex, multiple causation, and to make inferences in other ways.  We need to understand causation by numerous, small or even ephemeral statistical effects, without our current enumerative statistical methods of inference. In terms of the philosophy of science, doing that would require some replacement of the 400 year-old foundations of modern science, based on reductionistic, inductive methods that enabled science to get to the point today where we realize that we need something different.

The situation here is complicated relative to scientific revolutions in Copernicus', Newton's, Darwin's or even Einstein's time by the large, institutionalized, bureaucratized, fiscal juggernaut that science has become. This makes the rivalries for truth, for explanations that this time will finally, really, truly solve the complexity problem even more frenzied, hubristic, grasping, and lobbying than before.  That adds to the normal amount of ego all of us in science have, the desire to be right, to have insight, and so on.  Whether it will hasten the inspiration for a transforming better idea, or will just force momentum along incremental paths and make real insight even harder to come by, is a matter of opinion.

Sadly, the science funding system, including the role of lobbying via the media, is so entrenched in our careers, that dishonesty about what is claimed to the media or even said in grants is widespread and quietly acknowledged even by the most prominent people in the field: "It's what you have to say to get funded!", they say.  But where does dissembling end and dishonesty begin when it comes time to the design and reporting of studies (and, here, we're not referring to fraud, but to misleading results and over promising the importance of the work)?  The commitment to the ideology and the promises restrains freedom of thought, and certainly dampens innovative science.  But it's a trap for those who have to have grants and credit to make their living in research institutions and the science media.
Zip-line over rainforest canopy,
Costa Rica (Wikimedia)

But right now, scientists are like tropical trees, struggling mightily to be the one that reaches the sunlight, putting the others in their shade. What we need is a conceptual zip-line over the canopy.