Monday, April 8, 2013

Epigenetics isn't everything but it is something

'Epigenetics' is the new 'gene for.'  A good way to tell if a field in biology is hot is if it's become an -ome, with a Wiki page, and epigenetics has.  The 'epigenome' will now explain everything from why identical twins aren't, to why we get the diseases we'll get and why we behave as we do, and science studies and gender studies and social critics are using epigenetics to reconfigure their approach to understanding how biology and society are intertwined.

Epigenetics
There are a couple of issues here.  One is that some people use 'epigenetic' to refer to things like, say, obesity due to overeating because diet interacts with one's genetically based metabolism.  However, the current 'hot' meaning is that environmental factors directly affect gene expression, rather than just the result of normal gene expression.  So, a dietary component might lead to a gene being inactivated.

Now the fashionable aspect of this is, of course, that those in the area seem to want everything to be 'epigenetic' -- after all, one can get funded to do the epigenomics!  We mean not to disparage the field wholesale -- there is certainly something to it, but, like the human genome project, the epigenome can't possibly fill all the promises being made in its name.  It is, in this sense, a political ploy for funding and attention, as well as an enthusiasm for what seems new and possibly profound.

Still, epigenetic effects mean that DNA sequences alone do not determine what genes do, and the epigenetic modifications can have substantial effects, and be inherited.

The state of the art
A paper in the latest Trends in Genetics ("Bridging the transgenerational gap with epigenetic memory," Lim and Brunet) does an admirable job describing what's actually currently known about epigenetics, or 'non-Mendelian', non-genetic inheritance.  They point out that patterns of non-genetic inheritance have been described for almost 100 years, including Waddington's description of the inheritance of wing patterns influenced by heat shock in fruit flies in the 1940s -- it was Waddington who coined the term 'epigenetic', though his meaning was more general than ours today.

Parental imprinting, discovered during the 1980's, is another example of non-genetic inheritance.  This is when alleles from only one parent are inherited, the other set being silenced by DNA methylation and/or histone modification.  At the same time, Lim and Brunet point out, the discovery of transgenerational epigenetic inheritance (TEI) in mice was reported, in that case affecting coat color.  And many more instances have been reported since then, in many different organisms.  It became apparent as well that epigenetic modifications could last for at least several generations.

Lim and Brunet describe a recent experiment in which
mice with an insertion of LacZ into the Kit gene gave rise to genetically wild type offspring that still exhibited the tail and paw color phenotype characteristic of Kit mutants for at least two generations.  Genetically wild type descendants of ancestors that had the Kit mutant phenotype showed an altered pattern of Kit RNA expression, with RNA molecules of shorter size in brain and testis.  Microinjection of RNA from heterozygous mutant animals into one-cell embryos was sufficient to recapitulate the mutant phenotype in the following generation.
Our objections to the use of 'wild-type' and 'mutant' notwithstanding, the persistence of the traits from the altered mice into the next generations is of interest. The authors suggest that the transgene (the LacZ inserted into the Kit gene) disrupts a specific locus in the parental genome, which causes the production of abnormal RNA in sperm, which is then transmitted to at least the next two generations.  How this 'epigenetic memory' works is not yet clear but the apparent role of RNA in this process has been replicated in experiments with other traits.  One example is the injection of fertilized eggs with micro RNAs targeting specific enzymes that regulate cardiac growth, which had the capacity to slow the growth of the heart for several generations, though it seems that not all RNA has a similar capacity. 

Work has recently been done on transgenerational epigenetic inheritance in the model worm, C. elegans, as well, specifically on longevity and fertility.  The inheritance of sterility and longevity influenced by histone modifications -- demethylation -- has been documented, and observed to last for at least 5 generations. 

TEI and environmental stimuli
Another mode of TEI, which is perhaps more relevant to life as it's lived outside the lab, is that that might be induced by metabolic changes.  Over- or undernutrition of parents is the example we've heard most about, with respect, e.g., to the multi-generational consequences of widespread famine. 
Exposure to a chronic high-fat diet in rat fathers results in impaired insulin metabolism and pancreatic cell gene expression in female F1 offspring.  Female offspring from fathers fed a high fat diet mated with mothers fed a control diet exhibited an increase in blood glucose ... and a decrease in insulin secrettion compared with offspring with both parents fed a control diet.
Analysis of gene expression in islet cells showed differential levels of expression of various types, including signaling factors.  Other experiments have shown altered phenotypes and gene expression two generations after male mice were overfed, suggesting again that this is due to a TEI rather than genetic changes.  Some hypothesize that the epigenetic changes are to the contents of sperm and seminal fluid (which includes chromatin, RNA and metabolites).

The effects of undernutrition can also be inherited, including "increased expression of genes involved in fat and cholesterol biosynthesis" and genes involved in DNA replication.  Alterations in lipid metabolism and cell proliferation have been demonstrated in offspring of mice who've been deprived of food.  At least one study demonstrated epigenetic changes in the sperm of parental mice on restricted diets.

The effects of famine during World War II on a large family cohort in the Netherlands have been examined, and metabolic consequences shown to last for at least two generations.  A family cohort from 19th century Sweden shows much the same, as well as that food intake during adolescence of grandparents correlated with survival of grandchildren, suggesting that there may be a critical period for production of healthy gametes.

Further examples include the effects of heat shock in Drosophila, TEI of small RNAs from viruses in C. elegans with gene silencing consequences, TEI of behavior patterns such as depression, via exposure to psychological stress in utero, and olfactory imprinting behavior in C. elegans.

Transgenerational epigenetic inheritance can be via DNA methylation, histone modifications, noncoding RNAs, short RNAs and other aspects of RNA function.  Lim and Brunet point out that the mechanism or mechanisms behind these modes of inheritance are not yet well understood.  They list questions that remain to be answered, including how the signals are eventually erased, how the changes are maintained, and whether the strength of the environmental stimulus affects the number of generations the effect is maintained.

TEI and evolution
Finally, they suggest that because transgenerational epigenetic inheritance is known to occur in many organisms, it must have been selected.  They postulate that perhaps it was advantageous to pass on information about the environment -- but environments change so quickly that it seems more likely to us that the ability to adapt would have been what was selected for, rather than the ability to stay the same, and TEI is certainly one adaptive ability.

Or, they suggest, TEI might increase the evolvability or rate of evolution of an organism.  Perhaps it affects the accessibility of chromatin to DNA repair enzymes, thus making certain loci more and less mutable.

An active mechanism that can obscure
There are active genetically encoded mechanisms for applying and removing epigenetic changes such as methylation.  These can be very specific and differ between males and females.  Much of the genome is 'set' differently in the generation of sperm or egg cells.  But once modified, unless it is re-set each generation, the effect can appear to be DNA-encoded in epidemiological studies but non 'Mendelian': family members share the trait, but not because of specific DNA variants, since some may have inherited modified, and others unmodified copies of the same variant. This is one major reason why epigenetic factors can obscure studies, like genomewide scans, to find genes that contribute to important traits like disease.

Not non-Mendelian!
Note that the phrase 'non-Mendelian inheritance' is thoroughly wrong, but probably ineradicable from the current jargon.  Genes are inherited in a Mendelian way.  Each of us carries two instances of human genomes, and we more or less randomly transmit a copy of one of them to each sperm or egg cell.  This random transmission is what is Mendelian:  Mendel didn't know about genes, but used traits to signal the inheritance of these 'elements' or 'factors'.

But the traits are not--that is not--inherited!  Only the genes are inherited.  It is only if traits are tightly tied to specific genetic variants that the appeance of the trait is highly correlated with the genetic variant that was inherited that the trait seems to be 'Mendelian'.  Therefore epigenetic patterns of occurrence in families that are not 'Mendelian' refer to alleles that were inherited but that, because of epigenetic modification, their effect is not manifest.  The inheritance itself is in these instances not affected.

It is very sloppy to speak of Mendelian inheritance in regard to any phenotype, and we can lead ourselves into trouble if we aren't careful.  There is, in fact, a phenomenon of non-Mendelian inheritance (called segregation distortion) in which the two alleles a person carries are not transmitted with equal probability.

We also noted more verbal sloppiness in the literature, by terms such as 'non-genetic inheritance' and 'wild type'.  So, we quoted this earlier: "an insertion of LacZ into the Kit gene gave rise to genetically wild type offspring."  If the mouse is genetically altered, how can it have the wild type genotype?  What was meant, we think, was that the genetically modified animal had the phenotype of the unmodified animals.  But it inherited the modified not the wild type genome.

But we've probably said enough already about how self- as well as other-misleading such loaded terminology can be.

Friday, April 5, 2013

Harem-scare 'em? Is polygamy next? Why not?

The country is moving rapidly from cringing at gays, and gay marriage, to acceptance of any loving couple relationship as equally legitimate before the law, in terms of rights and responsibilities.  This post is triggered by the remarkably rapid turn of events by which even some Republicans are emerging from what seem like cretinous shells to accept, if not avow, the acceptability of gay coupling in our society.

The issue is not new, nor specific to our time and culture. These things are always based on cultural judgments. Homosexuality is related to mating, which is related to deep emotional issues of many kinds.  Mating-related issues include sex itself, procreation, resource acquisition and protection, group coherence and alliance, and so on.  But biologists favoring the relaxed laws are now quoted in the op-eds as lending Darwinian support to the argument, by claiming that because homosexual activity is widespread in nature, we should legalize it.  Of course, the do has nothing to do with the should.  The latter are cultural.

Unlike other species, religion is also part of any culture's coherence and solidarity, as well as resource control mechanisms.  In our culture we allow many religions, and they differ on almost any issue one can name.  Sex and marriage are among them.  In our society, religion cannot prevent homosexual feelings, but if one's religion prohibits gay activities or gay unions, then a believer must follow those tenets.  (Actually, other species have social norms and these can be locally variable.  Some may in some ways react to, or even try to suppress, homosexual activity)

Homosexuality is part of the naturally occurring variation in humans, and many if not most other species.  Because of its potential effects on our material interests of the above sort, and because humans are tribal and proud of it, society may care strongly and emotionally about it.  Policy may change this way or that over time.  That's because these are cultural responses.  They have nothing to do with the biology itself, and culture cannot dictate what is 'natural' (no matter that it may claim to do so). 

Biologists have no business (as biologists) chiming in here, no matter that we see their papers and op-eds  noting how widespread homosexuality is (and, presumably, therefore we should remove legal restrictions).  Comparative biology is absolutely irrelevant to the issues, and biologists provide zero relevant knowledge in this regard.  The issues are the same whether sexual preferences were hard-wired in our genomes or as totally free-willed and non-genetic as what brand of potato chips one chooses to eat. 

So we are moving towards an acceptance, even as unremarkable, of gay unions, and most of us think that's a good thing because it allows personal freedom when that freedom doesn't interfere with others' freedom.

There is no seriously new kind of precedent here.  Multiple sex partners, 'fornication', divorce, menages a trois, not to mention miscegenation, and so on have also histories of prohibition and acceptance.  And, of course, so do very delicate issues like the age of sexual consent, for which there is certainly widespread cultural variation and definition around the world.

Most educated people are aware of this cultural arbitrariness, but many tea heads may be so rooted in their fears and the here-and-now that they're not.

One of the arbitrary aspects of sexual behavior is polygamy.  This issue has been put back on the table in recent years, often by reactionaries objecting to gay rights.  We're not expert on religion, but certainly the sacred texts of Muslims and Mormons allow, if not encourage, polygamy.

Polygamy is common in cultures around the world, both contemporary and indigenous.  There are rules based on kinship, religion, group identity, social status and the like that, in any given polygamous culture, regulate it, as cultures regulate other forms of sexual union.

We're not here advocating such a change, nor opposing it, but just commenting upon it from a scientific (anthropological as well as biological) point of view.  It is worth thinking about, because it challenges the degree to which most people seem to believe there are ultimate rights and wrongs--even though we know these can change, as the move towards accepting gay marriage shows.

Is legalizing polygamy going to be on the front burner in the near future?  There are absolutely no objective reasons why we should prohibit polygamous unions any more than we do gay ones.  In both cases, as in normal heterosexual marriages, we have protections against abuse, limits on resource hoarding, rules about the care and education of children, and so on.

How far to go?
If we are scientists, we should also consider how far one should go in this kind of reasoning.  Is there any objective guideline for these sorts of thing?  Suppose it to be true, as so often joked about, that sex with animals, as in the proverbial shepherd-sheep escapades, is also part of the normal spectrum of human sexual behavior.  Should this bestiality be something to which the law is to be distanced in the name of freedom?  We would not need to deal with 'marriage' here, since that can objectively be defined as involving humans since it's about human rights, property and the like--but then, why can we leave property to dogs in our wills?

We're just musing here, but it is important to think about where our ideas and sense of meaning comes from, especially when the discussion can be made, or forced, to hinge upon claims about what is 'natural'.  There is nothing wrong with drawing lines around what is acceptable, but we should realize the culture-specific and in that sense highly arbitrary nature of those lines.

Thursday, April 4, 2013

When the crystal ball is cloudy: calling sequence data correctly

Here's a monkey wrench of a paper (O'Rawe et al.), just published in Genome Medicine.  We're all being sold on the idea that knowing our whole genome sequence is going to make us much healthier. The DNA sequencer cum crystal ball will tell us what we're likely to be in for, and this will give us plenty of lead time to prevent it -- by running, lowering our cholesterol intake, losing weight, or whatever -- or to prepare for it.

But, among many other assumptions, this assumes first and foremost that the data are being read correctly, no false positives or negatives.  And here's the clincher: O'Rawe et al. compared five different software packages that read and interpret DNA sequence data, and they report low concordance between results.  Discrepancies have been found before, but not when comparing reads of the same raw data.

This group sequenced whole exomes of 15 different individuals, in 4 families, and fed the raw data through 5 sequence analysis pipelines.  They also sequenced one whole genome.  Sequences were done at 20 - 154X coverage, 120X average, meaning each nucleotide was read at least 20 times, but most often more, and at least 80% of the target sequence was obtained.

They found that the 5 programs agreed on single nucleotide variants (SNVs) about 60% of the time.  That is, 40% of the time a SNV was called by fewer than 5 of the programs.  Each of the pipelines detect variants that the others do not, and they aren't necessarily all false positives.
This disagreement is likely the result of many factors including alignment methods, post alignment data processing, parameterization efficacy of alignment and variant calling algorithms, and the underlying models utilized by the variant calling algorithm(s).
That is, each step along the way potentially introduces errors.  Indel (insertion/deletions, segments of DNA one or more nucleotides in length) concordance rates were even lower, at 26% between three indel calling programs.  (The paper goes into much more detail about specific pipelines and error rates.) Using family data can help reduce inaccuracies when it is possible to determine which calls just cannot be correct.  But, otherwise, with current methods reducing false positives means increasing false negatives, and vice versa.

The authors write,
In the realm of biomedical research, every variant call is a hypothesis to be tested in light of the overall research design. Missing even a single variant can mean the difference between discovering a disease causing mutation or not. For this reason, our data suggest that using a single bioinformatics pipeline for discovering disease related variation is not always sufficient.
This somewhat understates the problem.  Serious level testing of a SNP (single nucleotide polymorphism) to see if it has an effect on disease risk--especially when these effects are typically very small in any case, and biased upwards in GWAS type data, is no joke.  What do you do?  Put that single change into a lab mouse or rat and see if it might be more likely to develop slightly higher blood pressure at old age?  Or have a slightly higher risk of some sort of cancer (again, to be a human model, it should be at older ages)?  Which mouse strain would you use?  If humans are to be used for validation, how would you do it?

The questions are serious because miscalls by sequencers go both ways.  A sequencer can miss a SNV call, so you don't identify one of the variants that you really want to be checking.  Or, it can give you a false positive, and lead you farther astray.  And if you must choose between hundreds of variants across the genome, with comparable estimated effects, you are already in a bit of a bind even if they are all perfectly called!

No technology, or medical test, will be correct 100 percent of the time, and sequencing technologies are likely to get better, not worse (though if MS Windows is any guide, that's not necessarily true!). But, when disease risk estimates depend on accurate DNA sequence, it is obvious that we are way premature in proclaiming findings so loudly and demanding that so much effort and resources be poured into doing more of the same.  Again, focused studies on problems more important, clearer, and less vulnerable to these kinds of errors is where the effort should be going.

And, some subtle manipulation, too?
By the way, the standard term for a single nucleotide variation in a population is SNP (single nucleotide polymorphism).  Now, some authors use SNV (single nucleotide variant), essentially doing two things.  First, they are rhetorically equating 'variant' with causal variant--that is, tactily, subtly, or surreptitiously planting in your mind that they are onto something causal.  And second, they are tacitly, subtly, or surreptitiously suggesting that one of the two is the 'good'  or 'normal' (i.e., health-associated) variant.  This perpetuates the 'wild type' thinking--see our earlier post 'walk on the wild-type side'.

These are ways in which the community of researchers inadvertently or intentionally (you decide which) cooks the books in your and journalists', and even their own minds, entrenching a de facto genetic-causation worldview into their and everybody's thinking.  That's good for business, of course.

Wednesday, April 3, 2013

Make a bee-line for truly important research!

You may think it's been too cold too long, but this was a really hard winter for honeybees.  Winters take their toll on bees even in a good year, with 5 - 10% mortality, but with the 'colony collapse disorder' (CCD) that has been affecting honeybees since it was first reported after the winter of 2006-7, mortality has risen to 20-30% and more.  If you like to eat, that's already pretty ominous news, since bees fertilize much of our food sources, but a story in The New York Times last week reports that this year 40 to 50% of all hives were wiped out (the accompanying video is worth a look), and no one is sure why.
“They looked so healthy last spring,” said Bill Dahle, 50, who owns Big Sky Honey in Fairview, Mont. “We were so proud of them. Then, about the first of September, they started to fall on their face, to die like crazy. We’ve been doing this 30 years, and we’ve never experienced this kind of loss before.”
This is of course devastating to beekeepers, but it's also going to be devastating to farmers who depend on bees to pollinate their crops -- almonds in California are a huge such crop.  A story at NBCNews.com reports that bee pollination is responsible for $15 billion in increased food value every year, perhaps a quarter of all foods.  And food losses will mean higher prices.  If this keeps getting worse, our own species itself could be in danger of starvation.  So of course we look hopefully to science to explain, and stop, the devastation of our buzzing friends.

According to the US Environmental Protection Agency, and this list of possible causes is fairly standard,
There have been many theories about the cause of CCD, but the researchers who are leading the effort to find out why are now focused on these factors: 
  • increased losses due to the invasive varroa mite (a pest of honeybees);
  • new or emerging diseases such as Israeli Acute Paralysis virus and the gut parasite Nosema;
  • pesticide poisoning through exposure to pesticides applied to crops or for in-hive insect or mite control;
  • bee management stress;
  • foraging habitat modification
  • inadequate forage/poor nutrition and
  • potential immune-suppressing stress on bees caused by one or a combination of factors identified above.
  • potential immune-suppressing stress on bees caused by one or a combination of factors identified above.
Additional factors may include poor nutrition, drought, and migratory stress brought about by the increased need to move bee colonies long distances to provide pollination services.
These possibilities have been proposed since the onset of CCD, and they are still live possibilities, but most have proven less explanatory than they'd seemed, and they don't explain why this winter was particularly hard.  Perhaps it was the drought in the midwest followed by a hard winter, though some beekeepers are reporting heavy losses despite good summer conditions.  The increase in pesticide resistant mites is another possibility, or viruses.  Or, perhaps it's a number of stressors in combination.

The explanation getting the most play these days is the increasing use of pesticides, fungicides and herbicides, although the EPA says there is no definitive evidence that pesticides are the cause ("To date, we’re aware of no data demonstrating that an EPA-registered pesticide used according to the label instructions has caused CCD."). And indeed, each of the chemicals now used on crops has been certified safe, but are our guardian officials being too lenient?  For example, any given combination may have unforeseen effects, and combinations haven't been tested.  Of particular concern is the only new class of pesticides developed in the last 50 years, neonicotinoids, derived from nicotine and developed in the 1980s and 90s.

A paper published online in Science March 29, 2012 (Whitehorn et al.) reported that neonicotinoids indeed do have a negative effect on bees.
Growing evidence for declines in bee populations has caused great concern because of the valuable ecosystem services they provide. Neonicotinoid insecticides have been implicated in these declines because they occur at trace levels in the nectar and pollen of crop plants. We exposed colonies of the bumble bee Bombus terrestris in the laboratory to field-realistic levels of the neonicotinoidimidacloprid, then allowed them to develop naturally under field conditions. Treated colonies had a significantly reduced growth rate and suffered an 85% reduction in production of new queens compared with control colonies. Given the scale of use of neonicotinoids, we suggest that they may be having a considerable negative impact on wild bumble bee populations across the developed world.
A second paper published in Science at the same time  (Henry et al.) tested the effects of a sublethal dose of a single one of these compounds on the homing behavior of honeybees, suspecting that it might affect the bee's ability to find its way home because of how it affects the insect nervous system.
They are highly potent and selective agonists of nicotinic acetylcholine receptors, which are important excitatory neurotransmitter receptors in insects.  Effects of sublethal neonicotinoid exposures in honey bees may include abnormal foraging activity, reduced olfactory memory and learning performance, and possibly impaired orientation skills.
They found that the neonicotinoid they tested affected forager survival, which may indeed have severe consequences for the survival of the hive.

Neonicotinoids are applied to the seed, and then travel through the sap to all parts of the plant as it grows.  They are said to be less toxic to mammals than other pesticides, and so have been used more liberally.  Because of the suspicion that they may be at least one of the agents responsible for colony collapse disorder, they've been banned in some European countries, and the ban may widen throughout Europe.  Indeed, Whitehorn et al. conclude their paper "...we suggest that there is an urgent need to develop alternatives to the widespread use of neonicotinoid pesticides on flowering crops wherever possible."

Several of these compounds are now under review by the EPA to determine whether they still meet requirements for certification.  If we were to bet, we'd bet they do.  Big agriculture relies heavily on chemicals to grow the food we eat.  Much less of this would be needed in more traditional, smaller-scale less corporately-tied agricultural practices, that many argue could still feed the earth.  Enough said.

Priorities when there's too much on our plate
We are currently pouring research resources into massive but mildly incremental topics like genomic disease and personalized genomic medicine (PGM).  Many, if not all, of these are the common diseases we get after living a long time in a sedate, well-fed (or over-fed) lifestyle.  These diseases are consequences of ease and privilege, and could clearly be  prevented, or greatly delayed, by basically painless changes in how we life.  They are not 'genetic' in any serious sense.  As a result, the payoff of these studies, in most cases, even if things were to work out as promised by PGM's advocates, would with some exceptions be exceedingly not-exceeding.  Indeed, it would be minor.  Minor relative to using the experience of relatives (heritability) rather than individualized genomes, minor relative to the baseline risk, minor relative to environmental exposures, and minor because genomic risk doesn't generally lead to gene-specific treatment.

Meanwhile, we really do have an important problem, one with orders of magnitude more potential for harm if not understood quickly and enormous potential for human good: colony collapse in bees.  A sane research policy would be aimed at solving societal problems in a rational priority order, rather than the vested-interest order that so predominates today.  These areas pale in importance compared to the problem of having adequate food.  That's even more important than climate change, though climate change may be a major threat to agriculture and our food sources as well.

Major problems often turn out to be complex and difficult to solve, and CCD may or may not turn out to be simple.  But it is an example, along with others like antibody resistance and overpopulation that are huge threats to our essential well-being.  Why aren't we pulling funds from what is sexy and media-exploitable research, but is entrenched and in many ways about problems of privilege, to areas that are much closer to the nitty-gritty of our very survival?

Tuesday, April 2, 2013

Are we closer to personalized genomic medicine yet?

A large study of 200,000 people with breast, ovarian or prostate cancer and controls was reported widely last week as one of the most comprehensive studies ever of these diseases.  The study was carried out by the  Collaborative Oncological Gene-environment Study, COGS, and the thirteen papers published in five different journals were collated by Nature and presented in an open access site, with commentary.

The aim of the study was to identify risk factors with which to stratify the population by risk status, and thus to be able to determine who to most intensively screen to prevent or detect cancers early. They identified 74 new genetic loci (chromosome regions, whether or not an actual causal site has been found, as usually it hasn't, to date) associated with risk; 49 with breast cancer, 8 for ovarian cancer and 29 for prostate cancer.  There's a lot of material, and we can't cover it all, but we'll give it a go.  Keep in mind that they aren't yet saying they've identified genes 'for' these cancers, but that their genomewide associate studies have located chromosomal regions that are statistically associated with disease, and the assumption is that they are involved in affecting the probability or risk of disease.  

Prostate
The genetic loci that increase risk of prostate cancer that have been reported to date have been common and 'low-penetrance' alleles, meaning that even if a man carries one or more of them his risk of getting prostate cancer is low.  The new study identifies 29 new loci associated with susceptibility to prostate cancer.  They are less common than the previously identified risk factors, presumably found by the COGS now because their study is larger than previous studies.

The consortium stratified their sample into aggressive and non-aggressive disease and found different though overlapping patterns of loci associated with each.  Of the new loci, 16 are associated with both forms of the disease.  The researchers say that each of the loci contains 'plausible' candidate genes.  That is, the causal genes in the chromosomal location the disease mapped to have not yet been identified.  It is a slight stretch to say 'plausible' rather than 'possible', but that temptation's not unusual, and identifying the true causal genes can be tricky (e.g., see our previous posts on mapping genes for traits, here and here.)  But ok.  The paper discusses these candidates, but the consortium can't (yet) confirm any of them.

More than 70 loci have now been reported for prostate cancer, which, according to the researchers, explains "~30% of the familial risk for this disease." That means that this is found in those who already have a family history of risk--that is, an enriched subsample most likely to be carrying genetic risk factors. This study reports that those at highest genetic risk of prostate cancer, the top 1% of the risk distribution in the population, are at 4.7-fold higher risk than average.  High risk was considered to be aggressive prostate cancer or prostate cancer in patients younger than 55.

Essentially this means that if we tote up your alleles in these 70 regions, and add up their separate independently estimated effects, the net risk is 4.7 times higher than that of a random person in the population.  These are subtle and often slippery issues when it comes to assessing how important the actual risk assessment is--even if the method of computing summed risk is appropriate.
 
Breast and ovarian
Twenty-seven loci associated with familial breast cancer risk have previously been identified, accounting for about 30% of familial risk.  This study reports 49 new susceptibility loci, and estimates that approximately 1000 genes will eventually be found to be associated with risk of breast cancer. 

With respect to genes already known to be major risk factors for disease, the study looked for additional genetic markers in women who carry BRCA1 or 2 mutations, which put them at higher risk of breast and ovarian cancers than the general population.  They replicated previous findings, and also identified new 'modifier genes' that increase risk of breast cancer in BRCA1 carriers, and 8 new genes that increase risk of ovarian cancer, bringing the total to 10 and 12 known genes associated with breast and ovarian cancer, respectively. They identified a new risk allele specific to BRCA2 carriers, as well.  They also found 21 genetic loci associated with risk in east Asian women. 

The consortium reports that these results will one day allow those BRCA1 and 2 carriers at highest and lowest risk of developing cancer to be identified.  Risk estimates for BRCA2 carriers now range from "21–47% risk of developing breast cancer by the age of 80 years for the 5% of the BRCA2 mutation carriers at lowest risk compared to 83–100% risk for the 5% at highest risk" (source).

It is a somewhat separate question what the same non-BRCA sites do, if anything, in those not carrying BRCA mutations.

Gene-environment interactions and breast cancer
The consortium made a stab at addressing the question of gene-environment interaction, acknowledging that risk alleles are rarely determinative (they do apparently believe they've found at least one set of genes that may increase risk to 100% for some breast cancer).  Even the risk associated with BRCA1 and 2, among the most clear cut genetic risk factors known, has been shown to vary considerably by year of birth, presumably because of the interaction between the gene and some environmental risk factor.

King et al., 2003; Science 302:643-646
The study looked at whether relative risk associated with other risk alleles was modified by 10 factors already thought to be associated with risk: age at menarche, parity, breastfeeding, body mass index, height, oral contraceptive use, use of hormone replacement therapy after menopause, alcohol consumption, smoking, and exercise.

They replicated previously reported interactions between specific loci and parity and alcohol consumption, with risk increasing with lower parity and, they assume, a genetic variant of the LSP1 gene, and with a variant of CASP8 and alcohol consumption.

Import
A commentary in Nature Genetics discusses the public health implications of the work.  Remember that the aim of the study was to facilitate the identification, on an individual basis, of those at highest risk of breast, ovarian and prostate cancer, in order to prevent and detect disease.  But how to translate this to population public health measures is another question. 

But we've got a ways to go on this. Tests for prostate cancer are very unreliable, and there is no screen for ovarian cancer.  Whether a risk prediction is accurate may depend on whether the cancer is found because it is symptomatic or by extensive testing.  That's an issue since it's known that many or perhaps even most prostate cancers, and at least some breast cancers, regress on their own.

The chromosome regions now identified to increase risk of ovarian cancer are estimated to double the average population risk to just 3%.  The cost-effectiveness of mammography is still the subject of intense debate, and genetic susceptibility to cancer, even for those alleles known to be major risk factors, is never 100% and usually much lower. And testing carries its own risk, as does being told you're at higher risk--so such information had better be accurate if this is about public health!

Further, risk is always estimated from past environmental exposures, and future environments cannot be predicted.

Family history and the simulation of mendelism
We already know that family history is a good--generally, clearly one of the best predictors of these kinds of diseases.  So far, genetic risk studies haven't done much to improve on that; whether the current batch does isn't clear to us, at least not yet.

The other major risk factors are sex and age.  Generally these are far more powerful than individual gene identification, but of course the reason would be at least in part because family history reflects the presence of those genetic variants.  We also know that sampling and studying families with multiple cases enriches for whatever factors, genetic or environmental or even tendency to have testing done by high-grade clinics, in those family members.  This can make the presence of the disease seem Mendelian, as if due to a single or very few genes.  Such 'simulation of Mendelism' can raise the apparent risk of an identified gene shared by the affected family members to values far above those the same would have in the population at large.

This kind of bias is but one of several that make actual risk estimation very difficult, and GWAS type case-control studies are known theoretically and empirically to generate inflated estimates of effects.

So there is a lot that must follow these studies if they are to be judged to satisfy the media noise that they have generated.

Monday, April 1, 2013

58th Carnival of Evolution

The 58th Carnival of Evolution is up, over at Synthetic Daisies. A nice long collection of evolution-related blog posts. 

How to tell a woodpecker from a monkey

Humboldt
The 1700-1800's were great times for intrepid naturalists and collectors of insect and plant specimens who traveled the world in search of exotica.  Alexander von Humboldt was one.  From 1799-1804 he traveled throughout the Amazon region of South America, collecting specimens and taking notes for his extensive 21 volume description of his travels. Humboldt brought scientific instruments and a belief that much could be understood through systematic observation to his travels.

Darwin's 5 year voyage
Charles Darwin and Alfred Russel Wallace were both inspired by Humboldt's work as they set off on their own travels.  Darwin, of course, spent 5 years as naturalist on The Beagle, traveling around South America, Africa, Australia and back again to England.  He collected thousands of specimens himself, sending them back to various contacts and museums throughout his journey.  (His handwriting on the labels he attached was so bad that curators curse him for it to this day.)

Wallace
Wallace first went to Brazil, but, not being of the upper class as Darwin was, his purpose was as much to make his living by selling specimens to interested collectors back in England as it was to document the flora and fauna he found there.  He spent 4 years in Brazil.  When, in 1852 he finally set sail to return to England with most of his collection, 26 days into the voyage the ship caught fire and all were forced to abandon ship.  His entire collection and most of his notes were lost.  Then, from 1854-1862 he traveled the Malay Archepelago, collecting specimens and making notes.  It was there that, in a malaria-induced fever, he first intuited his theory of evolution by natural selection, which motivated him to write to Charles Darwin suggesting his theory.  This in turn, as is well known, motivated Darwin to finally write the book he'd long been contemplating, his Origin of Species, describing the same theory.    

And these are only three of the best known explorers, people who set out for foreign shores knowing little to nothing of what they would find there.  We had occasion to think about this yesterday, as we were out walking in the woods on the first real spring day of the year.  Ahead of us by about 50 yards was a young couple.  He was American, she Chinese.

A pileated woodpecker called its eerie warlockian call in the distance.  If you don't know the sound, here it is.




We had caught up to the couple by then, and as we passed we heard him say, "No, it wasn't a monkey!  It was a bird!"  And they were laughing.

We laughed, too.  But not derisively.  It was lovely to imagine thinking we might be surrounded by monkeys in our local woods!  And indeed, how would she know?  Perhaps she came from a part of China where it was common to see monkeys, though whether they sound like pileated woodpeckers is a question we can't answer.

This incident of gentle innocence is in itself harmless of course, and even humorous.  But it shows how very easy it is for people from different, far-flung places, even intelligent educated people, to have deep misunderstandings of each other's world.  How might we do were we to take a walk in a wood somewhere in China?  And think about how much moreso this kind of naive misperception or expectations can be for each other's worldview. 

Were Humboldt, Darwin and Wallace better informed about the flora and fauna of the regions they explored when they first arrived than this woman about central PA?  In fact we don't know, but how could they have been?  They did work with various guides, ranging from European settlers or missionaries, or natives who were bilingual, and so on, so they were surely quickly disabused of any such innocent mistakes.  Until Europeans had explored these areas and settled more systematically, there were no guides to birds or insects or lizards -- or rocks, or trees, or flowers.  In a sense that was what they were trying to do, in good Enlightenment tradition, systematically observing and cataloguing what they saw. But they must have made a great many mistakes, until someone who lived there and knew better set them straight.

Where does ambiguity end and knowledge begin?
Indeed, in science we believe that we try to bridge understanding-gaps by being precise in our terminology, and stating our hypotheses and designs in unambiguous, logical, and specific terms. And, scientists at least tend to believe that we try to collect data objectively and without preconceived notions as to what we'll find.  After all, objectivity is our purported aim, is it not?  Or is it?

The extent and intricacies of science within areas like genetics or evolution are such that they can be most fully within the reach of only narrow specialists.  The rest of us have to assume we know what 'gene' or 'adapted' or 'function', or even 'next-generation sequencing'  actually mean.  And these are relatively simple compared to other issues that we face (like multiple-test correction, various other statistical niceties, and so on).  So someone coming in from another field could easily make the 'monkey' mistake, and be viewed as incredibly naive.

But there's more.  We are often, or even typically, clearly not trying to be 'objective' except in some rather technical senses of accurately describing our methods (e.g., the pH of our PCR reactions, our significance testing p-values, and so on).  When it comes to what 'hypothesis' we've chosen to 'test', what aspect of something we've chosen to study, our interpretations, or the issues we've stated (or not), what is clear is something rather different:  to a great extent all of us, even in science, have our preferences, assumptions, and predispositions to believe and hence prove (but not really test).

So one can wonder, when we as scientists are presenting their work, whether what's being heard is a monkey.....or just a bird.