Showing posts with label association studies. Show all posts
Showing posts with label association studies. Show all posts

Friday, August 29, 2014

Genomic cold fusion? Part II. Realities of mapping

Mapping to find genomic causes of a trait of interest, like a disease, is done when the basic physiology is not known—maybe we have zero ideas, or the physiology we think is involved doesn’t show obvious differences between cases and controls.  If you know the biology, you won't have to use mapping methods, because you can explore the relevant genes directly.  Otherwise, and today often, we have to go fishing, in the genome, to find places that may vary in association—statistical regularity—with the trait.

The classical way to do this is called linkage analysis.  That term generally refers to tracing cases and marker variants in known families.  If parents transmit a causal allele (variant at some place in the genome) to their children, then we can find clusters of cases in those families, but no cases in other families (assuming one cause only).  We have Mendel’s classical rules for the transmission pattern and can attempt to fit that pattern to the data—for example, to exclude some non-genetic trait sharing.  After all, family members might share many things just because they have similar interests or habits.  Even disease can be due to shared environmental exposures. Mendelian principles allow us, with enough data, to discriminate.

“Enough data” is the catch.  Linkage analysis works well if there is a strong genetic signal.  If there is only one cause, we can collect multiple families and analyze their transmission patterns jointly.  Or, in some circumstances, we can collect very large, multi-generational families (often called pedigrees) and try to track a marker allele with the trait across the generations.  This has worked very well for some very strong-effect variants conferring very high risk for very specific, even quite rate, disorders.  That is because the linkage disequilibrium—the association between a marker allele and a causal variant due to their shared evolutionary history (as described in Part I) ties the two together in these families.

But it is often very costly or impractical to collect actual large pedigrees that include many children each generation, and multiple generations.  Family members who have died cannot be studied and medical records may be untrustworthy, or family members may have moved, refuse to participate in a study, or be inaccessible for many reasons.  So a generation or so ago the idea arose that if we collect cases from a population we may also collect copies of nearby marker alleles in linkage disequilibrium—shared evolutionary history in the population—so that, as described in Part I, a marker allele has been transmitted through many generations of unknown but assumed pedigree, so that the marker will have been transmitted in the pedigree along with the causal variant.  This is implicit linkage analysis, called genomewide association analysis (GWAS), about which we’ve commented many times in the past.  GWAS look for association between marker and causal site in implicit but assumed pedigrees, and is another form of linkage analysis.

When genetic causation is simple enough, this will work.  Indeed, it is far easier and less costly to collect many cases and controls than many deep pedigrees, so that a carefully designed GWAS can identify causes that are reasonably strong.  But this may not always work, when a trait is ‘complex’, and has many different genetic and/or environmental contributing causes.

If causation is complex, families provide a more powerful kind of sample to use in searching for genetic factors.  The reason is simple: in general a single family will be transmitting fewer causal variants than a collection of separate families.  Related to this is the reason that isolate populations, like Finland or Iceland, can in principle be good places to search, because they represent very large, even if implicit, pedigrees.  Sometimes the pedigree can actually be documented in such populations.

If causation is complex, then linkage analysis in families will hopefully be better than big population samples for finding causal contributors, simply because a family will be segregating (transmitting) fewer different causal variants than a big population.  We might find the variant in linkage analysis in a big family, or an isolate population, but of course if there are many different variants, a given family may point us only to one or two of them.  For this reason, many argue that family analysis is useless for complex traits—one commenter on a previous Tweet we made from our course, likened linkage analysis for complex traits to ‘cold fusion’.  In fact, this was a mistake and is incorrect. 

Association analysis, the main alternative to linkage analysis, is just a combining of many different implicit families, for the population-history reason we’ve described here and in Part I.  The more families you combine, whether they are explicit or implicit, the more variation, including statistical ‘noise’, you incorporate.  The rather paltry findings of many GWAS are a testament to this fact, explaining as they have only a small fraction of most traits to which that method has been applied.  Worse, the greater the sample of this type, like cases vs controls, the more environmental variation you may be grouping together, again greatly watering down even the weak signal of many or, probably, by far most genetic causal factors.

In fact, if you are forced to go fishing for genetic cause, you may well be fishing in dreamland because you may simply be in denial of the implications of causal complexity.  In fact, all mapping is a form of linkage analysis.  Instead, one should tailor one’s approach to the realities of data and trait.  Some complex trait genes have been found by linkage analysis (e.g., the BRCA breast-cancer associated genes), though of course here we might quibble about the definition of 'complexity'. 

Sneering at linkage analysis because it is difficult to get big families, or  because even single deep families may themselves be transmitting multiple causes (as is often found in isolate studies, in fact), is often simply a circle-the-wagon defense of Big Data studies, that capture huge amounts of funding with relatively little payoff to date.

A biological approach?
Many linkage and association analyses are done because we don’t understand the basic biology of a trait well enough to go straight to ‘candidate’ genes to detect, prevent, or develop treatment for a trait.  Today, even though this approach has been the rule for nearly 20 years now, with little payoff, the defense is often still that more, more and even more data will solve the problem.  But if causation is too complex this can also be a costly, self-interested, weak defense.

If we have whole genome sequence on huge numbers of people, or even everyone in a population, or in many populations so we can pool data, that we will find the pot of gold (or is it cold fusion?) at the end of the rainbow.

One argument for this is to search population-wide genome sequenced biomedical data bases for variants that may be transmitted from parents to offspring, but that are so rare that they cannot generate a useful signal in huge, pooled GWAS studies.  This usually will still be in the form of linkage analysis if a marker in a given causal gene is transmitted with the trait in occasional families but the same gene is identified, even if via different families.  That is, if variation in the same gene is found to be involved in different individuals, but with different specific alleles, then one can take that gene seriously as a causal candidate.

This sometimes works, but usually only when the gene’s biology is known enough to have a reason to suspect it.  Otherwise, the problem is that so much is shared between close family members (whether implicitly or explicitly in known pedigrees) that if you don’t know the biology there will be too much to search through, too much co-transmitted variation.  Causal variation need not be in regular ‘genes’, but can be, and for complex traits seems typically to be, in regulatory or other regions of the genome, whose functional sites may not be known.  Also, we all harbor variation in genes that is not harmful, and we all carry ‘dead’ genes without problems, as many studies have now shown.

If one knows enough biology to suspect a set of genes, and finds variants of known effect (such as truncating a gene’s coding region so a normal protein isn’t made) in different affected individuals, then one has strong evidence s/he has found a target gene.  There are many examples of this for single-gene traits.  But for complex traits, even most genes that have been identified have only weak effects—the same variant most of the time is also found in healthy, unaffected individuals.  In this case, which seems often to be the biological truth, there is no big-cause gene to be found, or a gene has a big-cause only in some unusual genotypes in the rest of the genome.

Even knowing the biology doesn't say whether a given gene's protein code is involved rather than its regulation or other related factors (like making the chromosomal region available in the right cells, downregulating its messenger RNA, and other genome functions).  Even in multiple instances of a gene region, there may be many nucleotide variants observed among cases and controls.  The hunt is usually not easy even knowing the biology--and this is, of course, especially true if the trait isn't well-defined, as is often the case, or if it is complex or has many different contributors.

Big Data, like any other method, works when it works.  The question is when and whether it is worth its cost, regardless of how advantageous for investigators who like playing with (or having and managing) huge resources.  Whether or not it is any less ‘cold fusion’ than classical linkage analysis in big families, is debatable.  

Again, most searches for causal variation in the genome rest on statistical linkage between marker sites and causal sites due to shared evolutionary history.  Good study design is always important.  Dismissal of one method over another is too often little more than advocacy of a scientist’s personal intellectual or vested interests.

The problem is that complex traits are properly named:  they are complex. Better ideas are needed than what are being proposed these days.  We know that Big Data is ‘in’ and the money will pour in that direction.  From such data bases all sorts of samples, family or otherwise, can be drawn.  Simulation of strategies (such as with programs like our ForSim that we discussed in our recent Logical Reasoning course in Finland) can be done to try to optimize studies. 

In the end, however, fishing in a pond of minnows, no matter how it’s done, will only find minnows. But these days they are very expensive minnows.

Wednesday, December 11, 2013

Sometime geneticist Joe Terwilliger on genetics

We may recently have given the false impression that geneticist Joe Terwilliger gives less priority to science, or at least good science, than to other perhaps more frivolous pursuits (he is Abe Lincoln every February, for example, and tuba player the rest of the time -- unless he's cleaning up bean debacles as a diplomat, or being a basketball and language coach to Dennis Rodman), so we wanted to help correct any such misconceptions here.  Perhaps to that end, Joe (now known in South Korea, we're afraid, as "sometime geneticist Joe Terwilliger") suggested we republish a blog post he first posted on his own short-lived blog in 2008.  He recently dug this up again and says that few could disagree, even 5 years later.


Joe as Abe on the balcony (but not of Ford's Theater)

The point is that sometimes there is a lot of convenient hard-of-hearing even in science, which fancies itself to be an objective search for truth. Some of the details in Joe's post are out-of-date but we, and he, think that the basic thrust is not.  In a sense that makes the conclusion all the more cogent, because the same modes of thinking about genomic causation are still predominant, despite the vastly costly but essentially consistent results in the five years since 2008.  And, as Joe points out, he and Ken had much the same message in 2000.

One not-so-subtle change, we will note, is that promises by NIH Director Francis Dr Collins, and many others in presumably responsible positions, have steadily altered  their due date, which recedes into the distance like, say, an oasis as you grope for water, the fences if you want your pitchers to have a better earned run average, or a preacher's promises of ultimate salvation, if you weekly plunk coins into the basket.

So perhaps the lesson is that under these circumstances, rather than just dismiss critics, science--actual science as it's supposed to be--should feel a need to take stock of what it's doing.  But we leave it to you to judge.

And if you hear about Joe in other contexts in weeks to come, remember that he was a sometime geneticist here first:


The Rise and Fall of Human Genetics and the Common Variant - Common Disease Hypothesis
By Joe Terwilliger
Nov 2008

There is an enormity of positive press coverage for the Human Genome Project and its successor, the HapMap Project, even though within the field the initial euphoric party when the first results came out has already done a full 180 to be replaced by the hangover that inevitably follows such excesses.

For those of you not familiar with the history of this field and the controversies about its prognosis which were present from the outset, I refer you to a review paper I and a colleague wrote back in 2000 at the height of the controversy - Nature Genetics 26, 151 - 157 . The basic gist of the argument put forward for the HapMap project was the so-called common variant/common disease hypothesis (CV/CD) which proposed that "most of the genetic risk for common, complex diseases is due to disease loci where there is one common variant (or a small number of them)" [Hum Molec Genet 11:2417-23]. Under those circumstances it was widely argued that using the technologies being developed for the HapMap project, that one would be able to identify these genes using "genome-wide association studies" (GWAS), basically by scoring the genotype for each individual in a cross sectional study for each of 500,000 to 1,000,000 individual marker loci - the argument being that if common variants explained a large fraction of the attributable risk for a given disease, that one could identify them by comparing allele frequencies at nearby common variants in affected vs unaffected individuals. This point was contested by researchers only with regard to how many markers you might have to study for this to work if that model of the true state of nature applied. Many overly optimistic scientists initially proposed 30,000 such loci would be sufficient, and when Kruglyak suggested it might take 500,000 such markers people attacked his models, yet today the current technological platforms use 1,000,000 and more markers, with products in the pipelines to increase this even more, because it quickly became clear that the earlier models of regular and predictable levels of linkage disequiblrium were not realistic, something that should have been clear from even the most basic understanding of population genetics, or even empirical data from lower organisms.

Today such studies are widespread, having been conducted for virtually every disease under the sun, and yet the number of common variants with appreciable attributable fractions that have been identified is miniscule. Scientists have trumpetted such results as have been found for Crohn's disease, in which 32 genes were detected using panels of thousands of individuals genotyped at hundreds of thousands of markers - this sounds great until you start looking at the fine print, in which it is pointed out that all of these loci put together explain less than 10% of the attributable risk of disease, and for various well-known statistical reasons, this is a gross overestimate of the actual percentage of the variance explained. Most of these loci individually explain far less than half a percent of the risk, meaning that while this may be biologically interesting, it has no impact at all on public health as most of the risk remains unexplained. This is completely opposite to the CV/CD theory proposed as defined above. In fact, this is about the best case for any complex trait studied, with virtually every example dataset I have personally looked at there is absolutely nothing discovered at all.

At the beginning of the euphoria for such association studies, the example "poster child" used to justify the proposal was the relationship between variation at the ApoE gene and risk of Alzheimer disease. In an impressively gutsy paper recently, a GWAS study was performed in Alzheimer disease and published as an important result, with a title that sent me rolling on the floor in tears laughing: "A high-density whole-genome association study reveals that APOE is the major susceptibility gene for sporadic late-onset Alzheimer's disease" [ J Clin Psychiatry. 2007 Apr;68(4):613-8 ] - in an amazingly negative study they did not even have the expected number of false positive findings - just ApoE and absolutely nothing else... And the authors went on to describe how important this result was and claimed this means they need more money to do bigger studies to find the rest of the genes. Has anyone ever heard of stopping rules, that maybe there aren't any common variants of high attributable fraction??? This was a claim that Ken Weiss and I put forward many times over the past 15 years, and Ken has been making this point for a decade before that even, in his book, "Genetic variation and human disease", which anyone working in this field should read if they are not familiar with the basic evolutionary theory and empirical data which show why noone should ever have expected the CV/CD hypothesis to hold...

In many other fields, the studies that have been done at enormous expense have found absolutely nothing, and in what Ken Weiss calls a form of Western Zen (in which no means yes), the failure of one's research to find anything means they should get more money to do bigger studies, since obviously there are things to find but they did not have big enough studies with enough patients or enough markers - it could not possibly be that their hypotheses are wrong, and should be rejected... It is a truly bizarre world where failure is rewarded with more money - but when it comes to promising upper-middle-aged men (i.e. Congress) that they might not die if they fund our projects, they are happy to invest in things that have pretty much now been proven not to work...

While in a truly bizarre propaganda piece, Francis Collins, in a parting sycophantic commentary (J Clin Invest. 2008 May;118(5):1590-605) claimed that the controversy about the CV/CD hypothesis was "... ultimately resolved by the remarkable success of the genetic association studies enabled by the HapMap project." He went on to list a massive table of "successful" studies, including loci for such traits as bipolar, Parkinson disease and schizophrenia, and of course the laughable success of ApoE and Alzheimer disease. To be objective about these claims, let me quote from what researchers studying those diseases had to say.

Parkinson disease: "Taken together, studies appear to provide substantial evidence that none of the SNPs originally featured as PD loci (sic from GWAS studies) are convincingly replicated and that all may be false positives...it is worth examining the implications for GWAS in general." Am J Hum Genet 78:1081-82

Schizophrenia: "...data do not provide evidence for involvement of any genomic region with schizophrenia detectable with moderate [sic 1500 people!] sample size" Mol Psych 13:570-84

Bipolar AND Schizophrenia: "There has been great anticipation in the world of psychaitric research over the past year, with the community awaiting the results of a number of GWAS's... Similar pictures emerged for both disorders - no strong replications across studies, no candidates with strong effect on disease risk, and no clear replications of genes implicated by candidate gene studies." - Report of the World Congress of Psychiatric Genetics.

Ischaemic stroke: "We produced more than 200 million genotypes...Preliminary analysis of these data did not reveal any single locus conferring a large effect on risk for ischaemic stroke." Lancet Neurol. 2007 May;6(5):383-4.

And the list goes on and on of traits for which nothing was found, with the authors concluding they need more money for bigger studies with more markers. It is really scary that people are never willing to let go of hypotheses that did not pan out. Clearly CV/CD is not a reasonable model for complex traits. Even the diseases where they claim enormous success are not fitting with the model - they get very small p-values for associations that confer relative risks of 1.03 or so - not "the majority of the risk" as the CV/CD hypothesis proposed.

One must recall that in the intial paper proposing GWAS by Risch and Merikangas (Science 1996 Sep 13;273(5281):1516-7) - a paper which, incidentally, pointed out that one always has more power for such studies when collecting families rather than unrelated individuals - the authors stated that "despite the small magnitude of such (sic: common variants in)genes, the magnitude of their attributable risk (the proportion of people affected due to them) may be large because they are quite frequent in the population (sic: meaning >>10% in their models), making them of public health significance." The obvious corollary of this is that if they are not quite frequency, they are NOT having high attributable fraction and are therefore NOT of public health significance.

And yet, you still have scientists claiming that the results of these studies will lead to a scenario in which "we will say to you, 'suppose you have a 65% chance of getting prostate cancer when you're 65. If you start taking these pills when you're 45, that percent will change to 2". Amazing claims when the empirical evidence is clear that the majority of the risk of the majority of complex diseases is not explained by anything common across ethnicities, or common in populations... (Leroy Hood, quoted in the Seattle Post-Intelligencer). Francis Collins recently claimed that by 2020, "new gene-based designer drugs will be developed for ... ALzheimer disease, schizophrenia and many other conditions", and by 2010, "predictive genetic tests will be available for as many as a dozen common conditions". This does not jibe with the empirical evidence... In Breast Cancer for example, researchers claimed that knowledge of the BRCA1 and BRCA2 genes (which confer enormously high risk of breast cancer to carriers) was uninteresting as it had such a small attributable fraction in the population. Of course now they have performed GWAS studies and examined tens of thousands of individuals and have identified several additional loci which put together have a much smaller attributable fraction than BRCA1 and BRCA2, yet they claim this proves how important GWAS is. Interesting how the arguments change to fit the data, and everything is made to sound as if it were consistent with the theory.

I suggest that people go back and read "How many diseases does it take to map a gene with SNPs?" (2000) 26, 151 - 157. There are virtually no arguments we made in that controversial commentary 8 years ago which we could not make even stronger today, as the empirical data which has come up since then basically supports our theory almost perfectly, and refutes conclusively the CV/CD hypothesis, despite Francis Collins' rather odd claims to the contrary...

In the end, these projects will likely continue to be funded for another 5 or 10 years before people start realizing the boy has been crying wolf for a damned long time... This is a real problem for science in America, however, as NIH is spending big money on these rather non-scientific technologically-driven hypothesis-free projects at the expense of investigator-initiated hypothesis-driven science. Even more tragically training grants are enormously plentiful meaning that we are training an enormous number of students and postdocs in a field for which there will never be job opportunities for them, even if things are successful. Hypothesis-free science should never be allowed to result in Ph.D. degrees if one believes that science is about questioning what truth is and asking questions about nature, while engineering is about how to accomplish a definable task (like sequencing the genome quickly and cheaply). The mythological "financial crisis" at NIH is really more a function of the enormous amounts of money going into projects that are predetermined to be funded by political appointees and government bureaucrats rather than the marketplace of ideas through investigator-initiated proposals. Enormous amounts of government funding into small numbers of projects is a bad idea - one which began with Eric Lander's group at MIT proposing to build large factories for the sequencing of the genome rather than spreading it across sites, with the goal of getting it done faster (an engineering goal) instead of getting more sites involved so that perhaps better scientific research could have come along the way. This has led to a scenario years later in which the factories now want to do science and not just engineering, which is totally contrary to their raison d'etre, and leads to further concentrations of funding in small numbers of hands when science is better served, perhaps by a larger number of groups receiving a smaller amount of money so that more brains are working in different directions thinking of novel and innovative ideas not reliant on pure throughput. Human genetics has transformed from a field with low funding, driven by creative thinking into a field driven by big money and sheep following whatever shepherd du jour is telling them they should do (i.e. innovative means doing what they current trend is rather than something truly original and creative). This is bad for science, and also is bad science. GWAS has been successful technologically, and it has resoundingly rejected the CV/CD hypothesis through empirical data. If we accept this and move on, we can put the HapMap and HGP where it belongs, in the same scientific fate as the Supercollider, and let us get back to thinking instead of throwing money at problems that are fundamentally biological and not technological!


(most notably in terms of the big money NIH is sending into these non-scientific technologically-driven hypothesis-free studies, rather than investigator initiated hypothesis-driven science - one of the main causes of the "funding crisis" at NIH where a tiny portion of new grants are funded - get rid of the big science that is not working - like the supercollider! - and there is no funding crisis)

Monday, January 2, 2012

Those confounded links to the causes of asthma

We've blogged a number of times about asthma -- specifically, why has it increased in prevalence so dramatically over the last 30 years, and why hasn't epidemiology figured it out?  (Here's our most recent post on this, and here's one on the Hygiene Hypothesis.) A disease that goes from not so common to very common very quickly most likely has an environmental trigger.  It's unlikely to be some new genetic risk, because genes don't change that quickly, and it's likely to be an environmental cause with a major effect, since it's triggering the disease in a whole lot of people.  And, since epidemiology is best at figuring out major environmental effects -- smoking, infectious agents, um....  -- you'd think they'd have knocked this one long ago.  But no, the answer to why so many kids are getting asthma has been elusive.  Until maybe now.

But, let's back up.  Millions -- and millions -- of dollars have been spent on the genetics of asthma.  Numerous family studies, GWAS (genomewide association studies), admixture mapping studies, and so on, and nothing to explain any significant amount of variation in risk has yet been found.  The idea was -- we guess -- that while asthma was certainly increasing in prevalence, not everyone was getting it, and the explanation for that must be genetic.  Let's ignore the root cause, whatever had changed in the environment, and just explain why some people respond to whatever-it-is with asthma.  Beyond our current fetish with geneticizing everything that moves, the rationale was that druggable pathways would be identified this way, and so whether or not we understood the source of the epidemic, by golly, we would be able to sell a lot of drugs to treat it.

Ok, that's one approach.

But to be fair, environmental and observational epidemiologists did try their darndest to tease out the environmental cause, and all they got for their pains were confusing and contradictory results.  Breast feeding, bottle feeding; environments that were too clean or environments that were too dirty; lack of helminth infections, lack of air pollution (really -- this was based on observations such as that asthma was on the increase in the early days of the epidemic in places like West Germany, but not right over the wall in East Germany, where air pollution was high).  We got the Hygiene Hypothesis out of all this work, a still live conjecture about the importance of boosting our immune systems when we're young. Maybe true.

But now the asthma community has a new idea, and again to be fair, it's in part due to some good sleuthing by epidemiologists.  Actually, the idea isn't so new -- the first paper suggesting it was published in 1998 -- but the idea is just now getting legs.  It's probably the most likely possible solution to the question of what's causing this epidemic that has been offered to date.  (Though we're still disappointed that no one picked up on our suggestion in a 2006 paper that the use of plastic diapers grew right along with the asthma epidemic, a correlation that we thought might do with some looking into.) And if this current suggestion is true, it shows again the problem of correlation not proving causation.

A story in the New York Times reported on this recently. The idea is that acetaminophen causes asthma.  In the 1980's, when the epidemic began, doctors started recommending that parents treat their infants' pain and fever with acetaminophen rather than aspirin because aspirin had been found to cause Reyes' Syndrome.  So, Tylenol sales took off.  And, not too long after that, so did inhaler sales.  More than 20 studies that show this association have now been published, and it's starting to look strong, not only based on epidemiological measures such as the strength of the association (the strength of the association between air pollution and low risk was strong too, but more on that below), but with a fairly substantial understanding of the biological mechanism that could explain the risk.

Here's the abstract from a paper in the November issue of Pediatrics, by John McBride:
The epidemiologic association between acetaminophen use and asthma prevalence and severity in children and adults is well established. A variety of observations suggest that acetaminophen use has contributed to the recent increase in asthma prevalence in children: (1) the strength of the association; (2) the consistency of the association across age, geography, and culture; (3) the dose-response relationship; (4) the timing of increased acetaminophen use and the asthma epidemic; (5) the relationship between per-capita sales of acetaminophen and asthma prevalence across countries; (6) the results of a double-blind trial of ibuprofen and acetaminophen for treatment of fever in asthmatic children; and (7) the biologically plausible mechanism of glutathione depletion in airway mucosa. Until future studies document the safety of this drug, children with asthma or at risk for asthma should avoid the use of acetaminophen.
A somewhat more detailed description of the findings on JournalWatch is here.

So, some in the pediatric community are starting to take notice, and recommend that parents give their kids ibuprofen rather than acetaminophen or aspirin.  Though, some are still doubtful; parents give their children acetaminophen when they have fevers, often of viral origin, and it could be the virus that causes asthma, not the treatment, some say.  Others are dubious because many of the studies relied on parents recalling how much of the drug they gave their children sometimes years in the past.  Relying on recall is a common cause of iffy results in epidemiological studies.

In addition, according to the Times,
So far, only one randomized controlled trial has investigated the link. Researchers at Boston University School of Medicine randomly assigned 1,879 children with asthma to take either acetaminophen or ibuprofen if they developed a fever. The results, published in 2002, showed that children who took acetaminophen to treat a fever were more than twice as likely to seek a doctor’s care later for asthma symptoms as those who took ibuprofen.
Of course, these are children who already have asthma, which complicates the interpretation of causation, but it does suggest that the mechanism, glutathione depletion in airways, might indeed be involved (though see above).  What's really needed to confirm the role of acetaminophen is a prospective study of children from birth through childhood but the ethics of a study that involves giving acetaminophen to babies surely would be questionable at this stage.

So, what about these other associations that looked so strong not long ago?  The idea that helminth infections or air pollution might actually reduce risk of asthma?  They could still be true, it could be that asthma, which is as complex and variable a disease as almost any disease out there, could have many causes.  Or, if the acetaminophen link is true, then these other factors are either independent confounders or in some way actively interact with the acetaminophen-related mechanism.  In rural African villages where helminths are a common fact of life, or in cities where air pollution is high, acetaminophen use is not.  The Hygiene Hypothesis, even given that people have suggested possible biological explanations for it, may be on its way out.

If this association does explain the epidemic, there's still this question -- will geneticists stop looking for genes 'for' asthma?  We don't bet on it.  That's not to say that better treatment isn't needed, and that the triggers and responses need to be better understood because the biology is complex -- indeed, some say that every case is unique -- and asthma is often difficult to control.  But, continuing to search for genetic causation?  Bah, humbug.