Showing posts with label logical reasoning. Show all posts
Showing posts with label logical reasoning. Show all posts

Sunday, August 9, 2015

How many diseases does it take to map a SNP? Fifteen years on

Ken and I are here in Finland, preparing to teach a week of Logical Reasoning in Human Genetics with Joe Terwilliger and colleagues.  Not statistical methods, not laboratory techniques, not the latest way to analyze sequence data.  Concepts, logical reasoning.

Ken and Joe have been reasoning logically for a long time.  They've taught this course together in many places, and they wrote at least one logically reasoned paper 15 years ago.  That paper was published in Nature Genetics.  That journal shortly afterwards made an editorial policy decision to be the loudspeaker for genetic association studies (GWAS), and would be unlikely in the extreme to publish such a view today.  But Joe often says that it could, and probably should be published again, with very few wording changes.  (He also says that if overhead projectors were still available, he'd give the same talks he gave in 1995, since the issues in human genetics haven't changed.  We have lots more data, but no fundamentally new concepts or insights regarding SNP associations and complex traits.  In fairness, though, he does update his slides -- he adds photos of the latest places he has traveled.  Looking forward to photos of Crimea this week.)

The 2000 paper was called, "How many diseases does it take to map a SNP?"  They began:
There are more than a few parallels between the California gold rush and today's frenetic drive towards linkage disequilibrium (LD) mapping based on single-nucleotide polymorphisms (SNPs). This is fuelled by a faith that the genetic determinants of complex traits are tractable, and that knowledge of genetic variation will materially improve the diagnosis, treatment or prevention of a substantial fraction of cases of the diseases that constitute the major public health burden of industrialized nations. Much of the enthusiasm is based on the hope that the marginal effects of common allelic variants account for a substantial proportion of the population risk for such diseases in a usefully predictive way. A main area of effort has been to develop better molecular and statistical technologies often evaluated by the question: how many SNPs (or other markers) do we need to map genes for complex diseases? We think the question is inappropriately posed, as the problem may be one primarily of biology rather than technology.
Today, emphasis is more on ever larger sample sizes to find rare alleles, since common alleles turned out not to be the magical answer, but the issues are the same.  The problem is biological, rather than one of sample size. And not only do we have at least as much causal complexity due to environmental factors, but to the mix have been added the comparable complex 'genetic' causal factors as epigenetic modification of DNA affecting gene expression, and the potential contributions of highly complex microbiome.

The idea of mapping diseases from SNPs is that markers will be near the disease allele.  But, there are problems with this, as GWAS are successfully showing.
If traits do not strongly predict underlying genotypes, that is, if P(GP|Ph) is small, linkage and LD mapping may have very low power or may not work at all. As an extreme example, one's genotype cannot be reliably determined by merely stepping on the bathroom scale! But even if this could be done, there is a widespread but invalid belief that because something can be mapped (that is, P(GP|Ph) is high), the causal predictive power of the genotype (P(Ph|GP); Fig 1, blue arrow) will also be high. In fact, we have surprisingly little data on this latter topic, which requires extensive sampling from the general population, rather than patients. Note that the opposite can also be untrue—that is, if P(Ph|GP) is high it does not mean P(GP|Ph) will be high, as in genetically heterogeneous mendelian disorders such as retinitis pigmentosa. It is important to note that when we speak of P(Ph|GP) in this context, we speak of the marginal mode of inheritance, which is only valid for consideration of singletons, and relatives will not have independent and identically distributed penetrances (even without assuming epistasis or gene-environment interactions) because the other genetic and environmental factors are also correlated among them! Similar arguments can be made about detectance, P(GP|Ph), which must always be a function of the ascertainment, something that is often overlooked in the literature when investigators make comparisons of power for different study designs

 Figure 1. Schematic model of trait aetiology.
The phenotype under study, Ph, is influenced by diverse genetic, environmental and cultural factors (with interactions indicated in simplified form). Genetic factors may include many loci of small or large effect, GPi, and polygenic background. Marker genotypes, Gx, are near to (and hopefully correlated with) genetic factor, Gp, that affects the phenotype. Genetic epidemiology tries to correlate Gx with Ph to localize Gp. Above the diagram, the horizontal lines represent different copies of a chromosome; vertical hash marks show marker loci in and around the gene, Gp, affecting the trait. The red Pi are the chromosomal locations of aetiologically relevant variants, relative to Ph.
Other inconvenient biological issues, mentioned in the paper, include that linkage disequilibrium is stochastic, and this has implications for the use of SNPs in disease mapping, that regulatory rather than protein coding sites often affect disease risk, and these are generally impossible to identify (see below), that late-onset chronic diseases are much more complex than the clearly genetic pediatric disease, that the most effective disease mapping and association studies are done in "selective samples of individuals or families at high risk relative to the average risk in the population, and from populations with unusual histories" (hence, Joe's eclectic and interesting travelogue), etiology tends to be very heterogeneous, phenotype can't predict genotype and vice versa, environmental effects can be significant, but are unpredictable and often impossible to identify, and so on.

Whole genome sequencing will not be a general miracle cure.  Exome sequencing can sometimes find  coding variants that have strong effects because we know how to identify exomes and how to read their code.  But many if not most mapped sites for complex traits, as might be expected, are in regulatory regions.  Yet we are still quite inept at identifying regulatory regions, for many reasons not least having to do with their complexity and fluidity among individuals and populations.  So whole genome sequencing will likely have to be analyzed by using markers, as in GWAS, and that will not automatically show us where key regulatory affects are located or how they work.  If these are too heterogeneous, they'll vary hugely, so that mapping will still face the complexity problem.  Time will tell what transpires.
The problems faced in treating complex diseases as if they were Mendel's peas show, without invoking the term in its faddish sense, that 'complexity' is a subject that needs its own operating framework, a new twenty-first rather than nineteenth—or even twentieth—century genetics.
So, if the data are better and less costly now than 15 years ago, the basic issues haven't changed.  

Wednesday, September 3, 2014

Genomic cold fusion? Part III. Gene mapping: when minnows are whales

In the first two parts of this series we tried to outline the actual logic underlying the search for genes that affect a trait, disease or otherwise, that we might be interested in.  We titled this series ‘genomic cold fusion’ in response to a comment on a tweet made about our course Logical Reasoning in Human Genetics whose most recent offering was given by us a week or so ago in Helsinki, Finland.  The characterization was about the idea of ‘linkage analysis’—that is, in known pedigrees—to find genetic causal factors for complex traits. 

We tried to explain that evolutionary (population) history lies behind the logic of both family-based linkage and population-sample-based association approaches to genomewide mapping (such as in GWAS).  When causes are strong and not too numerous, mapping works in large families.  That’s because if something is genetic it must be familial—that in a sense is what ‘genetic’ means in this context—and one can trace transmission, following Mendelian principles, explicitly.

If causes are individually rare and there are many, pooling families doesn’t work very well, because getting large enough families to map individually is difficult and costly, but that is just what GWAS does in its implicit, unconstrained pooling of different families, where the family connections aren’t even known!

In the end, however, we concluded that if there were too many different causes, and they are weak or rare, and environmental factors are important, then the trait is basically the result of a mix of contributors, differing among individuals both within and between families.  Individually, we suggested, the causes are minnows, and fishing in a pond of minnows, no matter how it’s done, will only find minnows.  But there is more to the issues than this, and it deserves to be recognized.

When a minnow is a whale
There are tons of results in which a known genetic mutation identified as having a major effect is found to have lesser effects in some people.  Even family members sharing the variant may have different effects (more or less severe, for example, in regard to disease).  Some may have an essentially lethal phenotype, while others are only mildly affected.

The reason is that a variant’s causal effects depend fundamentally on its context.  This is true for environmental risk factors as much as genetic ones.  A causal minnow—a minor causal effect—can be major in some contexts.  Any approach to genetics that fails to take this basic fact seriously into account is, in a sense, amateurish.

A good illustration of this is that when a disease-causing genetic change is engineered into a laboratory mouse, it may or may not mimic the human trait.  Sometimes, perhaps most of the time, it will have a roughly similar effect in one strain of lab mice, but very different, or even no effect, in other strains.  Indeed, while this is very well-known to mouse workers (including ourselves when we were doing that sort of experiment), it is rarely taken seriously into account.  A transgenic effect is reported, but not checked in other strains of lab mice, or in other animal models, such as rats or dogs.  The reasons, especially for other species than mice, is that such testing is quite costly.  The bottom line is that we learn about the biology of the effect in one of its contexts, but extrapolate to other contexts, even humans, at our peril.  This, too, is well known.

This is why, among humans within or between populations, a mapping minnow can be a causal whale in some people, and vice versa.  It’s something that needs to be recognized more widely, but for which there really is no generic explanation.  It’s why risk estimates given for a genetic variation—such as by companies essentially practicing shell-game medicine without a license by advising customers about their ‘risk’ based on DNA analysis—are often not worth the electrons needed to send them.  Some risk factors are often very strong (one thinks of BRCA variation and breast cancer) but most are not, and some are very weak to start with and only strong in rare contexts.  Again, conscientious geneticists know this very well, or should.  It’s not secret.

Indeed, the fact that minnows can grow up to be whales or whales can shrink to minnows depending on the genomic and environmental pond they’re swimming in, is one of the important things genomicists should be directly addressing, rather than making the rather bold and expensive promises that they are making.

This isn’t an argument against doing genetics, but it is a reason to think differently, or at least carefully, before making very expensive promises that often are not very different from what preachers promise you if you’ll put some coin in the plate being passed.

Genetics is fundamental biology, and its challenges are great from the ground up.  At present, those challenges are typically whales, but are just as typically, and expediently, treated as if they are minnows. 

Wednesday, February 26, 2014

Godel's principle and respect for failure

In 1931 Kurt Godel shocked the mathematical world.  Math is the ultimate sanctuary for those who believe that some ultimately Platonic sense there is absolute, universal, unexceptionable--and understandable truth. The facts of geometry and mathematics are cosmically true (the Pythagorean theorem, the sum of angles =180 degrees, etc., 2+2=4, and the fact that the derivative of x-squared is 2x, etc). The way we do math may be a human or cultural convention, but the facts are not.  Our cultural convention of how we do it, or whatever it is, may for cultural or historical reasons simply overlook many equivalent truths, but it is working with at least some set of ultimate truths.  But how is it that Pythagorean theorem is true about right triangles, but there are no actual, perfect right triangles in the world??

2+2=4

At least, at the turn of the 20th century, for those studying such ultimate truths, it was widely thought that the principles of logic and logical reasoning were essentially the same as the principles of mathematics.  Not only were both equally true but logical reasoning could be expressed in the same kinds of terms as those of mathematics.  This would make the world a certain place, in a sense.  Truth is truth. Truth is internally consistent.  And truth is discoverable!

Infinity symbol in various typefaces; Wikipedia


There were some problems.  For example, what do we do with or about the notion of 'infinity'?  Nineteenth century mathematicians, notably Cantor, showed that there are even different levels of infinity.  The whole numbers are one level.  You can match even numbers up one for one with odd numbers.  But you can't match either up like that for the numbers between just 0 and 1.  That's because there are far more of the latter:  if we match 0, 1, and 2 with 0, 0.1, and 0.2 it might seem fine, but then how do we match up 0.001, 0.002, 0.00001, and so on?

The constant π is represented in thismosaic outside the mathematics building at the Technische Universität Berlin. Wikipedia

And then there is the little problem of what 'randomness' means.  For example, I recently learned that the digits in the value of  pi (relating radius to circumference in a circle) are randomly distributed.  Take any sequence, like 7623116, and you'll find it, on average every 10 million 7-digit sequencs in pi.  Or look at Stephan Wolfram's 'cellular automata'; these are simple rules for transitions in a string of (say) black or white boxes if each box produces a 'descendant' box and the rule says what color it is based on the array of colors in the current generation.  The resulting black/white pattern, determined by a simple rule, is, Wolfram says, indistinguishable from random: no pattern can be found along the string at any given time.  Yet 'random' seems such an obvious concept...until you think about them too much, and then they become quite disturbing.

And I have not mentioned the very problematic notion of probability.  We know how a cause can lead to an effect--well, we think we know that--but it is totally unclear how a cause can only lead to the future probabilistically.   Usually, the probabilistic nature of such results is attributed to measurement error, poor theoretical understanding, and so on, in a cosmos that, were we to know everything, would be purely and rigidly law-like.  But how could something cause something else 'with 27% probability'?  If you think about it, it is not at all clear what that means, in terms of actual causation, beyond errors and sampling effects.

And then there is 'chaos' theory.  Even in a purely deterministic, rule-bound process of cause and effect, where there is no uncertainty or probability involved, unless you have 100% measurement accuracy of things at some given time, you cannot predict with any accuracy what things will be like over the future.  Your predictive power, even with perfectly true theory, is zilch.  But how can you know what the underlying reason is?

Such things fly in the face of the views of the cosmos as a law-like place where certainty rules.  We want a knowable universe.  We spend our puny lives as scientists trying to understand it, and assuming it at least exists, even if it's hard to understand!

What we have learned
To general chagrin, what mathematician Kurt Godel showed halfway through the last century, was that even if it were true that the mathematical realm was all-of-perfection, an unknown fraction of it was unknowable.  That is, things that are true can't be proved and, worse, you could never know whether something you thought were true and were trying to prove it, was in reality untrue or just unprovably true.

Given all of this, it is surprising that in so many cases what we have learned is that the universe may not be law-like in the way we'd thought, true probability may or may not exist, not all things can be shown to be true even if they are true, and (in quantum mechanics and relativitiy and gravity at least) there are phenomena that seem truly to be unlike any of the above concepts or, as physicists often say, just are not consistent with 'common sense'....even if they're true.

Even physics and chemistry, not to mention biology and psychology or economics, are often swimming in uncertainty and claims of knowledge that, no matter how confidently asserted, simply don't hold up.

These various incarnations of indeterminacy, like probability, can shake our faith in the idea of a knowable universe whose causal nature we can pin down tight.  In ordinary sciences, and sometimes even in our daily lives, we don't know exactly how we should be viewing, much less approaching, causation.  What we end up doing is designing studies or experiments that we know how to design, using methods we know how to use, and crossing our fingers.  Whether we're being ostriches to our peril, or whether it doesn't matter and we should just carry on regardless, is unclear.

To the young and thoughtful, these provide things to think about, both in the practical sense of actually moving the science forward more than a millimeter at a time.....and in terms of our ultimate hope to understand life as it really is, not just in a statistical analysis.