Monday, January 16, 2012

Changing the goal posts: heritability lost and found

We interrupt our series on Probability and its meaning, for a post that we've been asked to write, related to a new paper on gene hunting that got some press last week, and will be stirring up controversy (and naturally, we can't resist including our usual editorializing, for better or worse):

Making complexity simple (again)?
Every age and every profession has its PT Barnums.  They're the slick-talking, fast-moving guys who will say anything to draw customers into the show they're running.  Truth, if it even exists other than ambiguously, is secondary in many ways to closing the deal.

One tactic that the pitchmen use in fields like politics is to change definitions so that problems never get 'solved' (that is, there is always something to keep you in office or to keep levying taxes, or to keep you afraid of some enemy or other).  This is a way of making the same facts serve new interests.

Science is rife with these kinds of self-interested maneuvers.  Changing the goalposts in genetics means redefining the objectives or the criteria for pushing ahead even in the face of contrary evidence.  This is a system we've built, step by step, once government funding largesse started flowing some time during the Cold War.  And today, in genetics, which plays on fear of death just as much as preachers do, we have to keep passing the plate.  Yet science is supposed to be objective and 'evidence based'.  So we have to change the goalposts to keep the game from ending.  In this case, there is a recent paper, with promise of more to come, by Eric Lander and colleagues (Zuk et al.).  Because of his prominence, skill, and salesmanship this will of course get a lot of attention.

The paper discusses at least one reason why GWAS have not been very good at accounting for heritability (something we ourselves have commented on in many posts).  Some, who are critical, say that this paper finally shows that GWAS and related big-scale approaches are proliferating even though they have themselves shown that they've reached diminishing returns.  Here's an example.  Of course there's the resistance that says no, Bravo! to the new paper, which shows that we're just getting started!

Naturally, everyone shares an interest in saying that to show their favorite view we need mega-GWAS, Biobank, gobs of wholegenome sequence, or other similar open-ended approaches, and whether this will lead to the ultimate small scale objective of personalized 'genomic' medicine.  Far too many vested interests are at play and, indeed, training in 'grantsmanship' and the whole research culture is about manipulating the system to get, keep, and increase funding. 

Zuk et al. ask why GWAS have failed to account for the heritability of so many traits, that is, to explain the correlation in risk among family members that reflects genetic effects.  The subject is complex, but here we can just say that if each gene adds a dose of effect to a trait, like nibbles on a given side of the caterpillar's mushroom added to Alice's stature, then even if the effect of nibbles is very small, if we sample enough mushrooms we can identify all of the effects.  Then, knowing them, we can tell which nibbles a given person has made, and hence predict his stature:  Voila!  personalized medicine!

Yet it hasn't worked that way.  Most heritability remains unaccounted for despite already very large studies.  Hence the demand for ever larger studies.  A glib commentary in Nature a few years ago coined the term 'hidden heritability' and made the search to find it akin to Sherlock Holmes' search for Moriarty.  That was a fantastic, if anti-science, marketing ploy on Nature's part, since it fed an ever-increasing demand for funds for genomic scavenger hunting....and that's good for the science journal business!

But the search for this hidden treasures has been frustrating, and Zuk et al. claim they now know that the search is in vain, and they provide a very sophisticated mathematical account of the reason why.  It is due to gene interactions, or 'epistasis'.  That means that a large part of the correlation among relatives can't be found by looking only for additive effects.  Here's roughly a basic underlying concept:

If trait T, such as stature or insulin levels (or their disease-risk consequences) is due to the effects of factor A plus those of factor B, then we can write

T = A + B

If your genotype includes a variant A-gene that gives you an additional level of factor A, then for you

T = A + A + B,   or 2A + B.

But if the factors interact, say in a multiplicative way, then

T = AxB

and if that's what's going on, and your A-gene genotype adds a second dose of A, your trait is

T = (2A) x B =  2AB

So, let's say a normal person has A=3 trait units and B=4 trait units.  In the additive case that person's trait would be A+B=7.  And if you have a mutation that doubles your A-dose, then your genotype makes you 2A+B=10 for your trait value.  But in the epistasis case, your trait would actually be 24.  So we expect you to have trait value 10, and conclude that more than half your trait value is unexplained.

Functional interaction
Nobody disputes a central fact in the Zuk et al. argument.  Life works by interactions among genes.  These are functional interactions in systems called 'networks' and other sorts of molecular interactions.  One molecule doesn't do anything by itself, nor does one gene as a rule.  Here, each gene-related component is subject to variation by mutation, and that will be inherited (its effects contributing to heritability).  So it is obvious from a cell biology point of view that single-factor explanations are not going to tell the whole story. But the fact of multiple factors doesn't tell the story we're interested in predicting states among variable traits.  That involves a different kind of interaction.

Quantitative interaction
If the true-fact is that biological traits are the result of functional interactions, what epidemiological risk estimation is about is not a list of pathways but the effects of variation in the pathways. When factors interact in a non-additive way, their net result is estimated from sample data using statistical techniques.  The A B example above showed conceptually how they work.  The additive contributions of variants in factor A are estimated from samples that compare those with and those without the variant in question.  You can estimate each factor's effects independently in this way and add up the estimates.

But if they interact, then in essence you have to have an additional estimate to make, of the average trait value in groups of people with each combination of the variants at the interacting factors.

Not only does this require more data, it won't show up in GWAS types of case-control or similar data.  You need to look in other ways.

Zuk et al. address this.  They build their idea that much of the observed heritability is estimated on a purely additive model, and yet at least some factors may be what in standard biochemical terms is known as 'rate limiting' effects.  At some level of concentration of such factors, they or what they interact with no longer works the same way if it works at all.  The authors outline a model which, under various assumptions about how many steps are rate-limiting such that variants in those steps define measures of heritability, might begin to explain familial correlations not accounted for by current additive effects.

It is already nigh impossible to get stable estimates of hundreds of additive effects, mostly very small (see our current series of posts on Probability does not exist!).  It's one thing to estimate the effects of, say, 100 additively contributing genes. Variants will be many in each, variable in their frequencies among samples.  If purely additive, the context doesn't matter: in each population you can estimate each factor's effect and add 'em up to get a given person's risk.  But if context matters, that is, if the effect of one factor depends on the specific variants in the rest of the genome (forgetting environmental effects!) then it's quite another to estimate those interactions.  Roughly, if 100 genes interact in pairwise fashion (2-way interactions like AxB only), that means 10,000 interaction effects to estimate. Zuk et al. certainly acknowledge this problem clearly.  But the authors suggest kinds of data on relatives of various degrees that might be practicably collected and could reveal discrepancies from expectations under additive models, and account for more or all of the heritability.

Zuk et al. promise that if we study 'isolated' populations we'll have a shot at the answer!  This is not new, and indeed studies in Finland and Iceland led a previous charge for similar reasons.  Perhaps it is a good idea....and the authors provide tests that could be done if there were adequate data collected.  It will be done and will expend more millions of research dollars in the process.  But, if complexity is real, the most we'll get is a few hits and a few weak statistical signals.  But we're in that place already, so in that sense this is another way of changing the goalposts, because we did not reap bonanza from those studies.

The answer
Like most such papers, Zuk et al. make gobs of assumptions about the model, the data, and the underlying basis of dichotomous disease (present or absent).  Facing such complex problems it's hard or impossible not to make simplifying or exemplifying assumptions.  As usual, if one probes these assumptions, there will be additional sources of variation and uncertainty that are being ignored or that must be estimated from data, so that even the rosy answers suggested by the authors, which are far from promising a complete understanding, will be overstated (and that is the clear-cut history of the senior author, and most of his peers in the profession, including Francis Collins).

The actual answer is well known, even if it's resisted because it's not very convenient: life is a mixture, or spectrum, of all of these kinds of components.  We know that additive models work very well in many domains, such as agricultural and experimental breeding (artificial selection), and that as GWAS sample sizes have increased steadily they have steadily been identifying more contributing genes.   That is, we have not plateaued to a point that bigger samples will not add more genes, the remaining recalcitrant effect due to interaction.  We think even the authors acknowledge that the interactions may comprise sets of individually weak contributors.

Further, because most variants in our population are in fact, and indisputably, rare to very rare, they form a large aggregate of potentially contributing genes that will vary from sample to sample and population to population.  This is, really, a fact of life.

Networks, the  sugar plum fairies of promised genomic medical miracles, involve tens of genes, many with multi-way interactions, like AxBxC. And what of higher-order interactions such as the square of one or more factor levels?  One might as well count stars in heaven as attempt to collect accurate data from the entire human species in order to get enough data to estimate these effects.   And that, of course, is foolish for many reasons, not least being that environmental factors change and vary and heritability is a measure of genetic relative to environmental effects.

Zuk et al. are promising a series of papers along this track, and have coined a new name for their idea (a standard marketing ploy). We'll be hearing a lot of me-tooism, users avidly diving into the LP (limiting pathway) model.  That certainly doesn't make the model wrong, but it does affect where the goalposts are. We're hopefully not hypocrites: we have our own simulation program, called ForSim (freely available to others) by which some of these things could be simulated, and we may do that.

Nobody can seriously question that context is very important, and that includes various kinds of interactions.  But the issue is not just what the mix of types of contribution are, how stable, how variable among samples, and so on.  Even if Zuk et al. are materially correct, it doesn't erase the problematic nature of trying to estimate, much less generally to do much about, the joint effects of the many genes, in their tens or hundreds, whose variation contributes to a trait of interest.

"Discovery efforts should continue vigorously"
We ourselves are far from qualified to find technical fault with the new model, if there is any.  We doubt there is.  But the point is not that this new paper is flim-flam, even if it simplifies and makes many assumptions that could be viewed as perhaps-necessary legerdemain, given the situation's complexity.  Or, perhaps more clearly, red-herrings to distract attention from the real point, which is  whether changing the goalposts in this kind of way changes the game.

One can ask--should ask--whether regardless of interaction or additivity, it's worth trying to document them, or whether Francis Collins' insistence on luxury medicine (personalized prediction of weak effects with lifelong treatment mainly available to paying customers) is a realizable goal.  The same funds could, after all, be spent in other ways on indisputably stronger and simpler genetic problems, that could be far more directly and sooner relevant to the society that's paying the bill.

Arguing either/or (additive or non-additive) and attempting to relate that to the desirability of keeping the Big Study funds flowing is a carnival barker activity.  The authors at one point subtly make the anti-Collinsian but obvious point that personalized gene-based prediction is generally going to be a bust.  But then they argue let's plow ahead to discover pathways. Let's have our goalposts everywhere at once!  There are other, better, logistically easier and probably less costly ways to find networks and if there aren't already, there ought to be research invested in figuring them out (e.g., in cell cultures).  Epidemiological studies are expensive and of low payoff in this kind of context.

As the authors clearly say: discovery efforts should continue vigorously" despite their points, acknowledging in fact and cleverly hedging all bets, that many variants remain to be discovered by current approaches.  This is a fair-grounds where PT Barnums thrive.  It keeps their particular circus's seats filled.  But that doesn't make it the best science in the public interest.

Friday, January 13, 2012

Probability does not exist. Part I. The very idea!

We have occasionally mused here on MT about what it means to talk about the risk of such-and-such -- 12% risk of heart attack, 30% risk of rain, 80% risk of a double dip recession.  For example, this could either mean everyone is at 12% risk (a fairly elusive concept), or that 12% of the group have 100% risk and the rest have 0%.

We are lead to post on this now because of a recent BBC Radio 4 program, More or Less, which is always about the meaning of statistics, but the Dec 30 program happened to mention the Italian statistician, Bruno de Finetti and his book, Theory of Probability, which begins thus:
                                    PROBABILITY DOES NOT EXIST
The abandonment of superstitious beliefs about the existence of the Phlogiston, the Cosmic Ether, Absolute Space and Time, . . . or Fairies and Witches was an essential step along the road to scientific thinking. Probability, too, if regarded as something endowed with some kind of objective existence, is no less a mis-leading misconception, an illusory attempt to exteriorize or materialize our true probabilistic beliefs. 
So, God is dead, but what does this mean, exactly?  A 2002 paper by Robert Nau on de Finetti's thesis explains that de Finetti meant that probability is nothing but a subjective analysis of the likelihood that something will happen, that probability does not exist outside the mind.  That is, it's the rate at which a person is willing to bet on something happening.  This is as opposed to the classicist or the frequentist's view of the likelihood of aparticular outcome of an event. That view depends on the assumption that the same event could be identically repeated many times over, and then 'probability' of a particular outcome has to do with the fraction of the time that outcome results from the repeated trials.  This example from Nau's paper clarifies the differences in approach.
For example, in the case of a normal-looking die that is about to be tossed for the first time, a classicist would note that there are six possible outcomes which by symmetry must have equal chances of occurring, while a frequentist would point to empirical evidence showing that similar dice thrown in the past have landed on each side about equally often. A subjectivist would find such arguments to be suggestive but needlessly encumbered by references to superfluous events. What matters are her beliefs about what will happen on the single toss in question, or more concretely how she should bet, given her present information. If she feels that the symmetry argument applies to her beliefs, then that is sufficient reason to bet on each side at a rate of one-sixth. But a subjectivist can find other reasons for assigning betting rates in situations where symmetry arguments do not apply and repeated trials are not possible.
In other words, the idea is that probability is not part of the real world, only of one's belief in the nature of that world.

What does this have to do with the risk of having a heart attack?  Well, how much are you willing to bet on your chances?  That is, if your chances are 20%, and you feel that's high, you might be willing to do whatever you can to reduce your cholesterol, you might take up going to the gym more regularly, quit smoking, or become a vegan.  Someone else, though, might feel that 20% is not so high, and do nothing at all to alter their (what are currently considered to be) risk factors. But the physical basis of that belief is far less clear than the idea of the belief. 

And, physicians advising us differ in how much they are willing to bet on the probability we'll get sick. Some are very diligent about cholesterol, some less so, some advise all men of a certain age to have a PSA test for prostate cancer, others none.  They are reading the same probabilities, but what they make of them differs. Indeed, the question of interpretation is secondary to the notion of probability itself.

Yet, how do we account for the role that probabilities do play in the real world, such as in the example from which much of probability was developed, when events can be repeated: gambling.  And 'bet' is the appropriate operational concept.  The formal theory was largely developed in the literal context of gambling, but the same idea applies to health.

In rolling dice, we have 6 outcomes and no reason to prefer any of them (see below!).  If we know about rolling, we might, in advance, decide that each possible outcome would be as likely to result.  We don't  know which, so we might say that in a large number of rolls, each face would come up the same number of times.

In fact, however gambling notions of probability first arose (scholars have some ideas, but we don't), by the time formal theories were being developed, there was extensive and systematic experience with past sets of rolls that actually did occur (apparently not obviously so in Roman times, where gambling with bones was thought to be related to things like how the gods viewed the gambler, etc.).  We don't personally know how extensive such data, experimental or otherwise, were but the notions of equal occurrences not only seemed intuitive at the time but backed up by experience.  '6' came up about 1/6th of the time in dice games, leading naturally to a theory that all sides had equal chances to arise -- fractions of the times it will arise.

Heart attack risks based on, say, cholesterol levels, are based on past experience, too.  But unlike dice, people are more than simple structures.  We have more than cholesterol levels.  So the fraction of people with cholesterol over some level, who had heart attacks in our studies, is used to estimate the fraction of people with such level will have a coronary in the future.  Yet we know very well that each person's 'ancillary' risk factors in the data on which the probability was based were different, and worse, that we simply cannot know about those risk factors in the future.  So what does the genetic-risk-perveryer's probability actually mean?

We also talk in probabilistic terms when we say such things as that God probably exists (or doesn't).  This clarifies the unclarity of such wording.  God either does or doesn't exist, clearly without any actual 'probability', so this is really just a statement, using a serious-sounding word, of strength of belief.  If you examine it closely, the same really applies to similar statements about whether we'll have a heart attack or not, or whether 5 & 6 will come up on the next roll of a pair of dice.  Even those sound more rigorous, or suggest experiments or relevance to actual data, even they are based on the assumption that multiple replicates of unique events can take place, and they are based on some idea, or 'model', of the process involved such as how we measure cholesterol, how we sample people and measure their cholesterol and diets and obesity and so on), and even how dice are rolled (see next installment!).

Some statisticians and books and lectures simply assume that we know what probability is and don't attempt to define it, or to do so in practical or frequentist terms.  Others try to wriggle out of this situation by saying that the frequentist terms can be disregarded and that there are other systematic ways to extract from the available data alone, some idea of what our best idea of the situation is.  These are called 'likelihood' and 'Bayesian' approaches (there may be others we're not aware of), but if they stay somewhat closer to actual data, they essentially are ways to strengthen belief, and belief is in ourselves rather than a physical property of the object of belief.  That is the subjectivist assertion.

Next time, we'll show how the meaning of  'probability' or the interpretation of repeated events in probability terms--even in seemingly simple cases like coin-flipping or dice-rolling--are far from clear.

Thursday, January 12, 2012

Do we still not know what causes cancer? Part III

This series of posts is about cancer but also about the nature of life itself.

Do we still not know what causes cancer?  We've discussed the origin of the current theory of cancer, the SMT or somatic mutation theory.  As we outlined in Part I, this 50 year old view, that cancer is a disease of diseased cells that have been screwed up by genetic mutation, has focused almost all cancer research in one way or another.  It is relevant as well to a proper understanding of life itself, as we suggested in Part II.  Is the lack of progress in cause-directed (that is, gene-based) therapy, a result of a badly misdirected effort, rather than just the heavy challenge of targeting genetically altered cells?

The competing 'tissue organization field theory' (TOFT) is that cancer is a tissue rather than cellular disease, and goes roughly like this, as outlined by Soto and Sonnenschein (hereafter, SS)  in the May 2011 BioEssays point-counterpoint that triggered this series of posts:

SS point to several aspects of cancer that do not reflect simple genetic causation, or simple one-cell-gone-bad-on-its-own  model of causation; the latter is what they, very inaptly in our view, imply is the heart of the SMT model.  Experiments show that a single transplanted cell cannot generate the multiple cell type architecture of the organ from which it was taken.  Sometimes cancers regress, or a transplanted cancer cell can integrate into a normal tissue architecture in the recipient organ without proliferating as a cancer.  That means that the cell is not, in these experiments, inherently abnormal as might be expected on a genetic model, assuming no experimental artifact.  Some tumors regress upon  hormone treatment, again showing that their abnormal behavior may be a matter of signaling and that it can be reversed or slowed by changing the signaling environment: the cell is not inherently mischievous.  Some normal cells can become cancerous if transplanted to some other tissue context, showing that mutations are not needed.  Environment may not be everything, but it's not nothing, either.

Further, SS point out that interactions among the different types of cell in a tissue cannot be reduced to individual cellular events.  That seems wholly correct, and shows the clearly relevant 3-dimensional architecture of tissues.  Most cancers arise in tissues that include both supportive (stromal) cells and actively dividing organ-specific functional (parenchymal, often epithelial) cells, and that these must interact in normal as well as abnormal tissues.



The figure is from the SS paper showing their idea of how context affects tumorigenesis.  Of course, they want this to be a tissue-architecture phenomenon in which the 'carcinogenic event' is not a genetic mutation.  They may well be right that many triggering events are not themselves mutations (but the SMT asserts that induced by the event to proliferate, the cells become vulnerable to mutation).  In any case, one could offer the same figure for events that are mutational, because of course once a cell does not respond to its environment properly it could be induced to grow in undisciplined ways for that tissue, as cancer cells do.

SS seem to criticize the genetic theory of cancer because tumors of the same organ from different patients seem to involve different sets of mutations, as if variation among cases (and imperfection of mutation detection methods) means that the SMT is an erroneous view.  Of course there are limits to what kinds of mutations can be detected in complex cancerous tissue (that also contains normal vessels, nerves, and so on), but in a polygenic view of cancer, as a complex trait involving many genes and signaling pathways, like other complex traits,  this variation and multiple gene involvement is not a reflection of a wiggling, erroneous theory, but is just what one would expect.  No geneticist we know thinks otherwise, even if they may want simpler answers (as many GWASers do).

SS focus their discussion only on 'sporadic' cancers, that is, ones without a family history of the same tumor type.  That is a completely false, indeed naive, dichotomy.  Most if not every cancer, being a polygenic trait, will involve some inherited risk components, even if GWAS or whole genome sequencing of tumor vs host-normal tissue can't detect weak effects.  SS are quite wrong that most 'inherited' cancers are early onset or pediatric--they are not including the multigenic effects, so theirs is a quite restricted view.

According to SMT, cancer is a disease of sick cells in a normal environment.  But in TOFT, it is a disease of normal cells in a sick environment.  The SMT is a special case of genetic evolution because modified genomes can be inherited but, since cancer is a disorder of tissue architecture, abnormal tissue architecture cannot--a fertilized egg has no tissue!

A recent and very informative installment of our favorite BBC Radio4 program In Our Time discusses macromolecules, and we posted on that separately, starting here.  But this installment just casually drops an observation about biomedical applications of macromolecules that is relevant here.  Webs of supporting tissue can now be constructed of synthetic macromolecules, and embedded with stem cells, then placed in context to repair skin, trachea (windpipe) or other types of tissue, where the cells flesh out the matrix which develops into normal tissue.  The casual comment is that each application requires a different artificial matrix, because stem cells respond differently to different substrates.  Clearly context matters!

We scientists are vain and we all want to be part of a major 'paradigm shift' in our respective fields--we want to be important, to live in important times, and to be the architect of a grand transformative event.  So it is common that we want to suggest (and, of course, name) new sweeping theories.  That may sometimes be correct, as it was for Newton, Darwin, Einstein and others of their fortunate and insightful ilk, but that's very rare, and it is usually uncalled for.  TOFT vs SMT oversimplifies what seem to be overlapping phenomena related to cancer--and, indeed life and its evolution itself.  But there is no conceptual revolution involved.  Cells induced by whatever means to misperceive or or respond wrongly to their signaling environmental context can go off on their own, until correct in some way, or in some instances can get out of any such control.  There's no reason to be surprised at that.

SS, hinting perhaps that they are paradigm-shifters themselves, conclude by citing some philosophers of science and using that to criticize the revisions that SMT advocates regularly make in their theory as the result of experiments that, SS argue, support TOFT instead.  There is an exchange of barbs in the September issue of BioEssays, but it adds nothing of substance to the discussion; there, SS again sneer at changes in the SMT theory as 'changing the goal posts', but indeed that is exactly what any valid theory of life (or any area of science) must do as more is learned.  It is only a valid criticism if the fundamental aspects of the theory are abandoned, but this is not at all the case.  Indeed, as the exchanges show, the SMT is only strengthened--especially if one takes context into account, as it must.

In our view, nobody can deny the importance of context or even that cells that are mis-informed about or that mis-interpret their environment can launch out-of-control growth.  But SS seem again to suggest that geneticists are holding essentially to single-gene concepts of causation, which as we noted above we think no sane geneticist does--even if some mutations may make individually strong contributions to risk. After all, BRCA mutations do that, yet nobody thinks the tumor waits 40 or more years to show up except because other events must also occur (indeed, BRCA genes are involved in mutation repair!).

But such philosophizing is irrelevant if not self-serving baloney!  Every science is always imperfect, and always under revision as new facts become known.  Whether the revisions in SMT are cogent is a separate question, but revision itself is not a fault.  Likewise, SS basically ignore the huge wealth of evidence of clonality and the many clearly known mutations relevant to cancers (in a sense, they don't even try to revise the TOFT).  None of this means we have 'the' answer, because there may be no single answer.  These are not dichotomous, incompatible views of abnormal cell behavior. 

In the end the SMT theory, that cancer can and usually does involve genomic mutations, is completely defensible, as Vaux very powerfully shows in his part of the BioEssays exchange.  This doesn't mean cancer is just a disorder of isolated cells!  Naturally, context must matter.  And, instances of cancer cells becoming normal, or not leading to cancer when transplanted into a normal context, shows that SMT can be oversimplistic.  But a polygenic view of cancer at the cell level is totally consistent with both.  Mutations are involved, but cancer is a disease, at least in part, of mis-cooperation--aberrant signaling or response to context.

We can't resist our own vanity in pointing out that in and around 1990, I had suggested in several papers and a book a somatic-polygenic etiology for cancer, in terms that for the specifics known in its time were essentially modern conceptually, and that were consistent with a contextual yet genetic idea of the nature of cancer.  The idea was compatible with what we know (and knew) of epidemiology,  genomic evolution and causation, and that cancer is a disorder of misbehaving cells, involving gene networks.  A similar view is the bottom line message of MT, the book.  A  mix of somatic and inherited genomic architecture involving multiple contributing genes, in a tissue context stimulated by non-genetic environmental factors such as mutagens and stimulants of cell division, provides a consistent if not simplistic view of cancer, because it puts cancer into the context of normal biology and its geological as well as somatic evolution. 

False dichotomies here, as so often, reflect yearning for simple explanations for complex phenomena.  In fact, what we clearly know about the nature of genomic action, and the essential role of cooperation in the making and maintenance of multicellular organisms, shows that we need no all-or-none 'theories'.  We just need to view cancer as a phenomenon in the kind of biology we already know very well.  That there will be variation in the trait and its cause is exactly what we expect.  So is the fact that causation can be difficult to attribute to individual factors.  That we cannot simplify polygenic phenomena is an apparent reality.  It's not a matter of one wrong or right theory.  It's a reflection of how life works!

Wednesday, January 11, 2012

Do we still not know what causes cancer? Part II

Part I in this series described the SMT, or somatic mutation theory of cancer.  The original theory was developed from some data on both the epidemiology and the known cases of inheritance of cancer susceptibility.  It lead to a focus on the idea that, at the cell level, cancer was a misbehavior disorder due to mutations--changes in DNA--in genes whose normal function was critical for the cell type in question--be it lung cells, intestinal cells, or other tissues.  The cell can't behave properly if its relevant genes have been changed.  The idea is then that a given case of cancer is due to the spread of a clone of cells, descended in the person's body from a single initial 'transformed'--misbehaving--cell, and cells in that clone then accumulate a diversity of subsequent mutations.

Tests of clonality and searches for mutations have been done, and these have been successful.  Similarly, genomewide tests have shown that cancer cells, relative to host normal cells, do reflect many mutational differences.  At least some of these are repeatedly found, and are in genes related to cell division and other relevant aspects of behavior.  Some of these changes can be inherited, leading to elevated risk, as we described briefly in Part I.  The picture was complex--essentially, polygenic, with different tumors manifesting different mutations.  Still, patterns of mutational change have been shown to be relevant to response to therapy and prognosis.  It all seemed consistent with the SMT.

However, there were some weak points in the data.  Normal cells also show mutations, and cataloging the differences from tumor cells is difficult.  After all, even under the SMT, normal cells would be expected to show mutations in the same genes found changed in tumors, because that's how the combinations of 'bad' changes accumulate.

Further, a 'theory' of the essential qualities of life, that goes beyond our contemporary obsession with Darwinian selection and genetic determinism based on competition, such as we try to discuss in MT (here, and in our book of the same name), stresses the role of complex multi-component cooperation in the nature of life among organisms, species, and cells.  Signaling interactions are a fundamental property of that cooperative aspect of life.  A cell's behavior is instructed by its current constituents (including the genes it's using at the time) and the conditions it detects in its environment.  What it detects alters the genes it will express or repress.  A stomach cell expresses appropriate genes for stomach-related behavior, but not genes involved in, say, liver, brain, or blood.  When the environment changes, the cell changes its gene expression and its behavior (thus, when a stomach stem cell detects the absence of adjacent differentiated stomach cells, it divides to replace the lost cells.

What we know about such complex phenomena is that they typically are polygenic, that is, are affected by many genes, and their variation can be due to many different combinations of variation in those genes.  This is what we write a lot about, for example, in the context of GWAS findings.  In our view, to this extent, cancer like other complex traits, is a polygenic phenotype involving cell-to-cell signaling as a determinant of the complex structure of organs like lungs, skin, brains, and ovaries.

Thus, a key feature of life is properly timed preparation, detection, and responsiveness of cells.  Once it becomes committed to the environment it has been prepared to detect, or to whose changes it detects, it is channeled in particular directions....and its set of expressed response detectors (signal receptors, for example) limit what it can do in the future.  In a polygenic view of tissue behavior, there would be many different ways to go awry.

Viruses and other cellular components, normal and from the outside, can also enter the genome or pop  copies of themselves elsewhere in the genome.  This can lead to abnormal effects on the regular genes in the vicinity of the genome where such a copy has, by chance, landed.  For example, a gene may be induced to be expressed abnormally as the result of such events.  This is not a mutation in the expressed gene itself, but in its anomalous usage.  But once the transposed bit of DNA is there, the cell and its descendants are doomed to obey its effects!  It is, in a sense, a kind of somatic mutation, but would never be detected in sequencing the affected gene itself.

The May 2011 BioEssays point-counterpoint includes one part, by Vaux, defending the SMT, but opposed by Soto and Sonnenschein who argue for a tissue organization field theory (TOFT).  One need not accept all-or-nothing combative 'theories' to ask whether we have somehow misinterpreted the SMT, leading to a research focus that will have only limited success, if based on the expectation of mutant genes as the cause of cancer.  That can be important, of course, in the research approach to effective therapies. But it could also be important in what we understand about genetics....and even about evolution and life itself.  That's because we might have been too deterministic in assuming genomes to be self-contained 'programs for life'.

This can be of fundamental importance for understanding life, far beyond its relevance to cancer.  That's because cooperation among genes and other cellular components, rather than gene structure itself, may provide the critical explanation of cellular behavior--even if genetic variation indubitably would be one way to affect that cooperation.  In part III of this series, we'll discuss the basic ideas underlying TOFT.

Tuesday, January 10, 2012

Do we still not know what causes cancer? Part I

Many theories have been proposed for the causes of cancer.  We were involved in some of this work, long ago, before molecular approaches were possible.  We were present when various immunological and other theories were being displaced by a genetic theory, and that theory has been developed over the years.  Cancer is an evolutionary as well as genetic phenomenon, but it's the evolution of genetic variation among cells within the body.

A few observations, too much to go into here but involving a fairly rare childhood eye cancer called retinoblastoma, led to the idea that one could inherit susceptibility mutations (variation in particular genes), but that that required waiting for other mutations to occur somatically, that is, in body cells as they divide throughout life.  Three of the key bits of evidence were, first, that frankly inherited susceptibility seemed rather rare (and it still does, for most types of cancer), even if inherited variation contributes to risk.  Second, it was shown by some clever early experiments that cancers are clones of cells within the affected person's body: the tumor began as a single 'transformed' cell.  Third, the risk of cancer rises with age in a way that seemed consistent with waiting time distributions; that is, a person had to 'wait' until some single cell was transformed, to become the progenitor of the tumor as it grew and sometimes spread around the body (metastasized).

Together, these facts suggested that cancer was a multihit, mutational disease.  It was a genetic disease because the progenitor cell was transformed by mutations.  It was clonal in that this single cell led to the entire descendant set of cells that comprised the tumor.  And it was multihit in the sense that many different mutations were required to transform a cell.  This was the somatic mutation theory (SMT) of cancer, which is what we ourselves worked on when we were in Texas long ago.  The idea is that cancer is a disorder of the cell itself, a damaged cell that did not behave properly in its context.

The age pattern of cancer--how fast risk increased with age--could be associated with the type of tissue.  Carcinomas grow in dividing tissue.  In most organs, partly differentiated stem cells divide and become terminally differentiated for their type of tissue (stomach, intestine, etc.).  When the differentiated cells died or were sloughed of, the stem cell would divide and produce more terminally differentiated cells.  The tissue maintained its integrity because the cells had the right genes expressed, receptors on their surface, and so on, to behave properly for their type of tissue.  Stem cells were normally quiescent until stimulated to divide.  Somatic mutation released that inhibition and disrupted the orderly responsiveness, leading to the relatively undisciplined proliferation that is cancer.

The more stem cells at risk in a given tissue type, and the more their natural pattern involved dividing and differentiating, the faster risk would accumulate.  If cells stopped dividing in a person's natural life-history, tumors became rarer and rarer as the person got older.  The pattern was consistent, and the age-pattern of onset suggested that many mutational 'hits'  were involved.

Work using techniques that became available mainly in the 1980s was consistent.  Evidence of mutation in cancer cells compared to normal cells from the same individual implicated multiple different genes with (as in other complex traits), different sets of mutations in different tumors of the same type (lung, intestinal, etc.).  At the same time, some of these mutations could be inherited, if the person had inherited a good copy of the gene along with a defective one.  Then, one might have to wait for the bad-luck mutation of the other copy of that gene, along with some other genes.  That's why even strong risk-affecting mutations don't cause cancer right away; instead, you have to wait less time for some other complement of mutations to arise.

It was clear that tumors were clones by and large, but as they grew and spread, new mutations, often involving unstable, multiple chromosomal changes, would arise so that the tumor itself then was comprised of a tree of varying descendant cells.  When the right (for the cell, though for the victim, the wrong) set of mutations arose, some cells gained the ability to spread more rapidly, to invade other types of tissue, and so on.  The more potent cells could out-compete the more sluggish ones, in a kind of selection.  This became a kind of natural selection when drug therapy is applied, as some cells could survive the drug, leading to resistant tumors.  The population of tumor cells were mutant but they were still the host's own cells, which explained why the immune system, structured to detect invading foreign cells, was not good at finding and removing them.  Modern genetic analysis of tumor cells and normal cells has generally found evidence consistent with these ideas.

In this view, cancer is always a genetic disease and perhaps thus always amenable (in principle) to a genetic therapeutic approach:  find the mutant gene and target cells expressing it in ways that are specific, so as to leave the same person's normal cells unaffected.

This was, then, an evolutionary theory of cancer involving genetic changes in somatic cells, complemented perhaps with some inherited mutations.  It had most of the elements of Darwinian organismal evolution, including its basis in genes.  It fit the epidemiology, including the role of environmental factors--largely being those that induced mutations or stimulated cell division.    It is a theory we worked on, wrote about, and believed seemed consistent with the evidence (including, yes, GWAS studies of cancer, and various cancer genome projects!).  We even suggested somatic polygenic models of evolution among body cells as consistent with the age-onset patterns.

But this theory has been questioned on a number of grounds, and there are, as usual, alternative explanations.  These have been aired in a recent point-counterpoint in the May 2011 BioEssays in which the protagonists discuss serious questions about the somatic mutation theory of cancer.  The facts discussed above may apply, but they may not account for all cases of cancer....or perhaps the data have been interpreted in the context of an assumed theory and hence seemed to be consistent with that theory.....a bias we often write about here on MT.  Perhaps there are other kinds of causation.

We'll discuss those, as raised in the article, in our next post in this short series.

Monday, January 9, 2012

Polymers: the not-so-secret Secret of Life. Part II. How it works

So, prompted by the recent BBC Radio 4 program, In Our Time, we blogged last week about polymers, long molecules made of strings of subunits.  Many naturally occurring polymers, and other physical structures like crystals, are built of large numbers of copies of the same subunit--the same type of atom, for example.  Even many of the substances produced by organisms, such as starches, are basically like that.

But the polymers of life, DNA, RNA, and proteins are different in a fundamentally important way.  They are constructed of chains of different subunits.

Life depends on the nature of heterogeneous polymers--made of arrangements of different chemical 'beads' in strings.  Life is a phenomenon of cooperation within the elements of and between, polymers.  One is DNA, with four types of bead, and another is proteins with 20 different types.  With just that knowledge alone, you couldn't tell whether the strings were from bacteria, fungi, plants, or animals.  The functional information in the string is based on the arrangement of the elements.

Conceptually, think of these as strings of pop-beads, as in this figure, that we just grabbed from the web:



The four colors correspond to the 4 nucleotide types in DNA.  These 'DNA-beads' can be strung together in any order, and for basically any length.  Now imagine a DNA molecule that is a pop-bead string that's 300,000,000 beads long, stretching for about 25 miles.  Now there will be various short strings (say, yellow-yellow-green, or blue-blue-blue) that are seen again and again, and correspond to some function.  That may be that they code for some particular amino acid (a 'bead' in a protein string of amino-acid pop-beads).  Proteins are similar but use 20 different bead colors, each with its own shape.  This is how DNA codes for the amino acid strings in proteins, and you can easily imagine that the order of individual beads, and of these small sequences of beads, carries the information in the DNA.  And thousands of different proteins, with their respective shapes are what carry out most of the jobs of life.  Not to be insulting, but the difference between you and a fungus is the set of these molecular shapes that you and it produce, and the sequence of the DNA that codes for them.

But there's more.  The beads of all types have the same knobs and sockets that allow them to chain together, but they are also of different shapes.  That's important.  So, for example, you can see that a green bead might fit into the depression in a blue bead.  So a string floating freely in a cell could fold up on itself, with blue-green 'bonds'.  The location of the blues and greens along the chain would determine where this folding occurred, and that in turn would give the self-folded chain a particular shape, and that shape could perhaps interact with the folded shapes of other, different strings.

Or, folded up molecules might patrol along a DNA-bead string, and wherever they found, say, blue-yellow-yellow-blue-red, they could stick to the shape of that bead sequence.  That could in turn attract other such molecules to the location.  This kind of bonding to DNA is what controls many of its functions: which nearby strings are made into RNA and then used to code for protein--that is, which of its genes a given type of cell uses, or into RNA that folds up upon itself as described above, giving it some function.

The point here is that with these shape-recognition features leading proteins and DNA to take shapes, and stick to parts of themselves or of other molecules based on these shapes, is the essence of how life works.  And why we say that life is a polymer phenomenon.  And it is all of these recognition-based interactions that give a cell or an organism its specificity.  And to do that, its countless components must be present in the right combinations at the right times.  That is, they must cooperate.

Messages are passed between cells that cause cells to change what they do, including which of their genes they use.  The message and the message-detector (signal and signal receptor molecules) are of the pop-bead type.  When the two join up, that triggers hundreds of other such interactions in the cell.  For the right messages to be passed, some cell somewhere must make and release the signal, and other cells must make the receptors and response molecules.  All of this is an exotic dance of similar cooperation.

To understand life this is what one must understand.  Variation in polymer arrangement leads to changes in function, and that can affect probabilities or rates of proliferation, and this of course is what evolution is all about.  A genetic change (mutation) is a change of the order of the DNA 'pop-beads', and if that leads to something that works, or works better under current circumstances, the change will be passed on to the next generation.  If some particular bead-order systematically confers an advantage, it may be transmitted more frequently than others in the population; that's what we call Darwinian natural selection.  But whether because of competitive advantage or just luck, different lineages of life accumulate different bead-orders, diverging more and more over time.

In this sense, that things have to be present together at the right time and place, in the right locations and combinations, makes life a logical phenomenon.  It's about presence and absence, combinations, arrangements, in space and time.  Since the same elements like signaling factors and their cell-surface receptors--the same substrings of pop-beads, can be used in different combinations to bring about different organs or structures, shows this rather clearly.  There's nothing physical about a signal factor or its receptor that resembles the physical structures of limbs or teeth, but the presence of the same factors, in different timing or combinations, is what makes these various structures.  And it is the differential use of these factors, triggered by signal-receptor combinations and regulatory protein-DNA binding, that makes different cell types (and, by extension, different species) what they are. Again, this shows that it is the logic rather than the specific physical traits that matter.  And this logic, combinations and arrangements, that is the essence of cooperation among the interacting elements that make an organism or a species or an ecosystem.  And all of this depends on, and is produced by, polymers made of non-identical 'beads'.

The discovery that evolution is a process of divergence from common ancestor, with the divergence increasing over time because of differential rates or patterns of proliferation of genomic variation--variation in the organisms among information-bearing polymers--has been perhaps more transformative than any other single realization in the history of science.   There are no 'paradigm shifts' involved, but there is a need for a rebalance of the central ideas or theory of life, to recognize that, yes, competition certainly does occur, but the essence of life is much more about cooperation, and it's the cooperation that comes first.

If you think of the huge numbers of cooperative interactions that must occur in the subset of  pop-bead strings that are involved in most individual biological traits, and that many different sets of strings can generate similar outcomes, then you can also see that, a cooperation-based view of life explains why even when natural selection is screening a function, there can be many ways to pass the screen, and any specific part of the system may be only very weakly affected by the selection.  That's why even when selection leads to specific functional adaptations, it's very hard to find solid evidence for it at the gene (pop-bead) level.

Again, understanding the nature of complex cooperation on which life is based, makes things fall into place more naturally than the oversimplified, competition-centered, often single-bead-focused view of life, that is so commonplace even among biologists (and certainly in the biomedical research world).  This is all ever so simple, and there's nothing secret or revolutionary about it. It's information available to anyone who wants to pay attention.  We think it's very important and that's largely why we wrote our book MT, to show its implications in detail.

Life is a polymer phenomenon.

Friday, January 6, 2012

In Memoriam: Jim Crow

On January 3rd, James F Crow died, at the ripe old age of 95.  Jim was one of the 20th century's most prominent population geneticists, training many other leaders in the field as well as providing much of the evolutionary theory we have today.  Madison, Wisconsin must be a healthy place to live because, among other things, Jim's predecessor and one of the ultimate founders of population genetics, Sewall Wright, lived and passed away there.  Wright lived til he was just shy of his 100th birthday, so perhaps Madison is just too tough a place to make it all the way.

I knew Jim from several meetings, though not as a close friend.  But all of us in our generation were weaned on Crow and Kimura (the latter, a founder of the 'neutral' theory of evolution as a counterweight to the prevailing strongly selectionist view, was a Crow student).  Many other of our most prominent population geneticists trained or worked with him.  His Wiki page stresses his role in teaching, suggesting it may have been his most important single contribution.  If one includes his books, and his very clear arguments about various subjects in his papers, then it would be hard to argue with that.

Crow may not be known for many specific major theorems or the like, but he worked extensively on the nature of natural selection and how to detect evidence of it, on the way interactions among genes were (or, he might insist) were not reflected in adaptive evolution, on the age effect of fathers on disease risk in their children, and on human population variation--these being things I knew him for (he also did experimental and fruit-fly work).

Jim Crow was personally a gentle man and a gentleman.  He was well-rounded personally and in his family.  I cribbed the picture, showing him playing the viola as he did for the Madison Symphony,  from John Hawkes' very fine post, where you can learn more about Crow (John is at Wisconsin).

As I personally witnessed, he could defend a point of view in discussions, but I never knew him to become aggressive about it, nor ad hominem, even when his protagonist was a bully (something I myself witnessed, the bully being alpha male Jim Neel).  He stuck to polite consideration of issues--even though he did have his points of views!

Like the other famous Jim (Watson), Crow did hold some views about human variation, that (I think) in naive ways conflated the facts that we're all genetically different, that genes affect our traits, with group (i.e., 'race') differences.  But we all have our blinders.

In his later, retired years, Jim continued to contribute and one noteworthy way was his editing of a column of retrospectives that ran in each issue of the prominent journal Genetics.  These looked back at major issues and historical figures in ways that might not be widely known among younger readers for whom history was not considered a very important part of science.

Jim had nowhere near the name-recognition of the great triumvirate of  population genetics, Wright, JBS Haldane, and RA Fisher.   In recent years, I've found that people in genetics often have not heard of him (to their knowledge).  But even the great troika, along with other giants in science of their time, are no longer remembered by name.  And Jim had the kind of career, and of kindness, that should satisfy anyone in or out of science.