Showing posts with label -omics. Show all posts
Showing posts with label -omics. Show all posts

Tuesday, January 8, 2019

Susumu Ohno: Accounting for Why Gene Counting Doesn't Account for Things

Gifts, gifts, gifts!  Every day in the media, often promoted by universities, journals, and NIH, we seem to be offered the imminent gift of immortality, if we but pony up for more and more 'omical' science (well, if you have to pay for it, even via taxes, I guess it's not exactly a gift!).

The promise that for nearly two decades has been the main course on the 'omicists' menus, is that by counting--adding up the contributions of a list of enumerated genome locations--all our woes will be gone!  The idea is simple: genes are fundamental to life because they code for proteins and stuff like that, which are the basis of life.  This, in a nutshell, is the justification for much of the Big Data endeavors being sponsored by the NIH these days, long driven for historical reasons by an obsession with genes.

But, at least partly, this obsession has revealed to us what we should--and could--already have known.  Genes are clearly fundamental to life, coding for proteins and other functions.  But the reason we're seeing increasing weariness with GWAS and other fiscally high but scientifically low yield approaches is not new.  It's not secret.  And it is not a surprise.  All we needed to do was to ask, where do genes come from?  It is not a new question, the genome has been intensively studied, and indeed the answer has been known for nearly 50 (that is, fifty) years.

Susumu Ohno (1928-2000), from Google images)
In 1970 Susumu Ohno published his deeply insightful Evolution by Gene Duplication.  This book should be a must-read for all life-science graduate students.  Instead, it has been casually forgotten--one might say conveniently forgotten, except that in our culpable ignorance of the history of our field  or to suit our self-serving careerism, it has not been deemed important to read anything published more than a few years ago.

So, what did Ohno say?

Where do 'genes' come from?
In his time, we didn't have much in the way of DNA sequencing.  We knew that genes coded for proteins, and were located on chromosomes.  We had learned a lot about how the code works, much of this from experiments, such as with bacteria.  We knew proteins were fundamental building blocks of life, and were strings of amino acids.  Watson and Crick and others had shown how DNA carries the relevant code, and so on.

But that did not answer the question: Where do all these genes come from?  I'm not an historian, and cannot claim to know the many threads leading to the answer.  But in essence, a point Ohno is credited for noting and whose importance he stressed, is that new genes largely arise from duplication events affecting existing genes.  He had noticed amino acid similarities among some known proteins (hemoglobins); this and other evidence suggested that chromosomal or individual gene duplication was a mechanism, if not the mechanism, for the origin of new genes.  Expecting random mutations in parts of DNA not already being used to code for RNA or DNA, to generate all the sequence aspects of a code for a new protein that would actually have some use, was too far-fetched.  Indeed, nowadays one can be skeptical if an 'orphan' gene is claimed--that is, one not part of a gene family, of which there are also other genes in the genome.

Instead, if occasionally a stretch of DNA or even a whole chromosome duplicates, the individual inheriting that expanded genome gains two potentially important attributes.  First, s/he has a redundant code; mutational errors in one gene that lead to a non-functional protein can be compensated for by the fact that an entirely different, duplicate gene exists and codes for the same protein.

Secondly, duplication is the basis of a much deeper, indeed fundamental aspect of life, going farther even than just gene: redundancy.

Evolution depends on redundancy: genomes are family affairs
By having redundant genes, the initial result of duplication, an individual is more likely to survive mutations.  And over the long haul, with lots of duplication, the additional copies of a needed gene can mutate and over time take on new function, without threat to the individual, who will still have one or more healthy versions of the gene.

Indeed, perhaps one of the far under-appreciated but even fundamental axioms of life is that it is built on redundancy: not only are genomes almost exclusively carriers of members of gene families whose individual genes arose by duplication events, but our tissues themselves are constructed by repeating fundamental units: multicellular organization generally; bilateral or radial symmetry; blood cells, intestinal villi, lobes and alveoli in lungs, nephrons in kidneys, and so on.

I think it is not easy to imagine a different evolutionary way for our very simple biochemical beginnings to generate the kinds of complex organisms that populate the Earth.  And this has deep consequences for those for whom dreams of omical sugar plums dance in their heads.

Why the 'omics' promises were always doomed to fail, or at least to pale
From the cell theory to Ohno to the very data that our 'omical dreams have yielded in extensive amounts, we have found that life relies on the protection of redundancy.  From genes on up, if one thing goes wrong, there's an ally to pick up the slack.  Redundancy means back-ups and alternatives.  It also provides individual uniqueness, which is also fundamental to the dynamics evolution.

Together, these facts (and they're facts, not just wild speculations) show that, and why, we can't expect to predict everything from individual genes or even gene scores.  There are many roads to the Promised Land.

It is important, I think, and entirely fair to assert that nothing I've said here has ever been secret, known only to a small, Masonic Lodge of biologists exchanging secret handshakes.  Indeed, these basic facts have been at the heart of our science since the advent of the cell theory, centuries ago.  Genomics has largely just added to what was already known as a generalization about life.

The implicit lesson, of Ohmo not Homer, is to Beware of Geneticists Bearing Gifts.

(updated to correct a spelling error in Prof. Ohno's name)

Thursday, October 18, 2018

When is a consistent account in science good enough?

We often want our accounts in science to be consistent with the facts.  Even if we can't explain all the current facts, we can always hope to say, truthfully, that our knowledge is imperfect but our current theory is at least largely true....or something close to that....until some new 'paradigm' replaces it.

It is also only natural to sneer at our forebears' primitive ideas, of which we, naturally, now know much better.  Flat earth?  Garden of Eden?  Phlebotomy?  Phlogiston?  Four humors?  Prester John, the mysterious Eastern Emperoro who will come to our rescue?  I mean, really!  Who could ever have believed such nonsense?

Prester John to the rescue (from Br Library--see Wikipedia entry)
In fact, leaders among our forebears accepted these and much else like it, took them as real, sought them for solace from life's cares not just because they were promised (as in religious figures) but as earthly answers.  Or, to seem impressively knowledgeable, found arcane ways to say "I dunno" without admitting it.  And, similarly, many used ad hoc 'explanations' for personal gain--as self-proclaimed gurus, promisers of relief from life's sorrows or medical woes (usually, if you cross their palms with silver first).

Even in my lifetime in science, I've seen forced after-the-fact 'explanations' of facts, and the way a genuine new insight can show how wrong those explanations were, because the new insight accounts for them more naturally or in terms of some other new facts, forces, or ideas.  Continental drift was one that had just come along in my graduate school days.  Evolution, relativity, and quantum mechanics are archetypes of really new ideas that transformed how our forebears had explained what is now our field of endeavor.

Such lore, and our more broad lionizing of leading political, artistic or other similarly transformative figures, organizes how we think.  In many ways it gives us a mythology, or ethnology, that leads us to order success into a hierarchy of brilliant insights.  This, in turn, and in our careerist society, provides an image to yearn for, a paradigm to justify our jobs, indeed our lives, make them meaningful--make them important in some cosmic sense, and really worth living.

Indeed, even ordinary figures from our parents, to the police, generals, teachers, and politicians have various levels of aura as idols or savior figures, who provide comforting answers to life's discomfiting questions.  It is natural for those burdened by worrisome questions to seek soothing answers.

But of course, all is temporary (unless you believe in eternal heavenly bliss).  Even if we truly believe we've made transformative discoveries or something like that during our lives, we know all is eventually dust.  In the bluntest possible sense, we know that the Earth will some day destruct and all our atoms scatter to form other cosmic structures.

But we live here and now and perhaps because we know all is temporary, many want to get theirs now, and we all must get at least some now--a salary to put food on the table at the very least.  And in an imperfect and sometimes frightening world, we want the comfort of experts who promise relief from life's material ills as much as preachers promise ultimate relief.  This is the mystique often given to, or taken by, medical professionals and other authority figures.  This is what 'precision genomic medicine' was designed, consciously or possibly just otherwise, to serve.

And we are in the age of science, the one True field (we seem to claim) that delivers only objectively true goods; but are we really very different from those in similar positions of other sorts of lore?  Is 'omics any different from other omnibus beliefs-du-jour?  Or do today's various 'omical incantations and promises of perfection (called 'precision') reveal that we are, after all, even in the age of science, only human and not much different from our typically patronized benighted forebears?

Suppose we acknowledge that the latter is, at least to a considerable extent, part of our truth.  Is there a way that we can better use, or better allocate, resources to make them more objectively dedicated to solving the actually soluble problems of life--for the public everyday good, and perhaps less used, as from past to today, to guild the thrones of those making the promises of eternal bliss?

Or does sociology, of science or any other aspect of human life, tell us that this is, simply, the way things are?

Tuesday, October 16, 2018

Where has all the thinking gone....long time passing?

Where did we get the idea that our entire nature, not just our embryological development, but everything else, was pre-programmed by our genome?  After all, the very essence of Homo sapiens compared to all other species, is that we use culture--language, tools, etc.--to do our business rather than just our physical biology.  In a serious sense, we evolved to be free of our bodies, our genes made us freer from our genes than most if not all other species! And we evolved to live long enough to learn--language, technology, etc.--in order to live our thus-long lives.

Yet isn't an assumption of pre-programming the only assumption by which anyone could legitimately promise 'precision' genomic medicine?  Of course, Mendel's work, adopted by human geneticists over a century ago, allowed great progress in understanding how genes lead at least to the simpler of our traits, with discrete (yes/no) manifestations, traits that do include many diseases that really, perhaps surprisingly, do behave in Mendelian fashion, and for which concepts like dominance and recessiveness been applied and that, sometimes, at least approximately hold up to closer scrutiny.

Even 100 years ago, agricultural and other geneticists who could do experiments, largely confirmed the extension of Mendel to continuously varying traits, like blood pressure or height.  They reasoned that many genes (whatever they were, which was unknown at the time) contributed individually small effects.  If each gene had two states in the usual Aa/AA/aa classroom example sense, but there were countless such genes, their joint action could approximate continuously varying traits whose measure was, say, the number of A alleles in an individual.  This view was also consistent with the observed correlation of trait measure with kinship-degree among relatives.  This history has been thoroughly documented.  But there are some bits, important bits, missing, especially when it comes to the fervor for Big Data 'omics analysis of human diseases and other traits.  In essence, we are still, a century later, conceptual prisoners of Mendel.

'Omics over the top: key questions generally ignored
Let us take GWAS (genomewide association studies) on their face value.  GWAS find countless 'hits', sites of whatever sort across the genome whose variation affects variation in WhateverTrait you choose to map (everything simply must be 'genomic' or some other 'omic, no?).  WhateverTrait varies because every subject in your study has a different combination of contributing alleles.  Somewhat resembling classical Mendelian recessiveness, contributing alleles are found in cases as well as controls (or across the measured range of quantitative traits like stature or blood pressure), where the measured trait reflects how many A's one has: WhateverTrait is essentially the sum of A's in 'cases', which may be interpreted as a risk--some sort of 'probability' rather than certainty--of having been affected or of having the measured trait value.

We usually treat risk as a 'probability,' a single value, p, that applies to everyone with the same genotype.  Here, of course, no two subjects have exactly the same genotype so some sort of aggregate risk score, adding up each person's 'hits', is assigned a p.  This, however, tacitly assumes something like that each site contributes some fixed risk or 'probability' of affection.  But this treats these values as if they were essential to the site, each thus acting as a parameter of risk.  That is, sites are treated as a kind of fixed value or, one might say 'force', relative to the trait measure in question.

One obvious and serious issue is that these are necessarily estimated from past data, that is, by induction from samples.  Not only is there sampling variation that usually is only crudely estimated by some standard statistical variation-related measure, but we know that the picture will be at least somewhat different in any other sample we might have chosen, not to mention other populations; and those who are actually candid about what they are doing know very well that the same people living in a different place or time would have different risks for the same trait.

No study is perfect, so we use some conveniently assumed well-behaved regression/correction adjustments to account for the statistical 'noise' due to factors like age, sex, and unmeasured environmental effects.  Much worse than these issues, there are clearly factors of imprecision, and the obvious major one, taboo even to think about much less to mention, that relevant future factors (mutations, environments, lifestyles) are unknowable, even in principle.  So what we really do, are forced to do, is extend what the past was like to the assumed future.  But besides this, we don't count somatic changes (mutation arising in body tissues during life, that were not inherited), because they'd mess up our assertions of 'precision', and we can't measure them well in any case (so just shut one's eyes and pretend the ghost isn't in the house!).

All of these together mean that we are estimating risks from imperfect existing samples and past life-experience, but treating them as underlying parameters so that we can extend them to future samples.  What that does is equate induction with deduction, assuming the past is rigorously parametric and will be the same in the future;  but this is simply scientifically and epistemologically wrong, no matter how inconvenient it is to acknowledge this.  Mutations, genotypes, and environments of the future are simply unpredictable, even in principle.

None of this is a secret, or new discovery, in any way.  What it is, is inconvenient truth. These things should have been enough, by themselves and without badgering investigators about environmental factors that (we know very well, typically predominate) prevent all the NIH's precision promises from being accurate ('precise'), or even to a knowable degree.   Yet this 'precision' sloganeering is being, sheepishly, aped all over the country by all sorts of groups who don't think for themselves and/or who go along lest they get left off the funding gravy train.  This is the 'omics fad.  If you think I am being too cynical, just look at what's being said, done, published, and claimed.

These are, to me, deep flaws in the way the GWAS and other 'omics industries, very well-heeled, are operating these days, to pick the public's pocket (pharma may, slowly, be awakening-- Lancet editorial, "UK life science research: time to burst the biomedical bubble," Lancet 392:187, 2018).  But scientists need jobs and salaries, and if we put people in a position where they have to sing in this way for their supper, what else can you expect of them?

Unfortunately, there are much more serious problems with the science, and they have to do with the point-cause thinking on which all of this is based.

Even a point-cause must act through some process
By far most of the traits, disease or otherwise, that are being GWAS'ed and 'omicked these days, at substantial public expense, are treated as if the mapped 'causes' are point causes.  If there are n causes, and a person has an unlucky set m out of many possible sets, one adds 'em up and predicts that person will have the target trait.  And there is much that is ignored, assumed, or wishfully hidden in this 'will'.  It is not clear how many authors treat it, tacitly, as a probability vs a certainty, because no two people in a sample have the same genotype and all we know is that they are 'affected' or 'unaffected'.

The genomics industry promises, essentially, that from conception onward, your DNA sequence will predict your diseases, even if only in the form of some 'risk'; the latter is usually a probability and despite the guise of 'precision' it can, of course, be adjusted as we learn more.  For example, it must be adjusted for age, and usually other variables.  Thus, we need ever larger and more and longer-lasting samples.  This alone should steer people away from being profiteered by DNA testing companies.  But that snipe aside, what does this risk or 'probability' actually mean?

Among other things, those candid enough to admit it know that environmental and lifestyle factors have a role, interacting with the genotype if not, usually, overwhelming it, meaning, for example, that the genotype only confers some, often modest, risk probability, the actual risk much more affected by lifestyle factors, most of which are not measured or not measured with accuracy, or not even yet identified.  And usually there is some aspect that relates to age, or some assumption about what 'lifetime' risk means.  Whose lifetime?

Aspects of such a 'probability'
There are interesting issues, longstanding issues, about these probabilities, even if we assume they have some kind of meaning.  Why do so many important diseases, like cancers, only arise at some advanced age?  How can a genomic 'risk' be so delayed and so different among people?  Why are mice, with very similar genotypes to humans (which is why we do experiments on them to learn about human disease) only live to 3 while we live to our 70s and beyond?

Richard Peto, raised some of these questions many decades ago.  But they were never really addressed, even in an era when NIH et al were spending much money on 'aging' research including studies of lifespan.  There were generic theories that suggested from an evolutionary theory why some diseases were deferred to later ages (it is called 'negative pleiotropy'), but nobody tried seriously to explain why that was from a molecular/genetic point of view.  Why do mice only live only 3 years, anyway?  And so on.

These are old questions and very deep ones but they have not been answered and, generally, are conveniently forgotten--because, one might argue, they are inconvenient.

If a GWAS score increases the risk of a disease, that has a long delayed onset pattern, often striking late in life, and highly variable among individuals or over time, what sort of 'cause' is that genotype?  What is it that takes decades for the genes to affect the person?  There are a number of plausible answers, but they get very little attention at least in part because that stands in the way of the vested interests of entrenched too-big-to-kill Big Data faddish 'research' that demands instant promises to the public it is trephining for support.  If the major reason is lifestyle factors, then the very delayed onset should be taken as persuasive evidence that the genotype is, in fact, by itself not a very powerful predictor.

Why would the additive effects of some combination of GWAS hits lead to disease risk?  That is, in our complex nature why would each gene's effects be independent of each other contributor?  In fact, mapping studies usually show evidence that other things, such as interactions are important--but they are at present almost impossibly complex to be understood.

Does each combination of genome-wide variants have a separate age-onset pattern, and if not, why not?  And if so, how does the age effect work (especially if not due to person-years of exposure to the truly determining factors of lifestyle)?  If such factors are at play, how can we really know, since we never see the same genotype twice? How can we assume that the time-relationship with each suspect genetic variant will be similar among samples or in the future?  Is the disease due to post-natal somatic mutation, in which case why make predictions based on the purported constitutive genotypes of GWAS samples?

Obviously, if long delayed onset patterns are due not to genetic but to lifestyle exposures interacting with genotypes, then perhaps lifestyle exposures should be the health-related target, not exotic genomic interventions.  Of course, the value of genome-based prediction clearly depends on environmental/lifestyle exposures, and the future of these exposure is obviously unknowable (as we clearly do know from seeing how unpredictable past exposures have affected today's disease patterns).

The point here is that our reliance on genotypes is a very convenient way of keeping busy, bringing in the salaries, but not facing up to the much more challenging issues that the easy one (run lots of data through DNA sequencers) can't address.  I did not invent these points, and it is hard to believe that at least the more capable and less me-too scientists don't clearly know them, if quietly.  Indeed, I know this from direct experience.  Yes, scientists are fallible, vain, and we're only human.  But of all human endeavors, science should be based on honesty because we have to rely on trust of each other's work.

The scientific problems are profound and not easily solved, and not soluble in a hurry.  But much of the problem comes from the funding and careerist system that shackles us.  This is the deeper explanation in many ways.  The  paint on the House of Science is the science itself, but it is the House that supports that paint that is the real problem.

A civically responsible science community, and its governmental supporters, should be freed from the iron chains of relentless Big Data for their survival, and start thinking, seriously, about the questions that their very efforts over the past 20 years, on trait after trait, in population after population, and yes, with Big Data, have clearly revealed.

Thursday, January 24, 2013

A 'paradigm shift' in science....or a manoever?

Thomas Kuhn's 1962 book The Structure of  Scientific Revolutions suggested that most of the time we practice 'normal' science, in which we take our current working theory--he called it a 'paradigm'--and try to learn as much as we can.  We spend our time at the frontiers of knowledge, and at some point we have to work harder and harder to make facts fit the theory.  Something is missing, we don't know what, but we insist on forcing the facts to fit.

Then, for reasons hard to account for but in a way that happens regularly enough that it's a pattern Kuhn could outline (even if rare), someone has a major insight, and shows how a totally unexpected new way to view things can account for the facts that had heretofore been so problematic.  Everyone excitedly jumps onto the new bandwagon, and a 'paradigm shift' has occurred. Even then, some old facts may not be as well accounted for, or the new paradigm may just explain issues of contemporary concern, leaving older questions behind.   But the herd follows rapidly, and an era of new 'normal science' begins.

The most famous paradigm shifts involve people like Newton and Galileo in classical physics, Darwin in biology, Einstein and relativity, and the discovery of continental drift.  Because historians and philosophers of science have in a sense glamorized the rare genius who leads such changes, the term 'paradigm shift' has become almost pedestrian:  we all naturally want to be living--and participating--in an important time in history, and far, far too many people declare paradigm shifts far too frequently (often humbly referring to their own work).  It's become a kind of label to justify whatever one is doing, a lobbying tactic, or a bit of wishful thinking.

Is 'omics' a paradigm shift?
The idea grew out of the Enlightenment period in Europe starting about 400 years ago, that empiricism (observation) rather than just thinking, was the secret to understanding the world.  But pure empiricism--just gathering data-- was rejected in the sense that the idea was for the facts to lead to theoretical generalizations, the discovery of the 'laws' of Nature, which is what science is all about.  This led to the formation of the 'scientific method,' of forming hypotheses based on current theory, setting up studies specifically to test the hypothesis, and adjusting the theory according to the results.

If the 17th-19th centuries were largely spent in gathering data from around the world, a first rather extensive kind of exploration.  But by the 20th century such 'Victorian beetle collection' was sneered at, and the view was that to do real science you must be constrained by orderly hypothesis-driven research.  Data alone would not reveal the theory.

With advances in molecular and computing technology, and the complexity of life being documented, things changed.  In the 'omic' era, which began with genomics, the ethos has changed.  Now we are again enamored of massive data collection unburdened by the necessity to specify what we think is going on in any but the most generic terms.  The first omics effort, sequencing the human genome, led to copy-cat omics of all sorts (microbiomics, nutrigenomics, proteomics, .....) in which expensive and extensive technology is thrown at a problem in the hope that fundamental patterns will be revealed.

We now openly aver, if not brag, that we are not doing 'hypothesis-driven' research, as if there is now something wrong with having focused ideas!  Indeed, we now often treat 'targeted' research as a kind of after-omics specialty activity.  Whether this is good or not, I recently heard a speaker refer to the  omics approach as a 'paradigm shift'.  Is that justified?

Before we could even dream about genomic-scale DNA sequencing and the like, we must acknowledge that our understanding of genetic functions and the complex genome had perplexed us in many ways.  If we had no 'candidate' genes in mind-no specific genetic hypothesis--for some purpose, such as to understand a complex disease, but were convinced for some reason that genetic variation must be involved, what was the best way to find the gene(s)?  The answer was to go back to 'Victorian beetle collection'.  Just grab everything you can and hope the pieces fall into place.  It was, given the new technology, a feeling of hope that this might help (even though we had many reasons to believe that we would find what we indeed did find, as some of us were writing even then).

The era of Big Science
Omics approaches are not just naked confessions of ignorance.  If that were the case, one might say that we should not fund such largely purposeless research.  No, more is involved.  Since the Manhattan Project and a few others, it did not escape scientists' attention that big, long, too-large-to-be-canceled projects could sequester huge amounts of funding.  We shouldn't have to belabor this point here: the way universities and investigators, their salaries and careers, became dependent on, if not addicted to, external grants, the politics of getting started down a costly path enabling one to argue that to stop now would throw away the money so-far invested (e.g., current Higgs Boson/Large Hadron Collider arguments?).  Professors are not dummies, and they know how to strategize to secure funds!

It is fair to ask two questions here:
First, could something more beneficial have been done, perhaps for less cost, in some other way?  Omics-scale research of course does lead to discoveries, at least some of which might not happen or might take a long time to occur.  After the money's been spent and the hundred-author papers published in prestige journals, one can always look back, identify what's been found, and argue that that justifies the cost. 

Second, is this approach likely to generate importantly transformative understanding of Nature?  This is a debatable point, but many have said, and we generally agree, that the System created by Big Science is almost guaranteed to generate incremental rather than conceptually innovative results.  (E.g., economist Tim Harford talked about this last week on BBC Radio 4, comparing the risk-taker science of the Howard Hughes Institutes with the safe and incremental science of the NIH.)  Propose what's big in scale (to impress reviewers or reporters), but safe--you know you'll get some results!  If you compare 250,000 diabetics to 500,000 non-diabetic controls and search for genetic differences, across a genome of 3.1 billion nucleotides, you are bound to get some result (even if it is that no gene stands out as a major causal factor, that is a 'result').  It is safe.

This is not providing a daring return on society's largesse, but it is the way things largely work these days.  We post about this regularly, of course.  The idea of permanent, factory-like, incremental, over-claimed, budget-inflated activity as the way to do science has become the way too many feel is necessary in order to protect careers. Rarely do they admit this openly, of course, as it would be self-defeating.  But it is very well-known, and almost universally acknowledged off the record, that this strategy of convenience seriously under-performs, but is the way to do business.

Hypothesis-free?
This sort of Big Science is often said to be 'hypothesis free'.  That is a big turn away from classical Enlightenment science in which you had to state your theory and then test it.  Indeed, this change itself has been called a 'paradigm shift'.

In fact, even the omics approach is not really theory- or hypothesis-free.  It assumes, though often not stated in this way, that genes do cause the trait, and the omics data will find them.  It is hypothesis-free only in the sense that we don't have to say in advance which gene(s) we think are involved.  Pleading ignorant has become accepted as a kind of insight.

For better or worse, this is certainly a change in how we do business, and it is also a change in our 'gestalt' or worldview about science.  But it does not constitute a new paradigm about the nature of Nature!  Nothing theoretical changes just because we now have factories that can systematically churn out reams of data.  Indeed, the theories of life that we had decades ago, even a century ago, have not fundamentally changed, even though they remain incomplete and imperfect and we have enormous amounts of new understanding of genes and what they do.

The shift to 'omics' has generated masses of data we didn't have before.  What good that will do remains to be seen, as does whether it is the right way to build a science Establishment that generates good for society.  However that turns out, Big Science is certainly a strategy shift, but it has so far generated no sort of paradigm shift.

Friday, January 4, 2013

Weighing in on a Weighty subject

Finally, definitive proof!
So, the latest (hottest, and certainly this time just must be true) report is that obesity (that is, Body Mass Index, or weight-for-height) isn't so clearly damaging to health and disease as billions of dollars and millions of pages of punditry and scientific hyperbole have suggested over a mere fifty years.  Whoopie!  We can eat again!  What a relief! 

Or is it?

The study we posted on yesterday seemed to say that.  But even forgetting our usual (of course always cogent and well-placed) reservations about science news bulletins, perhaps there is something else to note, that might cause at least a few milliseconds of thoughtful contemplation.

If this is a causal, material world, then as the argument goes, everything must be understandable  and predictable strictly in terms of molecules and energy--because that's all there is!  And since, the argument continues, evolution has molded life around DNA as the primary causal molecule, we simply must be the product of, and hence predictable from, our genes.

The BMI study was not a genetic report, and only concerned the predictive power of the net measure, BMI.  It was about the long-assumed health risks, or not, of obesity.  But there is a bit of slippage here:  BMI is easy to measure (your weight related to how tall you are), and so many different studies can collect comparable data, etc.  It is thus a convenient measure of choice for obesity.

Taking the current dogma of our time that everything simply must, obviously, necessarily be 'genetic', many studies you've paid for with your taxes have naturally done their best to find the genes 'for' this important health-risk trait.

No, not at all
Thus a major and very large GWAS on the genetics of BMI was published a couple of years ago (Nature Genetics, Nov. 2010).  This study of a mere 250,000 individuals found a small number of modest (statistically 'significant') locations in the genome, including one confirmatory gene (called FTO) that a blind person could find without using his hands.  Other 'known' obesity risk-factor genes weren't in this list, and of course there is the plethora of excuses--er, that is, alternative explanations--for why these genes didn't show up in the hit-list.

Now that is mysterious enough (unless you've been thinking critically about genetic causation, its evolutionary history, and the nature of such studies), but at least it's a large study that should illuminate at least the nub of the causal truth.

However, also in Nature Genetics in Nov. 2010 was another obesity GWAS.  This time the measure used was not BMI, but the waist-to-hip ratio (WHR).  This is another convenient, non-invasive, and cheaply measured index of obesity.  The study was the pooling of 61 studies of a total of a mere 114,000 participants.  Now this study essentially found no overlap in genome region 'hits' with the BMI study!  It also failed to show several genes well known to relate to obesity and related dieases, from many actually focused studies including mouse experimental work.

One can rationalize all one wants about this 'discrepancy' (to use a kind word for it).  But if 'obesity' is a meaningful trait with any sort of unitary causal nature, then measuring it in two ways should generate essentially the same result, after accounting for statistical vagaries.  Just as using a metric (Celsius) thermometer won't tell you anything more about water, ice, and steam than using a Fahrenheit scale.

237 traits linked to genomic loci by 1449 GWAS studies
(Source: www.genome.gov/GWAStudies); 2012
Clear and devastating indictment of the state-of-the-art
This issues seems not at all to have been noticed (openly, at least) by anybody. Instead, it should be seen as a clear and devastating indictment of the GWAS and related 'omics' grand-sample, meta-analysis, quick-and-dirty enterprise that we are investing so heavily and mechanically in.  It should be the miner's canary, telling us clearly that we are not going about this in a right way.

We can't blithely accept the current BMI and health finding as related to obesity in an interpretable way and are rather forced to recognize, as we said in our prior post, that BMI is a stand-in for some confounding factor(s) that may or may not have been measured.  That's because if different genes predict one measure of 'obesity' compared to another measure, there must be some seriously complex or heterogeneous causal variation in our data that we are not measuring, may not know about, but are not highly correlated with each other or, at least, are not consistently correlated with different ways we choose to define something as a trait, or risk factor.

'Obesity' is in some ways an obvious trait in its extremes (from skinny to very over-weight), and body weight is clearly related to health measures of various kinds.  But the Omics Way that is being taken is falling short, and this also means that the Epidemiological Way, of parsing a large plateful of variables into this or that correlation coefficient with various statistical significance levels, is also badly wanting.

We don't have the answers.  Indeed, the problem is not just that nobody has the answers, it's that the only reaction is to claim we need more and more, larger and larger, studies of essentially the same sort to get the answers!  But, bigger isn't always better!

Wednesday, August 22, 2012

The exactitude of -omics

On Exactitude in Science
Jorge Luis Borges, Collected Fictions, translated by Andrew Hurley.

…In that Empire, the Art of Cartography attained such Perfection that the map of a single Province occupied the entirety of a City, and the map of the Empire, the entirety of a Province. In time, those Unconscionable Maps no longer satisfied, and the Cartographers Guilds struck a Map of the Empire whose size was that of the Empire, and which coincided point for point with it. The following Generations, who were not so fond of the Study of Cartography as their Forebears had been, saw that that vast Map was Useless, and not without some Pitilessness was it, that they delivered it up to the Inclemencies of Sun and Winters. In the Deserts of the West, still today, there are Tattered Ruins of that Map, inhabited by Animals and Beggars; in all the Land there is no other Relic of the Disciplines of Geography.
         —Suarez Miranda,Viajes de varones prudentes, Libro IV,Cap. XLV, Lerida, 1658
We'd like to suggest that Borges' short story can be aptly applied to the current state of disease prediction.  Fifteen years ago or so we were being told that once we had the human genome (HG) sequenced we'd be able to predict the diseases people were going to get, prevent them, and everyone would live to older ages than we'd ever attained before.  Aside from the questionable ethics of enabling such a demographic catastrophe, not to mention the idea that "everyone" would surely be an exclusive club, this promise is not much closer to realization now than in pre-HG days.

The first HG sequence, such as it was, was published in 2001.  Since then the promises have been honed a bit--ok, so the sequence itself wasn't going to bring us as close to immortality as we'd hoped, but the Common Disease Common Variant project would.  That was the theory that was used to justify the HapMap project, to provide resources to use case-control comparisons to find causal variants; then we'd have the data in hand for disease prediction and prevention.  That project was itself fine-tuned and scaled-up over the years, eventually bringing us genomewide association studies (GWAS) which, depending on who you ask, are either justifiably dead because they're mainly finding genes with very small effects, or alive and well because there have been some successful studies (macular degeneration studies are always cited) and if we just fine-tune the method some more it will really work.  And think what we'll be able to do with more whole genomes.

The -omics boom was being born.  This is the era of 'hypothesis free' approaches.  When we don't know the cause or can't develop useful actual hypotheses, our 'hypothesis' is just that some element in the realm we're searching has causal effects.  The genome was the first such realm, and the idea was that the trait had to have some genetic cause and if we blindly search the entire genome it must be there, and so we'll find it (or instead of 'it', some tractable few numbers of such causal sites).

Genomics was driven by increasing technology and was addictive, because, it is not too cynical to say, it was thought-free, meat-grinder, factory science.  It was lucrative, did indeed teach us a lot about what genes and genomes do, and found a modest number of important causal genes.  Its success, at least in the fashion and funding senses, understandably spawned other hypothesis-free blind technological approaches, cashing in on the cachet of the 'omics' word and its rejection of the need for actual prior hypotheses to design studies:  nutriomics, connectomics, metabolomics, microbiomics, immunomics, epigenomics, and more.  How much of this was because the same people who were promising us that successful disease prediction with genetics was right around the corner realized that this just wasn't true, and needed to figure out ways to keep their labs running we can't say, but we certainly are a fad-following, money-following research culture and we know this is part of the story. To be fair, when other approaches hadn't solved any of the problems, there was natural appeal to a thought-free, safely factory-like turn.  In any case, many of the same people who were gung-ho about genetics are now equally gung-ho about the promise of the -omics boom to bring us disease prediction and prevention that will really work this time.

The current interest in the -omics of supercentenarians in order to figure how they lived to their ripe old ages, and thus how we can live to 120 is, we think, an example of this misguided fad.  One basic assumption of this work is that every cause is individually identifiable, predictable and replicable.  This is in fact true for causes with large effects--Mendelian diseases, e.g., or point source infections like cholera or malaria and so on--but there are many paths to heart disease or stroke.  When everyone's genome is unique and causes many and variable, however, too often each combination of environmental and genetic factors will be extremely rare if not singular, and impossible to identify with current statistical-sampling based methods, the identification of rarely replicated events will be next to impossible.  The idea that every cause can be identified is a reductionist approach to disease akin to the reductionist approach to evolution, which requires every trait to have an adaptive reason to have evolved when in fact sometimes it's just chance. 

But, once we venture into the quest to find environmental factors that influence longevity, we're necessarily identifying these factors retroactively, if they are even identifiable, and yet none of us is going to live in the past.  Future environments are unpredictable.  So, again, unless a factor has large effects--heavy radiation exposure, infectious agents, toxins, e.g.--it's unlikely to be useful in predicting individual cases of disease.

We can see the issues by the proliferation of ever-more 'omics' approaches.  Each omics-community advocates its realm as if it is the, or at least the critical, one.  Essentially, we always add but rarely reduce, the number of potential causes of the traits in the lives of individuals.  This adds to the combinatorial realm--number of possible combinations of factors (and their intensity)--through which we must search.  More causes, inevitably individually rare, means that to show that a combination is causal it has to be seen enough times. That means ever larger samples because 'seen enough times' means to enable us to rule out chance as the explanation for the association between the combination of risks and the outcome.  But when there are more reasonably plausible combinations than grains of sand on the earth's beaches (this is no exaggeration--it's if anything an understatement), there aren't enough people to get such results.  And subsequent generations will have different people with different combinations of risk factors.

We certainly wouldn't argue with the idea that what we eventually succumb to is likely to be the result of multiple -omics, that is, a combination of factors.  But, we do question the idea that they will be identifiable, or useful in prediction, which is presumably the point of all this work.  The current interest in documenting every possible factor that might have an effect on health and longevity is bringing us closer and closer to Borges' map of the Empire. 

Friday, February 18, 2011

Indecent Exposure

We thought we were parodying what's already absurd enough the other day in our post about the atomic bomb being the first Too-Big-To-Fail mega research project.  Actually, nobody suspected that this was the hydra whose heads couldn't be cut off fast enough, but the future was latent in its name: the atOmic bomb project.

Hydra, from Wikimedia Commons
Our past cynicism notwithstanding, Nature tops all that this week, reporting something we should have foreseen but even we in our cynicism did not.  It's the 'exposome'.  Researchers are proposing to wire people up and photograph or otherwise measure every breath they take, every bite they eat, every chemical they come into contact with, collecting both 'external' and 'internal' exposomes, to "reveal the effects of diet, toxins and other exposures", in order to determine which exposures cause which diseases.  

It's a bit of a technological challenge but if there's anything we excel at it's overcoming technological challenges, so we assume the measuring devices will be made (funds will  now likely be diverted to years of R and D to do that), and study subjects will be convinced to wear, swallow, implant or otherwise port cameras, breathalyzers, air quality monitors and so on for months or years at a time, and of course donating blood, urine and fecal samples at specified intervals for genetic and exposure analysis.  Ready your every orifice!

The challenge of creating and maintaining the huge databases this kind of research will yield will be met, and statisticians will figure out new ways to analyze the data and it will all keep many people busy for many years to come.  So much better of a grant bonanza even than the old-fashioned biobank idea, trivially small by comparison.   

Exposomics. Brilliant! 

The tiny little problem with all this is that it's already very easy to predict the results. Let's not even consider the problem that yesterday's risks, being all that we can estimate, are an unreliable predictor of tomorrow's risks. Just as with GWAS, these studies will find effects, but they will be small and explain little and will not be useful for predicting disease.  The few strong ones will be hyped to death (more material for Nature to trumpet, naturally), but most of them will or could have been identified by less exotic means.

Risk factors with major effects are not difficult to identify.  Any Joe Blow, even without a micro-array, can do it.  Single-gene diseases or major environmental risk factors like smoking or lead paint or cholera are readily revealed by current genetic and epidemiological methods.  As the late curmudgeon David Horobin, founder of the non-conformist journal Medical Hypotheses, once wrote, if you can't detect something in small samples (we think he may have said 30), then it's not worth detecting.

That may be going too far, but we already know that when there are multiple factors at play, each with a small effect, be it genes or environmental risk factors, we move into an arena where small samples won't do: either our methods fail, or the answers are unhelpful in any clinical or public health sense, for the same reason: if they are too small to be detectable they are too small, and ephemeral or fickle to be that useful.  Even if we can detect them, which GWAS, biobanks (and, yes, Exposomics) will occasionally do, it is far from obvious that the cost is worth the game.

When risks are very small, as for example, in dental x-rays, and we know that but can't really estimate them, by far the cost-effective approach is simply to restrain use to situations when something that is important is at stake.  That may not be the 100% best-in-principle approach, but in practice it will save far more than it costs.  And the research money can go to providing important dental x-rays for those who can't otherwise afford them.

Omics-itis is bound to spread.  We expect our local deli soon to have a placard outside saying:

Here now!  PeanutButterAndJellySandwichOmics!
Exposomics?  Really, now!

Monday, February 14, 2011

Omics and the atomic bomb

As we noted last week, Nature and Science, and surely many others involved in the 'genomics revolution', are feting the 10th anniversary of the human genome sequence, Nature in its Feb 10 issue, and Science all month.  It's petty at this juncture to point out that 'the' human genome sequence is still not complete, and that the anniversary being celebrated was a date on which it was politically expedient for all those involved in the rancorous sequencing race to declare victory....and go back to work.

But we point these things out anyway, as a reminder of just how much hype the human genome sequencing project has been from the start.  The hype ain't going away anytime soon. In fact, it is leading to an epidemic medical condition, of geneticists needing expensive physical therapy to repair their shoulders that have been dislocated in the latest round of exuberant patting themselves on the back.

DNA sequence,
Wikimedia Commons
This is not at all to say that nothing good (in addition to business for the physical therapists) has come of this whole endeavor.  We know a lot more about gene structure and function and so on than we used to, and indeed genetic technology as represented symbolically by the human genome sequencing has led to many new discoveries across the spectrum of the life sciences.

But when even the best believers in the project say, when confronted with more biological data than ever before amassed in the history of science, that what we need is ..... more data, or that the promised immortality due to genome-based cures will be decades in the future, one has to wonder what we're dealing with.

In fact, the completion of the human genome was quickly followed by the birth of a whole new -omics infrastructure; proteomics, biomics, nutrigenomics, cistronomics, epigenomics, microbiomics, metabolomics, connectomics and on and on, a recognition, tacit or otherwise, that genomics just wasn't going to be enough.  Surprise surprise.  And to some extent we owe it all to the atomic bomb.

The Manhattan project showed that mega-science with huge, long-term funding would be an employment boon unlike anything since the gold rush.  After the catastrophic end to WWII, a foundation was set up in Hiroshima to study the effects of radiation exposures to survivors.  That was about 65 years ago, and the Radiation Effects Research Foundation is still going strong.  Mega-science became the strategy du jour and that view has been growing ever since.

It's a little known fact that the origin of the term Omics is from the Greek, meaning either "Too-Big-to-Kill" or "No-Need-to-Think."  That's because once you start down the 'omics' pathway, that is, of using technology to document absolutely everything in everybody, rather than to think about what you're doing carefully and justify doing something selectively--that is, once you become an omicist, you'll never think again.  The project will involve so much investment that your Senator will not allow NIH to cancel the project--a good deal if ever there was one!  Or, if you're doing science but want to keep up with the Joneses, once you see the guy in the next department doing it, you have to have your own omics project, too.  After all, fair's fair!

We've written numerous times before about the exaggerated (some would make strong arguments for 'knowingly false') promises of the human genome project, and it's true that this is not being entirely overlooked in this celebratory time.  But feting this overwhelming mass of data, that we've just barely begun to make sense of, by calling for more data, to be collected at huge expense even as the cost of sequencing single genomes has plummeted, before we work out what we can do with the data we now have, is a cynical abuse of public trust.  The last set of promises is far from being filled, and we're now supposed to trust researchers with more money to answer the same questions they couldn't answer last time?

We have, say, 10 good candidate genes for effects on some disease, be it psoriasis or diabetes, and yet we continue to map, map, map to find even more genes--hundreds of them--that make ever more trivial contributions.  Why not stop spending on these larger-scale studies, and figure out what the reliably known genes, that may really do something, are doing and how to develop therapies and the like?  Of course some investigators are doing that, but more funds could be diverted to real problems if we but had the will....and hadn't set up so many Too-Big-to-Kill omics endeavors.

At least as serious is that scaling up means more money and longer-term research comfort, which is always easier than trying to think out serious problems to find more creative ways to understand them.  That is the situation we're in now.  New ideas come from the combination of data, genius, luck--and the struggle against a conceptually challenging problem.   Megafunding and megaprojects undermine the last, and most important of these, because they institutionalize science.  Many, including Darwin and Einstein and others of their stature, remarked that they couldn't have done what they did in the stultifying environment of universities, for example, and that was long before universities became the way they are today. 

The challenging problem is not a secret, and yet too many geneticists continue to dance around it -- complex traits are complex.  This is the single most consistent finding to come out of the last several decades of genetic research, including as we've also noted numerous times, from genomewide association studies (GWAS).  It's not a surprise, it's not a secret, and it's actually a positive finding, though too often ignored.  It is not that business as usual has brought no new findings, but the miniaturization of findings, the loss of focus on the really general, central issues we think is the problem.

Wednesday, December 29, 2010

Boondoggle-omics, or the end of Enlightenment science?

Mega-omics!
We're in the marketing age, make no mistake.  In life science it's the Age of Omni-omics.  Instead of innovation, which both capitalism and science are supposed to exemplify, we are in the age of relentless aping.  Now, since genetics became genomics with the largesse of the Human Genome  Project, we've been awash in 'omics':  proteomics, exomics, nutriomics, and the like.  The Omicists knew a good thing when they saw it:  huge mega-science budgets justified with omic-scale rhetoric.  But you ain't seen nothing yet!

Now, according to a story in the NY Times, we have the Human Connectome Project.  This is the audacious, or is it bodaceous, and certainly grandiose grab for funds that will attempt to visualize and hence computerize the entire wiring system of the brain.  Well, of some aspect of some brains, that is, of a set of lab mouse brains.  The idea is to use high resolution microscopy to record every brain connection.

This is technophilia beyond anything seen in the literature of mythical love, more than Paris for Helen by far.  The work is a consortium so that there will be different mice being scanned, and these will be inbred lab mice, and all that goes with their at least partial artificiality.  The idea that this, orders of magnitude greater complexity than genomes, will be of use is doubted even by some of the scientists involved....though of course they highly tout their megabucks project--who wouldn't?!

Eat your heart out, li'l mouse!
One might make satire of the cute coarseness of the scientists who, having opened up a living (but hopefully anesthetized) mouse, to perfuse its heart with chemicals to prepare the brain for later sectioning and imaging, occasionally munch on mouse chow as they do it.  Murine Doritos!  Apparently as long as the mouse is knocked out you can do what you want with them (I wonder if anyone argues about whether mice feel pain, as we now are forced to acknowledge that fish do?).

This project is easy to criticize in an era with high unemployment, people being tossed out of their  homes, undermining of welfare for those who need it, and in the health field itself.....well, you  already know the state of health care in this country.  But no matter, this fundamental science will some day, perhaps, maybe help out some well-off patrons who get neurological disease.

On the other hand, it's going to happen, and you're going to pay for it, so could there be something deeper afoot, something with significant implications beyond the welfare of a few university labs?

But what more than Baloney-omics might this mean?
The Enlightenment period that began in Europe in the 18th century, building on international shipping and trade, on various practical inventions, and on the scientific transformations due to people like Galileo and Newton, Descartes and Bacon, and others, ushered in the idea that empiricism rather than Deep Thought was the way to understand the world.  Deep Thought had been, in a sense, the modus operandi of science since classical Greek thought had established itself in our Western tradition.

The Enlightenment changed that: to know the world you had to make empirical observation, and some criteria for that were established: there were, indeed, natural laws of the physical universe, but they had to be understood not in ideal terms, but by the messiness of observational and experimental data.  A major criterion for grasping a law of nature was to isolate variables and repeatedly observe them under controlled conditions.  Empirical induction of this kind would lead to generalization, but this required highly specific hypotheses to be tested, what has since that time come to be called 'the scientific method'.   It has been de rigeuer for science, including life science, ever since.  But is that changing as a result of technology, the industrialization of science, and the Megabucks Megamethod?

If complexity on the scale of things we are now addressing is what our culture's focus has become, then perhaps a switch to this kind of science reflects a recognition that reductionism is not working the way it did for the couple of centuries after its Enlightenment launching.  Integrating many factors that can each vary, into coherent or 'emergent' wholes, may not be an effective approach, and enumerating the factors may not yield a satisfactory understanding.  Something more synthetic is needed, something that involves reductionistic concepts that the world is assembled from fundamental entities--atoms, functional genomic units, neural connections--but that to understand it we must somehow view it from 'above' the level of those units.  This certainly seems to be the case, as many of our posts (rants?) on MT have tried to show.  Perhaps the Omics Age is the de facto response, even a kind of conceptual shift that will profoundly change the nature of human approach to knowledge.

The Connectome project has, naturally, a flashy web site and is named 'human' presumably because that is how you hype it, make it seem like irresistible Disney entertainment, and get NIH to pay for it. But  the vague ultimate goal and the necessity for making it a mega-Project may be yet another canary in the mine, an indicator that, informally and even sometimes formally,  we are walking away from the scientific method, away from specific hypotheses, to a different kind of agnostic methodology:  we acknowledge that we don't know what's going on but, because we can now aim to study everything-at-once, the preferred approach is to let the truth--whatever form it takes, and whether we can call it 'laws', emerge on its own.

If that's what's happening, it will be a profound change in the culture of human knowledge, that has crept subtly into Western thought.