Showing posts with label human genome. Show all posts
Showing posts with label human genome. Show all posts

Thursday, March 7, 2013

What is genetic function? The ENCODE non-questions

The human genome is 3.1 billion nucleotides long.  If only 1-2% of it codes for proteins, what does the rest do?  And how do we figure that out?  One way is to commit multi-millions of dollars to a project dedicated to doing just that, and that's exactly what's been done by the ENCODE project, which recently published, with much fanfare, the results of years of work by a consortium of over 400 people and numerous labs. The comment by the ENCODE PR spokespeople that got the most attention during all the hoopla when the papers were published was the idea that while  98-99% of the genome was once called 'junk DNA', now it looks like 80% of the genome is in fact functional.

There is a heated, and indeed vitriolic debate about how misrepresentative and even highly wasteful, the ENCODE Megaproject was.  ENCODE is a cute acronym (we need that in science, after all, since much of what we're about is marketing, so we need a brand or trade-mark) for ENCyclopedia Of DNA Elements.  In a nutshell, the project was a large consortium whose objective was to identify as much of the functional elements in genomes as possible.

The interaction of tRNA and mRNA in protein synthesis:
Wikipedia
We know that DNA codes for protein, but that is only about 1-3% of the genome.  Another small fraction is transcribed into a variety of RNA molecules that do things on their own (that is, they don't just get translated into protein).  Examples are transferRNA, ribosomalRNA, and various others. Some of this, called microRNA, is used to affect gene usage, by interfering with messenger RNA and hence protein production.  Protein coding and these RNA other processes and interactions are what in our Mermaid's Tale book we called 'correspondence' codes in the genome, because DNA contains the code for -- corresponds to -- the RNA which is then used elsewhere in the cell. 

Then there are bits of DNA that contain what we called 'recognition' codes.  The DNA sequence is directly recognized by other molecules, such as proteins called transcription factors, that physically bind to the DNA sequence elements and, among other things, cause nearby protein-coding genes to be transcribed into messengerRNA.  These are codes, but they act locally on the DNA itself.

The middle and ends of chromosomes (centromeres and telomeres) contain DNA sequences used for protecting the integrity of the DNA molecule in the potentially hostile chemical environment of the cell, or in the process by which chromosomes are copied when the cell divides.  There are other codes of various kinds in DNA that affect how it is wrapped around proteins so it can fit into the nucleus and so on.

But various studies had shown that much or even most DNA is actually transcribed into RNA molecules of unknown (if any) function.  It is replicable--so not just chance or experimental trash.  Since the function isn't known, it is debatable whether this is truly 'functional' or not.

Meanwhile, 40% or even more of our DNA consists of repeat elements, short sequences that are found scattered all over the genome, and (among other ways) are copied from one location and inserted more or less randomly in some other location.  These relate to various processes, including errors in DNA replication (e.g., microsatellites) or use some viral-related mechanism on rare occasions, but enough over evolutionary time to proliferate in the hundreds of thousands.

By some accounts, especially since it seems to be transcribed into RNA, much or even most of the genome is 'functional'.  Such claims challenge well-established ideas that most of the genome has very little function--what was called 'junk' DNA--and that therefore only a small fraction really matters.

But what is function?
A lively, funny, but quite sharp--some would say vicious--attack on the excited reports of ENCODE by Dan Graur and colleagues was published recently.  First, even though the ENCODE authors, being good scientists, put lots of caveats in the original papers, they were not averse to the super-hyping given the report by the media.  Instead of saying that ENCODE had provided a very useful and accessible  data resource and some thought-provoking data, the usual hype about transformative new findings, mysteries uncovered, etc. was all over the media last year.  Graur et al. blasted such reportage as culpable, or even scientifically naive hype (or, perhaps, bovine droppings).  Indeed, the aspects of genome structure and use that were reported by the project were all to some extent or other already well-known, even if ENCODE provides a more systematic data resource and coverage of them than had been available before.

The controversy involves many different issues, some of them quite technical and methodological, but the core centered around ideas of 'function'.  The project investigators used various methods to find biochemical activity of different kinds to identify aspects of the genome that were functional by that standard.  Thus, for example, if a transcription factor protein stuck to a particular bit of DNA, that was activity and classified as function; it didn't have to be shown to affect a protein-coding gene's expression level.

From an evolutionary point of view, function only matters if it affects reproductive success--or 'fitness' in the Darwinian sense related to natural selection.  Why is this?  It's because if it doesn't affect fitness, then mutations will eventually disrupt the activity but with no loss to the organism's reproduction.  The bit of DNA will, over time, accumulate variation among individuals and between species.  By contrast, a bit of DNA that does have a fitness effect will have much less variation in the population, because mutational disruption will harm the individual, who won't reproduce, taking the variation out with it.  We say that relatively limited variation, or sequence conservation among or within species, indicates evolutionarily important function.  Indeed, even if the bit of DNA did have some function that affected a trait, say body shape, but not in a way that would be screened by natural selection--that is, not in a way that affected fitness--that function would sooner or later be erased by mutation.

In that sense the function might be real but evolutionary unimportant or irrelevant. The idea that one could have function but not be affected by mutation in this way is tantamount, the critics argued, to saying that organized structures could arise just by chance, without being molded by natural selection.  That is hard to justify (actually, there may be such reasons, but they're too much to go into  here).  But it's worth noting that Graur et al. do point out that such function could, under some circumstances, become relevant to natural selection, so that even highly variable bits of DNA may not be unrelated to evolutionary potential.  But looking at it at any given time can't tell you that, and doesn't warrant assigning function in the evolutionary sense to it.

Wasted electrons--debates over angels on pin-heads
This is a debate about many things, but in part centers around orthodoxy.   The discussion is over the question  "What fraction of the human genome is actually 'functional' in these latter senses?"  10%? 80%?  17.654382234887%? 

This is a thoroughly electron-wasting debate (using up electrons via the internet and airwaves), because we know very well, beyond any serious doubt, that the usefully used parts of the genome vary from person to person and, indeed, from cell to cell within each of us!  And if we take a broader evolutionary view, different parts and regions and fractions of genomes will be used over time. 

Among individuals in a species at any given time there are hundreds of dead or partly dead genes, regulatory regions with variable strength transcription factor binding, and so on, all across the genome.  These vary from person to person, as we have clear evidence to prove.  And, have you forgotten the hoopla over copy number variation, the hot recently-new finding that our numbers of genes and other parts of our genomes vary among by the thousands us and between the two copies of the genome each of us carries?

Since everyone differs, it is almost impossible in principle to ask this question of any single individual, because there just isn't enough information and unique observations that can't be tested with the statistical approaches needed to document it (needed for reasons that are not controversial).  Alternatively, we might come to some sort of average functional fraction for a species, but that is rather vague and perhaps misleading--misleading about how DNA functions.  For example, it's been estimated that a high fraction of our 'real' genes (protein-coding or regulatory regions) are individually dispensable if other well-working genes cover the same function.  In that sense, a high fraction even of those 'real' genes are dispensable. 

Whether something has a fitness effect is also a statistical question, since there is always a probabilistic aspect to reproductive success.  The amount of conservation in a DNA region is in principle an indicator of past history of natural selection, but it also involves other factors (population size, mutation rate, and so on), and it is inherently a relative measure.  Assessing what varies enough to be judged not to have a fitness-related function is also a statistical issue, not one with precise criteria.  If the relatively limited variation in protein-coding regions reflect real function, how much more variable reflects no, or 'less' function?  One might say that the evolutionary definition of function, based on more relative variation, is not entirely free of the subjectivity issues that plague the ENCODE definition of fitness based on having some biochemical activity.

This wastes trillions of electrons, because it stimulates the media streams of hot air, capitalizing on the flap, even though the issues themselves are debates over non-questions or even subjective issues, as we have tried to suggest here, that are not clear cut--and of course there is the strong vested interest of the investigators vigorously to defend the over-selling of yet another over-priced mega-project, so its funding won't be cut.  As usual, this electron stream misses much of the interesting and actually scientific aspect of the findings and their ambiguities.

For example, the issues rest on the tacit idea that genome functions can be enumerated at the nucleotide sequence level by essentially assuming that each function is independent of other functions, which is purely a fiction.  But interdependence makes these kinds of issues, that are based on differing criteria of 'function' and relative variation, very tricky.   And the evolutionary argument essentially assumes that non-conservation means no function which is also a misperception of the dynamic control and complexity of genetic mechanisms and evolutionary adaptation.  This is not the place for me to outline my view on that, but I do think that a proper understanding of genomes and their evolution can answer the perceived differences of point of view in the current food fight.

Of course, thinking seriously about evolution is harder, won't please Big Story seeking journalists, and makes less dramatic material for grant applications.  So the food fight is not at all surprising.

Tuesday, August 7, 2012

Ooops! The human genome does not exist! Part V. Why are we so comfortable with type representations?

On reference sequences and various loose ends
We started this series to explore some conceptual complexities and subtleties associated with the universally known, but perhaps less universally understood notion of 'the' human genome.  What seemed like a hard-won and very straightforward entity for human genetics turns out to be a lot less straightforward--even if it's exceedingly useful in many ways.

A type specimen in taxonomy can be a single individual.  If 'the' human genome sequence  in a future iteration is a diploid sequence from one person, it will be similar, as we noted in a previous part of this series.
Helixanthera schizocalyx type specimen
Kew Gardens
We have described the many ways in which such a sequence actually doesn't or never did exist.  In part this is because there will always be errors in the data.  But let's overlook all of that, and agree that the human genome as we know it is a representative or reference sequence--at least until we can get a single representation that digests all known variation in human genomes.  We suggested some ways that might be done, using what are known as gene 'logos' (visually reflecting the underlying data: the relative frequencies of each nucleotide at each position in the genome).

It is important to be aware that even as a reference, what it is a reference of is unclear.  In prior posts of this series we discussed some of the issues.  One general idea is that it's the sequence of the DNA from, at least, a 'normal' person.  But as we noted that is misleading.  The donor--when or if 'the' reference is from one person--is not in any way a particularly superior, nor even a particularly average, person.

A subtle point that seems very widely misperceived is that there is no 'normal' person.  Nobody is at the exact average value for glucose levels, stature, blood pressure, and so on.  Nobody has the 'average' DNA sequence (whatever that means).  No person has ever died at the exact life expectancy in his or her population!  An average, a 'norm', is a digest of variation. It is a Platonic concept.

If 37% of individuals in our species have an A, and 63% have a T at a given position in the genome, no person can be average.  The closest s/he can get is 50% A, 50% T since we each (if we're 'normal'!) have two copies of the genome, so a single person's allele frequencies can only be 0/1 or 50-50.  So it's hopeless even to think of a digest such as 'average' in this context.  The nucleotide frequencies at any give site will change, anyway, as samples increase (but we'll not know by how much, since people are leaving and entering the gene pool at every moment).  And, as a reminder, there is the further problem that the genome instances in each cell of an individual differ: 'the' genome sequence obtained from that individual alone is a reference, not a true description of his/her genomic contents.

So we are not really using a reference sequence to establish an ideal.  Yet, unless to make a point of something like political correctness, we would not use a person with, say, Tay Sachs, or who was congenitally blind, nor an indigenous Australian or an Inuit, as our reference.  This shows in yet another way that terms like 'reference' or 'representative' are potentially quite misleading, or even political.

And yet, an arbitrary reference is very useful!  Why is that?

'The' wildebeest; 'the' mongongo nut': Deep aspects of the mind?
Mongongo nuts
Whatever we eat, or whatever wants to eat us, we must recognize and digest its presence from sensory information--and sometimes we'd better do it fast or we'll be somebody's fast food!  We recognize our prey, and our predators, our dinner and our diners, by analyzing our sensory input and fitting it into categories.

However this works neurally, our ancestors knew what a wildebeest is, or a mongongo nut, even if no two of them are identical.  We have always created mental Platonic ideals and within rather broad limits we instantly fit sensory input into these categories and act on the result.

We use language--terms like 'the chicken' or 'the human genome', but the biological process involved in understanding the concept doesn't depend on language, and it certainly isn't anything particularly human!  Lions recognize wildebeests, bees recognize flowers, and ants recognize enemies.  The wiring process that leads to the ability quickly to characterize is long built-in to neural systems, however they work.

Using type specimens, including whatever we determine to use to represent genomes, is thus a manifestation of a deep mental characteristic.  Where we have trouble is at the edges of our categories, in the range of objects that we include, and this is true not just of taxonomy, or genomics, but in all of life.

Even when it comes to disease and gene mapping!
It is thus no surprise that in human genetics we  hunger for 'the' diabetes gene, or a gene 'for' diabetes--and have a definition of 'diabetes'.  That is an underlying reason for the fervor for enumerative approach to genome mapping of traits.  Categorization can be useful, since categories of disease allow us to develop protocols consisting of categories of treatments, tests, and identification.  For deep as well as pragmatic reasons, we are not easily accepting of complexity: we like to throw the word around perhaps, if it sounds fashionable and insightful, but in fact we do our damndest to reduce quantitative complexity to qualitative simplicity.  We simply can't stop wildebeesting.

Evolutionary typology, too
 We've stressed the central relevance of evolutionary population thinking that should be an omnipresent counterweight to typological thinking.  But we even do the very same Platonic thing in evolutionary biology, no matter how deeply we understand its quantitative nature.  We want there to be 'the' adaptive reason for a trait that we are interested in.  The adaptive 'fitness' of a trait is viewed in effect as an inherent, cosmically existant reality of which each individual's reproductive success is an imperfect reflection.  Indeed, often we do the same thing in defining 'the' trait in the first place: we want 'the' selective explanation for 'bipedalism', for example, as if that were a single thing and evolved as such.

Material Platonism
Categorical thinking is very hard to avoid, and population thinking devilishly hard to ingrain.  A new paper in PLoS One that a colleague, Charlie Sing at the University of Michigan, pointed out to us, looks at the current data from the '1000 Genomes' project, the effort to provide a large number of whole genome sequences, raises--or rather, demonstrates--problems that will be faced by the desire (or dream) of predicting disease from individual genomes.  The structure observed among known variants is used to identify causal variation by various mapping or association tests, as we've commented on many times (e.g., GWAS studies).  But while some variants identified in what might be called disease-specific 'reference' data will have strong effects, this reference even though based on many rather than one sequence, will not include many other relevant variants. 

How we should deal with rare variants is currently a hot topic in many labs, but whether one should be thinking in terms of such single-cause prediction is--or, rather, should be--at best an open question.  No composite of variation will be a perfect genome referent for the moving target that is a species.

Since the real world is at least as quantitative as it is qualitative, categorical type-specimen thinking may sometimes be helpful but also can lead to wasteful effort or even to harmful results.  Even war between 'us' and 'them' is an instance.  It's hard to be careful about how and when to use 'types' even as referents.  In a sense, the macro world has duality the way the micro world does in physics, where particles come in types, and 'the' electron is sometimes a particle, sometimes a wave....something, and yet not something.

Typological thinking is very deeply engrained.  For example, in what is almost laughably misleading yet ineradicably pervasive, almost every geneticist refers to 'mutations' when they mean 'alleles': mutation is a change due to a DNA 'mistake' between one cell and another, as between parent and offspring.  Every element of every sequence on earth today was, at one point, a 'mutation'.  'Allele' properly refers not to a change but to variant states present at any given time.  But we exceptionalize variants in this way by calling them mutations if, for example, they lead to a disease or trait under study.  This arises, conceptually, only because we think there is a 'normal'--a type--from which the variant is a 'mutant' (a world often with denigratory connotations). Yet we know that's at best inaccurate.

Worse, almost everyone in biology unstoppably refers to the 'wild type'--the wild type--when comparing to some variant ('mutant').  This has become ridiculously entrenched, because for reasons we've been discussing, there is no single type.  Even in  mouse research it is absolutely routine to refer to a strain, say C57BL/6 mouse as the 'wild type' relative to an experimental manipulation such as inducing a transgenic 'mutation' into one of those mice.  But there is absolutely nothing 'wild' about a C57BL/6 mouse--it is a laboratory produced inbred strain that couldn't last the proverbial 5 minutes out in the 'wild'.  It, like all inbred strains, have been very un-naturally produced.

So 'wild type' is a persistent, historically derived but obsolete term, a perversion even of the original type-specimen concept.  In daily practice, of course, this is jargon and geneticists know what someone is talking about when contrasting a wild type with a mutant.  But the habit reflects psychologically deep patterns of thinking that are far from what the real world is like, and we have tried to suggest, can be misleading, sometimes when it's important not to mislead.

A simple idea like 'the' human genome intuitively seems to exist, and yet on close inspection doesn't exist.  In a sense, one can say that humans are built for material Platonism.  The ideal, the abstract entity, is a figment of our individual imaginations, but our power of imagination was built by evolution.  Using ideals, based on material objects, comes natural to us.  We depend on it.

Plato's Allegory of the Cave by Jan Saenredam, 
according to Cornelis van Haarlem, 1604, Albertina, Vienna.
In a deep sense, Plato had things exactly backward.  He used the image of people in a cave, who were only able to see imperfect shadows of the real 'ideal' objects projected by firelight on the cave wall.  But that's backwards.  Ideals do exist, but not 'out there' in some Platonic immaterial realm unobservable by us.  They exist in our heads, and are material in that sense and because we each build ours from real-world observations.  It is thus the ideal that is imperfect relative to reality, not the other way round as Plato had it.  If there is anything mysterious about this, it is that we were designed by evolution to do things that way rather than to be better quantitative thinkers.

In the end, 'the' human genome sequence is imaginable, but doesn't exist, no matter how we might contrive to produce something less arbitrary than the essentially fabricated referent we now use.  The challenge is to understand the limits, as well as the strengths, of categories.

Monday, August 6, 2012

Ooops! The human genome does not exist! Part IV. What do we want to represent by 'a' genome sequence?

In the first post in this series, we said that the current Human Genome sequence is a kind of Platonic stereotype of something that does not actually exist.  This is true at least in the sense of its incompleteness, but considering it raises several issues--and by new things that are afoot in the human genomics community.

Any human varies at about 1 nucleotide per thousand of the simplest parts of 'the' genome, and much more than that for large subsets of it (e.g., 'repeat' regions, copy number variants, and others).  And we're diploid.  So listing only a single nucleotide per location is clearly misleading about any human person.  If the genome reference sequence is a composite of several donors, then of course nobody has the sequence.  How many people are represented has been shrouded in secrecy, to avoid misuse of the data, racist interpretations, and so on.  Even if it's only one person, a haploid version may not be accurate for the technical reason that in parts that vary, all variants on the same copy should be listed, whereas in the past my understanding is that when the individual was heterozygous (had a different nucleotide on each of his/her copies), one of them was arbitrarily chosen.  If we keep in mind that this is a reference sequence, a Platonic representation rather than a real sequence, this may be OK, but it would not be any actual sequence, even that of the donor/s.

Completely real salad.
What if, whoever or however many people were donors, sequence elements come from different parts of the world--different continental ancestry?  It is clearly a strange kind of biology to assemble bits and pieces from different places in that way. Even if from one person with mixed geographic ancestry, the person is real but the genome sequence would be difficult to interpret beyond identifying continents of origin (where that could even be done accurately).  That would be like taking a salad to represent vegetables, even if the salad is completely real.

Now, suppose it turns out that after all the DNA in the reference sequence was in fact entirely from one person and not 8 or so as have been many statements in the past.  And suppose that resolving difficult-to-sequence regions were not in any way complemented with known elements from other individuals.  Then, we can post the diploid sequence of that individual (and that is in the plans) for the same DNA that has currently been used as 'the' human reference sequence.  Once we have that diploid sequence, then except that it will be inevitably incomplete for technical sequencing reasons, it will at least be a sequence that actually exists!
Stuffed robin

That would seem like progress, but it really only gets us closer to the stuffed robin type specimen in a museum near you.  No one person represents our species except if we understand, clearly and consistently, that it is not 'the' species, or 'normal', etc. in any global sense.  Or, not in a way other than as an arbitrary reference, the way a map of  Pennsylvania represents where I-80 passes by State College.  If so, by what criterion and who decides?

In that sense, the diploid sequence isn't literally Platonic, but that is really almost a trivial technical point, since nobody else has the same sequence.  These may be philosophical questions, but at least 'the' reference person is a reference!  Or is it?

In fact this is not so clear.  The instances of the genome in every cell in every individual are different.  The fertilized egg by which your life began had 2 instances of a human genome.  But each subsequent cell, of which you are comprised of many many billions if not trillions, contains new mutations that arose when it was produced by a 'parent' cell dividing.  Mutations arising in the very first cell division will be shared by half your body, while the inherited instance will be in the other half.  But each subsequent time you make new cells even these split-genome personalities acquire new variation, that is faithfully inherited thereafter within your body (unless altered by subsequent mutation).

A reference sequence is obtained, basically, by taking a blood or other tissue sample, extracting the DNA, and cloning it, then sequencing the clone.  So 'your' sequence, and 'the' sequence is really only a composite of bits of DNA from one, or each bit from a different, cell.  We're back to square one: 'the' HG sequence, even a diploid one, really is essentially a Platonic ideal that doesn't exist and never did (or even if it did, will some day not, when you or the donor of 'the' HG sequence passes away).

It is very difficult to escape these problems.

Now, we suggested earlier the use of a 'gene logo' approach, that would represent variation in each position along the genome, as seen so far in the available sequences.  There are problems involved in rearranged parts of the genome, which we all have some of.  But forgetting those, one could make a frequency-based representation that would at least represent humans as a population, which is how we evolved.  We just use the available sequences and for each nucleotide position represent the observed variation.  Great!   Or.....not so.

The available samples of human genome sequences are not from a random sample of our species--whatever that would be.  They are from a few populations, selected for various usually arbitrary, political, or convenience reasons.  I know from direct experience that deep arguments were held to decide how to sample human genetic variation.  Haste and expedience (and a bit of cultural elitism and politics) ruled the day.  So, while we do have samples from various populations, they are not any sort of random sampling of our species.

Still, that's far better than nothing and, indeed, a major US genome browser (UCSC) that displays genome information (and from which you can download the data), does have an optional 'track' that that shows variation in a way similar to a gene logo, for some available population samples.  Here is the 'genome variants' track from 'the' human genome, in a region of the beta-globin gene, showing essentially the same kind of frequency portrayal of each site that varies in this region, in the available samples:


This is itself currently only tentative, and is just for some selected individuals rather than populations or over all known sequences, but it begins to make clear which sites vary along the genome at least in this very limited data (each vertical bar is split into proportions reflecting the variant nucleotides' frequencies--here, at most into two halves, since these are individuals who have just 1 or 2 alleles, AA, AT, or TT, say).  The legend above each track identifies the individual portrayed.  Note something, however: that the same sites often vary in multiple individuals.   One can dig around in the genome browser to identify details and get known population frequencies, but the data are not yet in a nice single logo-like visual track.

Such a track, for the '1000 Genomes' project, is under development here at Penn State as part of the various collaborative genome consortia, but not yet publicly available.  A separate frequency bar will be available for each population in that sample in the way individual data are in the above figure.  I understand that it, or some other similar visualization tracks will soon be available. There already are databases that portray allelic variation as little site-specific pie charts, too. 

This is progress, but these are ad hoc examples, and even if put into a single logo-like genome track pooling all available samples, will not be a completely good representations of human genetic variation, because the samples are a rather unsystematic representation.  That's in part just a temporary problem because new genome sequence data will rapidly accumulate, not just in population samples but also from countless unrelated individuals.  So perhaps the day is coming when we really could have a single logo-like presentation to give a more proper gestalt of 'the' human genome--if it is used as the reference representation.  It may be a long time before we have a 'representative' set of people, whatever that means, even if it will certainly be a  conceptual improvement, reminding everyone, not just specialists, of the population relevance of genomic variation.

But even this will still be stereotypical (even if every living human were included!) for the subtle reason that we discussed above, that each individual represented is only being sequenced in some way that essentially gets information from a single cell in that individual, even though the sequence in each cell differs.  No individual cell is much better than a Platonic ideal even of that one person's genomic variation! 

No sequenced individuals will be random or in obvious ways representative of the human species.  Indeed, how should or could we get such sample(s)?  What is the human species?  Thousands are born or die every hour, so the collection of human genomes constantly changes.  Representing genetic variation in a species is akin to representing a river with a snapshot of the water and waves (and fish) that happen to be there at the time.

Even biomedically, variation that's relevant changes all the time.  Depending on whether you are involved in public health (whole population level epidemiology), or a local isolated group, or a family, you might represent genetic variation differently.

In the end, we are perhaps stuck with Plato:  as limited human beings, we cannot grasp everything in our heads, and representations and reference guidelines are immensely useful.  Perhaps only at the edges of variation, or with exceptional cases, does the use of type specimens lead to problems, if we know what we're doing when we choose how to represent the variation.

There have been centuries of debate about the nature of Platonic ideas (even whether they somehow exist).  This debate occurs today even in mathematics and cosmology: do the rules and findings of mathematics somehow exist in a universal Platonic sense, or is that just how we can describe our own particular universe?  Is it anti-scientific mysticism to think of abstractions (like cosines, algebraic solutions, stuffed robins, or genome sequences) as having some sort of non-physical reality?

These are interesting and I would say profoundly important points.  Even if we can't understand the 'truth' of such ideals, they prove to be immensely useful and important.  For those familiar with mathematics, perhaps an analogy would be i, the 'imaginary' number equal to the non-existent square root of -1.  It doesn't exist, in the real world in the usual sense, but is fundamental to physics and how we understand the universe and much besides that.

Representations are fundamental to science.  The danger is when we don't understand them clearly and they become misrepresentations.  We'll have a few more finishing thoughts about this subject tomorrow.

Friday, August 3, 2012

Ooops! The human genome does not exist! Part III: a non-Platonic type specimen?

In the last couple of posts (here and here) we have been discussing the idea of 'type specimens', as that idea relates to genome sequences representative, in some way, of a species.  Stuffed robins and dodos seem to raise few scientific questions about the species they represent as type specimens, even though we know that each species varies.

Of course, we're engaging in a bit of tongue in cheek poetic license here, because there are lots of ways to view the notion of a type specimen (see Wikipedia on this, e.g.), and of course, every biologist knows that species vary.  The point here is that it is one or a few examples chosen in some way to represent a species, assuming adequate evidence that a type specimen does represent a proper species (such discussions are off-topic here).  People in genetics should be clear about what a reference sequence, an analogy to a morphological type specimen, is.  Not all of them are.

In the genomic age we now have whole genome sequences from one or more representative individuals, as DNA sequence type specimens of 'the' genome for that species.  Of prime interest of course is 'the' human genome sequence (HG), and we've pointed out that this must be viewed not as a real sequence carried by any individual, but as a sort of Platonic ideal: a reference or road map to the sequences we carry around but none of which is identical to--regardless of whether it does or ever did exist in any individual.  The latter is the central point.

Dodo

The reference sequence keeps changing as errors are found or problematic areas done again, to be clearer or sure about the sequence, yet it seems not to lose its status as 'the' reference. And now we have many sequences from different species--soon to be thousands for humans alone.  Since each differs, many in major ways but all in millions of minor ways, from any edition of the HG reference, we wonder if a single reference (or reference set) is the best way to characterize 'our' genome.  Since evolution works via populations, why not use a population-based way to represent our genome?

We suggested yesterday the idea of using a visual 'gene logo' to represent the variation found in existing sequences on file, augmenting that reference as more data become available. We show the example again here.  This is a 'population type' rather than a type specimen.  It would be 'the' referene instead of a single sequence, not as a supplement.  It shows the most common nucleotides as well as the variation, in every position along the genome.

This is not  a Platonic ideal, because it does not represent a single anchor sequence, but instead portrays genomic variation as it exists in available data--much more natural than a single reference and since genomes got here by evolution, and evolution is about variation,  a biologically more natural way to represent genomes.  There are problems, for example, that each instance of the HG has bits and pieces added, missing or rearranged, so not all the variation is in nucleotides in specific positions along the genome.  But ways could be found (there are, for example, graphical ways to show where segments of the genome have moved to and from).  And of course the underlying data are downloadable.  Still, could there be a better way?

An evolutionary Platonism?
Plato's idea of things like dogs or chairs or virtues is that there is an actual 'ideal' that really exists somehow but of which all actual ones we see are but imperfect instances.  Each new chair is an instance of 'chair', and one might liken this to being stamped out, never perfectly, by a chair-making machine.  Or perhaps every chair is a kind of real sibling produced by the same (ideal) parent.  In Plato's terms the parent may not even have physical existence.  That might give us an idea of how to handle biological types in a better, more physically real way.  The way is an evolutionary one.

Doll chair, The Children's
Museum of Indianapolis
The only reason we can even use a representative, reference HG sequence is that all humans are similar and the only reason that's true is that we evolved by descent with modification, Darwin's phrase, from common ancestry.  So could we represent a gene not by an arbitrary reference sequence of some person who happens to be in our lab, but in terms of ancestry?

If we look at large numbers of instances of a gene, we can make reasonable guesses as to what the common ancestral sequence was like.  If most people have an A in a given nucleotide position in the gene (even if, say, some have a T or G there), then A was likely the ancestral state of that position.   If we find that chimps also mainly have an A there, we can be even more confident.  Working site by site in this way, we can make an educated guess as to the ancestral sequence.  In fact, we do this all the time, and have routine, reasonably reliable software to do it on a large scale.

So why not use the ancestral sequence--which actually did exist, unlike some of the composite and hence artificial sequences we use today--as our reference?  Then, while nobody has the same exact sequence today, we can landmark what we do carry in reference to that 'ideal', a kind of real Platonism.

In many ways we do this in daily lab work and drawing evolutionary phylogenies.  We routinely use the idea of the common ancestral sequence, called the coalescent because all today's variation seems to 'coalesce' to the single ancestor as we look backward in time.  So the coalescent sequence of a gene could be our type specimen for that gene.

That seems like a sound way to 'Platonize' our genomic reference map without the arbitrary use of a randomly picked sequence today, or a sequence that has no material existence, since the coalescent is 'the' ancestor rather than just 'an' instance of the sequence.  But while this is evolutionary sensible, there are problems.

First, the ancestral gene may have been so different from today's that even if we reconstruct it accurately, using it as our reference would not work well from a functional, or perhaps even a structural point of view.

Second, while we can do this for a gene (even including inserted or deleted elements in the gene), what we need is the contextual arrangement of a whole genome sequence.  That means the sequence surrounding the genes (and that is about 95% of genomes!).  But when that is non-functional, or even when it has gene-regulating sequence elements, these can move around and evolve so fast that we often cannot easily guess their sequence.  Genes coding for proteins, are so highly conserved in sequence (because, evolutionarily, mutations were harmful and removed) that we can find and align them from other species.  But this doesn't work as well for surrounding sequence.

Third, the lack of such conservation in non-functional sequence might not matter so much if it weren't that the common ancestor of many genes existed long before there were humans and in some species often very distant from us--sometimes, even, an early vertebrate or even before that.  So the surrounding sequence can't even be guessed at.  And the ancestral gene sequence may have little interpretable relevance for human biology (or phylogeny).

Fourth, sequences are rearranged over evolutionary time, and the flanking bits aren't the same, so we can't really put together a very meaningful whole genome sequence (this has been attempted for mammals or vertebrates, but is very speculative and approximate, it never existed in an individual, and is essentially worthless as a reference to guide much of current work on genes and their functions).

Fifth, each gene's evolutionary-Platonic ancestral copy existed in a different individual, the ancestors for different genes living far apart and at very different times in history.  The animal (not even a person!) with 'the' original beta-globin gene would have been countless centuries and kilometers different from whoever had the ancestral blue-light receptor gene, and so on.  Unlike Platonic chairs, there may have been single ancestral genes, but not a single ancestral genome. 

So, while evolution is a better framework for discussing ideals than Plato's, because ancestors did actually exist, using ancestry to produce reference sequences--genome 'type specimens'--doesn't seem to lead to an obvious strategy.  This is strange, since a random instance of a genome sequence works as a reasonable reference because we're all so closely related--the same evolutionary reason that doesn't seem to work very well.

Our own view is that we should try some kind of variation-including method, such as the 'logo' approach above and that we discussed in Part II of this series, to represent our genomic 'type'.  It's just like using a single stuffed robin, or a single Neanderthal skeleton: imperfect, easily misleading, not actually 'representative' as a representative......but taking variation into account, perhaps the best we can do.

A few thoughts and updates about these topics will conclude this series, in Parts IV and V next week.

Thursday, August 2, 2012

Ooops! The human genome does not exist! Part II. What kind of reference is a reference?

We posted yesterday about the nature of type specimens, one individual (or sometimes a few more) chosen by scientists to represent an entire species.  How that choosing was done, how the type specimens were sampled, or what they actually represent is variable but always a bit (or more than a bit) arbitrary.

Nessie
When the age of genomics arrived, and we had large-scale sequencing, the Human Genome Project gave us something similar, 'the' Human Genome sequence (HG).  This, like previous type specimens--the stuffed carcases shipped back home by globe-trotting naturalists largely in the past 2 or so centuries--is an arbitrary instance of something we believe exists elsewhere in profusion.  But while the stuffed robin or Loch Ness Monster on display was an actual robin--or a grainy picture of an actual LNM--even this degree of specificity doesn't apply to the HG:  it has generally been said that it was not derived from a single individual, but 8 or so, whose identity or identities or ethnic origin etc.  are confidential.  In addition, it is under revision as errors are found and so on.

That means if you do a study today, for example, looking at some human genome region in a sample of people, you'll have to identify that region in terms of hg19, the current edition of the sequence. But if you want to compare your work to past work, or if those following up on your work want to compare to yours, they'll have to go back (or forward) to hg19.  Yesterday we noted that the starting position for the human beta-globin gene has shifted about 43,000 nucleotides between the current and previous versions of the genome.  This is not just due to error correction, but would happen if you compared any two instances of a human genome.

Thus not only do the sequence details of any given gene vary, but so does its location. The same is true for any other functional or even nonfunctional aspect of the genome.  The beta-globin or any other gene no more exists in reality than does 'the' HG itself.  What about other details? Just as every robin differs (we don't know about Nessies), so every instance of the HG differs.



Albino Robin, photo from www.seed-solutions.com
How many chromosomes do we have?  22 (two copies of each) plus X and Y?  Well not everybody has that number.  People with Down syndrome have 3 copies of chromosome 21.  Well, you may say that's abnormal, but are they not human?   In fact, every instance of a human genome varies in the number and location of substantial chunks of DNA (called 'copy number variation'), and there are shorter bits that vary in their number of copies in millions of different places across the genome.  There are tens or hundreds of defunct genes, different sets of them in each 'normal' person.  Further, of course, what is 'normal' is a largely arbitrary notion, based on what is average (the 'norm') or some range that we think of as 'normal' in some vague and subjective way.

We've discussed when or how we decide that your behavior is not 'normal' (e.g. here), and how disease is sometimes an arbitrary classification that we choose to make.  Are people 'abnormally' tall or short not human?  Is that a 'disease'? 

Somehow we need to decide how to categorize species--their traits and their DNA.  What would be a better way to deal with a Platonic ideal, a single kind of reference genomic sequence--an ideal  that seems to make sense, while at the same time there is no actual instance of that ideal?

Reference sets?
Rather than a continually revised type sequence, one possibility is to use some set of DNA sequences to determine what 'the' human genome is like.  But then what would be included in that set, and how many different sequences?  If we want a road map, strictly as a reference for comparing sequences, how would a set be different from a single type specimen?


If we had a set of sequences, they would somehow capture both the typical elements and their variability, and these are vital to understanding both gene function and evolution.  One way might be to use 'gene logos', such as in this figure:


Here, each position along a sequence is represented by a visual, color-coded, letter whose size is proportional to the the fraction of the  sample that contains an A, C, G, or T. So in the first position, every sequence had a C, but in position 7, G predominates.  The height of the letters sum to 1 (meaning 100%, unlike in this particular image).

This would provide a quantitative 'stereotype' of the human genome sequence that is not a rigid Platonic ideal, but would still be useful.  A given sequence, like the beta-globin gene, varies but a computer search could easily find it in any whole genome sequence (a search and compare method called BLAST does just that).

We could also include missing nucleotides by having a proportionally sized N (for None) in the location could be added.  Major copy variation or rearrangements would be harder to include, but probably ways could be found.

With computer technology, even a digest of all available sequences in a species could be stored in this way.  Since individuals will always die and be born, at any given time this would be a 'type species' rather than type or reference specimen, but that would perhaps be a much better way to think of things.

Then, it would not be necessary to use the concept of 'the human genome sequence' as a somewhat misleading guide--where we use something that doesn't exist to try to understand what does exist.  In fact, something like this approach is already available, though not standard, and will be discussed in Part IV of this series.

How we would do the same with robins (or Loch Ness monsters) is less clear.

Wednesday, August 1, 2012

Ooops! The human genome does not exist! Part I. The notion of a type specimen

It's sultry summer, and of course the science hype machine motors on unabated, but we want to ignore it briefly and relax at least a little bit before Penn State resumes it's serious academic activities--that is, football season.  Still, while you're basking in the sun getting your tan (and skin cancer), along with your mint juleps (and undoubtedly some other ugly diseases), here's something to think about:

The human genome doesn't exist
René Magritte: This is not a pipe.
It's a painting.  But a painting of 'a'
 pipe--but not  'the' pipe, which doesn't
 actually exist in the real world.  Or does it?
Despite many claims to the contrary, that The Human Genome project sequenced the human genome and thus set in motion the most exciting era of fundamental new scientific discovery since Galileo, it has turned out that the HG doesn't exist, after all.  It no more exists than does 'the chair' or 'the dog' as Plato once asserted.  He said we have dogs and chairs but they are only imperfect instances of the real true dog and chair.

Now what everybody raves on about, the HG, is like that.  It is not from one person, not even one copy from one person's two. It's not clear whether this is even from one person or from several donors, or if from one, who that person is in terms of his origins (his, because there's a Y chromosome in the sequence). 

In any case, it's not a 'normal' genome, because while we believe the sequenced individual(s) was/were healthy at the time they bled for the cause, they won't be healthy forever.  Since we are told every day by the press that every disease without exception must be genetic, and hence we should GWAS it endlessly,  the donors will eventually become genetically abnormal.

Worse, 'the' genome keeps changing!  We are now in version 19, released in 2009. It's either more details from the same donors, or now has bits that couldn't be sequenced easily from those donors' DNA and so was sequenced from additional donors (we don't know which is the case).  There will be future revisions.  Not only that, but it is a 'haploid' sequence, with only a single-nucleotide per position reference, whereas any human has two copies (except for X and Y chromosomes).  Yet in any one person 1% or more of sites actually vary.

This is not Van Gogh's
kitchen chair
Now, this means clearly that nobody actually has  'the' currently posted human genome sequence any more than your kitchen chair is 'the' chair.  Of course, what some biologists who are more savvy than most people who talk about it know is that the HG is a reference sequence.  Nobody has and nobody has ever had, this incomplete reference HG. All of us have some variants of that sequence.  It is wrong, but standard, to refer to what we each carry as 'copies' of the HG.  They are not copies!  There is no copy of the HG!  No one person ever had that sequence, so no person could copy (replicate) it to transmit to its offspring.

Probably, we should use a term like 'instance' rather than copy, in the same sense that each chair is an instance of the concept of chair, even if no Platonic actual ideal chair exists, except as a reference concept in our minds.

What is good for the goose is good for the gander.  If the HG doesn't actually exist, neither does a given gene say 'the' beta-globin gene, which is expressed in red blood cells and of which some variants are involved in anemia (like sickle cell anemia).  Here is hg19's version of a very small part of that sequence (click to see details):


Any one instance of this in an actual person, like you, may or may not have this exact sequence.  But unless it's been deleted from your genome, you'll have something very similar, not because it is a copy of some Platonic idea, but because they are descendant copies handed down through the generations since we shared a common ancestor, and that--evolution--is the crucial difference.  It's why we should be careful about treating a reference as being real, or how we relate the existing instances of the 'same' (homologous) sequence.

We use 'the' HG sequence as a way of referring to a standardized set of nucleotide locations along the chromosomes in the individuals who were sequenced.  Each of us varies from that sequence in millions of places (in each copy that we carry). To find similar functional elements--genes and so on--in my instances of the human genome, we use 'the' HG as a kind of map.  If I had your DNA sequence, I could identify which is 'the' beta-globin gene by its similarity to that in the figure, even if you didn't have the exact same sequence.

This is very useful, but it's important to be aware that not only does nobody have the same sequence, but even the structural details--the locations of the 'same' elements in you and us and the donors of 'the' HG--differ among people.  So the coordinates (the number at the top line) will not be the same.  For example, the beta-globin gene starts at chromosome 11 position 5,246, 696 in draft hg19, but at 5,203, 400 in draft hg18, a difference of around 43,000 nucleotides!  Where did they come from?  Which draft should we believe--if any?  What will hg20 say? What is 'the' HG?

Type specimens
American Robin
Think about this in another context.  For centuries we have had the concept of a type specimen.  There is 'the' robin, 'the' monarch butterfly (or even, perhaps, 'the' Loch Ness monster??).  'The' Neanderthal hominid is an arbitrary individual whose luck it was to die and not rot, so we could discover him around 30,000 years later in Germany.  The lucky old sod gets to represent all of his contemporaries and ancestors and descendants for hundreds of thousands of years.  And 'the' Neanderthal DNA sequence is partial, a composite of several individuals who lived hundreds of kilometers and thousands of years apart.  Think about that when you hear people pronouncing on how 'the' Neanderthal lived, or what color hair 'it' had, or whether it had religion!

Robin you'll see in the UK (image: RSPB)
The international organization of systematists--biologists who are concerned with cataloging, characterizing, and name species and their relationships--has decided what counts as a reference specimen of an organism, in the same way we decide what counts as a reference specimen of the genome of a species.  Indeed, and interestingly, the genome of 'the' mouse or 'the' horse, is likely rarely if ever to be from the same official physical type specimen on display at some authorized museum.  Thus some poor (former) robin sits rigidly on some display branch for all to see, representing all robin-hood individuals it never saw or dreamt of.

This leads to interesting questions about how we treat our subjects, biology and evolution.  Are there alternatives to type specimens?  So, while you nurse your mint julep in suspense, think about this--the whole concept of a reference specimen, be it genetic or physical, because tomorrow we'll give a least a bit of thought to this question.

Tuesday, May 1, 2012

Metaphysics in science, Part III. It's there because I say it's there....and I must be right!

This series is about the nature of reality and what is loosely referred to as metaphysics, the idea of reality above, apart from, or somehow different from material, real reality.  When abstractions about things of this sort, like 'the human genome,' are understood as no more than practical working baselines, there may be no problem.  But when what we have in mind, so to speak, are processes or theories rather than things, the situation is far less clear.

What is the reality of a theory?  When is a theory an assertion....or a belief?
A theory in our context is a general principle or 'law' of nature of some sort.  The theory of gravity states that objects have specific quantitative relationships with each other, that are universal no matter what the objects are or where they are--in the entire universe!  We may have a difficult time falsifying a theory like that, because if there is even one single minor exception that would show the theory is not literally true, we may never come across that exception.

On the other hand, there are so many direct tests, experimental and otherwise, of such a theory, that we have very good reason to believe it. At that point the theory is accepted as an abstraction of reality, essentially in some Platonic sense.  All we can directly see are specific manifestations of the theory, but we assume in essence that every single possible instance would obey the theory--even instances that have not yet arisen.

If a theory is just an approximation, what's it an approximation of?
In some instances we agree that the theory is an approximation of reality, at least insofar as we can measure it.  Thus, we may not know the exact speed of light, but we assume there is such a speed, whether or not we can measure it perfectly, and that this speed--this ultimate limit--somehow exists above and beyond any instances we might observe.  So even our imperfect measures are, in a sense, given metaphysical meaning.  They tell us about the real speed limit.

The human genome sequence may be an easily recognized abstraction rather than any sort of metaphysical (in the sense of unreal or mystical) view of existence.  But the speed of light is rather different.  So, what about the theory of evolution and, in particular for our purposes, the theory that it is due to divergence from common ancestry as a result of natural selection?

Darwin used barnacles, among other groups, to argue for these points. He found shared traits among different types of barnacles, and shared traits between barnacles and crustaceans.  He compared the organization of their bodies, their segments and so on.  He reconstructed an hypothetical ancestor based on the idea of common ancestry.  There is no problem about his reasoning from instance back up to principle.

Among the traits he studied were the various aspects of sexual reproduction among barnacle species.  In particular, the differences between separate sexes, females with separate but dependent degenerate males (well, most males are degenerate, you might say!), and those who were complete hermaphrodites (both sexes in one animal).  He inferred this as being due to selection clearly showing that he meant this to reflect instances of the concept (and we can ask if it is metaphysical) that applies more broadly.

But he also inferred, at least implicitly in what we're familiar with, that barnacles with partial hermaphrodism were on the way to complete hermaphrodism.  His reasoning was that he's seen some modern hermaprhodites, and these intermediates could be seen as....intermediates!  That builds into an instance, his metaphysical or Platonic ideal. Rather than the instance showing the idea, in the way your genome is an instance of 'the human genome', Darwin imposed the theory onto the data.  That, in many ways, makes it fundamentally metaphysical, something science doesn't like to acknowledge.

Darwin knew that his idea of inheritance (called pangenesis) did not fit the data, and he wriggled about that, but it made problematic his theory of ubiquitous natural selection.  His estimate of the age of geological formations, such as the valley near his home in Kent, was far too old for what  astronomers were estimating.  But he stuck to his guns!  He believed he was right.  That in many ways is a commitment to metaphysical truth in the absence of the required evidence.

Then he extended this to human evolution, human racial variation, and our relationship to other modern primates.  We agree with Darwin in general, but not with the kind of racial hierarchy he invoked as being due to natural selection.  So we would say Darwin was extending his theory too broadly (some argue still today that in many ways there is a racial hierarchy--again raising the same kinds of issues about whether natural selection is a real 'law' or just something that can happen and that has to be shown separately in each case).

One might say he knew he was right about the process (the Platonic ideal of evolution?), and just that his measurements were wrong. But was he right about barnacles being 'on the way' to hermaphrodism?  Or about racial hierarchies?

If the truth is that all we have are instances rather than universal law, how can we know that?  How can we know whether our knowledge is incomplete or our ideas (our theory) is wrong?  Or, if we accept that theory may only be approximate, how can we know that, as opposed to the theory being simply wrong?

When and why should a theory be abandoned?
This raises the metaphysical question in a somewhat new way.  If a theory, like gravity or evolution, is truly universal, then that is somehow a metaphysical concept of which we can only see instances.  It seems a bit less clear than the idea of 'chair' accepted to stand for the various actual chairs that we see.

And in science other questions arise here.  We can look at an object and see if it is an instance of chair or not, because chair is a human-defined symbol.  But universal principles or laws of nature are human-discovered and their universality is effectively assumed as an ideal.  If we assume a theory, and don't find the evidence for it, when do we abandon it and admit we were wrong, and when do we just say that the evidence simply is still flawed?

We see the problem in areas we write a lot about here on MT: evolution and genetic causation.   GWAS attempt to identify 'the' genes responsible for a particular trait, with the clear if sometimes unstated expectation that it's tractable in the number of genes and rather stable, replicable, and quasi-deterministic (high predictability).  What that's not what's found--as is routinely the experience--then why don't we abandon the idea of simple causation, or that we can use enumerative approaches to find 'the' causes?  Why do we cling to the theory, the Platonic abstraction of causation, and demand larger, longer, more expensive studies to find the elusive truth?

One can quibble about the terminology we're using relative to professional philosophy, but this to us shows that even modern science, that sneers so readily at metaphysics, is very metaphysical at its core.  Vested interests, beliefs, hunger for simple or tractable theories, unstated appeal to satisfying theories in other sciences, and so on, as well as occasionally supportive evidence, all lead us to generalize from instances to theory by forcing the theory on the instances.

In this sense, as we often say, rather than siding with those who complain about its lack of dramatic results, that GWAS have not been a failure.  They've been over-done for many unjustified reasons, and will be over-done even more in the future, but these studies have shown that some of our deeply wished-for ideas about nature, our abstractions, our Platonic ideas, simply aren't that way.  But we don't want to hear that message, and certainly not to be accused of delving in metaphysics. So we cling to the theory even in the face of clear evidence that it's inaccurate at best--or we reinvent the theory so that we water down our criterion for calling effects 'major'.

Is this the Zen of genomics:  when No means Yes?
There are some relevant issues here that make the story less clear than we've stated.  They have to do with the real, or metaphysical, nature of 'emergence' or 'interactions', and we'll comment on that next time.

Wednesday, February 22, 2012

How many genes can we live without?

It's well-known that 'the' human genome doesn't exist -- despite the hype about the complete sequencing of this thing (and despite the incompleteness of its sequencing).  In fact, each of us has a collection of DNA variants that, added together, means that our genome has never been seen before in the 3.8 billion year history of genomes, and will never be seen again.  And it's no mean collection; we all differ from each other at something like 3 million loci.

We need to understand first of all that 'the' human genome is not from one person, and even so, what it is, is a reference sequence, useful for comparing other data, but not definitive of our species.  The donors of the DNA were healthy at the time of donation, but that's about it.  They were not particularly special in any way.

Human genome by functions; Wikimedia Commons
In fact, as it turns out, not only do we differ at single nucleotides -- you have a T where I have an A, I have a G where you have a C, and so on, times 3 million -- but we are each carrying around a not insignificant number of variants that result in loss of function (LoF) of some of our protein coding genes.  And many of these are genes we think of as essential.

A paper published in last week's Science, "A Systematic Survey of Loss-of-Function Variants in Human Protein-Coding Genes," MacArthur et al., estimates that each of us has around 100 of these LoF variants, 20 of which result in complete loss of a gene. The authors looked at three pilot data sets from the 1000 Genomes Project, (58 Yoruba, from Nigeria, 60 European American from Utah, 30 Chinese individuals from Beijing and 30 Japanese from Tokyo) and a European genome, and, after filtering a larger initial set of candidate LoF variants, finally analyzed 1285 variants that they found to be likely to cause protein-coding genes to lose function.   

As MacArthur et al. point out, other recent studies of the complete DNA sequences of healthy individuals have also found many LoF variants -- from 200 to 800.  People walking around perfectly normal so far in their lives, but without the use of substantial numbers of their genes.  The specifics vary, and what you can do without, your genes that aren't working, likely depends on what is working.

Before the advent of complete genome sequencing, LoF variants were thought to be rare, largely associated with severe Mendelian disorders such as cystic fibrosis or Duchenne muscular dystrophy.  The finding that they aren't so rare after all suggests to MacArthur et al. "a previously unappreciated robustness of the human genome to gene-disrupting mutations and [this has] important implications for the clinical interpretation of human genome–sequencing data."
LoF variants found in healthy individuals will fall into several overlapping categories: severe recessive disease alleles in the heterozygous state; alleles that are less deleterious but nonetheless have an impact on phenotype and disease risk; benign LoF variation in redundant genes; genuine variants that do not seriously disrupt gene function; and, finally, a wide variety of sequencing and annotation artifacts. Distinguishing between these categories will be crucial for the complete functional interpretation of human genome sequences.
After they weeded out false positives (which were due to sequencing errors; of course, false positives are a problem in their own right in a clinical setting), the variants included indels (insertions or deletions of 1 or more nucleotides) that changed the splicing of the gene (and thus changed the amino acids that got strung together in the resulting protein), single nucleotide variants that introduced a stop codon into the gene sequence (that is, that caused transcription of the gene to halt prematurely), and large deletions that removed some of the coding sequence.  Some of the variants were found to affect all known protein-coding transcripts of the affected gene, and some affected only some of the coding transcripts.  That is, some transcripts were normal. 

The authors find that the common gene-disrupting variants described in this study are not a significant cause of complex disease.  Most of the LoF variants identified that are associated with complex diseases, all but one of which were heterozygous in these subjects (they had one functional copy, one not), are at low frequency, presumably due to purifying selection -- that is, selection against alleles that are severely harmful, thus preventing them from reaching high frequencies in any population. 

Individuals in this study have about 120 LoF variants, about 100 of these heterozygous and 20 of them homozygous -- both the person's copies are non-functional.  It's the homozygous LoF's that seem to have no effect that are of most interest.  The individuals in this study are healthy -- so either their LoF's truly have no effect, or they haven't yet had an effect.  If they are truly what MacArthur et al. are calling LoF-tolerant genes, and don't lead to disease, the authors suggest that they can be used to "define the functional and evolutionary characteristics that distinguish these genes from severe recessive disease genes."
We examined the 253 genes containing validated LoF variants that were found to be homozygous in at least one individual. These LoF-tolerant genes are significantly less conserved and have fewer protein-protein interactions than the genome average. They are also enriched for functional categories related to chemosensation, largely explained by the enrichment of olfactory receptor genes in this class (13.0% versus 1.4% genome-wide), and depleted for genes involved in embryonic development and cellular metabolism.
The finding that so many of the genes that are 'allowed' to vary are olfactory receptor (OR) genes isn't much of a surprise, as we all carry many OR genes that are pseudogenes, OR's that no longer function.  So, the researchers eliminated these from the set of genes that could lead to Mendelian disease, and then compared the remaining 213 LoF-tolerant genes with 858 known recessive disease genes.  They found these 2 categories were very different, and suggest that the characteristics of the recessive disease genes that are not shared with the LoF-tolerant genes could be used to prioritize candidate disease genes.

This can be important because we all have so many variants, and identifying which of these is or are contributing to a disease we may have is often impossible without data from other affected family members.  If likely candidates can indeed be prioritized based on the results of this study, this could be very helpful.  But, many genes are only deleterious in a given environment, or after years of exposure to environmental factors, and one person's LoF-tolerant gene might be another's disease gene.

To us, the fact that so many genes can apparently be disrupted with no discernible ill effect is further evidence of the adaptability that evolution has built in -- DNA replicating errors are common, the timing and locale of gene expression are often imprecise, and environments frequently change.  We're redundant, buffered, pretty hardy creatures!  The fact that all this can happen and we can live on with no ill effect is a beautiful fact of life.

Tuesday, February 14, 2012

Ptolemaic genetics: epicycles of lobbying

That was then...
Ibn al-Shatir's model for the
appearances of Mercury,
showing the multiplication of
epicycles in a Ptolemaic
enterprise. 14th century CE
(Wikimedia Commons).
Way back then, in the dark ol' days of science, the Roman astronomer Claudius Ptolemy (90-168AD) tried to explain the position of the planets in terms of divinely perfect circles of orbit around God's home (the Earth).  The idea that we were at the center of perfect celestial spheres was a standard 'scientific' explanation of the cosmos and our place in it.

But the cantankerous planets refused to play by the rules, and their paths deviated from perfect circles.  Indeed, occasionally the seemed to move backward through the skies!  Still, perfect circular orbits around Earth simply had to be true based on the fundamental belief system of the time, so astronomers invented numerous little deviations, called epicycles, to make the (we now know) elliptical orbital pegs fit the round holes of theory.

And then along came Nicolaus Copernicus (1473-1543 AD).  And the cosmos was turned inside out: the earth was not the center of things after all!

Thomas Kuhn famously described in The Structure of Scientific Revolutions how the best and the brightest scientists struggle valiantly to fit pegs into holes they don't really fit, until some bright person ccomes along and shows the benighted herd a better way to account for the same things.  Copernicus, Galileo, Newton, Einstein, and others were the knights in shining armor who inaugurated some of the most noteworthy of these occasional 'scientific revolutions.'  Darwin's evolutionary ideas are also a classic example.

The same kind of struggle is just what is happening now in genetics and evolutionary biology--indeed in many other fields in which statistical evidence runs headlong into causal complexity.  Whether, when, or what knightly change will occur is anyone's guess.

And this is now
Everyone remembers the hoopla the sequencing of the human genome was met with when it was announced (or rather, each time it was announced) -- we were promised that we would by now not only know why people were sick, but we'd be able to predict what we'd get sick with in future.  It was promised that this would be a silver-bullet reality by the early 21st century by no other than Francis Collins.  Others were promising lifespans in the centuries: all of us would be Methuselahs!

So, all those illnesses would now be treatable or preventable in the first place. How?  Well, the genome would allow us to identify druggable pathways, and common diseases must be due to common genetic variants (an idea that came to be known as common disease common variant, or CDCV), and if we could just identify them, we'd be in business.  After all, didn't Darwin show us that everything about everything alive was due to genetic causation and natural selection?  If that's the case, we should be able to find it, and our wizardry at engineering would take the ball and run with it.  Big Pharma jumped on the 'druggable' genome bandwagon and people running big sequencing labs jumped on the CDCV idea, and genomewide association studies (GWAS) were born.  And then the 1000 Genomes project, and all the -omics projects....  Big is better, of course!  Not that these efforts weren't questioned at the time, based on what everyone should have known about evolution and population genetics, but the powers-that-be plowed ahead anyway.

Well, we're no longer in a minority of naysayers.  It's widely recognized that GWAS haven't been very successful, relative to the loud promises being trumpeted only a few years ago.  And even the successes they have had -- and numerous genes associated with traits have been identified, it must be said -- typically explain only a small amount of the variation in disease, or any trait, in fact.  So now researchers are working on automating the prediction of disease from gene variants based on protein structure and other DNA-based clues.  But the assumption--the belief system, really--is still that the answer is in the DNA, and disease prediction is still going to be possible.

A piece in Feb 9 Nature describes a number of state-of-the-art approaches to predicting the effects of DNA variants, in part based on what amino acid changes do to proteins.  The idea now is that diseases are going to be found to be due to rare variants, and the challenge is to figure out what these variants do.  In part, evolution will help us to do this.
"Sequencing data from an increasing number of species and larger human populations are revealing which variants can be tolerated by evolution and exist in healthy individuals."
But, are we trying to explain a current disease, or predict the diseases someone will eventually get? These are different endeavors, though it may often be inconvenient to acknowledge that.  Rare pediatric diseases that are due to single genetic mutations, or genetic diseases that cluster in families (and, again, usually with young onset age and rare) are easier to parse than the complex chronic diseases that most of us will eventually get.  But, based on the comparison of the genomes that have already been sequenced, we now know that we all seem to differ from each other at something like 3 million bases.  That is, we all have a genome that has never existed before and never will again. Assigning function to all that variation is from daunting to impossible -- not least because a lot of it might not even have a function.  And the idea that we'll eventually be able to make predictions from those variants is based on questionable assumptions.

It's true in one sense that every disease we get is genetic -- everything that happens in our body is affected by genes -- but in another sense, much of what happens is a response to the environment, and so is environmentally determined--that is, is not due to genetic variation in susceptibility.  Predicting a disease from genes when it's due to combined action of genes and environment, therefore, is a very challenging problem.

Here is just one example of why: Native Americans throughout the Americas are about 65 years into a widespread epidemic of obesity, type 2 diabetes and gallbladder disease, diseases that were quite rare in these people before World War II.  There are a number of reasons to suspect that their high prevalence is due to a fairly simple genetic susceptibility.  But, if gene variants (still not identified) are responsible, they have been at high frequency in the descendants of those who crossed the Bering Straits from Siberia for at least 10,000 years -- which means that variants that are now detrimental were "tolerated by evolution and exist[ed] in healthy individuals" for a very long time.

If geneticists had wanted to predict 70 years ago what diseases Native Americans were susceptible to, these variants would have been completely overlooked, because they weren't yet causing disease.  And indeed these 'risk' genes, whatever they be, were benign -- until the environment changed.  We're all walking around with variants that would kill us in some environment or other, and since we can't predict the environments we'll be living in even 20 years from now, never mind 50 or 100, the idea that we'll be able to predict which of our variants will be detrimental when we're old is just wrong. In fact, we're each walking around with substantial numbers of mutant or even 'dead' genes, with apparently no ill effect at all -- but who knows what the effect might be in a different environment.

But, ok, some of us do have single gene variants that make us sick now.  Many of these have been identified, most readily when a family of affected individuals is examined (though the benefit of knowing the gene is rarely of use therapeutically), but many more remain to be.  The current idea is that this can be done by looking for mutations in chromosome regions that are conserved among species, and figuring out which of these change amino acids (and thus the protein coded for by the gene).  The idea is that unvarying regions are unvarying because natural selection has tested the variants that arose and found them wanting, thus eliminating them from the population.  They must, therefore, be functionally important!
A host of increasingly sophisticated algorithms predict whether a mutation is likely to change the function of a protein, or alter its expression. Sequencing data from an increasing number of species and larger human populations are revealing which variants can be tolerated by evolution and exist in healthy individuals. Huge research projects are assigning putative functions to sequences throughout the genome and allowing researchers to improve their hypotheses about variants. And for regions with known function, new techniques can use yeast and bacteria to assess the effects of hundreds of potential mammalian variants in a single experiment.
This is potentially useful, because for those with single gene mutations that cause disease -- 1 variant among 3 million other ways in which each person differs from everyone else -- homing in on the causative mutation is, again, difficult to impossible if you don't have a large family with similarly affected individuals in which to confirm the association of mutation and disease.

Well, if we can do with or without a protein (or other functional DNA element), depending on the variation we have across the genome, then even when the element is important its variation in a given individual may not be causal: there are many examples where that is clearly true.  Further, the same kind of evolutionary reasoning would say that centrally important -- and hence highly conserved -- parts of the genome probably cannot vary much without being lethal, largely to the embryo.  So, from that equally sound Darwinian reasoning, we would expect that disease-associated variation will be in the minor genes with only little effect!  So the 'evolutionary conservation' argument cuts both ways, and it's not at all clear which way its cut is sharpest.  It's a great idea, but in some ways the hope that searching for conservation will bail us out, is just more wishful thinking to save business as usual.

Methuselah (Della Francesca ca. 1550) 
To complicate things even more, not all amino acid changes cause disease, or even do much of anything.  And again, sometimes they will only be harmful in a given environment.  And, of course, not all diseases are caused by protein changing mutations -- sometimes they are caused by disturbances to gene regulation.

In fairness, the multitude of researchers trying to make sense of the limitless genetic variation that is pouring out of DNA sequencers recognize that it's complicated.  But then, why are they still saying things like this, as quoted in the Nature piece: “The marriage of human genetics and functional genomics can deliver what the original plan of the human genome promised to medicine.”

What's to the rescue?  Do we need another 'scientific revolution'?
We have no idea when or if our current model of living Nature will be shown to be naive, or whether our understanding is OK but we haven't cottoned on to a seriously better way to think about the problems, or indeed whether the hubris of computer and molecular scientists' love of technology will, in fact, be victorious.  If it comes, it could be.  But we are certainly in the midst of a struggle to fit the square truths about genetics and evolution into the round holes of Mendelian and Darwinian orthodoxy.

Perhaps the problem to be solved is how to back away from enumerative, probabilistic, reductionistic treatment of complex, multiple causation, and to make inferences in other ways.  We need to understand causation by numerous, small or even ephemeral statistical effects, without our current enumerative statistical methods of inference. In terms of the philosophy of science, doing that would require some replacement of the 400 year-old foundations of modern science, based on reductionistic, inductive methods that enabled science to get to the point today where we realize that we need something different.

The situation here is complicated relative to scientific revolutions in Copernicus', Newton's, Darwin's or even Einstein's time by the large, institutionalized, bureaucratized, fiscal juggernaut that science has become. This makes the rivalries for truth, for explanations that this time will finally, really, truly solve the complexity problem even more frenzied, hubristic, grasping, and lobbying than before.  That adds to the normal amount of ego all of us in science have, the desire to be right, to have insight, and so on.  Whether it will hasten the inspiration for a transforming better idea, or will just force momentum along incremental paths and make real insight even harder to come by, is a matter of opinion.

Sadly, the science funding system, including the role of lobbying via the media, is so entrenched in our careers, that dishonesty about what is claimed to the media or even said in grants is widespread and quietly acknowledged even by the most prominent people in the field: "It's what you have to say to get funded!", they say.  But where does dissembling end and dishonesty begin when it comes time to the design and reporting of studies (and, here, we're not referring to fraud, but to misleading results and over promising the importance of the work)?  The commitment to the ideology and the promises restrains freedom of thought, and certainly dampens innovative science.  But it's a trap for those who have to have grants and credit to make their living in research institutions and the science media.
Zip-line over rainforest canopy,
Costa Rica (Wikimedia)

But right now, scientists are like tropical trees, struggling mightily to be the one that reaches the sunlight, putting the others in their shade. What we need is a conceptual zip-line over the canopy.