The gaggle that continue to raid evolutionary biology blogs, patrolling for things that can be naively or intentionally misinterpreted as evidence for their theological views, specifically 'Intelligent Design' (ID), loves to concentrate on complex traits. They claim such traits cannot have evolved because the independent components won't function on their own and the whole breaks down without them. They call that Irreducible Complexity: since you can't take any components of complex traits away and still be viable, such traits could not have arisen gradually by natural selection. Therefore (the IDeologs say), Intelligent Design is true. But this is false on several grounds, not all of them even recognized by biologists, who often defend evolution by needlessly agreeing to do it on the IDeologs' turf.
First, it is IDiotic to argue that if an evolutionary claim is false, therefore creationism is true. That is simply a logical fallacy. If evolution as biologists see it were being misperceived, that in no way provides evidence for any specific counter explanation. Only an IDeolog would make such an argument. It would be just as sensible--that is, as nonsensical--to say that our misperception proved that life came to earth from a parallel universe in a spaceship made of banana peels. We get things wrong or understand them incompletely in evolutionary biology, which is why it remains an active science, but that is not evidence that evolution didn't happen.
Second, the major IDiotic argument about the need for completeness was one Darwin was aware of and even speculated on in regard to the eye, a favorite irreducible complexity example cited from that time to the present day. Darwin suggested ways that primitive light sensitivity could have evolved bit by bit. In what was really striking prescience, his basic speculations have been shown to be about right, because species alive today with 'partial' vision have been found, and genetic components of vision are shared among species with simple as well as complex light reception. Even saying 'partial' vision is a subtle misnomer, because each species uses what it has: the light sensitivity of a worm or bacterium is not partial for their uses, and to use the adjective suggests the IDiological view that humans are at an intended pinnacle, that our vision is somehow more complete or real than a clam's. That's an egocentric misperception of evolution.
Complexity is reducible! It always has been. It's a central aspect of life. Right here and now
Thirdly, and perhaps even more important than the first two reasons why the anti-evolutionary IDeology is just plain wrong is that complexity is typically reducible! The basic IDeologs' premise doesn't have to be refuted because it's not true.
What we know very well is that most traits of organisms are, in fact, the result of multiple interacting factors (gene networks, the polymeric, cooperative nature of DNA and proteins, signaling and receptors systems, gene regulation, and multipart proteins, etc.). And, eyes, too. That is a central fact, and a main point of MT (the blog and the book). We know from thousands of studies (yes, even the GWAS and other 'omics' studies whose excesses we love to point out) that complex traits really are complex at the gene level.
The same studies also show by their very nature--by the very fact that we are doing so many of them in the first place--that each person will have a different genotype, a different set of variants, involved--even if they have the 'same' trait, like stature, insulin levels, blood pressure, or behavior. That is why personalized genetic medicine is unlikely to work nearly as well as advertised. Personalized medicine almost assumes irreducible complexity: enumerate the parts and then any variation in the trait must be due to a broken part that can be identified. But that isn't how Nature works.
Reducible complexity is true even of vision: Color-blind people are people and they have vision, yet they are missing functional light-sensitive genes (e.g., genes that are used in red or green detection, or overall light sensitivity). Visual acuity varies in all sorts of ways among perfectly viable people.
This is typical of biological traits. And recent studies have clearly shown that each of us is walking around with numerous completely inactivated genes, whose 'damaged' sequence variants we have inherited--from parents who somehow had managed without them. One recent paper found that around 165 different genes were completely inactivated (both copies not working) in a typical person. And there are many others in which one of our two copies is not working normally. The combination of inactive genes would be different for each person, but the truth is that we do not normally need all the genes in our genome. That tolerance of variation is exactly the working material that biologists have known is at the basis for evolution from Darwin's own time.
Confirming this in another way, and also very clearly, is that it is routine that a gene experimentally inactivated in a laboratory animal, like a mouse, has serious effects in some strains but little or even no effect in others. A mutation causing a serious disease in humans may do nothing when the same mutation is tested in a mouse, or it may have similarly bad effects only in some strains. That's one of the notorious problems with mouse models for human traits: mice and people share many traits but we make them differently to various extents. There is more than one way to make the same trait. Complexity is reducible.
The reducibility of a trait, to put it in terms even an IDeolog could understand, depends on the combination of genes being viable, not on every gene having the most functionally efficient variants. The importance of component cooperation, a favorite MT word, is in part that various types of cooperation are viable. That aspect of redundancy and variation is one of the central reasons that complexity could evolve in the first place, exactly in the general fashion argued by Darwin and since. No biologist suggests that an eye just emerged wholesale from the primeval slime.
But there's more. Studies of the nature and evolution of genomes shows very clearly that genetic mechanisms arise largely by means that generate redundancy as well as alternative pathways to given outcomes, as cells respond to their local environment. Gene duplication occasionally leads to individuals with two copies of a gene where in their ancestors there was only one (this happens in species generally, not particular to humans in any way). That can provide redundancy, so that one of the copies can acquire mutations that alter what the gene does, while the other copy keeps plugging along with the original function. The new function can be due to mutations in the protein code of one of the copies, or the DNA sequences that regulate when and where the gene is used.
For these reasons, traits are the result of many different genetic contributions, all varying among individuals, each reaching similarly viable traits with different combinations of that variation. Those combinations that aren't functional don't survive or reproduce; those that have an advantage may do better. Over time, the mix of variation, including even the number and set of contributing genes, allow traits to evolve new or altered function.
This is how evolution works, gradually producing new or varied traits. We understand this because we are aware that complexity is often, or even typically, reducible. Although it hasn't been put this way before to our knowledge, this is nothing more than a modern understanding of classical evolutionary ideas.
The IDeologs claim that reduced complexity could not have existed in a stepwise, bit by bit, assembly of a new trait from parts that would not work on their own--that evolution couldn't get from there to here. But the deeper truth is that evolution is both there ('incomplete') and here ('complete') today and has been that way at any or even every time in the past. It isn't just that things have to be assembled over time by different steps, but that they exist at any given time in various steps or stages of 'completeness.' To a great extent, biological complexity is inherently reducible at any time as well as over time.
And one more reason: Of course, we needn't have gone through all of this to convince you that complexity was reducible, after all. That is because the IDeologs disprove their own irreducibility argument by their very existence: one can function as a human being even with a brain that allows you intentionally not to use it to recognize the realities of the world--by not using the thinking complexity they were born with! We would apply this to those who lead the movement, and do or should know better, but not those who they naively lure into adopting its know-nothing IDeology.
Finally, we may make sport of intentionally or willfully self-deluded critics of evolution. For any of those who are sincere but naive, one can only say that it's too bad, and poignant, too, that science shows the evolutionary nature of life, rather than the comforting existence of a benign divinity who graced the earth with our presence. How nice if that could be true! How hard it makes it to understand the injustices and suffering in the world. But science is about the real world, not the one we might wish for.
Wednesday, February 8, 2012
Tuesday, February 7, 2012
Doubt and dogmatism in science -- questioning natural selection
Through no fault of his own, a friend of ours finds his written words being used (or rather, abused) by the ID community. Again. This has happened to us from time to time as well, so we thought we'd address it here. Not the actual arguments, which we have no interest in, but the misconstruing of what scientists say, or clearly mean.
Adam Wilkins, a biologist and long-time very thoughtful editor of BioEssays, a leading biology journal, recently published a well-considered review of a new book by James Shapiro, Evolution: A View from the 21st Century. Adam reviews the book favorably in general as a thoughtful one that those seriously interested in the nature of evolution should read; but he takes widespread exception to the author's view, taking him to task for writing that natural selection may be less important than most biologists would accept.
Adam is not the first reviewer to take Shapiro to task for this. And, probably because he relegates natural selection to a minor role, Shapiro has been assumed to be an IDer by some, including some people in the ID community. This right here means there's a problem -- if a biologist can't question accepted wisdom in evolutionary theory, this makes evolutionary theory a dogma, just like ID. Science should always be questioning itself -- that's how knowledge is built and expanded upon.
But, giving succor to ID is not Shapiro's intention. This is clear from the 'debate' he has with IDers, which we won't even link to because it's tiresome, and really not much more than a clash of ideologies (you'll find it anyway, if you really must).
But, after this rather fruitless 'debate', Adam, or at least his review, gets pulled into the fray. Adam's piece was published in Genome Biology and Evolution in January. And the authors of the post believe it's a gotcha moment, saying that Wilkins admits something that few 'Darwinists' (and yes, that's a slur) will, which is that "a growing body of scientists" are starting to question the "alleged power of Darwin's natural selection to create the world of life that we see."
Yes, Adam does say that there are biologists who feel that the role of n.s. has been overstated (and yes, you've seen that here on MT, in fact). But, he absolutely does not include the 'therefore' that's implied -- therefore, if n.s. didn't do it, a designer did. The gotcha quote they pull from the review is this:
For biologists, for whom not a shred of evidence collected in the last 150 years has called into question the idea that all of life descended from a common ancestor that lived nearly 4 billion years ago, this kind of disagreement is arguing around the edges. But for an ID adherent, any kind of disagreement within the fold must mean that, therefore, evolution didn't happen.
We and other biologists don't question that natural selection can occur, or that it does occur, but ask when, where, how, how strongly, and how systematically it occurs--and how we can know which is which. We ask how it works in general or in specific instances relative to other factors that can lead to differential proliferation of variation, or of the way genetic and other transmissible variation arises and works. That is totally different from asking whether natural selection or Divine intervention account for life, which is not what legitimate science does. Science is only one way to know, but it rests on observable causation in the material world only.
We could put this another way. If it were somehow possible to show--to really show--that Divine intervention were the explanation, or that Jesus was divine, or that Mohammed really did get his inspiration from the Angel Gabriel, or that ants had souls, any sane scientist would love to be the one to do that. His or her reputation would dwarf even Darwin's! But that's simply not the message the material world gives us. Even if natural selection were somehow shown to be totally wrong, it would provide not a scintilla of evidence for creationist explanations. It would just say we've been accepting an incorrect theory and have more work to do. That is not a threat to science, even if many scientists do cling too tightly to simple explanations for complex things.
There are separate worldviews operating here, and, even if IDers actually understood the science, their fundamentalist view of the world still prohibits questioning their dogmatic view, creationism. And this is why something like Adam's review can be taken in vain -- IDers assume that if scientists disagree, that's a crack in the religion of evolution.
But evolution isn't supposed to be a religion. Scientists are supposed to question what they know. Jim Shapiro's questioning of the pre-eminence of natural selection in evolution is perfectly valid science, and will either stand the test of time, and questioning by other scientists, or it will fall. But it doesn't mean he sees the hand of a watchmaker behind every complex living thing. It means he thinks evolutionary theory hasn't yet explained everything.
Adam Wilkins, a biologist and long-time very thoughtful editor of BioEssays, a leading biology journal, recently published a well-considered review of a new book by James Shapiro, Evolution: A View from the 21st Century. Adam reviews the book favorably in general as a thoughtful one that those seriously interested in the nature of evolution should read; but he takes widespread exception to the author's view, taking him to task for writing that natural selection may be less important than most biologists would accept.
Adam is not the first reviewer to take Shapiro to task for this. And, probably because he relegates natural selection to a minor role, Shapiro has been assumed to be an IDer by some, including some people in the ID community. This right here means there's a problem -- if a biologist can't question accepted wisdom in evolutionary theory, this makes evolutionary theory a dogma, just like ID. Science should always be questioning itself -- that's how knowledge is built and expanded upon.
But, giving succor to ID is not Shapiro's intention. This is clear from the 'debate' he has with IDers, which we won't even link to because it's tiresome, and really not much more than a clash of ideologies (you'll find it anyway, if you really must).
But, after this rather fruitless 'debate', Adam, or at least his review, gets pulled into the fray. Adam's piece was published in Genome Biology and Evolution in January. And the authors of the post believe it's a gotcha moment, saying that Wilkins admits something that few 'Darwinists' (and yes, that's a slur) will, which is that "a growing body of scientists" are starting to question the "alleged power of Darwin's natural selection to create the world of life that we see."
Yes, Adam does say that there are biologists who feel that the role of n.s. has been overstated (and yes, you've seen that here on MT, in fact). But, he absolutely does not include the 'therefore' that's implied -- therefore, if n.s. didn't do it, a designer did. The gotcha quote they pull from the review is this:
…the book’s contention that natural selection’s importance for evolution has been hugely overstated represents a point of view that has a growing set of adherents. (A few months ago, I was amazed to hear it expressed, in the strongest terms, from another highly eminent microbiologist.) My impression is that evolutionary biology is increasingly separating into two camps, divided over just this question. On the one hand are the population geneticists and evolutionary biologists who continue to believe that selection has a ‘creative’ and crucial role in evolution and, on the other, there is a growing body of scientists (largely those who have come into evolution from molecular biology, developmental biology or developmental genetics, and microbiology) who reject it.Adam's following paragraph draws their scorn.
The arguments from paleontological evidence for the importance of natural selection largely concern the observed long-term trends of morphological change, which are visible in many lineages. It is hard to imagine what else but natural selection could be responsible for such trends, unless one invokes supernatural or mystical forces such as the long-popular but ultimately discredited force of “orthogenesis.”Obviously this draws scorn, because invoking the supernatural is exactly what IDers do. But, equally obviously, to a biologist such as Adam, that's not an explanation. Adam of course was writing for biologists, and for the overwhelming number of biologists evolution is a fact of life. He wasn't writing with creationists/IDers in mind. If he had been, he might have restructured his argument somewhat, but he still wouldn't have hidden the fact that there are disagreements among biologists about the strength or predominance of natural selection as a force in evolution. The disagreement doesn't make it false. It makes it science.
For biologists, for whom not a shred of evidence collected in the last 150 years has called into question the idea that all of life descended from a common ancestor that lived nearly 4 billion years ago, this kind of disagreement is arguing around the edges. But for an ID adherent, any kind of disagreement within the fold must mean that, therefore, evolution didn't happen.
We and other biologists don't question that natural selection can occur, or that it does occur, but ask when, where, how, how strongly, and how systematically it occurs--and how we can know which is which. We ask how it works in general or in specific instances relative to other factors that can lead to differential proliferation of variation, or of the way genetic and other transmissible variation arises and works. That is totally different from asking whether natural selection or Divine intervention account for life, which is not what legitimate science does. Science is only one way to know, but it rests on observable causation in the material world only.
We could put this another way. If it were somehow possible to show--to really show--that Divine intervention were the explanation, or that Jesus was divine, or that Mohammed really did get his inspiration from the Angel Gabriel, or that ants had souls, any sane scientist would love to be the one to do that. His or her reputation would dwarf even Darwin's! But that's simply not the message the material world gives us. Even if natural selection were somehow shown to be totally wrong, it would provide not a scintilla of evidence for creationist explanations. It would just say we've been accepting an incorrect theory and have more work to do. That is not a threat to science, even if many scientists do cling too tightly to simple explanations for complex things.
There are separate worldviews operating here, and, even if IDers actually understood the science, their fundamentalist view of the world still prohibits questioning their dogmatic view, creationism. And this is why something like Adam's review can be taken in vain -- IDers assume that if scientists disagree, that's a crack in the religion of evolution.
But evolution isn't supposed to be a religion. Scientists are supposed to question what they know. Jim Shapiro's questioning of the pre-eminence of natural selection in evolution is perfectly valid science, and will either stand the test of time, and questioning by other scientists, or it will fall. But it doesn't mean he sees the hand of a watchmaker behind every complex living thing. It means he thinks evolutionary theory hasn't yet explained everything.
Monday, February 6, 2012
Sweet tooth or sweet talk? The truth about the truth about sugar
You have to suspect any story with a title that begins "The truth about....", and the story in last week's Nature is no exception: "Public Health: The Toxic Truth about Sugar". Just the latest in a number of stories indicting sugar as the cause of all that ails us (almost literally), the piece's own summary is this:- Sugar consumption is linked to a rise in non-communicable disease
- Sugar's effects on the body can be similar to those of alcohol
- Regulation could include tax, limiting sales during school hours and placing age limits on purchase
The evidence is clear: even at Starbuck's it's hard to find some actual coffee (not to mention the impossibility of a 'small' coffee) in amongst the choco-banana-raspberry flattes. And then there are the Scots' deep-fried Mars bars. Who could doubt that our commercially whetted sweet tooth is the tooth that bites with poisoned fangs? But is the evidence actually so clear?
The conclusion looks suspiciously in need of the "correlation is not causation" reminder. Indeed, a reader who posted a comment on the paper in Nature, Geoff Russell, pointed this out with cogent examples. We'll reproduce his entire comment here, as it makes the point well.
- Geoff Russell said:
Australia provides a natural test of the sugar-is-the-evil-bullet theory. We don't produce much corn here, so continue to use cane sugar for most of our sweetening. In the 1960s we didn't have an obesity epidemic. How much sugar did we consume? According to the FAO, 52.3 kg per person per year in 1965 (of 55kg total sweeteners).
What about now, in the midst of our own obesity and type 2 diabetes epidemics?
We are down to 39.6 kg of cane sugar per person per year, with an additional 8kg of non-sugar sweeteners. Overall there has been a modest decline in all sugars despite a rise in obesity and diabetes. How has our food supply has changed over the past 4 decades? We have more Calories. If may be tempting to attribute the US obesity crisis to sugars, but obesity increases elsewhere demonstrate that more Calories and less exercise are a sufficient explanation.
Similarly, compare Cuba and Italy. Cuba consumes 500 kCal per day of sugar and Italy just 300 kCal, Italy has an obesity/type 2 diabetes problem while Cuba's rates are very low. Historically, Cuba has eaten even more sugar than she does now ... without the evil consequences that this article portends.There is no denying, of course, that obesity and what are usually thought of as its sequelae are public health problems in much of the world. Whether or not said sequelae are indeed sequelae of obesity, or whether obesity and the rest are, individually, consequences of fat consumption, or sugar consumption, or processed food in general, or simply of excess calories relative to energy usage (too much munching in front of the telly) has still not been determined, though many have their favorite candidates. Cholesterol, saturated fat, red meat, the non-Mediterranean diet, not enough exercise, and others.
There's a fundamental problem with this simple sugar analysis when it doesn't reliably predict on either the individual or population basis (as Geoff Russell's comment points out). Sugar, and/or what it's usually allied with, may well have detrimental effects on health, but clearly it's not as simple as is being said. There's a well-known issue in epidemiology called the ecological fallacy, whereby we paint individuals with a brush dipped in a population-based paint. That is, when we attribute generalizations about a group to causation at the individual level -- stereotyping is an example, but so is the (erroneous) assumption that because, say, risk of heart disease is higher among people who smoke, everyone who smokes will have a heart attack.
It simply can't be possible that sugar is the single or even primary cause of the obesity etc. epidemic. Too many healthy people consume a lot of sugar for this to be true. And surely too many unhealthy people don't. Trying to attribute this vast epidemic to a single dietary substance is denying the complexity of these diseases and of causation. Unfortunately (in our view), people are beginning to feel so fiercely about sugar as the root of all evil that they are (in our view) no longer able to assess the science. Instead, it's become such a strongly held belief that it's in danger of becoming a dogma that no longer needs to be tested, or the supporting studies examined with a skeptical eye. That's never a good thing.
This is relevant for an MT post not because we don't want it to be true because we both OD on sugar all the time (we don't), or that we only eat celery and carrots (we don't), but because it relates to the general problem of inferring causation, especially when we can't do definitive experiments. That is the common situation in human and evolutionary genetics. So, it's worth sitting contemplatively over a cup (small) of coffee discussing whether there are better ways to know about weak, gradual, or complex causation.
One lump, or two?
Friday, February 3, 2012
Ectopic thinking?
Our gene mapping project, which we first blogged about here, is starting to get interesting. And not necessarily for the reasons that we'd hoped. You might remember that our project is looking for genes involved in variation in specific craniofacial traits in the F34 generation of descendants of a cross between two inbred mouse lines. We've measured a handful of traits in over a thousand mouse skulls, and mapped their genetic effects by looking for genetic variation across the genome that might be associated with variation in the traits.
Sparing you the gory details, we'll just say that the chromosomal intervals that one or more of the traits mapped to span 30% of the genome, and 10% of all coding genes. That's 2400 genes or so that could potentially be of interest in affecting head shape in just these particular mice, with the restricted genetic variation they have (because they are descendants of only two inbred parental mouse strains). That's a lot of genes to wade through to figure out which might be most likely to be involved in the traits we're looking at.
One way to prioritize candidate genes from such a study is to look for the genes in every interval that you know from prior work to be involved in your trait of interest. Or to identify genes in families that include genes involved in your trait of interest -- these would then be considered guilty by association.
But this means most genes don't have a fighting chance of being considered, because you don't happen to know anything about them, or because nobody knows anything about them, or because what's known about them only partially represents what they do.
To try to minimize this, many people automate the search, with programs that cull the genes that the literature indicates might be of interest, or that seem to be expressed where you want them to be. So, this might solve the problem of no one knowing everything about all genes, but it doesn't solve the problem of nothing being known about so many genes, or that there's only partial knowledge. And it doesn't solve the problem of having to tell the program what to look for, which means you're constraining it in the same way you would if you were doing the search by hand, looking for specific families of genes. Nor, of course, does it solve the problem of what's happening in all the non-coding DNA that flanks all those genes.
Thus, we decided that the least biased way to comb the data was to go through all the genes in all the intervals by hand. We're still making sense of all that, not least because we are hoping not to be constrained by the usual ideas about statistical significance, but we've learned some interesting things along the way.
For example, one of the intervals of interest is loaded with olfactory receptor (OR) genes. Olfactory receptors reside on the cell surface of olfactory receptor neurons, and are involved in odorant detection. ORs form the largest family of genes in many genomes -- about 1000 different genes -- and they cluster in sets of genes in various locations on a number of chromosomes. ORs have a distinctive expression pattern, with only one expressed per neuron in the tissue lining the nose, where they each are sensitive to particular aspects of molecules the animal inhales, and hopes to smell. How expression of the remaining 999 genes in each cell is blocked is still not known.
ORs are an interesting example of something we've blogged about before, but that continually surprises us. One of the ways we're evaluating the possible role of all these genes in development of the traits we're looking at is to look at where they are expressed in the developing embryo. We initially thought this would be helpful for narrowing the search, but it turns out that about 95% or even more of genes (for which there are expression data) are expressed in the head (80% alone in the brain), so it's turning out that expression isn't all that helpful for narrowing the search. But it does mean we've looked at images of gene expression for around 2000 genes.
And ORs are a good example of how what we think we know can inhibit our understanding. Here, e.g., are the expression results for olfactory receptor 66 (Olfr66) in a developing mouse (at embryonic day 14.5). Just to orient you if you're not used to looking at such images, it's a single front to back section, the snout halfway down the image and pointing to the left, and the tail at the bottom. The dark blue is a stain showing cells where the gene is expressed at this particular stage of development. It's no surprise to see it in the olfactory epithelium in the snout, but notice that it's also in the axial skeleton (vertebral column), probably in cartilage cells that will soon become bone.
What's it doing there? These are olfactory receptors! You don't smell with your backbone! In fact, a lot of ORs are known to be expressed outside the olfactory region, particularly in the testes, but also in the spleen, the thyroid, salivary glands, the uterus, the skin, and other tissues. A 2006 paper is of interest in this regard, not only because it documents non-olfactory related expression, but because of its title -- "Widespread ectopic expression of olfactory receptor genes". Ectopic expression, meaning expression where it's not supposed to be.
But it's only not supposed to be expressed in the axial skeleton because that's not where its name says it will be, not because Nature says so! People named these genes! And, there is some discussion in the paper about how ORs might be involved in chemotaxis of sperm as they try to reach and penetrate the egg -- how they direct their movement, based on chemicals in their environment. Which is equivalent to assuming they are essentially carrying out their olfactory function in the testes, where a different form of molecular reaction than odorant-detection is going on. But, what about in cartilage, in the image above? It's hard to imagine chemotaxis has anything to do with OR function here.
Well then, maybe it's an experimental artifact -- maybe the experiment picked up expression of a gene sort of like Olfr66, but not quite, along with Olfr66? Maybe. But, then we'd have to explain away all the expression studies showing non olfactory expression of many ORs, and it's rather unlikely that it's all due to experimental artifact. This is how our own assumptions constrain what we know or even want to know about the function of so many genes. Maybe Olfr66 has a function we don't yet understand. As do other ORs. And, by extension, so many other genes.
But calling unexpected expression 'ectopic', or naming genes based on only a single role, or in their involvement in disease, when they have other perfectly normal functions, are ways of building in assumptions that, once accepted, can keep us from recognizing that there's a lot we don't yet understand about genes.
Sparing you the gory details, we'll just say that the chromosomal intervals that one or more of the traits mapped to span 30% of the genome, and 10% of all coding genes. That's 2400 genes or so that could potentially be of interest in affecting head shape in just these particular mice, with the restricted genetic variation they have (because they are descendants of only two inbred parental mouse strains). That's a lot of genes to wade through to figure out which might be most likely to be involved in the traits we're looking at.
One way to prioritize candidate genes from such a study is to look for the genes in every interval that you know from prior work to be involved in your trait of interest. Or to identify genes in families that include genes involved in your trait of interest -- these would then be considered guilty by association.
But this means most genes don't have a fighting chance of being considered, because you don't happen to know anything about them, or because nobody knows anything about them, or because what's known about them only partially represents what they do.
To try to minimize this, many people automate the search, with programs that cull the genes that the literature indicates might be of interest, or that seem to be expressed where you want them to be. So, this might solve the problem of no one knowing everything about all genes, but it doesn't solve the problem of nothing being known about so many genes, or that there's only partial knowledge. And it doesn't solve the problem of having to tell the program what to look for, which means you're constraining it in the same way you would if you were doing the search by hand, looking for specific families of genes. Nor, of course, does it solve the problem of what's happening in all the non-coding DNA that flanks all those genes.
Thus, we decided that the least biased way to comb the data was to go through all the genes in all the intervals by hand. We're still making sense of all that, not least because we are hoping not to be constrained by the usual ideas about statistical significance, but we've learned some interesting things along the way.
For example, one of the intervals of interest is loaded with olfactory receptor (OR) genes. Olfactory receptors reside on the cell surface of olfactory receptor neurons, and are involved in odorant detection. ORs form the largest family of genes in many genomes -- about 1000 different genes -- and they cluster in sets of genes in various locations on a number of chromosomes. ORs have a distinctive expression pattern, with only one expressed per neuron in the tissue lining the nose, where they each are sensitive to particular aspects of molecules the animal inhales, and hopes to smell. How expression of the remaining 999 genes in each cell is blocked is still not known.
ORs are an interesting example of something we've blogged about before, but that continually surprises us. One of the ways we're evaluating the possible role of all these genes in development of the traits we're looking at is to look at where they are expressed in the developing embryo. We initially thought this would be helpful for narrowing the search, but it turns out that about 95% or even more of genes (for which there are expression data) are expressed in the head (80% alone in the brain), so it's turning out that expression isn't all that helpful for narrowing the search. But it does mean we've looked at images of gene expression for around 2000 genes.
![]() |
| Olfr66, GenePaint, E14.5 |
What's it doing there? These are olfactory receptors! You don't smell with your backbone! In fact, a lot of ORs are known to be expressed outside the olfactory region, particularly in the testes, but also in the spleen, the thyroid, salivary glands, the uterus, the skin, and other tissues. A 2006 paper is of interest in this regard, not only because it documents non-olfactory related expression, but because of its title -- "Widespread ectopic expression of olfactory receptor genes". Ectopic expression, meaning expression where it's not supposed to be.
But it's only not supposed to be expressed in the axial skeleton because that's not where its name says it will be, not because Nature says so! People named these genes! And, there is some discussion in the paper about how ORs might be involved in chemotaxis of sperm as they try to reach and penetrate the egg -- how they direct their movement, based on chemicals in their environment. Which is equivalent to assuming they are essentially carrying out their olfactory function in the testes, where a different form of molecular reaction than odorant-detection is going on. But, what about in cartilage, in the image above? It's hard to imagine chemotaxis has anything to do with OR function here.
Well then, maybe it's an experimental artifact -- maybe the experiment picked up expression of a gene sort of like Olfr66, but not quite, along with Olfr66? Maybe. But, then we'd have to explain away all the expression studies showing non olfactory expression of many ORs, and it's rather unlikely that it's all due to experimental artifact. This is how our own assumptions constrain what we know or even want to know about the function of so many genes. Maybe Olfr66 has a function we don't yet understand. As do other ORs. And, by extension, so many other genes.
But calling unexpected expression 'ectopic', or naming genes based on only a single role, or in their involvement in disease, when they have other perfectly normal functions, are ways of building in assumptions that, once accepted, can keep us from recognizing that there's a lot we don't yet understand about genes.
Thursday, February 2, 2012
Triumph of the Darwinian Method, continued: genetics as function, function as history
By
Ken Weiss
So a powerful reason that the universal approach to biological questions was inspired by Darwin's theory of evolution (descent with modification), is that history leaves a trace in gene sequences, and gene sequences reveal history. Even a sequence that, statistically, is random by itself, is very non-random when compared to other sequences. That's because all DNA sequences are, in the history of life sense, related. They may be statistically random along the chain of nucleotides, but they are very much not random when compared to each other.
The same is true when we look into the structure of a DNA sequence, where once again history shows why what may seem random is anything but. Again, it is the Darwinian method, and the assumption of common ancestry, that makes it possible to understand this.
A century of work has shown us that DNA is related to the functions that go on in cells, in ways that are essentially common to all aspects of life. This has to do with how those functions are encoded in DNA, and once we know how to read the code, we can also see both persuasive evidence for natural and other forms of selection, but also that make sense strictly in light of life as history.
Here is a sequence of part of a human gene: (gastrin):
We picked this as a nice figure showing what we want, from a human genetics textbook by G. Moroni, 2001. The nucleotide sequence is in black, and the amino acid sequence, or protein code, is in brown, written under the nucleotide code. By itself, the DNA sequence appears just to be random nucleotides scrambled in a row. But here are labeled various parts that make it a gene: where messenger RNA is transcribed and the amino acids it codes for (brown), where the regulatory proteins bind to make this happen the TTATA in color), and a signal for where a string of A's will be attached (AAATAAA). These types of features, and others not shown, can be identified nowadays just by analyzing a naked DNA sequence.
You can go to what is called a genome 'browser' (the link is to the UCSC genome browser) and see the many structural and functional elements of DNA for any gene you can name in any species for which we have the DNA sequence. Here it is for the gastrin gene. The top lines are the sequence location, then the gene with its protein-coding parts (dark boxes), connected by a thin blue line for the noncoding parts (called introns), showing where the gene code begins and ends, and belowthat a grey bar whose darkness reflects the degree of sequence conservation, or similarity with other species, and below that black boxes showing the location of various short sequence elements that are 'motifs' found exactly or nearly repeated in many places in the human and other related species' genomes:
A browser can show many other features, but a snapshot showing them all would utterly clog this post (but, you can see the results for the gastrin gene here.) Note for example that the coding regions are areas where the conservation bars are darkest. That means that these areas are also very similar in sequence in other species for the 'same' gene. "Conservation", "same", "coding regions", etc., are all terms that refer to aspects of DNA we identify by comparison or that are similar because of history--shared ancestry. And the reason some areas in and around the gene have varied or changed less during that history, that is, that are conserved in sequence, we know from all sorts of data, is that they are functional parts of the DNA. As Darwin's theory would hold, they are important enough that mutational change in the DNA sequence probably did not work as well, and did not reproduce as well, as in less important areas: what Darwin called natural selection.
We know these things because, after many decades of research, we have learned that the general features, like the code for amino acids that make up the protein (here, gastrin), are essentially universal: that means that we can also find or identify these functional elements by various kinds of experiment, but the methods themselves derive from what's important: our ability to compare sequences of genes from any species or among species, and to compare sequences from the same gene in different species. Genes come and go, so we find the 'same' gene in sets of species that have diverged since the gene's origins.
Because of this, because of shared history and common ancestry, what by itself may seem to be an entirely random sequence of nucleotides, becomes understandable as an entirely non-random sequence when it comes to explaining what it does and why it exists. In fact, non-functional parts of the sequence may accumulate mutations randomly without selective constraint, but even they are transmitted faithfully (with occasional mutational change), and hence bear a trace of their history.
From the point of view of the Darwinian method, that is, the aspects of the scientific method that Darwin's insights set rolling, these tools that are related to life's nature as shared history are fundamental to our modern understanding of life. That is the triumph of the method.
Thus, a truly random sequence (such as in our previous post) would not only be statistically unpatterned itself (which is what 'random' means), but it would have none of the known functional structures that we know the evolution of life has produced in all its creatures, large and small. That's why we'd be truly spooked by an actual sequence that not only didn't fit anywhere on the known tree of life's creatures, but also didn't show any elements that we are know that evolution has made fundamental to the nature of living organisms.
There are still many things to be debated, about how to interpret various of these DNA-sequence factors, but not their nature as products of history. One can debate the nature and role of natural selection, or the strength of effect of individual sequence variants on the traits (like presence of disease) of the organism carrying them. We write all the time here on MT about over-stated claims about genetic causation and how easy it is to concoct adaptive Just-So stories. The Darwinian method is so powerful that it lures scientists to excess, to uncritical acceptance of scenarios and claims that go beyond what really is scientifically legitimate--in some sense, just as Adam Sedgwick accused Darwin of doing: assuming a theory which no facts could erode. That's ideology, not science. Scientists may indulge in such story-invention more, if anything, than Darwin himself did, so strong has a simplistic selectionist belief become in many quarters, either because a notion of Darwinian theory has been bought uncritically, or because as in many public arenas, education in biology has been a kind of Darwinian indoctrination.
Iron-clad theories are self-fulfilling, and too many in science and the media, buy into facile stories. There are many ways for differential proliferation of genetic variation and the traits it affects to occur or for variation to be distributed around the earth, of which natural selection is only one. Temptation to invent stories notwithstanding, however, the aspects of the Darwinian method that we've tried to explain are so pervasive that there is no serious doubt about the historical nature of life on earth, which applies to all the competing explanations, and the fact that to persist or proliferate systematically, DNA sequence elements must have tolerable, or advantageous, function. It should not have to be pointed out that this does not include creationist explanations, that do not require nor predict the kind of trace of history that we clearly observe. Biologists are not arguing about the fact of life as history nor that that is why DNA sequences are nonrandom in ways they are nonrandom.
Because evolution happened in the past, we must triangulate our approaches to understand it. This is why a mixture of repeated observation--induction--has been fundamental from Darwin to us today, and yet why deduction from a theory worked over the years by observation of DNA, leads us to predict what we'd find in some newly discovered sequence. And for the same reasons, why neither induction nor deduction would help us explain a truly novel DNA sequence. Scientific reasoning, especially when most things can't be proven by experiment and result from past events, must be a kind of social mix of various people taking various approaches, and combining their findings.
At the core of this mix is what is known as the Darwinian method. In a profound sense, whatever we in biology argue about, it isn't that!
The same is true when we look into the structure of a DNA sequence, where once again history shows why what may seem random is anything but. Again, it is the Darwinian method, and the assumption of common ancestry, that makes it possible to understand this.
A century of work has shown us that DNA is related to the functions that go on in cells, in ways that are essentially common to all aspects of life. This has to do with how those functions are encoded in DNA, and once we know how to read the code, we can also see both persuasive evidence for natural and other forms of selection, but also that make sense strictly in light of life as history.
Here is a sequence of part of a human gene: (gastrin):
We picked this as a nice figure showing what we want, from a human genetics textbook by G. Moroni, 2001. The nucleotide sequence is in black, and the amino acid sequence, or protein code, is in brown, written under the nucleotide code. By itself, the DNA sequence appears just to be random nucleotides scrambled in a row. But here are labeled various parts that make it a gene: where messenger RNA is transcribed and the amino acids it codes for (brown), where the regulatory proteins bind to make this happen the TTATA in color), and a signal for where a string of A's will be attached (AAATAAA). These types of features, and others not shown, can be identified nowadays just by analyzing a naked DNA sequence.
You can go to what is called a genome 'browser' (the link is to the UCSC genome browser) and see the many structural and functional elements of DNA for any gene you can name in any species for which we have the DNA sequence. Here it is for the gastrin gene. The top lines are the sequence location, then the gene with its protein-coding parts (dark boxes), connected by a thin blue line for the noncoding parts (called introns), showing where the gene code begins and ends, and belowthat a grey bar whose darkness reflects the degree of sequence conservation, or similarity with other species, and below that black boxes showing the location of various short sequence elements that are 'motifs' found exactly or nearly repeated in many places in the human and other related species' genomes:
A browser can show many other features, but a snapshot showing them all would utterly clog this post (but, you can see the results for the gastrin gene here.) Note for example that the coding regions are areas where the conservation bars are darkest. That means that these areas are also very similar in sequence in other species for the 'same' gene. "Conservation", "same", "coding regions", etc., are all terms that refer to aspects of DNA we identify by comparison or that are similar because of history--shared ancestry. And the reason some areas in and around the gene have varied or changed less during that history, that is, that are conserved in sequence, we know from all sorts of data, is that they are functional parts of the DNA. As Darwin's theory would hold, they are important enough that mutational change in the DNA sequence probably did not work as well, and did not reproduce as well, as in less important areas: what Darwin called natural selection.
We know these things because, after many decades of research, we have learned that the general features, like the code for amino acids that make up the protein (here, gastrin), are essentially universal: that means that we can also find or identify these functional elements by various kinds of experiment, but the methods themselves derive from what's important: our ability to compare sequences of genes from any species or among species, and to compare sequences from the same gene in different species. Genes come and go, so we find the 'same' gene in sets of species that have diverged since the gene's origins.
Because of this, because of shared history and common ancestry, what by itself may seem to be an entirely random sequence of nucleotides, becomes understandable as an entirely non-random sequence when it comes to explaining what it does and why it exists. In fact, non-functional parts of the sequence may accumulate mutations randomly without selective constraint, but even they are transmitted faithfully (with occasional mutational change), and hence bear a trace of their history.
From the point of view of the Darwinian method, that is, the aspects of the scientific method that Darwin's insights set rolling, these tools that are related to life's nature as shared history are fundamental to our modern understanding of life. That is the triumph of the method.
Thus, a truly random sequence (such as in our previous post) would not only be statistically unpatterned itself (which is what 'random' means), but it would have none of the known functional structures that we know the evolution of life has produced in all its creatures, large and small. That's why we'd be truly spooked by an actual sequence that not only didn't fit anywhere on the known tree of life's creatures, but also didn't show any elements that we are know that evolution has made fundamental to the nature of living organisms.
There are still many things to be debated, about how to interpret various of these DNA-sequence factors, but not their nature as products of history. One can debate the nature and role of natural selection, or the strength of effect of individual sequence variants on the traits (like presence of disease) of the organism carrying them. We write all the time here on MT about over-stated claims about genetic causation and how easy it is to concoct adaptive Just-So stories. The Darwinian method is so powerful that it lures scientists to excess, to uncritical acceptance of scenarios and claims that go beyond what really is scientifically legitimate--in some sense, just as Adam Sedgwick accused Darwin of doing: assuming a theory which no facts could erode. That's ideology, not science. Scientists may indulge in such story-invention more, if anything, than Darwin himself did, so strong has a simplistic selectionist belief become in many quarters, either because a notion of Darwinian theory has been bought uncritically, or because as in many public arenas, education in biology has been a kind of Darwinian indoctrination.
Iron-clad theories are self-fulfilling, and too many in science and the media, buy into facile stories. There are many ways for differential proliferation of genetic variation and the traits it affects to occur or for variation to be distributed around the earth, of which natural selection is only one. Temptation to invent stories notwithstanding, however, the aspects of the Darwinian method that we've tried to explain are so pervasive that there is no serious doubt about the historical nature of life on earth, which applies to all the competing explanations, and the fact that to persist or proliferate systematically, DNA sequence elements must have tolerable, or advantageous, function. It should not have to be pointed out that this does not include creationist explanations, that do not require nor predict the kind of trace of history that we clearly observe. Biologists are not arguing about the fact of life as history nor that that is why DNA sequences are nonrandom in ways they are nonrandom.
Because evolution happened in the past, we must triangulate our approaches to understand it. This is why a mixture of repeated observation--induction--has been fundamental from Darwin to us today, and yet why deduction from a theory worked over the years by observation of DNA, leads us to predict what we'd find in some newly discovered sequence. And for the same reasons, why neither induction nor deduction would help us explain a truly novel DNA sequence. Scientific reasoning, especially when most things can't be proven by experiment and result from past events, must be a kind of social mix of various people taking various approaches, and combining their findings.
At the core of this mix is what is known as the Darwinian method. In a profound sense, whatever we in biology argue about, it isn't that!
Wednesday, February 1, 2012
Triumph of the Darwinian Method, continued: genetics as history
By
Ken Weiss
Is a DNA sequence 'random'? How would one know? There are several tests for randomness that one can do on a DNA sequence, however it was derived -- there's a discussion of this issue here. For example, one can go down the sequence one nucleotide at a time and ask if that nucleotide can predict the next one in line. One way is to ask whether the next ones are A, C, G, and T each 25% of the time. That would mean that whatever the current nucleotide is, the next one is just unpredictable. Then one can ask the same for the nucleotide 2, 4, ..., etc 12,287 positions down the row. With some exceptions, such tests for predictability or periodicity would fail. That means, the sequence is random!
That is very weird, since DNA is responsible for organisms, and organisms seem to be anything but random! Or, could it be that at some higher level, organisms in this earth are, in some profound ways, 'random' in structure or behavior etc.?
If I give you a DNA sequence, and you go into some program on the web (of which there are many) that searches all known DNA sequences to see which are closest to the test sequence, the expectation is that it will be an exact or near match to something known. We have DNA sequence data from most branches of life, so we'd 'hit' a known species, or something similar in sequence. That would pin our test sequence on the tree of relationships which to a Darwinian is a tree of life, and the key aspect of that is that the tree is the result of a history.
If this seems obvious, it is at the same time a profound reflection of the only convincing hypothesis about life, and here our method assumes a 'tree' and fits a sequence on it, and the only reason we can do this is because life is history. If the sequence were from an individual already sequenced, there would be a complete match. If it were from that individual's sibling, there would be a very close match. If from the same population within a species, the similarity would be less but still very strong. If from a different but 'related' species (say, two different forms of cat) again similarities would be clear--way more than with a bird or lizard or maple sequence. This is only because of history, and it is only that fact that allows us to make sense of DNA similarities.
Suppose you were given the following sequence:
Now, go search the data bases for it, to see what known sequence it's closest to. Here is what you'll find if you use a common tool, called BLAST, for comparing sequences: "No significant similarity found."
Now, since we have sequence from basically every branch of life (though, at present, not whole genome sequence, to be sure, but this is a practical but not conceptual problem in our current context), how can our test sequence not fit the tree? If it really were unrelated to anything known, nor within the tree of known sequence, we would either have a sequence totally made up (which this one was!), or from a wholly unknown branch of life. We have so much data at present, that such a result would be very spooky and unlikely. Mutations arise randomly in DNA, relative to their effect on the organism, but even this kind of randomness is inherited, which is why even sequences that, by various statistical tests, seem to be random assemblages of nucleotides, fall into historical relationships with each other. They may be 'random' on their own, but not to each other.
A sequence truly unrelated to any other could deeply threaten our very Darwinian foundations! That is how strong and well-supported his hypothesis about life is. If that is a failure of induction, then so be it. We would go to very great lengths to find other explanations for our mysterious sequence, before we would even begin to question the hypothesis of evolution! Would it be less weird than current allegations of meteorite structures, to suggest that such a sequence came from Mars? Would it resuscitate arguments about spontaneous generation (see earlier post on this)?
And in another way such a sequence would be at least as profound, or perhaps much more profound than just a missing relationship. That is because, on its own DNA sequence can seem, statistically, to be a random string of nucleotides--once we know about history, and have experimental data (which we do, in profusion) we can see how utterly nonrandom DNA sequences are when it comes to what they do--that cannot by itself be 'read' off from the sequence alone, without this information. And this information is essentially connected to Darwinian ideas.
This gets us to consider not just the fact of the tree of life's history, but the functional roles of DNA, and Darwin's other idea, that the tree is built by natural selection. Comparative DNA sequences, viewed through the Darwinian method, also say something about that as well. That is for next time in this series....
That is very weird, since DNA is responsible for organisms, and organisms seem to be anything but random! Or, could it be that at some higher level, organisms in this earth are, in some profound ways, 'random' in structure or behavior etc.?
If I give you a DNA sequence, and you go into some program on the web (of which there are many) that searches all known DNA sequences to see which are closest to the test sequence, the expectation is that it will be an exact or near match to something known. We have DNA sequence data from most branches of life, so we'd 'hit' a known species, or something similar in sequence. That would pin our test sequence on the tree of relationships which to a Darwinian is a tree of life, and the key aspect of that is that the tree is the result of a history.
If this seems obvious, it is at the same time a profound reflection of the only convincing hypothesis about life, and here our method assumes a 'tree' and fits a sequence on it, and the only reason we can do this is because life is history. If the sequence were from an individual already sequenced, there would be a complete match. If it were from that individual's sibling, there would be a very close match. If from the same population within a species, the similarity would be less but still very strong. If from a different but 'related' species (say, two different forms of cat) again similarities would be clear--way more than with a bird or lizard or maple sequence. This is only because of history, and it is only that fact that allows us to make sense of DNA similarities.
Suppose you were given the following sequence:
ACGTCCAATCTGGGGTAAACCCGAGATCTGAGGCCTACCTGCAATTTCGGCCACACACAGGGTGTTACCCCGACTTCAGGGCA
Now, go search the data bases for it, to see what known sequence it's closest to. Here is what you'll find if you use a common tool, called BLAST, for comparing sequences: "No significant similarity found."
Now, since we have sequence from basically every branch of life (though, at present, not whole genome sequence, to be sure, but this is a practical but not conceptual problem in our current context), how can our test sequence not fit the tree? If it really were unrelated to anything known, nor within the tree of known sequence, we would either have a sequence totally made up (which this one was!), or from a wholly unknown branch of life. We have so much data at present, that such a result would be very spooky and unlikely. Mutations arise randomly in DNA, relative to their effect on the organism, but even this kind of randomness is inherited, which is why even sequences that, by various statistical tests, seem to be random assemblages of nucleotides, fall into historical relationships with each other. They may be 'random' on their own, but not to each other.
A sequence truly unrelated to any other could deeply threaten our very Darwinian foundations! That is how strong and well-supported his hypothesis about life is. If that is a failure of induction, then so be it. We would go to very great lengths to find other explanations for our mysterious sequence, before we would even begin to question the hypothesis of evolution! Would it be less weird than current allegations of meteorite structures, to suggest that such a sequence came from Mars? Would it resuscitate arguments about spontaneous generation (see earlier post on this)?
And in another way such a sequence would be at least as profound, or perhaps much more profound than just a missing relationship. That is because, on its own DNA sequence can seem, statistically, to be a random string of nucleotides--once we know about history, and have experimental data (which we do, in profusion) we can see how utterly nonrandom DNA sequences are when it comes to what they do--that cannot by itself be 'read' off from the sequence alone, without this information. And this information is essentially connected to Darwinian ideas.
This gets us to consider not just the fact of the tree of life's history, but the functional roles of DNA, and Darwin's other idea, that the tree is built by natural selection. Comparative DNA sequences, viewed through the Darwinian method, also say something about that as well. That is for next time in this series....
Subscribe to:
Posts (Atom)


