Showing posts with label 'junk DNA'. Show all posts
Showing posts with label 'junk DNA'. Show all posts

Wednesday, July 20, 2011

Epistemology and genetics: does pervasive transcription happen?

Pervasive transcription
RNA basepairing
It has been known for some time that only 1-5% or so of the mammalian genome actually codes for proteins. It does that by 'transcribing' a copy of a DNA region traditionally called a 'gene' into messenger RNA (mRNA) that was in turn 'translated' into an amino acid sequence (protein, in common parlance).  According to this established theory known as the Central Dogma of Biology (the problems with which we blogged about here), a gene had a transcription start and stop sites, specific sequence elements in DNA from which this canonical (regular) structure of a 'gene' could be identified from DNA sequence (but see below for problems with the definition of a gene).  From that, we got the estimate that our genome contains roughly 25,000 genes.

Not so long ago, the remainder of the genome, the non-coding DNA, was called 'junk DNA' because it wasn't known what, if any function it had.  Then, some of it, generally short regions near genes, was discovered to be sequence elements that, when bound by various proteins in a cell, cause the nearby gene to be transcribed.  So some of that DNA had a function after all.

But then, with various projects such as in 2007, the ENCODE project, a multi-institutional exploration of the function of all aspects of the genome in great detail, it was found that the majority of all the DNA in the genome is in fact transcribed into RNA in a process called 'pervasive transcription'.  As the ENCODE project reported,
First, our studies provide convincing evidence that the genome is pervasively transcribed, such that the majority of its bases can be found in primary transcripts, including non-protein-coding transcripts, and those that extensively overlap one another.  
What all this RNA did wasn't yet known, because it didn't have the structure that would be translated into protein, but that pervasive transcription happened seemed clear.  Some functions were subsequently discovered, such as 'microRNA' that codes for sequences complementary to mRNA that are used to inhibit the mRNA's translation into protein.  There were types of RNA that had their own functions within the classical idea of a gene, even if not translated into protein (these included ribosomal RNA and transfer RNA).

Or maybe not...
However, the pervasive transcription idea was challenged by van Bakel et al. in PLoS Biology, who said that most of the low-level transcription described by ENCODE was in fact experimental artifact or just meaningless noise, error without a function.

Oh, but it does!
Now, in a tit for tat, just last week a refutation of this refutation, a paper called "The Reality of Pervasive Transcription," appeared in PLoS Biology (on July 12).  Clark et al. confirm that pervasive transcription does in fact happen, and take issue with the van Bakel et al. results.
...we present an evaluation of the analysis and conclusions of van Bakel et al. compared to those of others and show that (1) the existence of pervasive transcription is supported by multiple independent techniques; (2) re-analysis of the van Bakel et al. tiling arrays shows that their results are atypical compared to those of ENCODE and lack independent validation; and (3) the RNA sequencing dataset used by van Bakel et al. suffered from insufficient sequencing depth and poor transcript assembly, compromising their ability to detect the less abundant transcripts outside of protein-coding genes. We conclude that the totality of the evidence strongly supports pervasive transcription of mammalian genomes, although the biological significance of many novel coding and noncoding transcripts remains to be explored.
Clark et al. question van Bakel et al.'s molecular technique as well as their 'logic and analysis'.
These may be summarized as (1) insufficient sequencing depth and breadth and poor transcript assembly, together with the sampling problems that arise as a consequence of the domination of sequence data by highly expressed transcripts; compounded by (2) the dismissal of transcripts derived from introns; (3) a lack of consideration of non-polyadenylated transcripts; (4) an inability to discriminate antisense transcripts; and (5) the questionable assertion that rarer RNAs are not genuine and/or functional transcripts.
They go into detail in the paper about how and why these are serious problems, and conclude that van Bakel et al.'s results are 'atypical' for tiling array data, their tissue samples were not sufficiently extensive, and that pervasive transcription is being detected by a variety of experimental methods.  (Tiling refers to the fact that sequencing is done one stretch at a time, and long stretches and their location on chromosomes from which they were copied is done by finding overlapping ends of these short stretches, that show how they 'tile' together relative to the chromosome as a whole.)

No, no, no! 
And finally (to date), van Bakel et al. respond, also in the July 12 PLoS Biology.
Clark et al. criticize several aspects of our study, and specifically challenge our assertion that the degree of pervasive transcription has previously been overstated. We disagree with much of their reasoning and their interpretation of our work. For example, many of our conclusions are based on overall sequence read distributions, while Clark et al. focus on transcript units and seqfrags (sets of overlapping reads). A key point is that one can derive a robust estimate of the relative amounts of different transcript types without having a complete reconstruction of every single transcript.
So, they defend their methods and interpretation and conclude that "a compelling wealth of evidence now supports our statement that 'the genome is not as pervasively transcribed as previously reported.'"

An epistemological challenge -- what do we know and how do we know it?
What is going on here?  Is most of the genome transcribed or isn't it?  We are intrigued not so much by the details of the argument but by the epistemology, how these researchers know what they think they know.  In the past, of course, a scientist had an hypothesis and set about testing it, and drew conclusions about the hypothesis based on his or her experimental results.  The results were then replicated, or not, by other scientists and the hypothesis accepted or not.  The theory of gravity allowed many kinds of predictions to be made because of its specificity and universality, for example, as do the theories of chemistry in relation to how atoms interact to form molecules.  This is a simplified description of course, but it's more accurate than not.

Molecular genetics these days is by and large not hypothesis driven, but technology driven.  There is no theory of what we should or must find in DNA.  Indeed, there is hardly a rule that, when we look closely, is not routinely violated.  Evolution assembles things in a haphazard, largely chance-driven way.  Rather than testing a hypothesis, masses of data are collected in blanket coverage fashion, mined in the hopes that a meaning will somehow arise.  That this is how much of genetics is now done is evident in this debate.  As first reported by ENCODE, complete genome sequencing seemed to be yielding a lot of DNA that was transcribed from other than protein coding regions, and so they speculated as to what that could mean.  Their speculation wasn't based on anything then known about DNA, or theory, but on results, results produced by then-current technology. That is the reason--and the only reason--that they were surprising.

And van Bakel et al. disagreed, again based on results they were getting and interpreting from their use of the technology.  Then Clark et al. disagreed with van Bakel et al.'s use and interpretation of molecular methods, and described their results as 'atypical'.  And both 'sides' of this debate attempt to strengthen their claim by stating that many others are confirming their findings.  And this is surely not the end of this debate.

We blogged last week about 'technology-driven science', suggesting that often the technology isn't ready for prime time, and thus that many errors that will be hard to ferret out are laced throughout these huge genetic databases that everyone is mining for meaning.  When interpretation of the findings is based on nothing more than whose sequencing methods are better, or whether or not a tissue or organism was sequenced enough times to be credible, or which sequencing platforms seem 'best' for a particular usage (by some criterion) -- meaning that the errors  of the platform are least disturbing to its objective -- rather than on any basic biological theory or prior knowledge, we're left with the curious problem of having no real way to know who's right.  If everyone's using the same methods, it can't be based on whose results are replicated most. If we haven't a clue what is there, even knowing what it means to be 'right' is a challenge!

But we don't even know what a gene is anymore!
These days, even the definition of a gene, our supposedly fundamental inherited functional unit, is in shreds and the suggested definitions in recent years have been almost laughably vague and nondescript.  Try this, by G. Pesole in a 2008 paper in the journal Gene, out for size, if you think we're kidding:
Gene: A discrete genomic region whose transcription is regulated by one or more promoters and distal regulatory elements and which contains the information for the synthesis of functional proteins or non-coding RNAs, related by the sharing of a portion of genetic information at the level of the ultimate products (proteins or RNAs).
Or how about this one, from 2007 (Gerstein et al., Genome Research):
Gene: A union of genomic sequences encoding a coherent set of potentially overlapping functional products.
Not very helpful!!

These definitions show how much we're in trouble.  Ken attended a whole conference on the 'concept of the gene', at the Santa Fe Institute, in 2009, but there was no simple consensus except that, well, there is no consensus and essentially no definition! 

The good old days
When microscopes and telescopes were invented, they opened up whole new previously unknown worlds to scientists.  If anyone doubted the existence of paramecium they had only to peer through the eyepieces of this strange new instrument to be convinced.  When cold fusion was announced by scientists in Utah some years ago, the reaction of disbelief was based on what was known about how atoms work.  That is, there were criteria, and theoretical frameworks, for evaluating the evidence.  Yes, there has always been a learning curve; when Galileo looked at the moon or planets, various optical aberrations actually were misleading, until they were worked out.

And now....
But they were building in an era when fitting theory was the objective.  In biology, we're in new territory here.

Friday, August 20, 2010

Genes playing possum!

A story by Gina Kolata in the NY Times reports a muscular disorder that is due to what was considerd a 'dead' gene, or 'junk' DNA.  A dead gene, usually called a pseudogene, is a DNA sequence derived from an incomplete copy of a functionally active gene, or a gene that was once actively used but whose transcription regulatory sequence mutated away.  Or a gene could suffer a mutation in its coding sequence that makes the resulting protein not work.

The smugness with which DNA sequence in between regular protein-coding genes was called 'junk' DNA is rapidly fading.  Much of our DNA has no known function, but even there we have evidence that there may be function--for example, some regions of such DNA have sequence that is conserved (basically the same) among species that haven't shared a common ancestor for many millions or more years.  If it had no function maintained by natural selection, why hasn't mutation simply erased the similarity among species of that sequence?

To find that a pseudogene region that was known to be transcribed into RNA can actually interfere with normal processes is interesting and biomedically important.  It's worthy of a story in the Times (and Kolata is a worthy person to write it) (see, we don't just criticize the popular science media!).

We can't make generalizations from this about 'junk' DNA however.  The degree to which a bit of non-coding DNA has some function, of some sort, is difficult to know simply because proving 'no' function is virtually impossible.  That's why evolutionary conservation is a persuasive indicator that something's still usefully active.

Francis Collins is quoted as saying that this is interesting and complex in its mechanism, which is not yet understood.  He rightly says that in genetics, whatever can go wrong will go wrong--a principle Ken called the Rusty Rule of life in his 1990 book on disease genes, because evolutionarily we know this had to be so, since mutation can strike anywhere in the genome.

There may be DNA with no function, not even spacing-function to keep other functional elements some proper distance apart.  Perhaps it could be called 'junk'.  But right now the problem we face in non-coding DNA is the opposite: there is so much of it, it's hard to understand how natural selection could be maintaining it.

When a sequence variant has very little function--and most of this DNA seems clearly in that category--then we expect genetic drift (chance) to determine how the frequency of the variant will change over time.  In that case, deep evolutionary conservation is not to be expected or at least should be less than that of really functional DNA.  But even saying 'less' is problematic, because we need a baseline for the rate at which variation will accumulate in truly nonfunctional DNA.

But if we can't be sure of what's really nonfunctional, where is our baseline?  We can try theory, try some experimental things (like watching bacteria over thousands of generations in a lab), but it's not easy to know.

Ironically, a common bit of DNA to use for that baseline is--you may have guessed it--pseudogenes!  Because a dead gene has no function!  Well, the Times example shows that some, at least, do have a function, and it is very possible that this disorder, though not lethal, could affect the reproductive success of those unfortunate enough to carry it.

Life is always playing tricks on it.  What you see may not be what you get.  Something may look pseudo, but only be playing possum.

Friday, June 4, 2010

A TUF nut to translate? The function of 'noncoding' RNA

One of the fundamental principles of life, which we've written about from time to time and include a lot about in our book, is chance. It's everywhere, from the randomness of which genes we inherit from each parent, to being the unlucky fish that get caught in the shark's maw as it sweeps through a school with an open mouth.

Another aspect of chance is sloppiness, which is also found everywhere. Echolocating bats are notably imprecise in their discrimination among prey, error-prone transcription of DNA into RNA is routine--it has been estimated, in fact, that perhaps 3% of a cell's energy is used to correct transcription errors.*

The environment presents us with the unpredictable, too, in the form of phenomena like hail storms or hurricanes, unusual heat or cold, volcanoes, etc., as well as the regular, predictable fluctuation of the seasons. (Of course, as Ken is a former meteorologist, he's inserting the self-defense caveat here that weather forecasting is not entirely unpredictable in the short run, though many details are, and weather may never be predictable very many days in advance--but then, for organisms other than humans, and then only recently, this is irrelevant).

Chance is so ubiquitous that being able to adapt to chance effects--within limits, of course; no amount of adaptiveness will help that fish escape being that marauding shark's dinner--is a characteristic of life that had to have evolved very early because all organisms can do it to some degree. The lineages that couldn't disappeared long ago.

Errors and chance are found at the molecular level, too.  It has been known for decades that only a small fraction of DNA is transcribed into messenger RNA and translated into proteins, and it was thought that the 98% of the genome that wasn't transcribed was 'junk', detritus from evolutionary trial and error. But recently, unexpected transcripts from non-coding DNA have been reported by a number of labs, leading to speculation about the actual role of all that 'junk DNA'. A recent paper in PLoS Biology by van Bakel et al., accompanied by a commentary, addresses this question. van Bakel et al. describe these excess transcripts this way:
Dubbed transcriptional “dark matter”, the “hidden” transcriptome, or transcripts of unknown function (TUFs), the exact nature of much of this additional transcription is unclear, but it has been presumed to comprise a combination of novel protein coding transcripts, extensions of existing transcripts, noncoding RNAs (ncRNAs), antisense transcripts, and biological or experimental background. Determining the relative contributions of each of these potential sources is important for understanding the nature and possible biological function of transcriptional dark matter.
To address the question of what these are, van Bakel et al. compared the usual method for identifying transcripts (tiling arrays) to a "single- and paired-end RNA-Seq" method, and found that the RNA-Seq method identified many fewer unexpected, or 'dark matter' transcripts (reminding us that high-powered technology can be error-prone, too). Most of the transcripts were identifably from intronic regions, suggesting that they were perhaps "fragments of pre-mRNAs", and were associated with open chromatin, that is, segments of DNA that are open for business, ready to be transcribed.
We conclude that analysis of data from tiling arrays leads to vast overestimates of the proportion of transcriptional dark matter. However, the mammalian transcriptome does contain thousands of unannotated transcripts, exons, promoters, and termination sites.
That is, they still found enough unidentified stuff to write home about, maybe 2% of the transcripts they analyzed. While that is considerably less than the dark matter that's given rise to so much speculation, it still suggests that even after 3 billion years, DNA copying enzymes make mistakes. As Richard Robinson says in the PLoS commentary,
The emerging picture of RNA polymerase is of an inherently imprecise, not to say promiscuous, copyist, one whose output includes some mistakes along with lots of valuable product. In this view, most dark matter transcripts are not signals emerging from a hidden universe within the genome, but instead simply the noise emitted by a busy machine.
If these latest results bear out, it will mean that at least some of the enthusiasm in recent years for the idea that there is a huge unknown realm of DNA functions uninvolved with direct protein expression may not be completely warranted, and that our standard ('old-fashioned'?) theory, including the error-proneness of the system, may be pretty accurate after all. 

-----------------
*Kurland, C., and J. Gallant. (1996). "Errors of heterologous protein expression." Current Opinion in Biotechnology 7(5):489-493.