Showing posts with label gene regulation. Show all posts
Showing posts with label gene regulation. Show all posts

Friday, June 5, 2015

Making a genomic Skype call

There is an interesting and, if we understand it adequately, important paper by Mifsud and others, in the June issue of Nature Genetics.  This relates to how genes are used by a cell, and how different parts of DNA functions are coordinated.

Regulatory sites are parts of the genome containing usually short sequences that are used to control when genes are used.  They can help them to be expressed, repress their expression, or prevent (insulate) one gene from being expressed when some other nearby gene has already been turned on.  Regulatory sites for a given gene are usually numerous, and can be upstream, internal to, or downstream of the gene.

As currently known, regulatory sites are usually found close to the gene they regulate, but not always.  There is currently only very fragmentary knowledge of regulatory sites and their location, but it's clear that they follow no general rule.  Additionally, many GWAS mapping studies have found 'hits', that is, DNA regions whose variation is associated with some trait, where the regions are not in or near any actual gene (that is, protein-coding region).  Further, many if not most GWAS hits for complex disease that have been confirmed affect regulation rather than the protein code itself.  This makes sense since most genes have multiple functions, and changing the coded protein's structure could affect many different traits and be quite damaging.  Altering its regulation will typically only affect one or some of its uses, and hence typically will be less detrimental.  That's the idea, at least.

Regulatory sites and their associated gene's transcription start sites (where RNA begins to be read off the DNA) are usually close together (or brought close together) because a complex of proteins is required to assemble at the start site and start the transcription process.  But how 'close'?

Various techniques have been developed to identify stretches of chromosomes that are physically close to each other in the nucleus of cells, usually from some cell culture or other source.  In short-hand, these are called Hi-C assays (there are various ways to do these).  The juxtaposed bits of chromosome are isolated from the cells, and then sequenced, and the sequences aligned to the human genome reference sequence to see where they are.  The analysis thus shows what parts of DNA are physically close in a given cellular context or cell type.  Remember that chromosomes insides cells are 3 dimensional structures, not just linear stretches of DNA.

The new paper uses a technique to identify parts of DNA that are where transcription starts (called 'promoter' sites) and regulatory sites ('enhancers', or other terms).  With this information, functional analysis can be done.  The new paper by Mifsud et al. looks at this issue.  Here is a figure from the paper that shows some of the points (I labeled some features for you):
Long range regulation can even skip over active genes.  From Misfud paper (modified to show features)
The authors use criteria based on chemical modification of DNA (by histones that package it, but are specifically informative for promoter or enhancer sites) to identify regulator and RNA transcription start sites (enhancers and promoters), and find that most regulatory sites are, as expected, near to the gene they thus appear to be regulating in these cells.  The figure also shows that genes actively being used may be in between the enhancer and promoter.

There are several particularly interesting points here. First, regulatory sites need not be near to a gene, but can be almost anywhere (or, at least, quite distant), so that we can't know a priori where the important sites are.  Second, as mentioned above, most GWAS 'hits' have been in regulatory sites.  Third, regulatory contacts between DNA bits can span actively used genes in between the sites; this raises the question of how those sites' enhancers and promoters are juxtaposed and/or stay open for business as spanning DNA parts are brought together in the nucleus.  Fourth, finding a mapping 'hit' in a non-coding region may tell us that some gene's activity is being affected and contributing to the measured trait (e.g., diabetes, stature, or whatever).

In a given cell thousands of genes (not to mention other regions that are transcribed into other sorts of RNA) are expressed differently in different contexts in the same cell (e.g., when it divides, when it is doing its normal business, when it responds to environmental changes).  And of course, each cell type will be using different combinations of genes.  This raises the question as to how the chromosomes all knot up in the orderly-appearing way that Hi-C methods identify, and then can re-knot as these or those genes go 'on' and 'off' (or 'higher' or 'lower' levels of transcription).   This would seem to be an intriguing 4-dimensional (space and time) geometric problem.  This analysis does not include trans connections, between enhancers on one chromosome and promoters on another (I thank senior author Cameron Osborne for clarifying this to me), yet a much larger kettle of fish as yet mainly unexplored.  So this is possibly, or probably, only the tip of the nuclear-interaction iceberg.

Genomic Skype calls
Regulation spanning very large distances effectively and rapidly is like making a complex multi-person international Skype call: instant communication from afar. This is remarkable, even if it confirms what we have suspected!  The finding raises the related question of how the conjoined parts of DNA 'find' each other.  Some data of this sort aggregates millions of cells at one go, but new methods have been applied to single cells, and they have found that there is stochastic variation among cells of the type in the same culture at the same time.  This paper seems to have been of aggregate data from many cells, so we don't know the role of variation among cells in the 'same' state, if that is really what can be said of the cell-source of these data.  So there are other issues yet to be understood (the authors don't claim otherwise!).

Genomic Skype calls may not just be across the country or across ocean, but maybe far out into space, figuratively speaking.  If the current limited technology is but the first opening of this sort of knowledge, then one can only wonder what far-reaching sorts of communication are going on within and maybe even among us.

It's of course one thing to document long-distance regulation in cell culture, and understand or even identify the related pairs from data on whole organisms--such as to find the relevant contributing genes to diabetes or some other trait.  Sometimes, experimental assays will be able to find the gene affected by a non-coding GWAS or other association-study 'hit'.  Other times, perhaps the vast majority, this won't really be possible or practicable.  And if hundreds of different genes are contributing, identifying them more accurately from mapping results won't necessarily simplify things. But it will help confirm those complex results, and will be interesting, potentially very important, new knowledge in its own right.

Wednesday, November 12, 2014

On cancer genetics

What 'causes' cancer?  This was a very mysterious disease for a long time, and there were many theories about it.  Prominently, in the 1970s or so, a major idea was proposed by Nobel laureate Macfarlane Burnet, an eminent Australian immunologist.  The idea was known as the 'forbidden clone' theory and was about autoimmune disease but, more generally, about somatic mutation.  The idea of cancer as a somatic mutational disease made sense if cancer arose from single founder cells, as accumulating evidence suggested, and yet was generally not inherited.  If it is 'genetic' in its etiological mechanism, what else could it be?  Viral causes were found, though I cannot recall when, relative to the rest of this history.

The idea of a mix of inherited and somatic mutations had appeal in the sense that if you inherited part of a mutational pathway to cancer, but not all of it, your parents would be unaffected but you would only have to 'await' complementary somatic mutation in order for some cell to be transformed to a cancer state.  This thinking led Al Knudsen in the early 70s to propose such a mechanism for the pediatric eye cancer retinoblastoma--a marvelous insight for which a Nobel prize would not have been inappropriate.  There, it has turned out that the major event is a second, somatic, mutational 'hit' in the RB gene itself, and the tumors occur so early in life that perhaps few other somatic events are needed to transform a retinoblast.  Also, retinoblasts may not divide much if at all after development, so if you escape the second event while the retina is developing, then you're safe.

The idea of cancer as a somatic mutational disease is widely acknowledged, though most of the ink is spilled lauding discoveries of inherited tumor variants, of which the best-known are variants in the BRCA1 and 2 genes (but there are others).  Virally induced cancers seem to be due to viruses incorporating into inappropriate locations in the genome, so while they are externally 'inherited', the cell-specific mechanism is consistent with other ideas.

It is still correct that, with a few exceptions like retinoblastoma, even those who inherit a high-risk variant such as in the BRCA genes typically do not get their cancer till much later in life.  And it is also true that inherited variants seem to need many subsequent complimentary mutations for a cell to be transformed.  Thus, even BRCA mutations are in themselves not a cause of cancer.  Indeed, if the story is correctly being understood, the BRCA genes are involved in mutation detection and repair, so that the associated breast and ovarian (and perhaps a few other) cancers are really due, at the cellular level, to other mutational changes that directly affect the cell's behavior.

Somatic mutations are generally hard to study, but even in cancer, a concentrated source of cells with such mutations, this is a challenge because a tumor grows rapidly and spreads, so even if all tumor cells are somatic descendants of the original transformed cell, these cells continue to acquire further mutations.  This accounts, in part at least, for the spread (metastasis) and evolution of drug resistance of tumors.

Most attention has been on protein causing changes--exome mutations--in the search for cancer-related  mutations.  But if cancer is a lineage of cells that do not constrain their processes or rate of cell division, then one might suspect that regulatory variation would be comparably or even more important than protein structure itself; that is, normal proteins related to cell behavior may cause problems if there are too many or too few of them in a cell under various conditions.  This has led to expanded, though more difficult, searches of DNA sequence in tumors.

Regulatory somatic mutations in cancer
A paper in the November 2014 issue of Nature Genetics, by Weihnold et al., reports on regulatory mutations found in cancer cells.  The authors used some existing cancer genome data bases that compared cancerous tissue to normal ('matched normal') control samples.  The samples were small and had other various limitations as the authors note, but the point is that in screening whole genome sequence they found a number of gene-regulating areas that had multiple mutations in the data and thus seemed to indicate regulatory somatic mutations.

This is interesting beyond even the tentative nature of the paper itself.  One might speculate that even when protein variation is responsible for the cell's initial transformation to found a tumor, the subsequent aspects of growth, metastasis, drug resistance and so on may well be due to changes in the regulatory behavior of the cells descendant from the original tumor.

It is theoretically obvious and well-documented specifically, that different parts of tumors contain different somatic-origin mutations.  This paper suggests that classical genes are not the only place to look for such variation.  Searching the 'noncoding' parts of the genome, which is the vast majority and is still largely not understood, will be daunting.  How complex, unique to individuals and tractable this approach will turn out to be is hard to predict.  But as we've noted recently here on MT, the evolution of cells within a given individual's lifetime is comparably (or more) complex than the evolution of individuals in a species,  Documenting this variation in adequate detail may require very different sorts of methods, but the story is surely going to be interesting.  How well it aids therapy is another story entirely.

Tuesday, September 11, 2012

What makes the traits of life? More on ENCODE

Our favorite tweet about ENCODE went something like this: "I have no idea what ENCODE is but I know it's making everyone mad."  There were valid reasons for all the to-do, but we want to go beyond that here and talk about the science.

The ENCODE project is a large-scale genomewide attempt to "build a comprehensive parts list of functional elements in the human genome."  The idea is that once we know what does, and what doesn't, have function, we can effectively home in on those sites that affect human traits and, of course, disease susceptibility in particular.  The new ENCODE  web page will be very helpful and welcome. 

Yes, the hoopla surrounding the release (or is it 'press release'?) of the ENCODE results was widely criticized, including by us, for its excesses and manifest if not blatant self-promotional advertizing.  The critiques included comments about possible over-acceptance of data that are less secure than the impression given, and other such issues.

These critiques are justified, and it will take an enormous almost open-ended effort to address or resolve all of the points that have been raised--if  that were even possible, given life's moving target.  And, of course, this must already have triggered a host of me-too demands by the mouse, fly, nematode, arabidopsis, maize, and other genome communities who will want to do all the same for their favorite species--and of course, they'll make the same claims about its urgent disease relevance.  Adjudicating the relative importance of what to fund may be tricky if budgets are limited.

As importantly, a sober evaluation of what and where the real value is may be contentious. This is because if causation is complex and much or most of the genome is involved in traits of interest, exhaustive enumeration of countless trivially-small or rare-variant effects may not have much point, even if each instance is really valid, which will be hard to prove.  Like trying to understand a changing beach by enumerating sand grains. Our earlier post dealt with some of these points.

However, there is one major actual scientific point that even in this sea of potential error or misinterpretation, seems rather safe--and particularly important.

Making life live
The traditional idea of genomes is that their purpose is to code for protein. We've long known that DNA has to replicate itself and that it strings along tens of thousands of protein-coding segments, and that these functions are fundamentally important to life.  Life is about proteins and the variation among cells and among species is largely due to variation among proteins.  Some DNA codes for RNA that is directly functional rather than being translated into protein.  We've known all this for many decades.

We've also known for somewhat fewer decades that part of DNA is used to determine which of our thousands of genes are expressed (used) in which of our cells and under what conditions.  This is called regulatory DNA.  A body is made of many cell types (skin, nerve, muscle, stomach, immune,....) and they are different because they use different sets of genes.  Likewise, cells change the genes they use, or their level of activity, based on circumstances.  That is why organisms can respond to changing environmental conditions and so on. This works because the regulatory DNA is recognized by proteins (transcription factors) that stick there, and that leads to other proteins transcribing nearby DNA into RNA.

We knew that regulatory DNA was just as important as coding DNA, so that there are two basic types of 'code'.  In Mermaid's Tale we called these correspondence  codes and recognition codes, respectively. The first term refers to DNA whose sequence corresponds to the protein (or functional RNA) that is copied from it and used elsewhere in the cell.  The second is a sequence code that is directly used--recognized by transcription factors--rather than only standing for something used elsewhere.  The transcription factors themselves are proteins, and so must be coded for by their own correspondence codes somewhere in the genome.

The dogma for a long time was that protein variation is the cause of evolutionary fitness differences and trait variation.  This followed essentially from Mendel's showing that what turned out to be protein variation--and hence variation in protein-coding genes--was responsible for trait differences in peas.  That led to a century of focus on genes as protein correspondence codes, the discovery of how that worked and how those genes were arranged on chromosomes.

But a few decades ago recognition codes were discovered, and that explained in principle how DNA also coded for the usage of its genes.  None of this was or is controversial.  But gene mapping to find genes 'for' (causally associated with) traits, and especially (from a funding point of view) disease, initially was intended to find the protein variants that were  responsible.  This was the focus and the prior belief of most investigators, but genomewide association studies, GWAS, were conceptually designed for that purpose, and it was obvious that regulatory variation would also be picked up by the same method.

Although the media and grant attention has focused on protein variation, even for years after we knew about the importance of regulation, it has turned out that many if not most well-documented mapping-identified common disease variants were not in protein coding regions.  And the important documentation of the ENCODE project shows rather persuasively and systematically, that the finding of regulatory rather than coding variation in the data that has been has been a valid indicator of how things are. 

The important impact of this is to make it clearer that most of our functional DNA is not about protein-codes directly, but is about how they are used.  It is the timing among cells, and regulation of that among cells, of the genes and their level of expression, that accounts for the production of organisms from fertilized eggs (single cells), and likewise that their variation and hence how they are seen by  natural selection is important to evolution.  In a sense, overall, evolution is a regulatory phenomenon.

What needs rethinking?
The ENCODE project did not raise any new points in this regard but the paper by Maurano et al. in the Sept 7 Science, "Systematic Localization of Common Disease-Associated Variation in Regulatory DNA," that shows the regulatory documentation from the project is an important one to be aware of.  And this paper only deals with classical 'regulatory' regions, but not with other kinds of DNA-encoded RNA molecules of various sorts that serve to regulate gene usage and dosage in other ways (others of the ENCODE papers in various places deal with some of these issues)

The finding that much or most of DNA function is regulatory casts a somewhat different light on the nature of organisms and on their evolution, but it is not a fundamental new theory of any sort.  Let's hear no talk of 'paradigm shifts' or any other such puffery!  Though we continue to learn of new DNA function and that more of the DNA has function that was thought at one time, the results are just a re-weighting of focus among things we've long known.  It's an incremental accumulation of knowledge.  Indeed, traits and their evolution and variation are still mainly about protein structure and variation, and that ultimately means genomic structure and variation. 

Conceptually, however, by trying to understand regulation, we gain a better way to view ourselves as 'emergent' phenomena, that is, that an organism and its highly complex organization is more than just a pile of genes, but 'emerges' from the way those genes are used.  This is similar to the idea that you cannot understand a building by enumerating the bricks, tiles, wires, and so on that it is built of--or sand grains and beaches.

And, importantly, the new project makes traits more, not less, complex and potentially problematic to deal with.  This is because it simply adds to the breadth and depth of factors whose variation contributes to the nature and variation of our traits.  Similarly, every new factor that we find contributing in subtle, and subtly variable ways to traits of interest, the harder it is to promise that we can enumerate risk gene by gene, or account for genome evolution by gene-focused selection models. Different ways of thinking will be called for.

So, filter out the blatant BS, and there is still importance in the findings. 

Wednesday, August 31, 2011

Controlling by cooperating

Here's another example of cooperation in nature that was largely unexpected and is an installment in our  life-as-cooperation campaign.  An article by Voss et al., in the August issue of the journal Cell, describes the way potentially competing proteins actually cooperate to affect gene expression.

The genome contains protein coding sequences, but these are only a small percent of our DNA.  Another part of DNA consists of generally very short sequences in the vicinity of a protein-coding gene, that are used to control when (in which cells) the nearby gene is to be transcribed into messenger RNA and then translated into protein.  These regulatory DNA sequences work by being bound (grabbed physically) by proteins called transcription factors (TFs).  A TF is a protein whose physical and chemical properties make it bind (find and stick to) specific regulatory elements (REs) like (to make one up) CCTGCA.

The idea has been that such regulatory elements are bound by a specific TF and if that TF is itself being produced by the cell, it will grab the RE (stick to CCTGCA) and cause the gene to be transcribed.  If the TF isn't being produced by the cell, the CCTGCA remains naked and the gene inactive.

We're oversimplifying greatly, but this is the general idea.  But what if there is more than one TF that recognizes the same RE?  Then, our hyperDarwinian friends might presume that the two TFs compete to bind the sequence, with some sort of consequence for gene expression.  This would then set up an opportunity for natural selection to choose a winner, for example, so that the loser TF would lose its function.



But instead, as shown schematically in the figure, TF A, on the left, binds to a specific RE (the blue box along the black DNA line, which helps open up the wrapped-up DNA (called chromatin--black line of DNA wrapped around some packaging proteins represented by the pink ball) near a particular gene, that event allows another complex of proteins (the remodeling complex in the figure) to modify the DNA so that TF A gives way to enable TF  B to get access to the same RE sequence.  The nearby gene is expressed.

This is just one experimental example of a particular laboratory setting, rather than an analysis of such activities generally, so we don't know how pervasive such cooperation is.  The idea of intricate cooperation is to us only an additional instance of such a phenomenon, which we believe is far more pervasive and important to understand, in terms of biology, than competition.  Competition may, of course, help establish cooperative interactions if they are useful--and that would be the standard Darwinian theory.  But every day powerful molecular technologies are finding new examples of the kinds of extensive cooperative interaction that occurs between the passive sequence of DNA and the very dynamic activities of organisms.  Without such cooperation, we wouldn't be here to write these posts....and you wouldn't be here to read them!

Monday, October 26, 2009

The genome in three dimensions

We all learn about DNA as a string of letters, of A's, C's, G's, and Ts, so it's not surprising that we tend to think of DNA as linear. But, inside the cell, where it really matters, DNA is actually wound up into a three-dimensional ball. That this is so has been known for a long time, but little has been known about the organization of that structure. A paper in the Oct 9 Science (Comprehensive Mapping of Long-Range Interactions Reveals Folding Principles of the Human Genome, Lieberman-Aiden et al., p 289-293) discussed on the BBC website here, begins to correct this. (The figure to the left is from the paper, via the BBC.)

Using a series of clever molecular techniques to cut and sequence neighboring pieces of DNA, Lieberman-Aiden et al. generated a compendium of interacting bits of DNA. That is, for each part of the genome, they were able to determine its neighbors in three dimensional space. Among other things, they show that distant DNA sequences interact with and regulate each other in ways that aren't easily envisioned when we conceive of DNA as linear sequence alone. In addition, they determined 'contact probabilities' for parts of the genome as a function of genomic distance (number of basepairs away), finding that intrachromosomal contact probabilities are greater than those for interchromosomal contact. Further, interchromosomal contact is most likely between small gene-rich pairs of chromosomes.

The current study is, of course, following up on earlier results of similar investigations, but based on a more powerful molecular method. Conventional wisdom has been that DNA has to be tightly wound in order to fit in the cell -- unwound, it's 2 meters long, so it has be be compacted somehow to fit into the cell, never mind the nucleus of eukaryotic cells.

But DNA seems to be packaged in a very orderly way, and reflects whether or not genes in a particular spot are open for expression. The packaging seems to be replicable, if the new paper is correct. And, this brings to mind a number of questions including whether this explains why it has been so difficult to track down regulatory regions for many specific genes. Is the pattern of inter- and intra-chromosomal contact the result of functional constraints, specific sequence-based interactions, or natural selection? Or is it just how DNA winds itself up, given its overall structure? How sensitive is the cell to its packaging? If someone's DNA doesn't roll up right, are they selected against?

If selection is important, there must be many opportunities for--that is, need of--cooperation that enable the DNA to fold up and around itself, as Lieberman-Aiden et al. demonstrate; not only does winding of DNA into this tight ball fit it into the cell, but it also facilities contact between pieces of chromosomes, enabling cooperative interactions such as gene regulation. If there are functional reasons why these alignments occur, then there would be co-evolutionary constraints that maintain their compatibility--what we refer to in our book as cooperation.

As we emphasize in our writing, cooperation is a fundamental principle of life, and these results reinforce that view. We'll have more to say on this subject after Ken and a colleague give a seminar on the Lieberman-Aiden et al. paper.

-Anne and Ken