Showing posts with label central tendency. Show all posts
Showing posts with label central tendency. Show all posts

Wednesday, April 18, 2012

Bulking up....or not: what chaos can tell us about order

A paper in last week's Science, (Using Gene Expression Noise to Understand Gene Regulation, Munsky et al.), as part of the special issue on computational biology, asks whether what looks like noisy gene expression can be informative about gene regulation.  Gene expression within a single cell is a topic of growing interest.  It has seemed to be fairly random, but Munsky et al. suggest if you look closely enough, the randomness can be indicative of some quite regular processes.  Mechanical engineer Brian Munsky and colleagues have used a similar approach to identify gene regulatory networks.

If molecules interact in a probabilistic way--bouncing randomly around the cell until they perchance (literally) bump into each other, and if each cell has countless zillions of molecules buzzing around in this way all the time, and if it's clear that the cells even in the same tissue in the same person (and hence the same genotype) are each a bit different....then how come we're so highly organized into tissues and organs that mostly work mostly correctly.....rather than being just a jiggling blob of formless jelly?

Sad to say, but Prairie Home Companion can't be true!
One obvious possibility is what one could call the law of large numbers, or a principle of central tendency.  All of these random motions have an average behavior, just like everybody's height or glucose levels vary but most of us are somewhere near the middle.  Unlike Minnesotans in Garrison Keillor's Prairie Home Companion, not all the children are above average!

With large numbers of cells in a given organ, there will be variation but only a small percent of cells will behave very differently from the average, and even if they are very naughty indeed, their effect on the organ as a whole--and, say, on the person's health--is trivial: the vast majority of well-behaving cells cover for the wayward ones.  And, indeed, we have bodily systems to detect cells that are far too misbehaving.  When they fail we can get nasty conditions such as cancers.

So how is this stochastic (probabilistic) buzz-fest made manifest at the level of individual genes and their levels of expression (use) by cells?   As the authors of this paper note, because of the vagaries of Brownian motion, two cells, even those produced by the same progenitor cell, will never be identical at the molecular level.  Thus, things like the number of transcription factor molecules per cell, needed to cause specific other genes to be expressed, are unlikely to be identical, and this cell-to-cell variability will affect gene expression levels among cells, and ultimately can lead to phenotypic variability as well.
Consider a single mother cell dividing into two daughter cells of equal volume. During the division process, all the molecules in the mother cell are in Brownian motion according to the laws of statistical mechanics. The probability that each daughter cell inherits the same number of molecules is infinitesimally small. Even in the event that the two daughter cells receive exactly one copy of a particular transcription factor, each transcription factor will perform a Brownian random walk through its cellular volume before finding its target promoter and activating gene expression. Because Brownian motion is uncorrelated in the two daughter cells, it is statistically impossible for both genes to become activated at the exact same time, further amplifying the phenotypic difference between the two daughter cells.
Munsky et al. use cell-to-cell variability in gene expression as a way to understand gene regulation, and they present quantitative models for this. Genes can be expressed all the time, or their expression can be episodic or timed -- this is 'constitutive' vs 'regulated' expression.  Taking into account copy number of transcripts of genes of interest, they determine that when transcript births and deaths are not related, and seem to follow a Poisson distribution (that is, they are independent events that occur at a known average rate, but with different probabilities of any specific rate); this indicates constitutive expression.  Deviation from the Poisson distribution--too many rare or too many common copies, for example--suggests regulated or episodic expression; expression of a gene within a cell can be more or less tightly regulated over time, and can switch states.

Documenting and making sense of transcript levels in single cells, the authors write, can be informative about gene regulation as well as gene networks in ways that looking at gene expression in multiple cells or tissues can't be.  This is because average statistics, such as of the number of transcripts of a particular gene among cells, mask the distribution within each cell, and so regulatory mechanisms can't be inferred. Yet most cellular studies are of test-tubes full of the 'same' kinds of cells, analyzed in aggregate, masking this informative, underlying variation.

Hounds and Hares
If, as the authors predict, sequencing of all the transcribed genes in a single cell becomes routine, understanding gene networks will be easier.  And it will account for our orderliness as organisms.  This can be seen at higher levels, in ordinary experience, too, as this example may help make clear:

Hounds chase hares, but each hound and each hare live individually different lives.  If we want to understand the overall organization of the hound-hare part of the ecosystem, such as how their respective populations vary over space and time, we can look at the aggregate behavior: the chance a hound will sniff a hare, the chance it will catch the hare and so on.  But if we want to understand details, we might have to follow a number of individual hounds and hares, because not all will be equally successful in their hunt, and the circumstances of the hunts will vary.

Bulking up....or battening down?
These ideas apply to situations when there are many 'identical' cells, as in a given organ.  The central tendency would seem to provide safety in numbers.  But if that's the case, how big do the numbers have to be to protect the organism from its fraction of unusually behaving cells?  This will depend on many things, including the 'variance' (variation from cell to cell) of the process, how many cells are in sync at each time, and so on.

There's another important issue. In complex tissues, gene expression is changed via various processes that we can call 'signaling'.  Cells sense their surroundings and respond to them.  They send out signal molecules appropriate for their location.  This reinforces cells in a given tissue to do what's appropriate for the tissue.  And, importantly, signaling can be homeostatic:  some signals are called 'activators' because cells detecting the signal's presence activate that same gene, or some set of response genes, as a result.  Other signals are 'inhibitors' and induce cells to do the opposite.  Waves of expression can generate waves of tissues (like hairs or scales) in an embryo, but activation and inhibition interactions can also generate stability, as signal levels can lead cells too far out of line, so to speak, to fall back into line--to batten down and stay within acceptable limits.  Could this be relevant for the question at hand?

In organs with lots of cells, say in the millions, perhaps that is bulked up enough to bar the door against cellular chaos.  But what about smaller organisms, of which there are many, in whom, like small fleas  biting the backs of larger fleas, the organs can be very small indeed?  Is there any evidence that their organ stability is less, or is more vulnerable to vagrant cells?  Does natural selection work differently in such organisms (e.g., they have to reproduce more or faster, to stay viable as a population)?  We haven't thought about this directly, though we did refer to the issue of embryonic selection as contrasted with competitive Darwinian selection in our book.  Perhaps there's no issue here, or perhaps there is something interesting to follow up.


If we're understanding this paper correctly, it's an interesting application of physics to biology.  This might seem to resemble the ideal gas law in chemistry and physics.  There, for a given container and number of molecules, the pressure or temperature of any gas follows the same law.  But this does not require tracking any individual molecules.  In a way central tendencies of similar cells are like that.  But since each cell is different, each gene is different, and gene expression can affect other gene expression, cells are not like containers of oxygen or hydrogen.  Still, when large numbers of molecules are involved, the distribution of traits among similar cells do seem to follow orderly statistical properties.

Thursday, March 22, 2012

Random events result in order -- how?

Development is ultimately very organized and predictable -- children look like their parents, legs are generally where they belong, and a lion never gives birth to a whale -- but yet another paper describes the randomness of the processes at the cellular level.  How can this be?

The paper is in the April BioEssays; "Genes at work in random bouts", Alexey Golubev.  Golubev says that things that go on inside cells are generally thought to be determined by the interaction of different molecules, which is itself determined by the concentration of those molecules in the cell.  Ordinary differential equations (ODEs) describing all this can be written, and, Golubev says, "ODE solutions may be consistent with oscillatory and/or switch-like changes in molecule levels and, by inference, in cell conditions."  This begins to make intercellular processes sound determined and law-like.

But, the article is basically about the stochastic (random, or probabilistic) events occurring in cells that affect their gene expression patterns, and hence the cycle between cell divisions or the time it takes the cell to express the genes related to its particular tissue  This variation, the author notes, makes stem cells--cells not committed to just one cell-type--plastic and flexible. 

But, as he points out, the idea of molecular concentrations is only true at the level of populations of cells, not in single cells themselves.  There's a lot of randomness in terms of what's going on in single cells, in cell differentiation and cell proliferation, particularly with respect to when genes are turned on or off, and thus which proteins are available, and what happens when. 

The question becomes, then, given all this stochasticity in cellular activity, how development is so organized.  The apparent problem is that once one reaction has taken place, it affects the next reaction, and this includes hierarchical changes such as changes in gene expression in the cell.  Thus, the cell is not just a mix of things, each in large numbers, that will 'even out' over time.  Differences that can be occasioned by chance in a cell can add up. Of course, if the cell continues to detect the same external conditions, its response may adjust so that things do even out.  But it doesn't need to happen. 

On the other hand, most tissues in most organisms are comprised of many cells of the same type.  Each may be experiencing stochastic changes, but their tissue-specific behavior may usually 'even out' because the variation will be slight and in different directions among the cells, so that on average they are doing the same, appropriate, thing.  In unusual circumstances, if this doesn't happen, the organism may be very different from its peers....or  it may not survive.

This perhaps reflects a fundamental property of populations, known as the 'law of large numbers'.  The theory behind this (Ken was just realizing from reading a book called The Taming of Chance (1990, by Ian Hacking, Cambridge Press)), comes from the study of populations of individually differing individuals, whose aggregate behaviors have regular distributions: the 'normal' or bell-shaped--or at least orderly--distributions of stature, incomes, and so many other things.  In another common phrase, they have 'central tendencies.'  Normal meant that most were near the norm.  This statistical idea was worked out over the 18th and 19th century, and raises interesting questions about causation.  Hacking's book shows how people had to learn that causation was not about precisely fore-ordained laws, but about probabilities, and that this applied to society.

The classic cases had to do with things like suicide.  One can't predict who will commit suicide in a given year, or by what means.  But the numbers, and the number who do it by each method, are very similar from year to year in a given population.  Likewise, the life expectancy is an average, and nobody lives exactly that long: some die younger, some older.  So are many social facts like political affiliations and so on.

Why is this?  It is the net, end result of many different individuals each with slightly varying characteristics.  There were many explanations of why this was, that are beyond this post, but in essence there are many contributing factors of diverse kinds, that mostly aren't known, so that a few individuals are exposed to many, others to only a few, but most of us to some 'average' amount of these factors.  The fraction exposed to many such factors is the fraction of individuals who are taller, more intelligent, ...., or who commit suicide.

The law of large numbers is a statistical fact that can be proven mathematically under rather general conditions.  This leads to central tendencies.  That is why population statistics took on a central role in social sciences, where often the underlying causal factors and their specific effects are unknown or hard to estimate accurately.  Social sciences can 'understand' society--at least predict some things about it--without understanding causation in the strict sense.  And in some situations, these things don't work very well--economics is one, in which stability of population outcomes occasionally, at least, takes a quick left-turn.

Probably the same applies to the populations of cells that make up a tissue, and if so this would make the high amount of probabilistic events in cells that would make each cell different, which can lead to a central tendency for the kidney to filter blood is similar ways, and so on.  Because of local differences among cells, different genotypes, and different life-experiences, kidney functions differ among people.  Some are at the extremes and we call that 'disease', but most are roughly near the norm.

Biologists routinely speak of chance, but often act as if they believe that genes 'determine' the organisms the way a program determines what a computer does.  They know about variation in populations, and how, for example, polygenic traits like stature or blood pressure (the darlings of the GWAS world) vary, even if they are driven to enumerate all the underlying causes that vary among individuals in the population.  In a sense, the population concept applied to tissue is of the same sort, and provides another source of variation between genotype and trait.