Showing posts with label heritability. Show all posts
Showing posts with label heritability. Show all posts

Wednesday, December 16, 2015

Let's be intelligent about intelligence

A lot of confusion reins over assertions about whether a physical or even behavioral trait is  'genetic'. There are several reasons for this.  One is the difference between mechanism and variation. Every human trait is genetic in the first sense: an organism develops from a fertilized egg because it has genes, and without its genes it could do or even be nothing.  So every trait is 'genetic' in the mechanism sense. But the other meaning of 'genetic' has to do with variation, and that is where the difficulty and often the contention lies.  The assertion that a trait is 'genetic' in this sense means that some people with a trait, or a particular trait measure, have it because of some particular genotype. That is, we all differ in the trait because of causal genetic differences.  Identifying genetic mechanisms or demonstrating that genetic variation is responsible for variation in a trait are genuine challenges.

Searching for genetic mechanisms responsible for, say, heart disease is one of those challenges.  It's difficult scientifically, but unlike with some other traits, the scientific question isn't politically loaded. Many people fervently want to stress the genetic role in intelligence, for example, and it's often for thinly disguised racist or elitist reasons.  A common response to almost any suggestion that an individual's intelligence might not be inborn, due to variants in his/her inherited genotype (meaning built-into the person's DNA sequence), is an accusation that the person is in denial of reality (but see our Dec 14 post about genetics and dialectics).  But who is really denying reality in such cases?  In our view, it is those who misperceive or misuse measures like heritability and have deep, emotional commitment to inborn destiny.

And, again, it's pretty clear that just slightly beneath the surface is often a racist or other discriminatory agenda: "let's identify 'them' and do something about it, to 'improve' them or prevent them from harming everybody else" (Trump's throw the Muslims out campaign, or the reluctance to invest 'our' resources in groups with inferior IQ, or in the worst case, eliminate them). If it's important to understand why people behave as they do (intelligence being just one aspect of behavior; there are of course many others), the argument goes, then one needs to know if it's genetic, that is, built into the genome at conception!  Again, then depending on who such knowledge is important to, individuals in the population can (should, must) be tested.

Of course, it's worth asking carefully whether what's really being looked for are individual differences, or group differences.  Why 'we' (those in power) 'need' (that is, want) to know which of 'their' behaviors are built-in, is unclear, but seems frequently to justify acting in discriminatory ways, favoring some and neglecting others.  In other words, of course intelligence is the result of gene action, but the argument is really about variation rather than mechanism.

But before we address these issues, it is worth providing a quick description of the core of the 'scientific' basis of the argument, which typically rests on a measure called 'heritability' (denoted here by H but typically written h-squared).

Heritability: simple-sounding word, but a slippery measure
When the genetics of intelligence, or most other behavioral traits for that matter, is considered, the proof that they are genetic is usually that their heritability is high.  Heritability has been known for decades to be a rough indirect indicator of genetic mechanistic cause, but it's a very elusive measure. The usual measure of H is basically the ratio of the amount of variation in genes (G) divided by the amount of variation in genes + variation in environment, G/(G+E), all within a particular sample at a particular time.  This is estimated typically by comparing the trait measure in relatives, since close relatives share specifiable fractions of their respective genetic variants.

This figure schematically shows the scatter of genetic similarities, each dot being values of the measure in an offspring compared to the average of its mother and father.  The figure shows the difference in such correlations if environmental effects are great and genetic variation accounts for only 10% of the similarity (left panel), or small where the environments contribute only 10% (right).

From Wikimedia images, taken from Nature


H in itself measures no specific genes or gene-variants, nor any specific environmental variants.  To avoid some confounding or confusing contributors to the trait, various additional types of sample are often studied or comparisons made, such as between adoptees vs biological children, or dizygous vs monozygous twins. Heritability studies also often try to remove correlations among relatives that are due to shared family environments that could, in the computation, falsely appear as genetic.  While these strategies are not useless, they are well-known to be imperfect.

Since the measure H is a ratio that depends on the particular conditions in your particular sample, if one of the terms (G or E) were to change, even within that same sample, the H value would also change. In other words, let the same population (the exact same set of genotypes) experience changed environments, and H will change. In that sense H is not an absolute measure of how genetic a trait is, but of how relatively important it is.  Let us repeat that--heritability is not a definitive measure of the genetic contribution to a trait.  It is about its context in a particular sample.

Every study of traits like IQ test scores, used as hopeful stand-ins for 'intelligence', shows that there is substantial heritability, though usually far below 1.0.  That means that environmental effects are important, usually predominant, even if genetic variation is contributing as well.  That's about all that H measures show.  'Environment' in this sense tells us nothing in itself about what the specific individual contributing factors might be, because they don't behave the way genetic factors do, thanks to the rules of genetic transmission from parent to offspring; environmental factors don't have theoretically specifiable patterns of clustering among people or even among relatives. The apparent environmental component estimated in H studies can also include things like chance, testing inadequacy, measurement error and so on.

The undeniable bottom line is that variation in traits like intelligence test performance is certainly affected by genetic variation because the trait itself is mechanistically affected by genes. But that is a crude and almost useless fact because the genetic component is generally polygenic, meaning that it is affected by large numbers of varying genomic elements, each making very small individual contributions. Here, we conveniently ignore whether current fad factors such as microbiomic or epigenetic effects are relevant, because each of them is variable, in each population or sample, and over time--even in each individual over time--and could in principle be inherited and hence appear in families as being 'genetic'.

What this means is that even each individual's inborn genetic component will be very different, that is, each of us will have different combinations of variants at tens or hundreds (or more) of contributing gene regions.  The predictability of achieved results from genomes, much less individual variants, will be correspondingly small, practically useless, as we've clearly seen for so many other complex traits (GWAS results, for example, even of IQ test scores). If we could measure environments the way we can measure genomic variation, they would be similarly complex with many individual factors involved, most with individually weak effects.  As with genotypes, the complexity of these environmental factors would mean each person is unique and predictions are weak, and that changing circumstances and imprecision in the risk estimates would have a large potential effect on each person's achieved results.  We've discussed these limitations (and the overselling) of genetic association studies many times here.

But, if one is determined to pry into everyone's inherent worth, here's how to do it properly:
Here's an idea: Let society decide that we want to know the real genetic truth about behaviors, not just the mechanisms but the effect of variation among individuals.  To do that, we must pass legislation to ensure that all environmental factors that contribute to behavior--all of them!--are exactly the same for everyone, from conception onward.  Once that is done, variation in test performance will be entirely due to genes, since the environmental variance, E, would be zero, so that H would be 1.0.  Now we can see how strongly genes in general, or individual genetic variants, determined results.  However, we assert with confidence that the result of individual genetic prediction would still be hopelessly complex in most cases (excepting, for example, clearly pathogenic genetic variants, which we know to be rare, and even they are usually not simple).

But this is of course a fantasy: making environmental effects uniform for everyone is obviously impossible, for at least two reasons.  First, we can't make the climate in Maine like that in Florida or California.  We can't have identical schools everywhere, or the same number of books in every home, or the same number of words spoken to each infant at each developmental stage.  And so on.  So maybe a more realistic idea would be to make the environmental variation the same everywhere, so that in a sense it was a kind of uniformly distributed 'error' term in measuring genetic effects.  Of course that can't be done either, for the same sorts of reason.

Secondly, genes don't work on their own, but interact with 'environments' in almost every imaginable way, and certainly in the development of the brain.  That means that separating G and E (as in G+E) is clearly an oversimplification of something very poorly understood.  Even fixing the same environment everywhere would not have the same effect on every genotype.

The bottom line, in reality, is that arguments, usually by those in privilege, that behaviors (and hence their societal value) are inherent, are almost inevitably working some other form of self-advantaging agenda.  Racism is right beneath the surface in much of this, but so are xenophobia and class differences.  So are hopes of producing babies with some desired property.  That's clear from the history of the subject.

Since it's impossible to think that society could make environments uniform for everyone so all that's left is genetic variation, the next most salubrious thing a society could do would be to provide the best environmental conditions for all of its members to thrive in, not expecting everyone to achieve the same but at least to have safe, satisfactory lives.  More socioeconomic equity by the elimination of poverty and privilege would be a solution if such equity were the real objective. Of course since the beginning of history this has been the stated goal of those who bemoan the unfairness of society (though less so of others who say we're inherently unequal and we ought to reward the privileged).  We gain little by peering into individual genomic 'souls' and condemning those found genetically wanting to fates that we, in the elite, decide is best for them (inevitably making sure we stay at the top).

This doesn't seem too cynical a view of the subject: If what those who assert the deep importance of genetics of behavior really want is for society to be fair, the first thing is to understand the environmental effects that obviously are the predominant causes of behavioral variation, and rectify the inequities.  Let society ensure that everyone has the same conditions: no upper class advantages in schools, ballet lessons, Kaplan prep courses for SATs to get them into Princeton, no jobs to get through family or parents' contacts, same number of books in every house, no corner drug dealers nor rats in the hallways in poor neighborhoods......  Or, how about broader 'intelligence' testing ideas, to include smarts like the ability to read defenses in basketball while flying through the air, or work a fork-lift efficiently, or fix one of today's complicated cars....

H is a complex factor that is misused as much as it is used, because there are too many reasons to interpret its computational subtleties in ways that conveniently favor one's own social agenda. Not everyone interprets these issues in this way, but behaviors like intelligence are too juicy for those with such intentions to resist.  But, yes, let's be scientific, and commit to a concerted effort to make H approach 1.0, so that we can really understand the genetic contributions--that is, to make test-score differences really 'genetic'!  Then we could make sense of 'genetic' causes.  But, would any serious thinker believe it would be very useful?

Tuesday, November 5, 2013

Smarten up....but do it quickly!

To what extent is  'intelligence' genetic?  That is, to what extent is your intellectual ability a native ability rather than something learned or developed?  This has long been a very heated 'discussion', not because it's about individuals with truly impaired intellectual abilities, which can often be due to known genetic mutations.  Instead it's largely been about groups, that is, 'races', and it's therefore inextricably related to society at large.

One argument that IQ (here, we let that stand for whatever is measured, without making any supportive judgments that, or when and where it is an appropriate measure of something).   Heritability of IQ is not-trivial, generally estimated to be around 80%.  The remaining 20% is 'environment', and measurement error and the like.

Idiocracy
xkcd: Idiocracy

Previous studies had shown the tremendous advantage that early exposure to language and complex concepts has on later development, and that means later school success, and that means higher IQ test achievement.  Prior work showed that by age 3, children from privileged professional families had heard millions more words spoken than children from unpriviliged families.  They also had heard many different words spoken and learned their use.  Now, the NY Times reports that this difference can be detected even by or before age 2.  Samples were small but the study reinforces the idea that very early experience is telling throughout later life.

So what?
Those who see important group differences in IQ are going to defend their viewpoint by saying that they know very well about environmental variation and take that into account but that early experiences cannot obscure the entire group difference.  We don't happen to agree, and this is clearly a matter of personal politics all round, but there are a couple of questions that are fair to ask.

First, is it possible that the actual heritability of the measure is much lower than its estimates?  If social class is correlated in families, say by neighborhood, race, or education etc., then this can inflate heritability estimates, showing similar values for similar reasons, in professional as well as lower SES families.  This is why adoption studies are often used to show the true heritability, but even there there has been evidence of SES correlation in adoptions.

Secondly, if SES inequality were removed from the picture, the overall heritability might stay the same but there would be no difference of the average and far less variance (variation among individuals) among what were previously very different SES groups.

Third, education policy strives and presumably would strive even harder, to standardize what children are taught from birth on up.  They'd be taught or exposed to what our society values, be it vocabulary or mathematics or music or sports.  IQ test scores could increase steadily, as they have done for the past several decades, by making the most of everyone's inherited abilities.

There will always be those who are unusual on this or any other kind of value-score system.  There will be those who are seriously impaired or seriously gifted, however the neural mechanism works.  But the issues of group differences would largely if not entirely disappear; there will always be some average difference between any two groups that are compared on almost any measure of attributes, but that doesn't make the difference 'important', which is a social judgment.

The current study doesn't take us all the way back to Freudian ideas that the first glimpse an infant has of the world is hugely transformative, though who knows what further studies might find.  In any case, the important fact is not about the IQ controversy, because there is no reason to doubt that an enriched environment throughout life is an enriching fact of life.

Monday, January 16, 2012

Changing the goal posts: heritability lost and found

We interrupt our series on Probability and its meaning, for a post that we've been asked to write, related to a new paper on gene hunting that got some press last week, and will be stirring up controversy (and naturally, we can't resist including our usual editorializing, for better or worse):

Making complexity simple (again)?
Every age and every profession has its PT Barnums.  They're the slick-talking, fast-moving guys who will say anything to draw customers into the show they're running.  Truth, if it even exists other than ambiguously, is secondary in many ways to closing the deal.

One tactic that the pitchmen use in fields like politics is to change definitions so that problems never get 'solved' (that is, there is always something to keep you in office or to keep levying taxes, or to keep you afraid of some enemy or other).  This is a way of making the same facts serve new interests.

Science is rife with these kinds of self-interested maneuvers.  Changing the goalposts in genetics means redefining the objectives or the criteria for pushing ahead even in the face of contrary evidence.  This is a system we've built, step by step, once government funding largesse started flowing some time during the Cold War.  And today, in genetics, which plays on fear of death just as much as preachers do, we have to keep passing the plate.  Yet science is supposed to be objective and 'evidence based'.  So we have to change the goalposts to keep the game from ending.  In this case, there is a recent paper, with promise of more to come, by Eric Lander and colleagues (Zuk et al.).  Because of his prominence, skill, and salesmanship this will of course get a lot of attention.

The paper discusses at least one reason why GWAS have not been very good at accounting for heritability (something we ourselves have commented on in many posts).  Some, who are critical, say that this paper finally shows that GWAS and related big-scale approaches are proliferating even though they have themselves shown that they've reached diminishing returns.  Here's an example.  Of course there's the resistance that says no, Bravo! to the new paper, which shows that we're just getting started!

Naturally, everyone shares an interest in saying that to show their favorite view we need mega-GWAS, Biobank, gobs of wholegenome sequence, or other similar open-ended approaches, and whether this will lead to the ultimate small scale objective of personalized 'genomic' medicine.  Far too many vested interests are at play and, indeed, training in 'grantsmanship' and the whole research culture is about manipulating the system to get, keep, and increase funding. 

Zuk et al. ask why GWAS have failed to account for the heritability of so many traits, that is, to explain the correlation in risk among family members that reflects genetic effects.  The subject is complex, but here we can just say that if each gene adds a dose of effect to a trait, like nibbles on a given side of the caterpillar's mushroom added to Alice's stature, then even if the effect of nibbles is very small, if we sample enough mushrooms we can identify all of the effects.  Then, knowing them, we can tell which nibbles a given person has made, and hence predict his stature:  Voila!  personalized medicine!

Yet it hasn't worked that way.  Most heritability remains unaccounted for despite already very large studies.  Hence the demand for ever larger studies.  A glib commentary in Nature a few years ago coined the term 'hidden heritability' and made the search to find it akin to Sherlock Holmes' search for Moriarty.  That was a fantastic, if anti-science, marketing ploy on Nature's part, since it fed an ever-increasing demand for funds for genomic scavenger hunting....and that's good for the science journal business!

But the search for this hidden treasures has been frustrating, and Zuk et al. claim they now know that the search is in vain, and they provide a very sophisticated mathematical account of the reason why.  It is due to gene interactions, or 'epistasis'.  That means that a large part of the correlation among relatives can't be found by looking only for additive effects.  Here's roughly a basic underlying concept:

If trait T, such as stature or insulin levels (or their disease-risk consequences) is due to the effects of factor A plus those of factor B, then we can write

T = A + B

If your genotype includes a variant A-gene that gives you an additional level of factor A, then for you

T = A + A + B,   or 2A + B.

But if the factors interact, say in a multiplicative way, then

T = AxB

and if that's what's going on, and your A-gene genotype adds a second dose of A, your trait is

T = (2A) x B =  2AB

So, let's say a normal person has A=3 trait units and B=4 trait units.  In the additive case that person's trait would be A+B=7.  And if you have a mutation that doubles your A-dose, then your genotype makes you 2A+B=10 for your trait value.  But in the epistasis case, your trait would actually be 24.  So we expect you to have trait value 10, and conclude that more than half your trait value is unexplained.

Functional interaction
Nobody disputes a central fact in the Zuk et al. argument.  Life works by interactions among genes.  These are functional interactions in systems called 'networks' and other sorts of molecular interactions.  One molecule doesn't do anything by itself, nor does one gene as a rule.  Here, each gene-related component is subject to variation by mutation, and that will be inherited (its effects contributing to heritability).  So it is obvious from a cell biology point of view that single-factor explanations are not going to tell the whole story. But the fact of multiple factors doesn't tell the story we're interested in predicting states among variable traits.  That involves a different kind of interaction.

Quantitative interaction
If the true-fact is that biological traits are the result of functional interactions, what epidemiological risk estimation is about is not a list of pathways but the effects of variation in the pathways. When factors interact in a non-additive way, their net result is estimated from sample data using statistical techniques.  The A B example above showed conceptually how they work.  The additive contributions of variants in factor A are estimated from samples that compare those with and those without the variant in question.  You can estimate each factor's effects independently in this way and add up the estimates.

But if they interact, then in essence you have to have an additional estimate to make, of the average trait value in groups of people with each combination of the variants at the interacting factors.

Not only does this require more data, it won't show up in GWAS types of case-control or similar data.  You need to look in other ways.

Zuk et al. address this.  They build their idea that much of the observed heritability is estimated on a purely additive model, and yet at least some factors may be what in standard biochemical terms is known as 'rate limiting' effects.  At some level of concentration of such factors, they or what they interact with no longer works the same way if it works at all.  The authors outline a model which, under various assumptions about how many steps are rate-limiting such that variants in those steps define measures of heritability, might begin to explain familial correlations not accounted for by current additive effects.

It is already nigh impossible to get stable estimates of hundreds of additive effects, mostly very small (see our current series of posts on Probability does not exist!).  It's one thing to estimate the effects of, say, 100 additively contributing genes. Variants will be many in each, variable in their frequencies among samples.  If purely additive, the context doesn't matter: in each population you can estimate each factor's effect and add 'em up to get a given person's risk.  But if context matters, that is, if the effect of one factor depends on the specific variants in the rest of the genome (forgetting environmental effects!) then it's quite another to estimate those interactions.  Roughly, if 100 genes interact in pairwise fashion (2-way interactions like AxB only), that means 10,000 interaction effects to estimate. Zuk et al. certainly acknowledge this problem clearly.  But the authors suggest kinds of data on relatives of various degrees that might be practicably collected and could reveal discrepancies from expectations under additive models, and account for more or all of the heritability.

Zuk et al. promise that if we study 'isolated' populations we'll have a shot at the answer!  This is not new, and indeed studies in Finland and Iceland led a previous charge for similar reasons.  Perhaps it is a good idea....and the authors provide tests that could be done if there were adequate data collected.  It will be done and will expend more millions of research dollars in the process.  But, if complexity is real, the most we'll get is a few hits and a few weak statistical signals.  But we're in that place already, so in that sense this is another way of changing the goalposts, because we did not reap bonanza from those studies.

The answer
Like most such papers, Zuk et al. make gobs of assumptions about the model, the data, and the underlying basis of dichotomous disease (present or absent).  Facing such complex problems it's hard or impossible not to make simplifying or exemplifying assumptions.  As usual, if one probes these assumptions, there will be additional sources of variation and uncertainty that are being ignored or that must be estimated from data, so that even the rosy answers suggested by the authors, which are far from promising a complete understanding, will be overstated (and that is the clear-cut history of the senior author, and most of his peers in the profession, including Francis Collins).

The actual answer is well known, even if it's resisted because it's not very convenient: life is a mixture, or spectrum, of all of these kinds of components.  We know that additive models work very well in many domains, such as agricultural and experimental breeding (artificial selection), and that as GWAS sample sizes have increased steadily they have steadily been identifying more contributing genes.   That is, we have not plateaued to a point that bigger samples will not add more genes, the remaining recalcitrant effect due to interaction.  We think even the authors acknowledge that the interactions may comprise sets of individually weak contributors.

Further, because most variants in our population are in fact, and indisputably, rare to very rare, they form a large aggregate of potentially contributing genes that will vary from sample to sample and population to population.  This is, really, a fact of life.

Networks, the  sugar plum fairies of promised genomic medical miracles, involve tens of genes, many with multi-way interactions, like AxBxC. And what of higher-order interactions such as the square of one or more factor levels?  One might as well count stars in heaven as attempt to collect accurate data from the entire human species in order to get enough data to estimate these effects.   And that, of course, is foolish for many reasons, not least being that environmental factors change and vary and heritability is a measure of genetic relative to environmental effects.

Zuk et al. are promising a series of papers along this track, and have coined a new name for their idea (a standard marketing ploy). We'll be hearing a lot of me-tooism, users avidly diving into the LP (limiting pathway) model.  That certainly doesn't make the model wrong, but it does affect where the goalposts are. We're hopefully not hypocrites: we have our own simulation program, called ForSim (freely available to others) by which some of these things could be simulated, and we may do that.

Nobody can seriously question that context is very important, and that includes various kinds of interactions.  But the issue is not just what the mix of types of contribution are, how stable, how variable among samples, and so on.  Even if Zuk et al. are materially correct, it doesn't erase the problematic nature of trying to estimate, much less generally to do much about, the joint effects of the many genes, in their tens or hundreds, whose variation contributes to a trait of interest.

"Discovery efforts should continue vigorously"
We ourselves are far from qualified to find technical fault with the new model, if there is any.  We doubt there is.  But the point is not that this new paper is flim-flam, even if it simplifies and makes many assumptions that could be viewed as perhaps-necessary legerdemain, given the situation's complexity.  Or, perhaps more clearly, red-herrings to distract attention from the real point, which is  whether changing the goalposts in this kind of way changes the game.

One can ask--should ask--whether regardless of interaction or additivity, it's worth trying to document them, or whether Francis Collins' insistence on luxury medicine (personalized prediction of weak effects with lifelong treatment mainly available to paying customers) is a realizable goal.  The same funds could, after all, be spent in other ways on indisputably stronger and simpler genetic problems, that could be far more directly and sooner relevant to the society that's paying the bill.

Arguing either/or (additive or non-additive) and attempting to relate that to the desirability of keeping the Big Study funds flowing is a carnival barker activity.  The authors at one point subtly make the anti-Collinsian but obvious point that personalized gene-based prediction is generally going to be a bust.  But then they argue let's plow ahead to discover pathways. Let's have our goalposts everywhere at once!  There are other, better, logistically easier and probably less costly ways to find networks and if there aren't already, there ought to be research invested in figuring them out (e.g., in cell cultures).  Epidemiological studies are expensive and of low payoff in this kind of context.

As the authors clearly say: discovery efforts should continue vigorously" despite their points, acknowledging in fact and cleverly hedging all bets, that many variants remain to be discovered by current approaches.  This is a fair-grounds where PT Barnums thrive.  It keeps their particular circus's seats filled.  But that doesn't make it the best science in the public interest.