Showing posts with label DNA sequencing. Show all posts
Showing posts with label DNA sequencing. Show all posts

Wednesday, September 25, 2013

Incidentally,...... (Interpreting incidental findings from DNA sequence data)

Genome sequencing yields masses of data.  It's one of the founding justifications of the current cachet term Big Data.  The jury is still out on how much of it is meaningful in any sort of clinical way as opposed to other sorts of data; that's fine, it's early days yet.  But too much of current thinking seems to rest on ideas that are outdated in fundamental ways.  It would be helpful to get beyond this.

A new paper in The American Journal of Human Genetics, ("Actionable, Pathogenic Incidental Findings in 1,000 Participants’ Exomes", Dorschner et al.) reports on a study of gene variants in 1000 genomes of participants in the National Heart, Lung, and Blood Institute Exome Sequencing Project.  How many variants associated with genetic conditions that might be undiagnosed does each individual carry?  This is addressing the issue of how much individuals should be told about "incidental findings" in their genome or exome sequences.

The investigators looked at single nucleotide variants in 114 genes in 500 European Americans and 500 African American genomes.  They found 585 instances of 239 unique variants identified by the Human Gene Mutation Database as disease-causing.  Of these, 16 autosomal-dominant variants in 17 people were thought to be potentially pathogenic; one individual had 2 variants.  A smattering of other variants not listed in HGMD were found as well.  The paper reports a frequency of ~3.4% and ~1.2% of pathogenic variants in individuals of European and African descent, respectively.

The 114 genes were chosen by a panel of experts, and pathogenicity determined by the same. 
“Actionable” genes in adults were defined as having deleterious mutation(s) whose penetrance would result in specific, defined medical recommendation(s) both supported by evidence and, when implemented, expected to improve an outcome(s) in terms of mortality or the avoidance of significant morbidity.
Variants were classified as pathogenic, likely pathogenic VUS (variant of uncertain significance), VUS and likely benign VUS.  Classification criteria included allele frequency of the variant (if low in the healthy population, it was considered to be more likely pathogenic than if high, relative to disease frequency), segregation evidence, number of reports of affected individuals with the variant, and whether the mutation has been reported as a new mutation or not. The group decided not to return VUS incidental findings to the individual if the variant was in a gene unrelated to the reason they were included in the study in the first place.

Reviewers of the data followed stringent criteria to classify alleles.  For example, variants were considered suspect if they were identified by the HGMD as disease-causing.  But they were not considered disease-causing if the allele frequency was common enough that this meant, relative to disease frequency, that the allele alone couldn't be causal.

This raises a point that we've made before, too often to a deaf audience; the variant is not 'dominant' and the 150 year old term due to Mendel should be dropped from usage, in favor of a more accurate conception of probabilistic causation (see below)If the allele was more common than the disease, other alleles or factors must also be involved.  Or, perhaps the allele was improperly classified as causal in the first place.
Punnett square showing results of crossing yellow and green peas; Wikipedia

And,"maximum allowable allele frequencies for each disease were calculated under a very conservative model, including the assumption that the given disorder was wholly due to that variant."  Disease frequencies were overestimated when they weren't known. But, is dominance a function of disease frequency?  Suggesting such a thing should raise big red flags about semantics and the conceptual working frameworks being used.
The eight participants with confirmed pathogenic (versus likely pathogenic) mutations included three with increased risk of breast and ovarian cancer (MIM 604370, caused by BRCA1 mutations, or MIM 612555, caused by BRCA2 mutations), one with a mutation in LDLR, associated with familial hypercholesterolemia (MIM 614337), one with a mutation in PMS2, associated with Lynch syndrome (MIM 614337), and two with mutations in MYBPC3, associated with hypertrophic cardiomyopathy (MIM 115197), as well as one person with two SERPINA1 mutations, associated with the autosomal-recessive disorder alpha-1-antitrypsin deficiency (MIM 613490).
Fewer actionable alleles were found in African Americans than European Americans, presumably because fewer studies have been done in this population and fewer causal alleles identified.  Dorschner et al. did not have access to phenotypes of the people included in this study, so can't know their health status with respect to these variants.  Nor, of course, can they know whether individuals might have a condition for which a causal allele was not found. 

They report few pathogenic alleles, though the fact that they looked only for alleles associated with adult onset conditions could partially explain this.  And, their criteria were stringent.  And, of course, they were looking only for single gene disorders, so it's not a surprise that they identified so few potentially pathogenic alleles, in fact. Of course, single gene disorders are only a small subset of conditions that might affect us.

A 2011 Cell paper we've mentioned before ("Exome sequencing of ion channel genes reveals complex variant profiles confounding personal risk assessment in epilepsy," Klassen et al.) looks at this question from a different angle.  Klassen et al. compared the exomes of 237 ion channel genes (known to be associated with epilepsy) in affected and unaffected people.  They found rare variants in Mendelian disease genes at equivalent prevalence in both groups.  That is, healthy people were as likely to have purportedly causal variants as those with sporadic, idiopathic epilepsy.  They caution that finding a variant is only a first step.

The unjustified dominance of 'dominance'
We feel compelled to comment again, and further, on the terminological gestalt involved in papers such as these, as we commented yesterday (and have done on earlier posts).

The concept of dominant single-locus causation goes back to Mendel, who carefully chose traits in peas that worked that way.  He knew other traits didn't.  He learned that not all 'dominant' traits showed 'Mendelian' inheritance and even came to doubt his own theory later in life.

There are traits that seem to 'segregate' in families in classical Mendelian fashion.  There is, for such traits, a qualitative (e.g., yes/no) relationship between genotype and phenotype (trait).  When a dominant allele is present, you always get the trait.  This means we see (in the case of dominance) about 1/2 of offspring of an affected parent who are affected.   Much of the 20th century in human genetics was spent trying to fit patterns of inheritance to such single-gene models.  But it was very clear that dominance (even when there was evidence for it) wasn't dominance!  The traits were not always black and white (or should we say green and yellow?).  And the segregation proportion wasn't 50%.  What to do?

The conviction that the trait was due to the effects of a single locus seemed to have strong support, so the concept of 'penetrance' was introduced.  This is not the first time that a fudge factor has been used to force a model to fit data when it didn't really.  The idea in this case is that the inheritance of the allele (variant) from the parent had a 50% probability, but that the allele, once present, did not always cause the trait.  If you can add a factor of 'incomplete penetrance', that can vary from 0 to 1, then you can fit a whole lot of data that otherwise wouldn't support Mendelian causation.

What we now know is that there are many variants at genes associated with single-gene traits, that other genes almost always also contribute (along with environmental factors as well, most of the time), and that the trait itself is quantitative: the same 'A' dominant allele doesn't always cause the same degree of severity and so on.  In other words, the trait is mainly (in a statistical sense) due to the presence of variation in a single gene, but the effect depends on the specific allele that is involved in a given case and also is affected by the rest of the genome.

In other words, there is a quantitative relationship between genotype and phenotype.  This is a general, accurate description of the pattern of causation.  The pattern is not 'Mendelian' and the causation is not dominant.  Or, to be clearer, the extreme cases are close to, or even exactly, single-allele dominant, but this is the exception that proves (tests) the quantitative-relationship rule.

We should stop being so misled by constraining legacy terminology. Mendel did great work.  But we don't have to keep working with obsolete ideas.

Wednesday, August 28, 2013

Cystic fibrosis, genetic variation and gene function

We wrote last week about the difficulty in assessing gene function, including determining whether a given genetic variant is in fact responsible for a disease or trait. The ever-decreasing cost of DNA sequencing will add to the rapidly growing set variants of undetermined function -- some will be associated with a disease, but some will have no effect.  This is not only a challenge to human geneticists trying to characterize gene function, but also to clinicians trying to explain, predict or treat disease. Last week's post was in the context of X-linked intellectual disorders, but the problem is ubiquitous.  Today we write about cystic fibrosis.

More than 2000 variants in the cystic fibrosis transmembrane conductance regulator gene (CFTR) have been linked to the disease.  About 70% of Europeans with cystic fibrosis have 2 copies of a 3 base deletion, called ΔF508 because the deletion occurs at position 508 in the gene.  This deletion causes a particular piece of the CFTR protein not be made when the protein is being synthesized, and this leads to an abnormal protein, and disease. That there can be one predominant allele in this 'recessive' disease is consistent with population genetics theory and also with the early discovery of the gene--because a high fraction of cases have two copies of the allele (or at least one copy, along with some other variant) made it findable with the techniques of a generation ago.

Why this deletion causes the particular symptoms of cystic fibrosis is well-understood, but why most of the remaining 2000 variants cause disease has not been demonstrated. Indeed, most of them are quite rare and since this disease is "recessive" cause little if any effect on their own.  Now, a new paper in Nature Genetics ("Defining the disease liability of variants in the cystic fibrosis transmembrane conductance regulator gene," Sosnay et al.) addresses this question.

Data collected for the study; Sosnay et al., Nature Genetics, 8/25/13
Sosnay et al. collected phenotype and genotype data from nearly 40,000 people with cystic fibrosis in Europe and North America.  Phenotypes vary in how severe they are and whether, for example, they involve just lung and breathing issues or also obstruct pancreatic ducts causing other problems, and so on.

In this sample, 159 CFTR variants were at a frequency greater than 0.01%, accounting for 96% of the cases--that is, of people known from symptoms to be affected, even if the severity is not uniform.  Of these, they confirmed that 80% met the clinical and functional definition of cystic fibrosis.  Information from about 2000 fathers of people with CF allowed the investigators to determine that a handful of the remaining variants seemed to have no effect, and the effect of another handful is still undetermined.  That the investigators looked for effect in cases' fathers is somewhat curious if the disease is recessive, because one would expect no effect--but see below.

This study contributes new and useful information to knowledge of the genetics of the disease.  Understanding whether, or better, how, a variant causes disease is important for determining carrier status in prospective parents, and can be an important part of understanding how to treat the disease. But the task is never-ending, as new variants arise with every meiosis -- most will be benign, but some will not.

Some 'meta' explanation of this approach
Once the ΔF508 allele was found, it was shown to have a strong effect in carriers of two copies.  This was consistent with the idea that CF is "recessive", and in a kind of circular strategy this is what was found because the variant is so common.

But then other patients were found who had only one copy of that allele, plus a not-normal sequence in their other copy.  Because the trait was assumed to be "recessive", this other variant gets incorporated into the data base as if it's causal.  This is a kind of circular reasoning gone another step: if the disease is assumed to be recessive, then the variants in both the patient's copies of the gene must be causal!

In some cases, the nature of both mutations in the gene, in terms of where in the protein structure the variant occurred was able to confirm a likely causal effect.  This is consistent with the current paper.  However, it was obvious that the phenotypes were not all the same, and many if not most individual had two different variants.  That is not what "recessive" is classically supposed to mean--that is, two copies of the 'bad' variant.  Instead, we have a quantitative relationship between the diploid genotype (that is, both copies of the gene) and the severity of the trait.

If this assumption is wrong, it could be that one or even both variants the patient has are not themselves causal, but only causal in the context of some other site(s) in the genome that interact with the CFTR gene, or even that have similar effects on their own.  So to continue to use the term we esssentially get from Mendel and his peas ("recessive") when what we really have is a quantitative relationship between genotype and phenotype, that may not even always involve the same gene, is an example of a theory lasting beyond the evidence, because investigators 'want' the trait to fit the simple model.  The persistence of the simple model shows our addiction to it and the historical legacy by which terms and concepts cling on when they should be modified--or abandoned.  This is a common iassue that applies to many purportedly single-gene traits.  It's why we put the word in quotes in this post.

 Of course when a variant is very rare, and is generally seen in patients who also carry one of the strong-effect alleles, we have almost no way to test whether that variant is causal or not.  If it hits a known major part of the gene, some severity correlations can be identified--and this has been known since about 1990.  The current paper provides some larger-scale documentation, but really is confirming what has long been well known.  Naturally, it is no surprise that the authors could not attribute specific mechanism to the almost 2000 genes in their study, many of which will be singletons, and why they'll be so hard to confirm.  This requires enough samples, or some clear form of experimental evidence--difficult to obtain for many different alleles of various types--to get statistical evidence to show tha the specific variant is involved.

So here we see an example of the interplay between sampling, theory, inference, and the challenging problem of understanding genetic causation--even in one of the classical 'simple' diseases.

Thursday, April 4, 2013

When the crystal ball is cloudy: calling sequence data correctly

Here's a monkey wrench of a paper (O'Rawe et al.), just published in Genome Medicine.  We're all being sold on the idea that knowing our whole genome sequence is going to make us much healthier. The DNA sequencer cum crystal ball will tell us what we're likely to be in for, and this will give us plenty of lead time to prevent it -- by running, lowering our cholesterol intake, losing weight, or whatever -- or to prepare for it.

But, among many other assumptions, this assumes first and foremost that the data are being read correctly, no false positives or negatives.  And here's the clincher: O'Rawe et al. compared five different software packages that read and interpret DNA sequence data, and they report low concordance between results.  Discrepancies have been found before, but not when comparing reads of the same raw data.

This group sequenced whole exomes of 15 different individuals, in 4 families, and fed the raw data through 5 sequence analysis pipelines.  They also sequenced one whole genome.  Sequences were done at 20 - 154X coverage, 120X average, meaning each nucleotide was read at least 20 times, but most often more, and at least 80% of the target sequence was obtained.

They found that the 5 programs agreed on single nucleotide variants (SNVs) about 60% of the time.  That is, 40% of the time a SNV was called by fewer than 5 of the programs.  Each of the pipelines detect variants that the others do not, and they aren't necessarily all false positives.
This disagreement is likely the result of many factors including alignment methods, post alignment data processing, parameterization efficacy of alignment and variant calling algorithms, and the underlying models utilized by the variant calling algorithm(s).
That is, each step along the way potentially introduces errors.  Indel (insertion/deletions, segments of DNA one or more nucleotides in length) concordance rates were even lower, at 26% between three indel calling programs.  (The paper goes into much more detail about specific pipelines and error rates.) Using family data can help reduce inaccuracies when it is possible to determine which calls just cannot be correct.  But, otherwise, with current methods reducing false positives means increasing false negatives, and vice versa.

The authors write,
In the realm of biomedical research, every variant call is a hypothesis to be tested in light of the overall research design. Missing even a single variant can mean the difference between discovering a disease causing mutation or not. For this reason, our data suggest that using a single bioinformatics pipeline for discovering disease related variation is not always sufficient.
This somewhat understates the problem.  Serious level testing of a SNP (single nucleotide polymorphism) to see if it has an effect on disease risk--especially when these effects are typically very small in any case, and biased upwards in GWAS type data, is no joke.  What do you do?  Put that single change into a lab mouse or rat and see if it might be more likely to develop slightly higher blood pressure at old age?  Or have a slightly higher risk of some sort of cancer (again, to be a human model, it should be at older ages)?  Which mouse strain would you use?  If humans are to be used for validation, how would you do it?

The questions are serious because miscalls by sequencers go both ways.  A sequencer can miss a SNV call, so you don't identify one of the variants that you really want to be checking.  Or, it can give you a false positive, and lead you farther astray.  And if you must choose between hundreds of variants across the genome, with comparable estimated effects, you are already in a bit of a bind even if they are all perfectly called!

No technology, or medical test, will be correct 100 percent of the time, and sequencing technologies are likely to get better, not worse (though if MS Windows is any guide, that's not necessarily true!). But, when disease risk estimates depend on accurate DNA sequence, it is obvious that we are way premature in proclaiming findings so loudly and demanding that so much effort and resources be poured into doing more of the same.  Again, focused studies on problems more important, clearer, and less vulnerable to these kinds of errors is where the effort should be going.

And, some subtle manipulation, too?
By the way, the standard term for a single nucleotide variation in a population is SNP (single nucleotide polymorphism).  Now, some authors use SNV (single nucleotide variant), essentially doing two things.  First, they are rhetorically equating 'variant' with causal variant--that is, tactily, subtly, or surreptitiously planting in your mind that they are onto something causal.  And second, they are tacitly, subtly, or surreptitiously suggesting that one of the two is the 'good'  or 'normal' (i.e., health-associated) variant.  This perpetuates the 'wild type' thinking--see our earlier post 'walk on the wild-type side'.

These are ways in which the community of researchers inadvertently or intentionally (you decide which) cooks the books in your and journalists', and even their own minds, entrenching a de facto genetic-causation worldview into their and everybody's thinking.  That's good for business, of course.

Friday, February 11, 2011

The human genome birth month

Ten years ago this week, Nature and Science published results of the first pass at sequencing the human genome.  And now Nature and Science are feting the event with commentaries from people who were central to the initial effort, as well as assessments of where we are now. Science in fact will have 4 issues celebrating the birth month of the human genome sequence, not just a single issue. We will blog about a few of these commentaries in the next few days.

A News Focus from the Feb 4 issue of Science asks how all we've learned in the last 10 years is being translated into clinical practice.  Titled "Waiting for the Revolution", the message is that if physicians saw some practical use for genetic information, they'd be using it, but they don't and they aren't.  One consistent message from those who defend the commitment of time and money on genetic research is that it's early days yet, and we can't expect miracles overnight.  The practical use of genetics that genetics enthusiasts frequently cite is sequencing of tumors to allow an informed choice of treatment.

There are a few noted clinical decisions that can be made based on genetic analysis, but they are far fewer than had been hoped (or promised) 10 years ago.  The amount of genetic information that can be collected on any individual continues to far outweigh its usefulness for disease prediction or clinical applicability.  For example, should individuals with a history of deep-vein blood clots be tested for known genetic risk factors, factor V Leiden and prothrombin gene variants?  According to the News piece,
Both genes influence such clotting. People who have had such clots should be treated with anticoagulants anyway, regardless of genetic status, the panel concluded. And in a second group—relatives of people who have had clots but who themselves have not—the panel judged that it would be too risky to treat preemptively with anticoagulants (which can cause hemorrhaging) based on genetic status alone.
And the same is true for many other conditions.  Treat it when it happens, but preventive treatment is not useful, or is even risky, so genotyping is an unnecessary added cost.  Researchers have been looking for genes for diabetes for decades, with little success, for example.  A colleague of ours recently told Ken that although he's been working in the field of heart disease and lipid research genetics for decades, focusing on ApoE (one of the best candidate genes that has been studied ad nauseum), he still cannot put any of the knowledge about ApoE into his internal medicine practice.  

But let's say they find genes for whatever their favorite disease (an outcome that isn't at all guaranteed).  In what sense will that help prevent or treat it?   Does it really matter as a rule?  If you just eat better, exercise and don't smoke, your risk of heart attack will drop by far more than even the most enthusiastic gene-O-hyper can claim.  Enthusiasts will say that genetics will transform medicine to target treatment, itself based on molecular technologies, to each individual's particular risk.  The truth will be a mix, and each person has his or her own view on where the balance, and the distribution of resources, should lie.

Stepping back, we can say that it's always true that most work in a field is pedestrian, unremarkable, journeyman-like but not spectacular.  A century (or more) later we may remember the 'genius' but forget the unheard-of peons in a given field.  Our era is distinguished by its materialistic greed and immodesty, but is otherwise the same as it's always been.  There will be expensive trails of cow-droppings, but there will also be gains, some of them important.  They may or may not correlate to the bravado of those today who boast about their successes, lobby for funds, and so on.  Most of our work will be tiny, incremental, incidental, or irrelevant.

A century from now there may or may not be blogs in which reflection can be seen to look back on the early 21st century to point out where the remarkable insights were.  But we can be sure that then, as now, there will be both the remarkable minority, and the pedestrian majority.

Monday, August 31, 2009

Gutting it out

Only a small fraction of bacteria can be grown in laboratories; apparently nobody understands what they need in their environment well enough. This can be a problem for microbiologists trying to identify the bacterium causing a new infectious disease, but it also means that it has not been possible to know all of the little bugs to which we our bodies are willing or unwilling hosts. The same would be true for other animals, wild ones as well as our pets and farm species. We say 'willing or unwilling' because, of these pathogens, some are presumed to be harmless commensals, and others are necessary to our survival (such as the E. coli in our intestine, that we depend on for digestion, and similarly for other mammals such as grazers like cows and goats who need bacteria in their rumens to digest cellulose).

One of the characteristics of bacteria is that they can exchange structures (e.g., plasmids) that contain some actively used genes. This is where many if not all of the genes are that lead the bacteria to resist nasty things in their environments such as antibiotics that bacterial targets have evolved to protect themselves. A vulnerable strain of bacteria can acquire a gene that makes them antibiotic resistant. This is of course a very important current problem in farm animals and humans, driven by the amount of antibiotics we ingest. And farm animals are reservoirs of pathogens that can affect humans, so humans and the animals they live with are part of large bacterial-host ecosystems.

Since we can't culture most bacteria, our knowledge of who's where, and the characteristics of many bacterial species, has been quite limited. But DNA sequencing technology has opened the way to identifying our visitors. By extracting all the DNA from a sample of some tissue, fragmenting the DNA and sequencing the fragments, their owners can be identified, and their genetic makeup and function characterized. This is done by comparing the sequence fragments against all sequences currently known (in Genbank). Even if we don't identify an exact match, we can find a known species that is close enough to our tissue-sampled sequence to identify its place in the bacterial tree of life. Antibiotic resistance genes can be identified in the same way, too, since many are already known.

A new paper in Science (Functional Characterization of the Antibiotic Resistance Reservoir in the Human Microflora, Sommer et al., Aug 28, 2009, 1128-1131) reports on a detailed analysis of the antibiotic resistance genes found in the microbiome, the resident bacteria, of healthy individuals. The study was done in an effort to learn more about how antibiotic resistance genes are acquired by pathogens that infect humans.

This study found that, indeed, many of the resistance genes in multidrug resistant bacteria were acquired by lateral gene transfer--that is, the gene hopped into the pathogen on a plasmid from a different bacterium. Bacteria in the wild usually live in large communities of many different species, where they can promiscuously exchange plasmids, thus antibiotic resistance can spread rapidly.

The authors wondered if the extensive history of antibiotic use in humans might mean that microflora in the human gut might be a reservoir of resistance genes readily transferable to pathogens. They isolated DNA from saliva and fecal samples from two healthy individuals who had not taken antibiotics for at least a year, and sequenced fragments of bacterial DNA that were resistant to all the antibiotics they tested.

They found that there was some, but by no means complete overlap between the two people who were sampled, and that nearly half the resistance genes they identified in this way were identical to resistance genes in human pathogens. This doesn't tell them whether gene transfer went from the commensal microflora to the pathogen or vice versa, but Sommer et al. suggest that it's quite plausible that our gut microflora are a reservoir of resistance genes just waiting to jump into pathogens which are now controllable with drugs. The remaining genes, although not yet found in pathogens, were functional when transferred to E. coli, suggesting that if they do eventually find their way into pathogens, they will be active, although there seems to currently be a barrier to lateral gene transfer which isn't yet understood.

The authors conclude:
Many commensal bacterial species, which were once considered relatively harmless residents of the human microbiome, have recently emerged as multidrug-resistant disease-causing organisms. In the absence of in-depth characterization of the resistance reservoir of the human microbiome, the process by which antibiotic resistance emerges in human pathogens will remain unclear.
This study provides interesting ecological information about bacterial dynamics, and of course warns us that antibiotic resistance may be more complex, challenging, and difficult to predict than we have thought. It's evolution in action.