Showing posts with label weather. Show all posts
Showing posts with label weather. Show all posts

Wednesday, August 15, 2018

On the 'probability' of rain (or disease): does it make sense?

We typically bandy the word probability around, as if we actually understand it. The term, or a variant of it like probably, can be used in all sorts of contexts that, on the surface seem quite obvious and related to some sense of uncertainty; e.g., "That's probably true," or "Probably not."  But is it so obvious?  Are the concepts clear at all?  When are they, actually, more than just informally and subjectively, meaningful?

Will it rain today?  Might it?  What is the chance of rain?
One of the typical uses of probabilistic terms in daily life has to do with weather predictions.  As a former meteorologist myself, I find this a cogent context in which to muse about these terms, but with extensions that have much deeper relevance.

Here is an episode of a generally very fine BBC Radio 4 program called More or Less, whose mission is to educate listeners on the proper use and understanding of numbers, statistics, probabilities and the like.  This episode deals, somewhat unclearly and to me quite vaguely, unsatisfactorily, and even somewhat defensively, about the use and interpretation of weather forecasts.

So what does a forecast calling for an x% chance of rain mean?  Let's think of an imaginary chessboard laid over a particular location.  It is raining under the black, but not under the white squares.  There is nothing probabilistic about this.  50% of people in the area will experience rain.  If I don't know where you live, exactly, I'd have to say that you have a 50% chance of rain, but that has nothing to do with the weather itself but rather with my uncertainty of where you live.  Even then it's misleadingly vague since people don't live randomly across a region (they are, for example, usually clustered in some sub-regions).

Another interpretation is that I don't know where the black and white squares will be exactly, at any given time, but my weather models predict that in about half of the region, rain will fall.  This could be because my computer models, necessarily based on imperfect measurement and imperfect theory, are therefore imperfect--but I run them many times, making small random changes in various values to account for that imperfection, and I find that among these model runs, 50% of the time at any given spot, or 50% of the entire area under consideration, experiences rain.

Or, is it that there is an imaginary chessboard moving overhead and so the 50% of the land will be under the black and hence getting rain at any given time, and thus that any given area will only get it 50% of the time, but every area will certainly get rain at some time during the forecast period, indeed every area will be getting rain half of the period?  Then the best forecast is that you will get wet if you stay outside all day, but if you only run out to get the mail you might not?  Might??

Or is it that my models are imperfect but theory or experience tell me that there is a 50% chance of any rain in the area--that is, my knowledge can tell me no more than that.  In that case, any given place will have this guesstimated chance of rain.  But does that mean at any given time during the forecast period, or at every time during it?  Or is it that my knowledge is very good, but the meteorological factors--the nature of atmospheric motion and so on--only probabilistically form droplets that are large enough not just to be clouds but to fall to earth?  That is, is it the atmospheric process itself that is probabilistic--at least based on the theory, since I can't observe every droplet.

If a rain-generating front is passing through the area, it could rain everywhere along the front, but only until the front has moved past the area.  Thus, it may rain with 100% certainty, but only 50% of the specified time, if the front takes that amount of time to pass through.

I've undoubtedly only mentioned some of the many ways that weather forecast probabilities can be intended or interpreted as meaning.  It is not clear--and the BBC program shows this--that everyone or perhaps even anyone making them actually understands, or is thinking clearly about, what these probability forecasts mean.  Even meteorologists themselves, especially when dumbing down for the average Joe who only wants to know if he should carry his brolly with him, are likely ('probably'?!) unclear about these values.  Probably they mean a bit of this and a bit of that.  I wonder if anyone can know which of the meanings are being used in any given forecast.

Well, fine, everyone knows that nobody really knows everything about the weather.  Anyway, it's not that big of a deal if you get an unexpected drenching now and then, or more often haul your raincoat to work but never need it.

But what about things that really matter, like your future health?  My doc takes my blood pressure and looks at my weight, and may warn me that I am at 'risk' of a heart attack or stroke--that without taking some preventive measures I may (or probably will) have such a fate.  That's a lot more important than a soaked shirt.  But what does it mean?  Isn't everybody at some risk of these diseases?Does my doc actually know?  Does anybody?  Who is thinking clearly about these kinds of risk pronouncements?

OK, caveats, caveats: but will I get diabetes?
In genomics 'precision' genomic medicine is one of the genomics marketing slogans of the day, the very vague (I would say culpably false) promise that from your genotype we can predict your future--that's what 'precision' implies.  The same applies even if weaseling now would include environmental factors as well as genomic ones.  And the idea implies knowledge not just of some vague probability, but by implication it means perfection--prediction with certainty.  But to what extent--if any at all--is the promise, or can the promise be true?  What would it mean to be 'true'?  After all, anyone might get, say type 2 diabetes, mightn't they?  Or, more specifically, what does such a sentence itself even mean, if anything?

We know that, today at least, some people get diabetes sometime in their lives, and even if we don't know why or which ones, that seems like a safe assertion.  But to say that any person, not specifically identified, might become diabetic is rather useless.  We want a reason--a cause--and if we have that we assume it will enable us to identify specifically vulnerable individuals.  Even then, however, we don't know more than to say, in some sense that we may not even understand as well as we think we do, that not all the vulnerable will get the disease: but we seem to think that they share some probability of getting it.  But what does that mean, and how do we get such figures?

Does it mean that among all those with a given GWAS! genotype, (1) a fraction f will get diabetes?(2) a fraction f will get diabetes if they live beyond some specified age? (3) a fraction f will get diabetes before they die if they live the same lifestyle diet as those from whom the risk was estimated? (4) a net fraction f will get diabetes, pro-rated year by year as they age; (5) a net fraction related to f will get diabetes, but that is adjusted for current age, sex, race, etc.?

What about each individual consulting their Big Data genomic counselor?  Are these fractions f related to each individual as a probability p=f that s/he will get diabetes (conditional on things like items 1-5 above)?  That is, is every person at the same risk?

Only if we can equate our past sample, from which we estimated f by induction to the probability p used by deduction to assert for each new individual might this, even in principle, lead to 'precision genomic medicine'.  It is prediction, not just description that we are being promised.  Even if we were thinking in public health terms, this is essentially the same, because it would relate to the fraction of individuals who will be affected in the future, because each person is exposed to the same probability.

Of course, we might believe that each person has some unique probability of getting diabetes (related, again, to the above items), and that f reflects the mix (e.g., average) of these probabilities.  But then, we have to assume that all the genotypes and lifestyles and so on in the current group whose future we're offering 'precision' predictions is exactly like the sample from which the predictions were derived, that this mix of risks is, somehow, conserved.  How can such an assumption ever be justified?

Of course, we know very well that no current sample whose future we want to be precise about will be exactly the same as the past sample from which the probabilities (or fractions) were derived.  Obviously, much will differ, but we also know that we simply have no way to assess by how much it will differ.  For example, future diets, sociopolitical, and other factors that affect risk will not be the same as those in the past, and are inherently unpredictable.  So, on what meaningful basis can 'precision' prediction be promised?

Just for fun, let's take the promise of precision genomic medicine at its face value.  I go to the doc, who tells me
"Based on your genome sequence, I must advise you of your fate in regard to diabetes."
"Thanks, doc.  Fire away!"
"You have a 23.5% chance of getting the disease."
"Wow!  That sounds high!  That means I have a 23.5% chance that I won't die in a car or plane crash, right?  That's very comforting.  And if about 10% of people get cancer, then of my 76.5% chance of not getting diabetes, it means only a 7.65% chance of cancer!  Again, wow!"
"But wait, Doc!  Hold on a minute.  I might get diabetes and cancer, right?  About a 7.65% percent chance of that, right?"
"Um, well, um, it doesn't work quite that way [to himself, sotto voce: "at least I think so..."].....that's because you might die of diabetes, so you wouldn't get cancer.  Of course, the cancer could come first, but it would linger, because you have to live long enough to experience your 23.5% risk of diabetes.  That would not be good news.  And, of course, you could get diabetes and then get in a crash.  I said get diabetes, not die of it, after all!"
I gather you, too, can imagine how to construct many different sorts of fantasy conversations like this, even rashly assuming that your doctor understood probability, had read his New England Journal regularly when not too sleepy after a day's work at the clinic--and that the article in the NEJM was actually accurate.  And that NIH knew in sincerity what they were promising in the way of genomic predictability promises.  But wait!  The medical journals, and even the online genotyping scam companies--you can probably name one or two of them--change your estimated risks from time to time as new 'data' come in.  So when can I assume case-closed and I (well, the Doc) really knows the true probabilities?

I mean, what if there are no such true probabilities, because even if there were, not just knowledge, but also circumstances (cultural, not to mention mutations) continually change, and what if we have no way whatever to know how they're gonna change?  Then what is the use of these 'precision' predictions?  They, at best, only apply to a single, current instance.  So what (if anything at all) does 'precision' mean?

It only takes a tad of thinking to see how precisely imprecise these promises all are--must be, except very short-term extrapolations of what past data showed, and extrapolations of unknown (and unknowable) 'precision'.  Except, of course, the very precise truth that you, as a taxpayer, are going to foot the bill for a whole lot more of this sort of promises.

Unlike the weather, we don't have anything close to as rigorous an understanding of human biology and cultures as we do of the behavior of gases and fluids (the atmosphere).  We might want to say, self-protectingly and more honestly modest, that our use of 'probability' is very subjective and really just means an extrapolated rough average of some unspecifiable sort.  But then that doesn't sound like the glowing promise of 'precision', does it?  One has to wonder what sort of advice would make scientifically proper, and honorable, use of the kind of probabilistic, vague, ephemeral evidence we have when we rely on 'omics approaches, or even when it's the best we can do at present.

In meteorology, it used to be (when I was playing that game) that we'd joke "persistence is the best forecast".  This was, of course, for short range, but short range was all we could do with any sort of 'precision'.  We are pretty much in that situation now, in regard to genomics and health.

The difference is, weather forecasters are honest, and admit what they don't know.

Wednesday, January 27, 2016

"The Blizzard of 2016" and predictability: Part III: When is a health prediction 'precise' enough?

We've discussed the use of data and models to predict the weather in the last few days (here and here).  We've lauded the successes, which are many, and noted the problems, including people not heeding advice. Sometimes that's due, as a commenter on our first post in this series noted, to previous predictions that did not pan out, leading people to ignore predictions in the future.  It is the tendency of some weather forecasters, like all media these days, to exaggerate or dramatize things, a normal part of our society's way of getting attention (and resources).

We also noted the genuine challenges to prediction that meteorologists face.  Theirs is a science that is based on very sound physics principles and theory, that as a meteorologist friend put it, constrain what can and might happen, and make good forecasting possible.  In that sense the challenge for accuracy is in the complexity of global weather dynamics and inevitably imperfect data, that may defy perfect analysis even by fast computers.  There are essentially random or unmeasured movements of molecules and so on, leading to 'chaotic' properties of weather, which is indeed the iconic example of chaos, known as the so-called 'butterfly effect': if a butterfly flaps its wings, the initially tiny and unseen perturbation can proliferate through the atmosphere, leading to unpredicted, indeed, wildly unpredictable changes in what happens.
  
The Butterfly Effect, far-reaching effects of initial conditions; Wikipedia, source

Reducing such effects is largely a matter of needing more data.  Radar and satellite data are more or less continuous, but many other key observations are only made many miles apart, both on the surface and into the air, so that meteorologists must try to connect them with smooth gradients, or estimates of change, between the observations.  Hence the limited number of future days (a few days to a week or so) for which forecasts are generally accurate.

Meteorologists' experience, given their resources, provide instructive parallels as well as differences with biomedical sciences, that aim for precise prediction, often of things decades in the future, such as disease risk based on genotype at birth or lifestyle exposures.  We should pay attention to those parallels and differences.

When is the population average the best forecast?
Open physical systems, like the atmosphere, change but don't age.  Physical continuity means that today is a reflection of yesterday, but the atmosphere doesn't accumulate 'damage' the way people do, at least not in a way that makes a difference to weather prediction.  It can move, change, and refresh, with a continuing influx and loss of energy, evaporation and condensation, and circulating movement, and so on. By contrast, we are each on a one-way track, and a population continually has to start over with its continual influx of new births and loss to death. In that sense, a given set of atmospheric conditions today has essentially the same future risk profile as such conditions had a year or century or millennium ago. In a way, that is what it means to have a general atmospheric theory. People aren't like that.

By far, most individual genetic and even environmental risk factors identified by recent Big Data studies only alter lifetime risk by a small fraction.  That is why the advice changes so frequently and inconsistently.  Shouldn't it be that eggs and coffee either are good or harmful for you?  Shouldn't a given genetic variant definitely either put you at high risk, or not? 

The answer is typically no, and the fault is in the reporting of data, not the data themselves. This is for several very good reasons.  There is measurement error.  From everything we know, the kinds of outcomes we are struggling to understand are affected by a very large number of separate causally relevant factors.  Each individual is exposed to a different set or level of those factors, which may be continually changing.  The impact of risk factors also changes cumulatively with exposure time--because we age.  And we are trying to make lifetime predictions, that is, ones of open-ended duration, often decades into the future.  We don't ask "Will I get cancer by Saturday?", but "Will I ever get cancer?"  That's a very different sort of question.

Each person is unique, like each storm, but we rarely have the kind of replicable sampling of the entire 'space' of potentially risk-affecting genetic variants--and we never will, because many genetic or even environmental factors are very rare and/or their combinations essentially unique, they interact and they come and go.  More importantly, we simply do not have the kind of rigorous theoretical basis that meteorology does. That means we may not even know what sort of data we need to collect to get a deeper understanding or more accurate predictive methods.

Unique contributions of combinations of a multiplicity of risk factors for a given outcome means the effect of each factor is generally very small and even in individuals their mix is continually changing.  Lifetime risks for a trait are also necessarily averaged across all other traits--for example, all other competing causes of death or disease.  A fatal early heart attack is the best preventive against cancer!  There are exceptions of course, but generally, forecasts are weak to begin with and in many ways over longer predictive time periods they will simply approximate the population--public health--average.  In a way that is a kind of analogy with weather forecasts that, beyond a few days into the future, move towards the climate average.

Disease forecasts change peoples' behavior (we stop eating eggs or forego our morning coffee, say), each person doing so, or not, to his/her own extent.  That is, feedback from the forecast affects the very risk process itself, changing the risks themselves and in unknown ways.  By contrast, weather forecasts can change behavior as well (we bring our umbrella with us) but the change doesn't affect the weather itself.


Parisians in the rain with umbrellas, by Louis-Léopold Boilly (1803)

Of course, there are many genes in which variants have very strong effects.  For those, forecasts are not perfect but the details aren't worth worrying about: if there are treatments, you take them.  Many of these are due to single genes and the trait may be present at birth. The mechanism can be studied because the problem is focused.  As a rule we don't need Big Data to discover and deal with them.  

The epidemiological and biomedical problem is with attempts to forecast complex traits, in which most every instance is causally unique.  Well, every weather situation is unique in its details, too--but those details can all be related to a single unifying theory that is very precise in principle.  Again, that's what we don't yet have in biology, and there is no really sound scientific justification for collecting reams of new data, which may refine predictions somewhat, but may not go much farther.  We need to develop a better theory, or perhaps even to ask whether there is such a formal basis to be had--or is the complexity we see is just what there is?

Meteorology has ways to check its 'precision' within days, whereas biomedical sciences have to wait decades for our rewards and punishments.  In the absence of tight rules and ways to adjust errors, constraints on biomedical business as usual are weak.  We think a key reason for this is that we must rely not on externally applied theory, but internal comparisons, like cases vs controls.  We can test for statistical differences in risk, but there is no reason these will be the same in other samples, or the future.  Even when a gene or dietary factor is identified by such studies, its effects are usually not very strong even if the mechanism by which they affect risk can be discovered.  We see this repeatedly, even for risk factors that seemed to be obvious.

We are constrained not just to use internal comparisons but to extrapolate the past to the future.  Our comparisons, say between cases and controls, are retrospective and almost wholly empirical rather than resting on adequate theory.  The 
'precision' predictions we are being promised are basically just applications of those retrospective findings to the future.  It's typically little more than extrapolation, and because risk factors are complex and each person is unique, the extrapolation largely assumes additivity: that we just add up the risk estimates for various factors that we measured on existing samples, and use that sum as our estimate of future risk.  

Thus, while for meteorology, Big Data makes sense because there is strong underlying theory, in many aspects of biomedical and evolutionary sciences, this is simply not the case, at least not yet.  Unlike meteorology, biomedical and genetic sciences are the really harder ones!  We are arguably just as likely to progress in our understanding by accumulating results from carefully focused questions, where we're tracing some real causal signal (e.g., traits with specific, known strong risk factors), as by just feeding the incessant demands of the Big Data worldview.  But this of course is a point we've written (ranted?) about many times.

You bet your life, or at least your lifestyle!
If you venture out on the highway despite a forecast snowstorm, you are placing your life in your hands.  You are also imposing dangers on others (because accidents often involve multiple vehicles). In the case of disease, if you are led by scientists or the media to take their 'precision' predictions too seriously, you are doing something similar, though most likely mainly affecting yourself.  

Actually, that's not entirely true.  If you smoke or hog up on MegaBurgers, you certainly put yourself at risk, but you risk others, too. That's because those instances of disease that truly are strongly and even mappably genetic (which seems true of subsets of even of most 'complex' diseases), are masked by the majority of cases that are due to easily avoidable lifestyle factors; the causal 'noise' that risky lifestyles make genetic causation harder to tease out.

Of course, taking minor risks too seriously also has known potentially serious consequences, such as of intervening on something that was weakly problematic to begin with.  Operating on a slow-growing prostate or colon cancer in older people, may lead to more damage than the cancer will. There are countless other examples.


Life as a Garden Party
The need is to understand weak predictability, and to learn to live with it. That's not easy.

I'm reminded of a time when I was a weather officer stationed at an Air Force fighter base in the eastern UK.  One summer, on a Tuesday morning, the base commander called me over to HQ.  It wasn't for the usual morning weather briefing.....

"Captain, I have a question for you," said the Colonel.

"Yes, sir?"

"My wife wants to hold a garden party on Saturday.  What will the weather be?"

"It might rain, sir," I replied.

The Colonel was not very pleased with my non-specific answer, but this was England, after all!

And if I do say so myself, I think that was the proper, and accurate, forecast.**


Plus ça change..  Rain drenches royal garden party, 2013; The Guardian


**(It did rain.  The wife was not happy! But I'd told the truth.)

Tuesday, January 26, 2016

"The Blizzard of 2016" and predictability: Part II: When is a prediction a good one? When is it good enough?

Weather forecasts require the prediction of many different parameter values.  These include temperature, wind at the ground and aloft (winds that steer storm systems, and where planes fly), humidity on the ground and in the air (that determines rain and snowfall), friction (related to tornadoes and thunderstorms), change over time and the track of these things across the surface with its own weather-affecting characteristics (like water, mountains, cities).  Forecasters have to model and predict all of these things.  In my day, we had to do it mainly with hand-drawn maps and ground observations--no satellites, basically no useful radar, only scattered ship reports over oceans, etc.), but of course now it's all computerized.

Other sciences are in the prediction business in various ways.  Genetic and other aspects of epidemiology are among them.  The widely made, now trendy promise of 'precision' medicine, or the predictions of what's good or bad for you, are clear daily examples.  But as with the weather, we need some criteria, or even some subjective sense of how good a prediction is.  Is it reliable enough to convince you to change how you live?

Yesterday, I discussed aspects of weather prediction and what people do in response, if anything.  Last weekend's big storm was predicted many days in advance, and it largely did what was being predicted.  But let's take a closer look and ask: How good is good enough for a prediction?  Did this one meet the standard?

Here are predicted patterns of snowfall depth, from the January 24th New York Times, the day after the storm, with data provided by the National Weather Service:



And now here are the measured results, as reported by various observers:




Are these well-forecast depths, or not?  How would you decide?  Clearly, the maximum snowfall reported (42") in the Washington area was a lot more than the '20+"' forecast, but is that nit-picking?  "20+" does leave a lot of leeway for additional snowfall, after all.  But, the prediction contour plot is very similar to the actual result. We are in State College, rather a weather capital because the Penn State Meteorology Department has long been a top-rated one and because Accuweather is located here as a result.  Our snowfall was somewhere between 7 and 10 inches.  The top prediction map shows us in the very light area, with somewhere between 1-5" and 7-10" expected, and the forecasts were for there to be a sharp boundary between virtually no snowfall, and a large dump.  A town only a few miles north of us had very few inches.

So was the forecast a good one, or a dud?

How good is a good forecast?
The answer to this fair question depends on the consequences.  No forecast can be perfect--not even in physics where deterministic mathematical theory seems to apply.  At the very least, there will always be measurement errors, meaning you can never tell exactly how good a prediction was.

As a lead-up to the storm's arrival in the east, I began checking a variety of commercial weather companies (AccuWeather, WeatherUnderground, the Weather Channel, WeatherBug) as well as the US National and the European Weather Services, interested in how similar they were.

This is an interesting question, because they all rely on a couple of major computer models of the weather, including an 'ensemble' of their forecasts. The local companies all use basically the same global data sources, and the same physical theory of fluid dynamics, and the same resulting numerical models.  They try to be original (that's the nature of the commercial outfits, of course, since they need to make sales, and even the government services want to show that they're in the public eye).

In the vast majority of cases, as in this one, the shared data from weather balloons, radar, ground reports, and satellite imagery, as well as the same physical theory, means that there really are only minor differences in the application of the theory to the computed models.  Data resources allow retrospective analysis to make corrections to the various models and see how each has been doing and adjust them.  For the curious, most of this is, rightly, freely available on the internet (thanks to its ultimately public nature).  Even the commercial services, as well as many universities, make data conveniently available.

In this case, the forecasts did vary. All more or less had us (State College) on a sharp edge of the advancing snow front.  Some forecasts had us getting almost no snow, others 1-3", others in the 5-8" range.  These varied within any given organization over time, as of course it should when better models become available.  But that's usually when D-day is closer and there is less extrapolation of the models, in that sense less accuracy or usefulness from a precision point of view.  At the same time, all made it clear that a big storm was coming and our location was near to the edge of real snowfall. They all also agreed about the big dump in the Washington area, but varied in terms of what they foresaw for New York and, especially, Boston.  Where most snow and disruption occurred, they gave plenty of notice, so in that sense the rest can be said to be details.  But if you expected 3" of snow and got a foot, you might not feel that way.

If you're in the forecasting business--be it for the weather or health risks based on, say, your genome or lifestyle exposures--you need to know how accurate forecasts are since they can lead to costly or even life-or-death consequences.  Crying wolf--and weather companies seem ever tempted to be melodramatic to retain viewers--is not good of course, but missing a major event could be worse, if people were not warned and didn't take precautions.  So it is important to have comparative predictions by various sources based on similar or even the same data, and for them to keep an eye on each other's reasons, and to adjust.

As far as accuracy and distance (time) is concerned, precision is a different sort of thing.  Here is the forecast by our local, excellent AccuWeather company for the next several days:

This and figure below from AccuWeather.com

And here is their forecast for the days after that.



How useful are these predictions, and how would you decide?  What minor or major decisions would you make, based on your answers?  Here nothing nasty is in the forecast, so if they blow the temperature or cloud over on the out-days of this span, you might grumble but you won't really care.

However, I'm writing this on Sunday, January 24.  The consensus of several online forecasts was all roughly like the above figures.  Basically smooth sailing for the week, with a southerly and hence warm but not very stormy air flow, and no significant weather.  But late yesterday, I saw one forecast for the possibility of another Big One like what we just had.  The forecaster outlined the similarities today with conditions ten days ago, and in a way played up the possibility of another one like it.  So I looked at the upper-air steering winds and found that they seem to be split between one that will steer cold arctic air down towards the southern and eastern US, and another branch that will sweep across the south including the most Gulf of Mexico and join up with the first branch in the eastern US, which is basically what happened last week!

Now, literally as I write, one online forecast outfit has changed its forecast for the coming week-end (just 5 days from now) to rain and possibly ice pellets.  Another site now asks "Could the eastern US face more snow later this week?" Another makes no such projection.  Go figure!

Now it's Monday.  One commercial site is forecasting basically nothing coming.  Another forecasts the probability of rain starting this weekend.  NOAA is forecasting basically nothing through Friday.

But here are screenshots from an AccuWeather video on Monday morning, discussing the coming week.  First, there is doubt as to whether the Low pressure system (associated with precipitation) will move up the east coast or farther out to sea.  The actual path taken, steered by upper-level winds, will make a big difference in the weather experienced in the east.

Source: AccuWeather.com

The difference in outcomes would essentially be because the relevant wind will be across the top of the Low, moving from east to west, that is, coming off the ocean onto land (air circulates as a counter-clockwise eddy around the center of the Low).  Rain or possibly snow will fall on land as the result.  How much, or how cold it will be depends on which path is taken.  This next shot shows a possible late-week scenario.

Source:  AccuWeather.com
The grey is the upper-level steering winds, but their actual path is not certain, as the prior figure showed, meaning that exactly where the Low will go is uncertain at present.  There just isn't enough data, and so there's too much uncertainty in the analysis, to be more precise at this stage.  The dry and colder air shown coming from the west would flow underneath the most air flowing in from offshore, pushing it up and causing precipitation.  If the flow is more eastward of the alternatives in the previous figure, the 'action' will mainly be out at sea.

Well, it's now Monday afternoon, and two sites I check are predicting little if anything as of the weekend....but another site is predicting several days in a row of rain.  And....(my last 'update'), a few hours later, the site is predicting 'chance of rain' for the same days.

To me, with my very rusty, and by now semi-amateur checking of various things, it looks as if there won't be anything dropping on us.  We'll see!

The point here is how much things change and how fast on little prior indication--and we are only talking about predicting a few days, not weeks, ahead.  The above AccuWeather video shows the uncertainty explicitly, so we're not being misled, just advised.

This level of uncertainty is relevant to biology, because meteorology is based on sophisticated, sound physics theory (hydrodynamics, etc.).  It lends itself to high-quality, very extensive and even exotic instrumentation and mathematical computer simulation modeling.  Most of the time, for most purposes, however, it is already an excellent system.  And yet, while major events like the Big Blizzard this January are predictable in general, if you want specific geographic details, things fall short.  It's a subjective judgment as to when one would say "short of perfection" rather than "short but basically right.".

With more instrumentation (satellites, radar, air-column monitoring techniques, and faster computers) it will get inevitably better.  Here's a reasonable case for Big Data.  However, because of measurement errors and minor fluctuations that can't be detected, inaccuracies accumulate (that is an early example of what is meant by 'chaotic' systems: the farther down the line you want to predict, the greater your errors.  Today, in meteorology, except in areas like deserts where things hardly change, I've been told by professional colleagues who are up to date, that a week ahead is about the limit.  After that, at least under conditions and locations where weather change is common, specific conditions today are no better than the climate average for that location and time of year.

The more dynamic a situation--changing seasons, rapidly altering air and moisture movement patterns, mountains or other local effects on air flow, the less predictable over more than a few days. You have to take such longer-range predictions with a huge grain of salt, understanding that they're the best theory and intuition and experience can do at present (and taking into account that it is better to be safe--warned--than sorry, and that companies need to promote their services with what we might charitably call energetic presentations).  The realities are that under all but rather stable conditions, such long-term predictions are misleading and probably shouldn't even be made: weather services should 'just say no' to offering them.

An important aspect of prediction these days, where 'precision' has recently become a widely canted promise, is in health.  Epidemiologists promise prediction based on lifestyle data.  Geneticists promise prediction based on genotypes.  How reliable or accurate are they now, or likely to become in the predictable future?  At what point does population average do as well as sophisticated models? We'll discuss that in tomorrow's installment.

Monday, January 25, 2016

"The Blizzard of 2016" and predictability: Part I--the value of prediction

Mark Twain famously quipped, "Everybody talks about the weather but nobody does anything about it." But these days, that's far from accurate.  At least, an army of specialists try to predict the weather so that we can be prepared for it.  The various media, as well as governmental agencies, publicize forecasts.  But how good are those forecasts?

As a former meteorologist myself (back--way back--when I was an Air Force weather officer), I take an interest, partly professional but also conceptual, in how accurate forecasting has become in our computer and satellite era.

Last week, a storm developed over the southwest, and combined with atmospheric disturbance barreling down from the Canadian arctic, to cause huge rain and wind damage across the south and then veered north where it turned into "The Blizzard of 2016", dubbed by the exaggeration-hungry media.  How well was it forecast and did that do any societal good?

Here is a past-few-days summary page of mapped conditions at upper air (upper left), surface (upper right) and other levels.  On a web page called eWall ( http://mp1.met.psu.edu/~fxg1/ewall.html ) you can scroll these for the prior 5 days.  The double Low pressure (red L's) on the right panel represent the center of the storm, steered in part by the winds aloft (other panels).



If you followed the forecasting over the week leading to the storm's storming up the east coast to wreak havoc there, you would say it was exceedingly well forecast, and many days in advance. Was it worth the cost?  One has to say that probably many lives were saved, huge damage avoided, and disruption minimized: people emptied grocery store shelves and hunkered down to watch the Weather Channel (and State College's own Accuweather).  Urgent things, including shopping for supplies in case of being house-bound, were done in advance and probably many medical and other similar procedures were done or rescheduled and the like.  Despite the very heavy snowfall, as predicted, the forecast was accurate enough to have been life-saving.

Lots of people still don't do anything about it!
And yet....
Despite a lot of people talking about the weather, on all sorts of public media, masses of people, acting like Mark Twain, don't do anything about it, even with the information in hand.  At least 12 people died in this storm in accidents, and others from coronaries while shoveling, and this is just what I've seen in a quick check of the online news outlets.  Thousands upon thousands were stranded for many hours in freezing cold on snow-sodden highways.  There were things like 25-mile-long stationary lines of vehicles on interstates and thousands of car and truck accidents.  That's a lot of people paying the price for their own stubbornness or ignorance.  This is what such a jam looks like:

A typical snowstorm traffic jam (www.breakingnews.com)
People were warned in the clearest terms for days in advance.  Our fine National Weather Service, in collaboration with complementary services in other countries, scoped out the situation and let everyone know about it, as is their very important job.  Some states, like New York,  properly closed their roads to all but necessary traffic. Their governments did their jobs.  Other states, like Kentucky, failed to do that.  So then, how is it that there was so much of what seems like avoidable damage?

Let's put the issue another way: My auto insurance rates will reflect the thousands of costly claims that will be filed because of those who failed to heed the warnings and were out on the highways anyway. So I paid for the forecasts first through my taxes, and then through the purchase prices of goods whose makers pay to advertise on weather channels, but then I also have to pay for those whose foolhardiness led to the many accidents they'll make claims for.  That's similar to people knowingly enjoying an unhealthy lifestyle, and then expecting health insurance to cover their medical bills--that insurance, too, is amortized over the population of insured including those who watch their lifestyles conscientiously.  That's the nature of insurance.

Some people, of course, simply can't stay home.  But many just won't.  Countless truckers were stranded on the roads.  They surely knew of the coming storm.  Did commercial pressure keep them on the road?  Then shame on their companies!  They surely could have pulled over or into Walmart parking lots to wait out the snowfall and its clearance--a day or so, say.  Maybe there aren't enough parking lots for that, but surely, surely they should not have been on the Interstates!  And while some people probably had strong legitimate reasons for being out, and a few may not have seen the strong, repeated forecasts over the many preceding days, most and I would say by far the most, just decided to take their trips anyway.

Nobody can say they aren't aware of pileups, crashes, and hours-long stalls that happen on Interstates during snowstorms.  It is not a new phenomenon!  Yet, again, we all will have to pay for their foolhardiness.  Maybe insurance should refuse to cover those on the road for unnecessary trips. Maybe those who clog the roads in this way should be taxed to cover the costs of, say, increased insurance rates on everyone else or emergencies that couldn't be dealt with because service vehicles couldn't get to the scene.

The National Weather Service, and companies who use their data, did a terrific job of alerting people of the coming storm, and surely saved many lives and prevented damage as a result.  Just as they do when they forecast hurricanes and warn of tornadoes.  Still, there are always people who ignore the warnings, at their own cost, and at cost to society, but that's not the fault of the NWS.

But what about predictability? Did they get it right?  What is 'right'?
It is a fair and important question to ask how closely the actual outcome of the storm was predicted.   The focus is on the accuracy in detail, not the overall result, and that leads one to examine the nature of the science and--of course in our case here on this blog--to compare it with the state of the art of epidemiological, including genetic, predictions.  Not all forecasts are as dramatic and in a sense clear-cut as a major storm like this one.

I have been in the 'prediction' business for decades, first as a meteorologist and subsequently in trying to understand the causal relationships, genetic and evolutionary, that explain our individual traits.  Tomorrow, we'll discuss aspects of the Big Storm's forecasts that weren't so accurate and compare that with the situation in these biological areas.