Given a Storm
Severe-storm environment analysis is excellent at saying what a storm can become, poor at saying whether one will form, and its sharpest tools quietly assume the storm already exists.
2026-09-15
The first tornado forecast in history, issued at Tinker Air Force Base in March 1948, is the origin of severe convective environment analysis: read the air around a storm and say what it will do. This episode argues that the method has become excellent at one question, what kind of storm forms if one forms, and remains poor at another, whether one forms at all, and that the two are routinely confused. It walks through the proximity-sounding climatologies that settled the first question, the sensitivity and field-campaign results that limit the second, and the 2024 and 2025 storm-motion papers showing that even the best tornado discriminators depend on the storm already existing. It closes with the 1999 re-analysis of the 1948 charts, which found the two Tinker days were not alike after all.
Transcript
Follows the audio as it plays — tap any sentence to jump there.
On the evening of the twentieth of March, nineteen forty eight, the forecaster on duty at Tinker Air Force Base in Oklahoma had promised gusty winds and nothing worse. Just before ten at night a tornado crossed the base and wrecked more than fifty aircraft. It was the most expensive tornado Oklahoma had ever seen. Nobody had forecast it, and the official inquiry concluded that nobody could have.
Two officers at the base weather station, Major Ernest Fawbush and Captain Robert Miller, did not accept that. They spent the next few days pulling charts from past tornado days, looking for a common pattern. On the morning of the twenty fifth they thought the new charts looked like the old ones, and they told the base commander, General Fred Borum. In the afternoon the general asked whether they believed tornadoes were likely near the base. When they said the probability was high enough to justify a warning, his entire reply was: do it. At three o'clock they issued the first operational tornado forecast in history, valid from four to six. At about six a second tornado hit the base. Miller later put the odds of that at twenty million to one.
That is the founding story of severe convective environment analysis: the idea that you can look at the air around a storm, before the storm, and say what it will do. Seventy eight years on, the same idea runs the Storm Prediction Center, most forecast offices, and a good share of the research literature. My argument today is that this method became very good at one question and is still poor at a second, and that the two are confused constantly, including in the founding story.
The first question: given that a storm forms here, what kind of storm will it be? The second: will a storm form here at all? Call them the conditional question and the unconditional one. Here is the picture.
A balloon rises through the atmosphere and reports temperature, moisture and wind at every height. Think of it as a core sample of the air. From that column you compute a few numbers. One is CAPE, convective available potential energy. Lift a bubble of surface air. If it stays warmer than its surroundings it keeps rising, and CAPE is the size of the hill it gets to roll down. Another is wind shear, how much the wind changes speed and direction as you climb the column. Shear is what makes a storm tilt and rotate instead of raining out on itself. A third is the height of the cloud base, a measure of how moist the air near the ground is.
The conditional question was largely settled around the turn of the century. In nineteen ninety eight Erik Rasmussen and David Blanchard took every evening balloon sounding from the United States in nineteen ninety two, matched each to the storms nearby, and sorted them into ordinary thunderstorms, supercells, and supercells with significant tornadoes. Combining CAPE with low level rotation did the best job of separating supercells from ordinary storms, though with plenty of overlap. And the single best number for picking out the tornadic supercells was neither CAPE nor shear. It was cloud base height. Half the tornadic soundings had cloud bases below eight hundred metres. Half the non tornadic supercells had them above twelve hundred.
Five years later Richard Thompson and colleagues at the Storm Prediction Center repeated the exercise with hourly model analyses instead of balloons, on four hundred and thirteen storms. Deep layer shear separated supercells from ordinary storms with no overlap between the bottom tenth of the supercells and the top tenth of the others. Low level shear and low level moisture separated the tornadic supercells. They folded the best ingredients into two products, the supercell composite parameter and the significant tornado parameter, which now sit on every severe weather desk in the country.
Now notice what every one of those studies did. They started from storms. A proximity sounding is, by definition, a sounding near a storm that happened, so every number in those tables is conditional on a storm existing. Rasmussen and Blanchard said so themselves: a sounding can only tell you whether conditions are generally favourable, and the forecaster must then look for the mesoscale features, the boundaries, that decide where a storm actually appears. Thompson warned against using his parameters as magic numbers.
The unconditional question is a different animal, and the reason is a nineteen ninety six paper that I think is still underrated. Andrew Crook ran a cloud model over a heated, moist boundary layer and asked how much he had to change the low level air to switch a storm on or off. The answer was one degree of temperature, or one gram of water vapour per kilogram of air. Those are the ordinary measurement errors of a summer afternoon. The difference between a clear evening and a severe storm can live below the resolution of the observing network. Crook concluded that initiation has limited predictability with the observations we have. Six years later the IHOP field campaign sent aircraft and radars to the boundaries where storms were expected to form, and most missions came back with no storm. Colliding boundaries do not always fire, and dry lines were not even the leading trigger that summer.
Put the halves together and you get a base rate problem. Reanalysis climatologies count how often the air over a place looks favourable, and favourable hours vastly outnumber storms. A high CAPE, high shear afternoon in Oklahoma in May is usually just a hot afternoon. The parameters are honest about severity and silent about occurrence. That is also why they travel badly. Mateusz Taszarek's ERA5 climatology found the tail of the CAPE distribution in the United States runs to twice the European values. A study of Chinese supercells published last year found that the thermodynamic ingredients of the American tornado parameter, CAPE and cloud base among them, barely discriminate there at all, and refitted the whole thing around the lowest three hundred metres of shear. A threshold is a property of a sample, not of the atmosphere.
Chuck Doswell and David Schultz wrote the sharpest version of this critique in two thousand and six. A diagnostic variable describes the atmosphere now. A forecast parameter is one whose current value predicts the weather later, verified on cases it was not built from. Almost none of the parameters in daily use, they argued, had ever been tested that way. Their reviewers, who included the man who built the significant tornado parameter, pushed back hard, and the exchange is printed with the paper. The reviewer says the parameters mark where ingredients overlap, and nobody expects a signal twelve hours ahead. The authors reply: then say so, and stop treating them as forecasts. Both were right. They were answering different questions.
Now the part I find genuinely new. For twenty years the field argued over which layer of low level rotation matters most for tornadoes: the lowest five hundred metres, or the lowest three kilometres. Field project balloons said deep. Model based climatologies said shallow. In twenty twenty four Michael Coniglio and Richard Thompson, and in twenty twenty five Brice Coffer and colleagues, found the reason, and it was not the atmosphere. It was the storm motion you plug in. Every rotation parameter is computed relative to how the storm moves, and without a storm you estimate that from the wind profile with a formula. Swap the formula for the observed motion and the ranking flips. Deep layers win, and storm relative winds that had shown no skill suddenly show a lot. Tornadic supercells move noticeably slower than the formula says and a little to its right. Non tornadic ones move a little faster and well to its left.
The best discriminator between a tornadic and a non tornadic supercell depends on how the storm actually moves, which you learn only once the storm exists. Coffer's paper says it outright: deviant motion may be a cause of tornado potential or an effect of it, and the question is left open. It adds that once a storm has modified its own surroundings, it is not certain what a correct background environment even is. The most refined conditional tool of all is conditional on the storm twice over.
One more layer, from Daniel Chavas and Daniel Dawson in twenty twenty one. They built two soundings with the same CAPE and the same low level shear. One produced a long lived supercell in the model. The other produced a cell that died within an hour. The difference was how the moisture was distributed with height, which a bulk number integrates away. So what is environment analysis for?
It is a sieve, not a switch. It tells you which days cannot produce a strong tornado. It tells you, if storms form, roughly what they can become. It does not tell you whether they will form, and its sharpest instruments quietly borrow from the storm they are meant to anticipate. The Storm Prediction Center made the split explicit this March, when it added conditional intensity groups to its outlooks: one set of numbers for how likely severe weather is, and a separate set for how bad it will be if it happens. Two questions, finally drawn on the same map as two questions. Which brings me back to Tinker.
In nineteen ninety nine Robert Maddox and Charlie Crisp went back to the nineteen forty eight charts. Miller had always said the two days looked alike. They did not. The twentieth of March was, by modern standards, an unremarkable day. The twenty fifth was a textbook severe pattern, one that Miller himself later listed in his forecasting manual as one of the two most dangerous in the country. Fawbush and Miller had correctly diagnosed a favourable environment. What they could not have known was that a storm would form on it and cross the same square mile of Oklahoma inside a two hour window. Maddox and Crisp call the forecast fortuitously successful, with luck both good and bad. The environment was right. The rest was the unconditional question, and it still is.
Check your understanding
Try answering before revealing — these are the points the episode turns on.
1. Rasmussen and Blanchard's sample held hundreds of soundings and Thompson's later ones thousands. Why can neither, however large, tell you how often a favourable environment actually produces a storm?
Both were built from proximity soundings, which are collected on the condition that a storm occurred nearby. The favourable days that stayed quiet are absent by construction, so the denominator of 'how often' never enters the dataset. Estimating occurrence needs the full population of environments, which is what reanalysis climatologies provide, and there favourable hours outnumber storms by a wide margin. Discrimination between storm types and prediction of storm occurrence are estimated from different samples, and one cannot stand in for the other.
2. Crook found that one degree of temperature or one gram of water vapour per kilogram in the boundary layer could switch initiation on or off. Why does that make initiation a measurement problem more than a theory problem, and what kind of observation would change it?
The model knew the physics exactly and still could not tell which case it was in, because the deciding difference was smaller than the uncertainty in its input. That is a predictability limit set by the observing network rather than by understanding. It moves only if low-level temperature and moisture are measured at the scale of their natural variability, which is why campaigns like IHOP invested in water-vapour profiling near boundaries, and why forecasters still watch radar fine lines and satellite cumulus, not parameters, for the initiation decision.
3. Coniglio and Thompson, and then Coffer and colleagues, found that replacing estimated storm motion with observed motion changes which depth of storm-relative helicity discriminates best. What does that imply about using the winning parameter in a forecast issued before storms exist?
Storm-relative quantities are properties of the environment paired with a storm motion, not of the environment alone. Before the storm exists the only motion available is the estimate, so the operationally relevant ranking is the one computed with the estimate, and on that basis near-ground helicity stays ahead. The observed-motion result explains the physics and reconciles the balloon and model datasets, but it is diagnostic: it uses information that arrives with the storm. If the deviant motion is partly an effect of tornado potential rather than a cause, part of the improvement is circular, and Coffer's paper leaves exactly that open.
4. Chavas and Dawson produced two soundings with the same CAPE and low-level shear that gave a long-lived supercell and a one-hour dud. What does that tell you about what a composite parameter can encode, and what it cannot?
A bulk parameter is an integral, and an integral discards the shape of what it integrates. Two different vertical distributions of moisture or buoyancy can yield the same number, so the parameter can only say the column lies somewhere in the favourable region of a high-dimensional space, not where. Composites compress further by multiplying several integrals together. They mark where ingredients overlap, which is useful, but they cannot carry the profile details that decide whether a storm persists. Recovering that detail is the point of Chavas's analytic sounding and of the open question on vertical-structure attribution.
5. Thompson's STP threshold of one and the Chinese refit come from the same physics. Why should a threshold that works in Oklahoma not be expected to transfer to Europe or China?
The threshold is calibrated to the joint distribution of ingredients in the developmental sample. In the United States storms sit at the high end of a CAPE distribution whose tail reaches twice the European values, and cloud base height separates tornadic from non-tornadic cases. In China the thermodynamic ingredients barely differ between the two classes, so the American product's thermodynamic terms add noise and the signal lives in the lowest three hundred metres of shear. The physics of rotation is universal; the climatology of which ingredient is limiting is not, and a threshold encodes the climatology.
6. The 25 March 1948 forecast verified. In what sense was it a good forecast, in what sense luck, and how would you score it?
The environmental diagnosis was sound: by modern standards 25 March was a textbook severe pattern, later canonised as the Miller type-B setup. What could not be forecast was that a storm would form on it and cross the same square mile inside a two-hour window. Scored as a conditional forecast of storm type over a region and a window, it deserves credit. Scored as a point forecast of a tornado at Tinker, the long-run hit rate of such forecasts would be near zero and the verification is luck. Maddox and Crisp also found the twentieth, which prompted the whole effort, was the less favourable day. A good forecast and a lucky one look identical on a single case; only a population of forecasts separates them.
Further reading
- Doswell and Schultz (2006), On the Use of Indices and Parameters in Forecasting Severe StormsFree, Electronic Journal of Severe Storms Meteorology. The diagnostic-variable versus forecast-parameter argument, with the verification criteria in section 5. The reviewer exchange printed after the paper, including Richard Thompson's defence of STP, is the best part.
- Maddox and Crisp (1999), The Tinker AFB Tornadoes of March 1948Free PDF on Chuck Doswell's site, from Weather and Forecasting. Re-analysis of the 20 and 25 March 1948 charts; the source for General Borum's 'Do it!' and for the finding that the two days were not alike.
- Thompson et al. (2003), Close Proximity Soundings within Supercell Environments Obtained from the Rapid Update CycleFree PDF hosted by the Storm Prediction Center. The 413-sounding climatology that defined the supercell composite and significant tornado parameters; the 'magic number' caveat is in section 4d.
- Coffer, Parker, Coniglio and Homeyer (2025), Supercell environments using GridRad-Severe and the HRRRFree preprint; published in Weather and Forecasting 40, 1405 to 1428. Shows that swapping estimated for observed storm motion reverses which depth of storm-relative helicity discriminates tornadic supercells, and leaves open whether deviant motion is cause or effect.
- Chavas and Dawson (2021), An Idealized Physical Model for the Severe Convective Storm Environmental SoundingFree preprint of the Journal of the Atmospheric Sciences paper. Two soundings with the same CAPE and low-level shear, one giving a long-lived supercell and one a cell that dies in an hour.
- Zhang et al. (2025), Environments of tornadic and non-tornadic supercells in China and optimized significant tornado parameter for China regionQuarterly Journal of the Royal Meteorological Society; may be paywalled depending on your access. Finds the thermodynamic terms of the US parameter do not discriminate in China and refits STP around 0 to 300 m shear and helicity.
- Storm Prediction Center, Conditional Intensity in SPC Convective OutlooksFree. SPC's own description of the 3 March 2026 change: outlook probabilities unchanged, with intensity now stated conditionally on a report occurring.
Against the library
How this episode sits against the listener's own paper library. Coverage: covered.
Adds: Doswell and Schultz 2006 is not in the library. It supplies the diagnostic-variable versus forecast-parameter framework and concrete verification criteria (independent developmental and verification datasets, contingency-table skill against climatology or persistence, accuracy as a function of lead time, lagged correlation) that gap-composite-parameter-out-of-sample asks for, plus the printed reviewer exchange in which Richard Thompson states SCP and STP are short-term overlap aids and that he does not expect a signal 12 h before an event. Maddox and Crisp 1999 add a synoptic re-analysis of 20 and 25 March 1948: the days differed markedly aloft, 20 March was unremarkable by modern standards, and the first tornado forecast is described as fortuitously successful. Zhang et al. 2025 (QJRMS) add that in China CAPE, LCL, low-level RH and CIN do not discriminate tornadic from non-tornadic supercells while 0 to 300 m SRH and shear do, with a refitted STP300cn using most-unstable parcels. SPC's conditional intensity groups, effective 3 March 2026, are the operational adoption of the conditional versus unconditional split.
Differs: No source contradicts a Claim; three qualify them. claim:thompson-2003-soundings#C7 (STP of 1 as a reasonable guideline) is discrimination within a dependent proximity sample, and Doswell and Schultz argue this establishes diagnostic rather than forecast value, the same limitation claim:thompson-2003-soundings#C8 already records. claim:rasmussen-1998-climatology#C7 (EHI best for supercell versus ordinary) is reframed by Doswell and Schultz via Rasmussen 2003: EHI above 0.5 gives roughly a fifty percent chance of a significant tornado conditional on a supercell, from about thirty cases per class. claim:coffer-2025-supercell#C4 (observed motion reverses the SRH-depth ranking) is read here as narrowing forecast-time applicability, consistent with claim:coffer-2025-supercell#C7 and synthesis S3 of the 2026-09-06 brief; Zhang 2025 finds a shallow-layer (0 to 300 m) preference in China with estimated motion, which does not test the observed-motion flip. Miller's account that the two 1948 days were alike is contradicted by Maddox and Crisp, but that is not a library Claim.
Relates: Directly on gap-composite-parameter-out-of-sample: Doswell and Schultz define the independent-verification standard the gap names, and Zhang 2025 is an out-of-region refit that answers the gap's threshold-transfer subquestion with a no. Touches gap-ci-intrinsic-predictability-limit through claim:crook-1996-sensitivity#C1 and #C6 and claim:weckwerth-2006-ihop#C6 and #C11; gap-severe-null-case-completeness through the base-rate argument; gap-analysis-sounding-offsite-fidelity through claim:thompson-2003-soundings#C10; and gap-vertical-structure-outcome-attribution through claim:chavas-2020-sounding#C2 and #C3. Cross-domain: the critique of single-scalar composites applies to mcgovern-2026-ewb's use of a composite convective-risk proxy in ml-weather-forecasting, and the lagged-correlation criterion Doswell and Schultz propose is a time-series-analysis question.
Read into the library after publication:
- On the Use of Indices and Parameters in Forecasting Severe Storms https://ejssm.com/ojs/index.php/site/article/view/4
- The Tinker AFB Tornadoes of March 1948 10.1175/1520-0434(1999)014<0492:TTATOM>2.0.CO;2
- Environments of tornadic and non-tornadic supercells in China and optimized significant tornado parameter for China region 10.1002/qj.5027