How Swelter's WBGT Is Validated
Swelter is validated against the world's largest WBGT monitoring network. This page explains the data we compared against, how the comparison worked, the results, and the limits of what they show.
Why we ran the comparison
Swelter estimates WBGT from weather data. It does not read from a physical instrument, and no estimate can match a calibrated device. We wanted to know, with real numbers, how close the estimates come. So we compared Swelter's engine against the largest set of reference values we could find.
The method
WBGT is computed as the standard outdoor blend that has been used for decades:
WBGT = 0.7 × Tw + 0.2 × Tg + 0.1 × Ta
where Tw is the wet-bulb temperature, Tg is the globe temperature, and Ta is the air temperature (dry-bulb).
Swelter estimates WBGT from forecast data, provided by Apple Weather. Wet-bulb temperature is calculated with the Stull (2011) formula, sun strength is estimated from your location's sun angle and cloud cover, using NOAA's solar position equations and the Kasten–Czeplak cloud model, and the standard 70/20/10 outdoor blend combines them with air temperature. We also apply our own wind cooling credit, which is capped and fades in very humid or very sunny conditions. The final number is classified with ACSM activity flags (82 / 87 / 90 °F, or 27.8 / 30.6 / 32.2 °C).
Stull, R. (2011). Wet-Bulb Temperature from Relative Humidity and Air Temperature. Journal of Applied Meteorology and Climatology, 50(11), 2267–2269.
What goes into the estimate
- Wet-bulb: measures what the moisture actually does. It is the coolest temperature that evaporation can reach in the current air, so it depends on both moisture and air temperature. This is the wet-bulb in WBGT, and it carries most of the weight in the blend.
- Sun strength: measures solar radiation, the part of sunlight that heats you and everything around you. Swelter calculates the sun's angle in your sky for the hour you are viewing, then reduces its strength based on cloud cover. A high sun in a clear sky adds the most heat, and a low sun or thick clouds add very little.
- Air temperature: enters the estimate in two places. It is an input to the wet-bulb calculation, and it also stands alone as the smallest share of the 70/20/10 blend.
- Wind: subtracts a small cooling credit from the estimate. The credit is capped, and it fades as the air gets more humid or the sun gets stronger. In very humid air or extremely sunny conditions, even strong wind cools very little.
Assumptions worth knowing
Swelter's estimates make a few assumptions worth knowing.
- Forecast inputs: the engine was validated on measured station inputs, but live readings are computed from Apple Weather forecasts, so forecast error passes through. Readings close to a flag boundary say so with the app's "Near …" line.
- Direct sun: readings assume you are out in the sun, the way outdoor WBGT is defined. If you are in the shade, actual heat stress is lower than shown.
- Individual differences: the flags are general guidance, not personalized. Heat affects everyone differently, and it takes time to adapt. If you are new to being active in the heat, treat each flag with extra caution.
The reference network
Japan operates a national heat-stroke monitoring system, and it is the world's largest WBGT monitoring network. The network publishes hourly WBGT values for stations across the country. Those values are computed from measured weather-station data, including measured solar radiation.
The published values are public. Anyone can check our comparison against the same source.
How the comparison worked
We ran Swelter's engine on the same hours of measured weather inputs, at 20 stations across Japan, from July into August 2026. For each hour, we compared the engine's WBGT and risk flag against the network's published value for that hour. The comparison covered 21,370 station-hours.
The results
- 21,370 station hours
- 0.7 °F avg. error
- 90% flag agreement
- 99.96% within one flag level
90% is exact flag-for-flag agreement. Nearly every difference was a single boundary step: the same borderline readings the app calls out with its “Near …” line. Within one adjacent flag level, agreement was 99.96%.
The misses overwhelmingly lean toward over-warning rather than under-warning.
Three ways we test
The headline numbers above come from the strictest test we can run: the engine on measured station inputs, compared against the reference network's published values. We also run two other checks.
- Cross-model comparison: we compare the engine's output against the National Weather Service's own WBGT estimates for stations across the United States. This is a model-to-model comparison rather than a comparison against a measured reference, and the agreement is consistent with the Japan results.
- Full-pipeline comparison: the app does not run on measured station inputs. It runs on Apple Weather forecast data, so we also test the complete pipeline the app actually ships, forecast data in and WBGT out, against the same reference network. Forecast error passes through this test, so pipeline agreement is naturally lower than engine agreement and moves with the weather. In our testing so far, the pipeline has still matched the network's flag for the large majority of hours, with the differences concentrated in the same borderline readings the Near line marks. That is also why we treat forecasts as planning guidance rather than measurements.
The comparisons are repeatable. The reference network's hourly values are public, and anyone can run the same check against the same source. The validation archive keeps growing across seasons, and the app's Sources & Algorithm screen shows the full methodology, the sources, and a live demonstration of the math running on the real engine.
Why not a physical model?
Research-grade WBGT models, such as Liljegren's, compute the wet-bulb and globe temperatures from first principles. They are the right tool when the inputs come from instruments, including measured solar radiation, and they are what agencies like the National Weather Service use with instrumented data. Forecast data is different. It carries no measured solar radiation, and its values arrive with forecast error, so a sensitive physical model running on forecast inputs inherits noise it was never designed to absorb.
Swelter takes the other path: simpler component estimates, chosen for how they behave on forecast inputs, calibrated together as one system against measured reference values. The headline numbers above are the result of that calibration.
Limitations
The comparison used measured station inputs. Day-to-day readings in the app come from forecast data and can vary more.
The comparison covered one summer season in one country's climate. WBGT from any app is always an estimate from weather data, not a reading from a calibrated device.
Swelter is a planning tool, not a medical device. Guidance in the app comes from published, publicly available materials, and the app is not affiliated with, endorsed by, or sponsored by the organizations behind them.
Sources
- NWS Heat Safety
- NIOSH Criteria: Occupational Exposure to Heat
- U.S. Army TB MED 507: Heat Stress Control
- ACSM Position Stand: Exertional Heat Illness
- Stull (2011): Wet-Bulb from Temperature & Humidity
- NOAA: Solar Position Calculations
- Kasten & Czeplak (1980): Solar Radiation and Cloud Cover
- Liljegren et al. (2008): Modeling WBGT
- Japan Ministry of the Environment: Heat Illness Prevention (WBGT Network)
Disclaimer
Swelter is a planning tool, not a medical device. Its readings are estimates from weather data, not measurements from a physical device. Conditions can change quickly. When in doubt, choose the cooler option and use your best judgment. Guidance in the app and on this website comes from published, publicly available materials, and Swelter is not affiliated with, endorsed by, or sponsored by the organizations behind them.
Swelter is a WBGT app for iPhone, built by RBRN Labs.