I Ran Last Year's Typhoon Data Through DeepMind's Weather Forecast Model
Yesterday I saw DeepMind publish a Nature paper claiming their WeatherNext model extended hurricane forecast lead times by more than a day. My first reaction was: here we go again, hyped up in the paper, but probably full of pitfalls in practice.
However, I was genuinely curious because last week I had spent two whole days doing typhoon path prediction using traditional numerical weather prediction models (the kind that run physics equations on supercomputers). So I decided to pull down WeatherNext's public code and test it against several typhoons in the Pacific last year.
First, the preparation phase. My environment was Python 3.10 plus PyTorch 2.1, with an RTX 4090 GPU having 24GB of VRAM. The model weights were about 6GB, and downloading them was quick. But installing dependencies was a headache; it required specific versions of libraries like einops, xarray, and dask. I initially tried installing the latest versions via pip, which resulted in errors about dimension mismatches. Eventually, I strictly followed requirements.txt, which took over an hour.
Then came data preparation. WeatherNext uses ERA5 reanalysis data (high-quality meteorological datasets fusing historical observations with numerical models), which I needed to convert into the format required by the model. There was a pitfall here: the data needs to be resampled onto a latitude-longitude grid. I didn't pay attention to the resolution at first and used the raw data, resulting in a predicted typhoon path off by about 200 kilometers. Later I realized I needed to interpolate to a 0.25-degree grid first, and after reprocessing, it worked normally.
For the first run, I chose a super typhoon from September last year, covering the period from formation to landfall—a total of 9 days. Traditional models usually predict landfall points accurately up to 5 days ahead, with large errors beyond 7 days. After running WeatherNext, I checked the output: the 7-day-ahead path prediction had an error of within 80 kilometers compared to the actual path I looked up later. This number was better than I expected.
But what really surprised me wasn't the path, but the intensity prediction. Predicting typhoon intensity has always been a tough nut for traditional models, often severely overestimating or underestimating it. The maximum wind speed change curve provided by WeatherNext matched the trend of actual observations quite well, especially during the rapid intensification phase, where the timing was captured accurately. I compared it with the European Centre's high-resolution forecasts, and at least in this case, the AI model held its own.
However, there were plenty of pitfalls. While model inference only took tens of seconds, the time spent on preliminary data processing and format conversion was no less than traditional methods. It requires boundary conditions for the coming days—in other words, it still relies on traditional models to provide the "background field," so it's not entirely independent. I'm unsure about the model's generalization ability for extreme cases, given that training data contains few samples of such super typhoons.
DeepMind stated in their blog that WeatherNext gave forecasters "an extra day" of warning time. After running this case, I think the claim isn't exaggerated, but the prerequisite is getting the data pipeline right.
I also ran it on two ordinary tropical storms, which lacked the drama of rapid intensification, but the path predictions were still within reasonable ranges. Comparatively, I tend to believe the model performs quite stably under normal conditions, with the disruptive potential mainly lying in intensity prediction.
The conclusion depends on who you are.
If you're a researcher or looking to build AI applications related to meteorology in enterprises, this model is worth spending time on. The code and weights are open source, and the documentation is relatively clear, making it easier to get started than I expected. But if you want to plug it directly into business systems for typhoon warnings, I suggest caution, because the engineering complexity of data preprocessing and boundary conditions might be higher than the model itself. Moreover, its interpretability is currently a black box; explaining "why this prediction" to clients would put you in a difficult position.
My takeaway is: the future of AI forecasting lies not in paths, but in intensity. But in terms of engineering, the data pipeline is the real barrier.
Personally, this model helped me redefine the boundaries of end-to-end learning in physical simulation. Previously, I thought claims about AI replacing physical models were nonsense, but my view has changed—at least in this niche area, AI can compete with traditional methods. If asked about "AI implementation in scientific computing" in future interviews, this should be a good talking point.
📌 This article is compiled from ArsTechnica. Original text: https://arstechnica.com/science/2026/08/deepminds-hurricane-model-bought-forecasters-an-extra-day/
Copyright belongs to the original authors. This is a compilation and independent analysis based on public reports.
Physix Frontier