AHALabAHALab · Science and Us
Back to Slides ↗

After the weather model is created,
what other evidence is needed?

The Atria team demonstrated a case where AI participated in weather model development. A working demo allowed us to see the results of the work; where and under what weather conditions it can be reliably used still needs further testing.

The full weather demonstration interface embedded in the Atria report, click to enlarge
The embedded original image used in Figure 2(a) on page 5 of the original report. The interface has a historical timestamp from 2022; it is not the current weather. The author notes in the caption that this interface image does not report forecast accuracy; one cannot conclude the model is reliable just because two maps appear similar.

What does this report record?

The author introduces that in a case without network search, AI processed more than 100GB of meteorological data, implemented a network with over 400 million parameters, ran 45,000 training steps, involving 69 meteorological variables.

These are case records provided by the author. This course has not independently reproduced the training, nor has it used this model to release real-time forecasts.

Source: report pages 4–5. The author notes that these are selected demonstrations and cannot be taken as an average task success rate.

How to continue testing?

  1. Clarify who produced the data, and what time periods, locations, and variables it covers.
  2. Conduct final testing using data not involved in training or tuning, to prevent answers from leaking in advance.
  3. Compare with appropriate references and baselines under the same tasks and conditions; check forecast time, region, variables, and extreme events separately.
  4. Confirm what the user needs. Global weather forecasting, mountain road trips, and flood warnings require different information and testing.

These are the verification requirements raised in class and do not indicate that these checks have already been completed for the demonstration. The reference data itself also has errors and applicable ranges.

If the weather map has already been generated, which item would you check first before deciding whether to use it?

How do humans and AI work together?

The team also analyzed 769 task records from 56 participants. In this set of records, the AI proposes plans, executes, and modifies; humans participate in selecting goals and methods, supplementing context, and judging results.

These materials can help us understand the specific collaboration process. They come from this one R&D team and cannot prove that all research tasks adopt the same division of labor, nor can they prove that human judgment is always correct.

Basis: Report pages 8, 10–11; the scope and denominators of tasks differ for various percentages and cannot be mixed.

What cannot be concluded here for the time being?

“One-minute prediction for the next week” cannot be interpreted as completing the entire model development in one minute. The corresponding sections in the report consulted this time do not provide detailed results enough to verify this time configuration or compare it with FourCastNet.

AI can participate in developing another AI system, but that does not mean it has been proven to continuously and reliably improve itself autonomously. The report still lists this as an open question.

Basis: Report pages 7–8, 12–13. The class does not use promotional headlines or rankings as evidence of practicality.

Source and reading records

Local report · Page 5 ↗Official report ↗Official project ↗

Based on a 23-page report snapshot obtained on September 15, 2026. GitHub main may continue to update; images are not redrawn, and authors and the original sources are preserved. Class discussion is organized by AHALab.

View the original image source record