File size: 4,440 Bytes
929e312 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 | ---
frameworks: PyTorch
language:
- en
license: apache-2.0
tags:
- OneScience
- Earth Science
- Ocean Simulation
- Global Ocean Forecasting
- OM4
- ConvNeXt
tasks: []
datasets:
- M2LInES/Samudra-OM4
---
<p align="center">
<strong>
<span style="font-size: 30px;">Samudra</span>
</strong>
</p>
# Model Introduction
Samudra is a global ocean emulator developed by the M2LInES team.
Paper: Samudra: An AI Global Ocean Emulator for Climate
https://doi.org/10.1029/2024GL114318
# Model Description
Samudra predicts global ocean states on an approximately one-degree grid with a five-day time step. It is designed to emulate the evolution of the OM4 ocean circulation model with a deep neural network.
# Use Cases
| Scenario | Description |
| :---: | :--- |
| Global ocean simulation | Train the model on OM4 data following the 77-state-channel and 4-forcing-channel Samudra protocol. |
| Local quick validation | Use synthetic NPZ data to check training, inference, and ocean-field visualization. |
| ModelScope / OneCode execution | Download the standalone model package, install dependencies, and run the scripts directly. |
| Multi-GPU training | Launch PyTorch DistributedDataParallel with `torchrun`. |
# Usage Guide
## 1. OneCode Usage
Experience intelligent one-click AI4S programming through the OneCode online environment:
[Click to Experience Intelligent One-Click AI4S Programming](https://web-2069360198568017922-iaaj.ksai.scnet.cn:58043/home)
## 2. Manual Installation and Usage
**Hardware Requirements**
- A GPU or DCU is recommended.
- CPU can be used for import and small-scale connectivity verification; full training and inference will be slow.
- DCU users must install DTK in advance. DTK 25.04.2 or above, or the OneScience recommended version matching your cluster, is recommended.
### Download the Model Package
```bash
hf download OneScience-Group/Samudra --local-dir ./Samudra
cd Samudra
```
### Install the Runtime Environment
**DCU Environment**
```bash
# Please activate DTK and CONDA first
conda create -n onescience311 python=3.11 -y
conda activate onescience311
# uv installation is supported
pip install onescience[earth-dcu] -i http://mirrors.onescience.ai:3141/pypi/simple/ --trusted-host mirrors.onescience.ai
```
**GPU Environment**
```bash
# Please activate CONDA first
conda create -n onescience311 python=3.11 -y libstdcxx-ng=12 libgcc-ng=12 gcc_linux-64=12 gxx_linux-64=12
conda activate onescience311
# uv installation is supported
pip install onescience[earth-gpu] -i http://mirrors.onescience.ai:3141/pypi/simple/ --trusted-host mirrors.onescience.ai
```
### Training Data Introduction
The official training data is generated by NOAA/GFDL OM4. The M2LInES project provides the dataset and documentation:
https://huggingface.co/datasets/M2LInES/Samudra-OM4
The dataset must be converted to the native NPZ layout expected by this repository. When real OM4 data is unavailable, generate a synthetic fixture for pipeline validation:
```bash
python scripts/fake_data.py
```
The synthetic fixture contains 77 prognostic state channels and 4 boundary-forcing channels and is not suitable for scientific evaluation.
### Training
Single GPU:
```bash
python scripts/train.py
```
Multi-GPU:
```bash
torchrun --nproc_per_node=8 scripts/train.py
```
The default checkpoint is saved to `data/checkpoints/model_bak.pth`.
### Training Weights
This repository provides a `weight/` directory for Samudra checkpoints. The weight files will be uploaded soon and are expected to be available in the near future.
### Inference
Inference performs an autoregressive rollout from `data/test.npz` and reads `data/checkpoints/model_bak.pth` by default:
```bash
python scripts/inference.py
```
Predictions are written to `result/output/prediction.npz`.
### Evaluation and Visualization
```bash
python scripts/result.py
```
The default outputs are `result/forecast_maps.png` and `result/temperature_profile.png`.
# Official OneScience Resources
| Platform | OneScience Main Repository | Skills Repository |
| --- | --- | --- |
| Gitee | https://gitee.com/onescience-ai/onescience | https://gitee.com/onescience-ai/oneskills |
| GitHub | https://github.com/onescience-ai/OneScience | https://github.com/onescience-ai/oneskills |
# Citation and License
- Paper: https://doi.org/10.1029/2024GL114318
- This repository is an independent reproduction of the original Samudra paper.
|