File size: 4,440 Bytes
929e312
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
---
frameworks: PyTorch
language:
- en
license: apache-2.0
tags:
- OneScience
- Earth Science
- Ocean Simulation
- Global Ocean Forecasting
- OM4
- ConvNeXt
tasks: []
datasets:
- M2LInES/Samudra-OM4
---
<p align="center">
  <strong>
    <span style="font-size: 30px;">Samudra</span>
  </strong>
</p>

# Model Introduction

Samudra is a global ocean emulator developed by the M2LInES team.

Paper: Samudra: An AI Global Ocean Emulator for Climate

https://doi.org/10.1029/2024GL114318

# Model Description

Samudra predicts global ocean states on an approximately one-degree grid with a five-day time step. It is designed to emulate the evolution of the OM4 ocean circulation model with a deep neural network.

# Use Cases

| Scenario | Description |
| :---: | :--- |
| Global ocean simulation | Train the model on OM4 data following the 77-state-channel and 4-forcing-channel Samudra protocol. |
| Local quick validation | Use synthetic NPZ data to check training, inference, and ocean-field visualization. |
| ModelScope / OneCode execution | Download the standalone model package, install dependencies, and run the scripts directly. |
| Multi-GPU training | Launch PyTorch DistributedDataParallel with `torchrun`. |

# Usage Guide

## 1. OneCode Usage

Experience intelligent one-click AI4S programming through the OneCode online environment:

[Click to Experience Intelligent One-Click AI4S Programming](https://web-2069360198568017922-iaaj.ksai.scnet.cn:58043/home)

## 2. Manual Installation and Usage

**Hardware Requirements**

- A GPU or DCU is recommended.
- CPU can be used for import and small-scale connectivity verification; full training and inference will be slow.
- DCU users must install DTK in advance. DTK 25.04.2 or above, or the OneScience recommended version matching your cluster, is recommended.

### Download the Model Package

```bash
hf download OneScience-Group/Samudra --local-dir ./Samudra
cd Samudra
```

### Install the Runtime Environment

**DCU Environment**

```bash
# Please activate DTK and CONDA first
conda create -n onescience311 python=3.11 -y
conda activate onescience311
# uv installation is supported
pip install onescience[earth-dcu] -i http://mirrors.onescience.ai:3141/pypi/simple/  --trusted-host mirrors.onescience.ai
```

**GPU Environment**
```bash
# Please activate CONDA first
conda create -n onescience311 python=3.11 -y libstdcxx-ng=12 libgcc-ng=12 gcc_linux-64=12 gxx_linux-64=12
conda activate onescience311
# uv installation is supported
pip install onescience[earth-gpu] -i http://mirrors.onescience.ai:3141/pypi/simple/  --trusted-host mirrors.onescience.ai
```

### Training Data Introduction

The official training data is generated by NOAA/GFDL OM4. The M2LInES project provides the dataset and documentation:

https://huggingface.co/datasets/M2LInES/Samudra-OM4

The dataset must be converted to the native NPZ layout expected by this repository. When real OM4 data is unavailable, generate a synthetic fixture for pipeline validation:

```bash
python scripts/fake_data.py
```

The synthetic fixture contains 77 prognostic state channels and 4 boundary-forcing channels and is not suitable for scientific evaluation.

### Training

Single GPU:

```bash
python scripts/train.py
```

Multi-GPU:

```bash
torchrun --nproc_per_node=8 scripts/train.py
```

The default checkpoint is saved to `data/checkpoints/model_bak.pth`.

### Training Weights

This repository provides a `weight/` directory for Samudra checkpoints. The weight files will be uploaded soon and are expected to be available in the near future.

### Inference

Inference performs an autoregressive rollout from `data/test.npz` and reads `data/checkpoints/model_bak.pth` by default:

```bash
python scripts/inference.py
```

Predictions are written to `result/output/prediction.npz`.

### Evaluation and Visualization

```bash
python scripts/result.py
```

The default outputs are `result/forecast_maps.png` and `result/temperature_profile.png`.

# Official OneScience Resources

| Platform | OneScience Main Repository | Skills Repository |
| --- | --- | --- |
| Gitee | https://gitee.com/onescience-ai/onescience | https://gitee.com/onescience-ai/oneskills |
| GitHub | https://github.com/onescience-ai/OneScience | https://github.com/onescience-ai/oneskills |

# Citation and License

- Paper: https://doi.org/10.1029/2024GL114318
- This repository is an independent reproduction of the original Samudra paper.