File size: 7,335 Bytes
bf314e8 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 | <p align="center">
<strong>
<span style="font-size: 30px;">UMA</span>
</strong>
</p>
# Model Introduction
UMA (Universal Materials Interaction Model) is a universal machine-learning interatomic potential model for materials and catalysis systems, built on equivariant graph neural networks. It predicts energy and forces for atomic structures.
# Model Description
UMA is based on an equivariant graph neural network architecture and is trained on diverse materials and catalysis datasets such as OC20, OC22, OC25, OMat, OMOL, ODAC, and OMC. It performs energy and force prediction and structure optimization for catalytic adsorption, inorganic materials, molecular systems, MOFs, and related scenarios.
# Applicable Scenarios
| Scenario | Description |
| :---: | :---: |
| OC20 energy and force fine-tuning | Train the UMA model using the standard configuration to read OC20 fine-tuning data |
| OC22 oxide catalysis fine-tuning | Train the UMA model using the standard configuration to read OC22 fine-tuning data (1P2 only) |
| OC25 (electro-)catalysis fine-tuning | Train the UMA model using the standard configuration to read OC25 fine-tuning data (1P2 only) |
| OMat inorganic materials fine-tuning | Train the UMA model using the standard configuration to read OMat fine-tuning data |
| OMOL molecules + polymers fine-tuning | Train the UMA model using the standard configuration to read OMOL fine-tuning data |
| ODAC MOFs fine-tuning | Train the UMA model using the standard configuration to read ODAC fine-tuning data |
| OMC molecular crystals fine-tuning | Train the UMA model using the standard configuration to read OMC fine-tuning data |
| Training workflow pre-check | Check configuration, data paths, run scripts, and checkpoint placement |
| Catalytic adsorption system modeling | Refer to OC20/OC22/OC25 tasks for adsorption/catalytic surface system training and inference |
| Inference script reference | Use upstream inference examples for crystal relaxation, adsorption system relaxation, or molecular MD adaptation |
| Custom data migration | Convert ASE-readable structures into UMA fine-tuning data and replace the training and validation paths |
# Usage Instructions
## 1. Using OneCode
You can try out intelligent one-click AI4S programming in the OneCode online environment:
[Try intelligent one-click AI4S programming](https://web-2069360198568017922-iaaj.ksai.scnet.cn:58043/home)
## 2. Manual Installation and Usage
**Hardware Requirements**
- GPU or DCU is recommended for fine-tuning training.
- CPU can be used for configuration and data-path pre-checks; formal UMA training is not recommended on CPU.
- DCU users need to install DTK in advance. DTK 25.04.2 or above, or the OneScience-recommended version matching the current cluster, is suggested.
### Download the Model Package
```bash
modelscope download --model OneScience/UMA --local_dir ./UMA
cd UMA
```
### Install the Runtime Environment
**DCU Environment**
```bash
# Please activate DTK and CONDA first
conda create -n onescience311 python=3.11 -y
conda activate onescience311
# uv installation is also supported
pip install onescience[matchem-dcu] -i http://mirrors.onescience.ai:3141/pypi/simple/ --trusted-host mirrors.onescience.ai
```
**GPU Environment**
```bash
# Please activate CONDA first
conda create -n onescience311 python=3.11 -y libstdcxx-ng=12 libgcc-ng=12 gcc_linux-64=12 gxx_linux-64=12
conda activate onescience311
# uv installation is also supported
pip install onescience[matchem-gpu] -i http://mirrors.onescience.ai:3141/pypi/simple/ --trusted-host mirrors.onescience.ai
```
### Training Dataset Description
This repository does not include built-in training data. The following uses **OC20 fine-tuning** as an example; the procedures for OC22, OC25, OMat, OMOL, ODAC, OMC, and other tasks are the same—just replace `--uma-task` and the data path accordingly.
```bash
modelscope download --dataset OneScience/oc20 --local_dir ./data
```
**Data Format Conversion**
The downloaded raw data is usually in `.extxyz` format and must first be converted to ASE-lmdb format using `scripts/create_uma_finetune_dataset.py`, which also computes `elem_refs` and `normalizer_rmsd`. The script supports the following tasks:
| Task | Description |
| --- | --- |
| `oc20` | Catalysis (example) |
| `oc22` | Oxide catalysis (1P2 only) |
| `oc25` | (Electro-)catalysis (1P2 only) |
| `omat` | Inorganic materials |
| `omol` | Molecules + polymers |
| `odac` | MOFs |
| `omc` | Molecular crystals |
```bash
python scripts/create_uma_finetune_dataset.py \
--train-dir data/oc20/s2ef_200k_uncompressed \
--val-dir data/oc20/s2ef_val_id_uncompressed \
--uma-task oc20 \
--regression-tasks ef \
--output-dir data/oc20_finetune \
--num-workers 8
```
After conversion, the following are generated:
```text
data/oc20_finetune/
├── train/ # ASE-lmdb training data
├── val/ # ASE-lmdb validation data
└── data/ # Generated data configuration yaml
└── uma_conserving_data_task_energy_force.yaml
```
Then use `scripts/update_demo_config.py` to update the generated `elem_refs`, `normalizer_rmsd`, and data paths into the demo configuration file:
```bash
python scripts/update_demo_config.py --demo-config demo/configs/oc20_ef_4dcu.yaml
```
`demo/run.sh` automatically uses the repository root as `ONESCIENCE_DATASETS_DIR`, so relative paths in the configuration file will match automatically.
### Training Weights
This repository already includes the rotation basis file `weight/Jd.pt`. UMA pre-trained checkpoints (e.g., `uma-s-1p1_converted.pt`) need to be downloaded from the fairchem official repository, converted to UMA format, and placed at (upload coming soon):
```text
weight/uma-s-1p1_converted.pt
```
- fairchem official repository: https://github.com/facebookresearch/fairchem
`demo/run.sh` and the example scripts under `inference/` will automatically detect `weight/Jd.pt` and set `ONESCIENCE_UMA_JD_PATH`.
### Fine-tuning
```bash
bash demo/run.sh --config demo/configs/oc20_ef_4dcu.yaml
```
### Inference
```bash
python inference/run_molecular_md.py
```
You need to specify the pre-trained weights path yourself.
## OneScience Official Information
| Platform | OneScience Main Repository | Skills Repository |
| --- | --- | --- |
| Gitee | https://gitee.com/onescience-ai/onescience | https://gitee.com/onescience-ai/oneskills |
| GitHub | https://github.com/onescience-ai/OneScience | https://github.com/onescience-ai/oneskills |
---
## Citation and License
- The UMA example code comes from the matchem example implementation in the OneScience project and references the upstream fairchem project (https://github.com/facebookresearch/fairchem). The upstream fairchem repository software is released under the [MIT License](https://fair-chem.github.io/core/install.html#license); fairchem model checkpoints and datasets may carry their own separate licenses, so please follow the corresponding instructions when using them.
- If you use UMA fine-tuning or inference results in scientific research, we recommend citing the UMA/related universal materials interaction model methods, fairchem/OneScience project information, and the sources of the datasets actually used.
|