M2RL-RL_Coding / README.md
Jackwang111's picture
Update README.md
22df4d6 verified
|
Raw History Blame Contribute Delete
1.95 kB

To Mix or To Merge?
Toward Multi-Domain Reinforcement Learning for Large Language Models

arXiv Hugging Face ModelScope GitHub COLM 2026

Haoqing Wang†, Xiang Long†, Ziheng Li†, Yilong Xu, Tingguang Li, Yehui Tangβœ‰
Samsung Research, Beijing, China   Β·   Peking University


πŸ“° News

  • [2026.09.07] πŸŽ‰ The model checkpoints are now open-sourced on Hugging Face and ModelScope! Feel free to use our checkpoints for your post-training research (e.g., weight merging, multi-teacher on-policy distillation)!
  • [2026.07.09] πŸŽ‰ Our paper is accepted to COLM 2026!

πŸ“š Citation

If you find this work useful, please consider citing:

@inproceedings{
wang2026to,
title={To Mix or To Merge: Toward Multi-Domain Reinforcement Learning for Large Language Models},
author={Haoqing Wang and Xiang Long and Ziheng Li and Yilong Xu and Tingguang Li and Yehui Tang},
booktitle={Third Conference on Language Modeling},
year={2026},
url={https://openreview.net/forum?id=jP7j5XkG8J}
}