3PO family methods trained on DapoMath-17k using Olmo3-IVON-SFT-7B and Qwen2.5Math-IVON-SFT-7B
BayesRL
non-profit
AI & ML interests
None defined yet.
Recent Activity
View all activity
A collection of three models trained on the Nemotron Post Training Dataset for reasoning tasks with IVON
-
BayesRL/Llama3.1-IVON-SFT-8B
Text Generation ⢠8B ⢠Updated ⢠4.23k -
BayesRL/Qwen2.5Math-IVON-SFT-7B
Text Generation ⢠8B ⢠Updated ⢠1.1k -
BayesRL/Olmo3-IVON-SFT-7B
Text Generation ⢠7B ⢠Updated ⢠763 -
Parameter Exploration for RLVR via Variational Learning
Paper ⢠2608.09805 ⢠Published ⢠7
3PO family methods trained on DapoMath-17k using Olmo3-IVON-SFT-7B and Qwen2.5Math-IVON-SFT-7B
A collection of three models trained on the Nemotron Post Training Dataset for reasoning tasks with IVON
-
BayesRL/Llama3.1-IVON-SFT-8B
Text Generation ⢠8B ⢠Updated ⢠4.23k -
BayesRL/Qwen2.5Math-IVON-SFT-7B
Text Generation ⢠8B ⢠Updated ⢠1.1k -
BayesRL/Olmo3-IVON-SFT-7B
Text Generation ⢠7B ⢠Updated ⢠763 -
Parameter Exploration for RLVR via Variational Learning
Paper ⢠2608.09805 ⢠Published ⢠7