Nemotron-Pre-Training-Datasets Collection Large scale pre-training datasets used in the Nemotron family of models. • 15 items • Updated Aug 11 • 191
view article Article Scaling Pedagogical Pre-training: From Optimal Mixing to 10 Billion Tokens codelion • Mar 6 • 7
david-thrower/HelixLM-41M-IT-v2-20260608-2335-d512-h8-nl3-s512 Text Generation • 40.1M • Updated Jul 10 • 28
david-thrower/HelixLM-20260529-2240-d512-h8-nl3-ffn2-s512-1500MT-ep3 Text Generation • 40.1M • Updated Jul 10 • 28
david-thrower/HelixLM-41M-IT-v2-20260610-0430-d512-h8-nl3-s512-x Text Generation • 40.1M • Updated Jun 10 • 18
david-thrower/HelixLM-41M-IT-v2-20260610-0430-d512-h8-nl3-s512-x Text Generation • 40.1M • Updated Jun 10 • 18
david-thrower/HelixLM-41M-IT-v2-20260610-0008-d512-h8-nl3-s512-x Text Generation • 40.1M • Updated Jun 10 • 15
david-thrower/HelixLM-41M-IT-v2-20260610-0008-d512-h8-nl3-s512-x Text Generation • 40.1M • Updated Jun 10 • 15
david-thrower/HelixLM-41M-IT-v2-20260609-2209-d512-h8-nl3-s512 Text Generation • 40.1M • Updated Jun 9 • 16
david-thrower/HelixLM-41M-IT-v2-20260609-2209-d512-h8-nl3-s512 Text Generation • 40.1M • Updated Jun 9 • 16
david-thrower/HelixLM-41M-IT-v2-20260609-1902-d512-h8-nl3-s512 Text Generation • 40.1M • Updated Jun 9 • 19
david-thrower/HelixLM-41M-IT-v2-20260609-1902-d512-h8-nl3-s512 Text Generation • 40.1M • Updated Jun 9 • 19
david-thrower/HelixLM-41M-IT-v2-20260609-1749-d512-h8-nl3-s512 Text Generation • 40.1M • Updated Jun 9 • 15
david-thrower/HelixLM-41M-IT-v2-20260609-1749-d512-h8-nl3-s512 Text Generation • 40.1M • Updated Jun 9 • 15