IndustryCorpus Collection 多行业预训练语料库,由 Xiaofeng Shi(BAAI)策划并维护。 A multi-industry pre-training corpus curated and maintained by Xiaofeng Shi at BAAI. • 19 items • Updated 2 days ago • 9
Infinity Instruct Collection Scaling Instruction Selection and Synthesis to Enhance Language Models • 17 items • Updated Feb 4 • 12
Dolma Collection allenai's Dolma dataset as native Hugging Face datasets • 12 items • Updated Sep 8, 2025 • 8
SimPO Collection This collections contains a list of SimPO and baseline models. • 49 items • Updated Mar 16, 2025 • 24
Chinese Llama-3 series Collection This collection hosts the LLMs of Chinese-LLaMA-Alpaca-3 project, including Llama-3-Chinese, Llama-3-Chinese-Instruct, etc. • 12 items • Updated Aug 23, 2024 • 15
Faro Series Collection Faro chat models are fine-tuned on Fusang, focusing on practicality and long-context modeling. • 6 items • Updated Apr 11, 2024 • 3