AI & ML interests

Agents make mistakes. Those mistakes are the most valuable training data there is. Only While. turns them into a dataset, post-trains an open model on it with SFT and RL, and proves the agent stopped repeating them on a held-out test.

Recent Activity

whileai  updated a dataset about 5 hours ago
while-ai/brand
whileai  updated a Space about 5 hours ago
while-ai/README
View all activity

Organization Card

wai, the While whale. Models improve while they work.

MID-TRAINING AND POST-TRAINING FOR LANGUAGE MODELS

Models improve while they work. A scientific RL and SFT post-training library.

While is the mid-training and post-training platform for language models. Open models, the latest SFT and RL algorithms, and an open-source SDK that builds and optimizes datasets from production traces. Every gain is measured on a held-out set.

Everything here was made with the SDK and published with the numbers behind it. A model ships when it beats its base on a held-out set with an interval that clears zero. If it does not, it does not ship.

pip install whileai
import whileai as wai

wai is While's whale and the alias of the whileai SDK.

Docs | Platform | SDK | Recipes

Collections group this org by use case.