Spaces:
Running
title: HuggingEnvs
emoji: π€
colorFrom: yellow
colorTo: purple
sdk: static
pinned: false
license: mit
π€ HuggingEnvs: Open RL Environments
HuggingEnvs is a home for end-to-end RL environment recipes, built to make it easier to explore, reproduce, train, and evaluate agent systems.
Explore complete and reproducible environment projects from us and the community, including:
- π Open RL environments
- π§© End-to-end environment recipes
- π» Complete implementations
- π¦ Models, datasets, and artifacts
- π§ͺ Training and evaluation setups
- π Demos and Spaces
- π Tutorials and guides
All the reproducible code β environments, rollouts, training configs, notebooks, article and slide sources β lives in one repo: github.com/adithya-s-k/HuggingEnvs. The artifacts those produce live here on the Hub.
HuggingEnvs Projects
A growing collection of open projects, environments, resources, and artifacts.
| Project | What it is | Explore |
|---|---|---|
| HuggingEnvs Academy | Articles, guides, tutorials, slides, and hands-on resources for learning how to build RL environments and agent systems. | Explore β |
| Data Agent | Training SLMs for data science with multi-harness RL environments. | Explore β |
Articles & Talks
| What it covers | Read / Watch | |
|---|---|---|
| π The Ultimate Guide to RL Environments | Building and scaling RL environments in the LLM era β how frameworks are built, how rewards are wired, how they scale to thousands of concurrent sessions. | Read β |
| ποΈ RL Environments 101 | From "what is an env?" to training your own: RL fundamentals β environment anatomy β OpenEnv β training with TRL. | Watch β |
| π Scaling RL for LLMs | RL environments and RL training β what an environment is, how reward hacking happens, how to train against your own. AMD AI Dev Day. | Watch β |
| π Multi-Harness Training | OpenEnv Γ Harbor β why an environment's failure model decides whether it can be trained against. | Watch β |
Environments
Three reference environments, each implemented across six frameworks β openenv, ors, nemo_gym, verifiers, skyrl_gym, gem. Same logic, six dialects. Source β
| Environment | Tools | OpenEnv | ORS | NeMo Gym |
|---|---|---|---|---|
| Jupyter agent β real code execution in an E2B sandbox | 4 | Space | Space | Space |
| Wordle β multi-turn, pure Python, no backend | 1 | Space | Space | Space |
| Desktop β computer-use, vision-driven Linux desktop | 19 | Space | Space | β |
Build your own
Five agent skills turn a plain-English description into a runnable RL environment across four frameworks β works with Claude Code, Cursor, Codex, OpenCode, Gemini CLI and others.
npx skills add adithya-s-k/HuggingEnvs
We're looking for new end-to-end recipes β a task, an environment, a training run, and honest results. Contributing guide β
