LittleLearner: Language Models Under Pedagogically Controlled Knowledge Exposure
Abstract
A curated elementary-grade pretraining corpus and 5B-parameter model create a controlled sandbox for studying knowledge acquisition, representation, and bounded capability growth via post-training and in-context learning.
Modern language models are trained on heterogeneous web-scale text corpora. Consequently, studying knowledge and skill acquisition is difficult, as prior exposure to related content is hard to characterize. To address this challenge, we introduce LITTLECURRICULUM, a curated 88B-token pretraining corpus tailored to U.S. elementary school material, explicitly excluding concepts, facts, and vocabulary taught above Grade 5. Training a 5B-parameter LLM from scratch on LITTLECURRICULUM yields LITTLELEARNER, a model with sufficient language competence for open-ended evaluation, yet with clear knowledge and capability boundaries mapped to interpretable curriculum guidelines. We release LITTLECURRICULUM and LITTLELEARNER as a developmentally restricted sandbox to study how models acquire, represent, and use data under a well-defined training scope. We illustrate the sandbox's utility in a first suite of experiments on injecting new knowledge through post-training and in-context learning. These methods let LITTLELEARNER better utilize existing knowledge, but do not raise out-of-scope capabilities. Our findings underscore the value of this controlled environment for future investigations.
Community
What happens when an LLM never sees material beyond fifth grade?
The 5B LittleLearner model trained from scratch on LittleCurriculum, a corpus restricted to K–5 material, answers this question and allows to test the effect of post-training, prompting and scaling on moving beyond the pretraining knowledge boundary.
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- Understanding Reasoning from Pretraining to Post-Training (2026)
- Knowledgeless Language Models: Suppressing Parametric Recall for Evidence-Grounded Language Modeling (2026)
- Co-LMLM: Continuous-Query Limited Memory Language Models (2026)
- Scalable Visual Pretraining for Language Intelligence (2026)
- Information Abundance Paradox: Long-Context Training Undermines Parametric Knowledge (2026)
- Knowledge Distillation from Large Reasoning Models to Compact Student Models: A Case Study on the John O Bryan Mathematics Competition (2026)
- Structured Thoughts For Improved Reasoning And Context Pruning (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2608.13545 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 15
littlelearner/littlelearner-5b-chatty
Datasets citing this paper 1
littlelearner/LittleCurriculum
Spaces citing this paper 0
No Space linking this paper