Papers
arxiv:2609.13356

ZGCM-1: A Fully Open and Extremely Efficient Foundation Model for Math and Agentic Search

Published on Sep 11
Β· Submitted by
Yifei Shen
on Sep 15
Authors:
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,

Abstract

ZGCM-1 is a 7B open foundation model that combines internal reasoning with external tool use, trained via efficient architecture-system co-design, progressive long-context scaling, and autonomous agent workflows to achieve strong reasoning and efficiency.

In this work, we present ZGCM-1, a fully open 7B dense foundation model trained from scratch with extreme data, system, and algorithmic efficiency. ZGCM-1 is founded on a core premise: compact models cannot passively memorize the open web, but can overcome parametric capacity limits by coupling deliberate internal thinking with active external tool use. To support this paradigm across a 256K context, we develop an end-to-end, high-efficiency open training recipe: Architecture & System Co-design: interleaved gated sliding-window and full attention, and a stable FP8 Muon optimizer; Progressive Curriculum & MDP Mid-Training: context scaling across 16K, 64K, and 256K, and the reformulation of interaction traces into Markov Decision Processes. Furthermore, we establish an AI-native R&D workflow where agent swarms autonomously manage cluster operations, data curation, and rapid diagnostic evaluation. Extensive evaluations show that ZGCM-1-7B is competitive across 7B model family on general benchmarks. On several challenging mathematical reasoning and agentic search suites, it remains competitive with frontier models orders of magnitude larger, such as Qwen3-235B-A22B and GLM-5.1. We also show that our pre-training design offers a ~4.2x efficiency improvement in 16K pre-training time-to-loss. Across the full development lifecycle, we distill eight actionable empirical findings-spanning architectural scaling, SFT quality pruning, long-context generalization, and agentic co-training dynamics. To facilitate community research, we open-source model weights from the pre-training, mid-training, and post-training stages, intermediate checkpoints, training code, per-stage data and data recipes, and W&B logs.

Community

Paper submitter

A fully open 7B LLM, including model, training infra, data, and wandb. Match Qwen3-8B on general datasets and competitive with frontier models orders of magnitude larger, such as Qwen3-235B-A22B and GLM-5.1 on math and agentic search datasets. Its highlight includes FP8 pretraining, MDP midtraining, and AI4AI.

Code: https://github.com/zgcagi/ZGCM-1
Model: https://huggingface.co/zgcagi/ZGCM-1-7B
Data: https://huggingface.co/datasets/zgcagi/ZGCM-1-Data
ZGCM

A big step toward recursive self-improvement (RSI) πŸ”₯πŸ”₯πŸ”₯

The "AI4AI" part is the real headline β€” agent swarms running cluster ops and data curation means the model partially built itself. RSI era loading… πŸš€

So cool!!!!! Really surprising work! A big step to RSI&AGI.

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2609.13356
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 1

Datasets citing this paper 1

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2609.13356 in a Space README.md to link it from this page.

Collections including this paper 1