NCP OLMo 3 Stage1 - Step 100,000

NCP_Olmo3_Stage1_Step100000 is the intermediate checkpoint from the NCP OLMo 3 Stage1 training run. It was converted from Megatron iteration 100,000 into a pure Hugging Face release with sharded SafeTensors and bundled Transformers remote code.

This is a research checkpoint for the Next-Concept-Prediction (NCP) architecture. Its token-level Transformer design is aligned with OLMo 3, while NCP adds dense residual connections and discrete concept representations. The model encodes token sequences into concepts, models them at the concept level, and decodes them back into token space.

Evaluation Results

The table below reports the final Stage1 checkpoint and is included only as a reference; it is not an evaluation of this intermediate checkpoint.

Model GSM8K HumanEval MMLU
OLMo-Stage1 39.88 26.94 62.25
NCP-Olmo3-StepLast (ours-hf) 46.78 (+6.90) 31.27 (+4.33) 64.77 (+2.51)

Model Details

Field Value
Architecture NCPOlmo3ForCausalLM
Training stage Stage1
Source checkpoint Megatron iteration 100,000
Parameter scale 7B-class
Hidden size 4,096
Token backbone layers 32 (16 encoder + 16 decoder)
Concept-special layers 8
Attention heads 32
FFN hidden size 11,008
Vocabulary size 100,278
Maximum sequence length 8,192 tokens
Attention pattern 4K sliding-window attention with periodic full attention
Concept chunk size 4 tokens
Discrete concept codebooks 32 codebooks x 128 entries
Weight dtype BF16
Weight format 5 sharded SafeTensors files (842 tensors)
Weight key format Pure Hugging Face state-dict keys
Pure Transformers backend Yes
ConceptLM vLLM backend Yes
External Megatron runtime for HF loading Not required
Package validation Pure-HF export and standalone package validated
Unified HF 4K+1K parity Not run for this Stage1 release batch

The release stores one shared set of SafeTensors with dedicated Hugging Face projection keys such as q_proj, k_proj, v_proj, gate_proj, up_proj, and down_proj. No native Megatron-key weight copy is required.

Requirements

  • A CUDA-capable NVIDIA GPU with enough memory for this 7B-class checkpoint.
  • PyTorch with CUDA and bfloat16 support.
  • A Transformers version compatible with the bundled remote code.
  • trust_remote_code=True when loading through Transformers.
  • No external ConceptLM or Megatron checkout is required for the pure-HF backend.

Download and Load

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

repo_id = "ArchSpace-Collection/NCP_Olmo3_Stage1_Step100000"

tokenizer = AutoTokenizer.from_pretrained(repo_id)
model = AutoModelForCausalLM.from_pretrained(
    repo_id,
    trust_remote_code=True,
    torch_dtype=torch.bfloat16,
    device_map="cuda",
)
model.eval()

Generation

prompt = "The role of hierarchical representations in language modeling is"
inputs = tokenizer(prompt, return_tensors="pt").to("cuda")

with torch.inference_mode():
    output_ids = model.generate(
        **inputs,
        max_new_tokens=128,
        do_sample=True,
        temperature=0.8,
        top_p=0.95,
        use_cache=True,
    )

print(tokenizer.decode(output_ids[0], skip_special_tokens=True))

Files and Provenance

  • config.json defines the pure-HF architecture and auto_map entries.
  • model.safetensors.index.json maps all HF state-dict keys to five shards.
  • conversion_manifest.json records conversion provenance and shard hashes.
  • standalone_manifest.json records standalone HF/vLLM readiness.

Intended Use and Limitations

This checkpoint is released for research and evaluation. It is a base language model, not a safety-tuned assistant. Users should review the bundled custom code before enabling trust_remote_code=True, validate outputs for their application, and follow the terms of the training data and downstream deployment environment.

The bundled vLLM readiness refers to the ConceptLM vLLM integration; compatibility with an unmodified upstream vLLM installation is not implied.

Downloads last month
14
Safetensors
Model size
9B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Dataset used to train ArchSpace-Collection/NCP_Olmo3_Stage1_Step100000