Correct Answers, Invalid Traces

The models of Correct Answers, Invalid Traces: What Verifiable Grade-School Math Reveals About Chain-of-Thought Traces (Puduppully, Misra, Iyer, Kalwar, Palod and Kambhampati, 2026). Each is a 12-layer GPT-NeoX model (124M parameters) trained from scratch on iGSM for 100k steps at batch 512. They differ only in the trace that followed the problem during training. Code and the evaluation protocol are at https://github.com/ratishsp/igsm-trace-validity.

folder training trace
clean-run-a minimal valid trace (the paper's main model)
clean-run-b the same recipe, second run
clean-run-c-200k the same recipe, 200k steps
swapped trace of a different problem
swapped-op-matched trace of a different problem with the same op count
shuffled-10, -30, -50, -75, -100 the first 10 to 100 percent of the trace's tokens shuffled
no-trace answer only
non-minimal valid trace with unnecessary steps
non-minimal-90 non-minimal for 90 percent of problems, minimal otherwise
seed43/* seed-43 replicates of swapped, swapped-op-matched, shuffled-30 and non-minimal
from transformers import GPTNeoXForCausalLM
model = GPTNeoXForCausalLM.from_pretrained("ratishsp/igsm-trace-validity", subfolder="clean-run-a")

The tokenizer is GPT-2's, with iGSM's special tokens 222, 223 and 224 for [PROB], [SOL] and [ANS]. evaluate.py in the code repository runs the paper's evaluation on any of these folders.

Citation

@article{puduppully2026correct,
  title   = {Correct Answers, Invalid Traces: What Verifiable Grade-School Math Reveals About Chain-of-Thought Traces},
  author  = {Puduppully, Ratish and Misra, Pranabendu and Iyer, Paarth and Kalwar, Durgesh and Palod, Vardhan and Kambhampati, Subbarao},
  journal = {arXiv preprint},
  year    = {2026}
}
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support