TMT 200M Base
TMT 200M Base is scaled up and PyTorch pretrain only version of TMT-4.5M(https://github.com/jrz97619761/test-model-thing). TMT was originally designed and implemented in MLX by jrz97619761. Credit for the original design belongs to that creator. This PyTorch scaling experiment uses a different training procedure and its checkpoints are not compatible with the MLX runtime.
Model details
| Property | Value |
|---|---|
| Architecture | 21 parallel recurrent trace layers, width 3,072 |
| Vocabulary | 256 UTF-8 byte values |
| Parameters | ~200M |
| Checkpoint | model.pt, TMT version-1 model weights and configuration |
| Framework | Custom PyTorch implementation |
| Training stage | Base pretraining, 24,502 optimizer updates |
Each trace layer reads the same byte embedding, maintains a decaying recurrent memory, and contributes to a residual sum before byte prediction. Training uses backpropagation through 1,024-byte sequences. The released checkpoint contains model weights only, without optimizer or resume state.
Training
The sampling weights were 60% FineWeb-Edu-Dedup and 20% Cosmopedia v2 from SmolLM-Corpus, plus 20% CodeParrot-Clean. These weights do not mean the full source datasets were used. Exact source revisions and shard paths are in training_sources.json.
Training used one RTX 3090, sequence length 1,024, batch size 4, gradient accumulation 4, and peak learning rate 1e-4. Only 12 hours of GPU time was used and if a week or month was given much better results would come.
Evaluation and limitations
At step 24,500, diagnostic validation cross-entropy was 1.2864 nats per byte (1.8558 bits per byte). Validation excluded exact document hashes seen in training, but this small source-specific sample is not an independent benchmark and does not rule out near duplicates. No general generation, safety, or downstream benchmark score is available.
This is a research checkpoint for this novel design. Output may be incoherent, repetitive, inaccurate, biased, or unsafe. English was the primary training language; performance in other languages is unknown. The model has not been shown to match Transformer models of similar parameter count.
Run locally
Install a compatible PyTorch build, then the package requirement:
pip install -r requirements.txt
python -m tmt_torch.inference model.pt --prompt "The history of computing began" --max-bytes 256
The script chooses CUDA when available and otherwise uses CPU. --device cpu, --temperature, --top-k, and --seed control inference. It prints the prompt followed by generated text. This is a custom PyTorch model; transformers.AutoModel.from_pretrained() and the standard Hugging Face text-generation widget cannot load it.
Attribution
The original Test-Model-Thing architecture and MLX project were created by jrz97619761. This release scales that design and changes the framework and training method. The original creator's MIT copyright and permission notice are retained in this release. Please preserve that notice and credit when sharing derivatives.