File size: 7,140 Bytes
aeb0a98 b5d930b 4786222 3a4a043 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 | ---
license: apache-2.0
datasets:
- HuggingFaceTB/smollm-corpus
- HuggingFaceFW/fineweb
- HuggingFaceFW/fineweb-edu
- bigcode/starcoderdata
- HuggingFaceTB/finemath
language:
- en
tags:
- byte-level
- custom-architecture
- linear-attention
- long-context
- hierarchical-radial-transformer
- sub-quadratic
---
# HRT-v7 (Hierarchical Radial Transformer) β 148M Base
**HRT (Hierarchical Radial Transformer)** is an experimental, tokenizer-free, sub-quadratic neural network architecture designed for ultra-long context modeling (up to **128k context**) with extreme memory efficiency on consumer-grade hardware.
This checkpoint is the **Base Pretrained Model (~148M parameters)** trained on ~2.5 billion raw UTF-8 bytes using a multi-stream balanced curriculum of web text, synthetic textbooks, code, and mathematics.
---
## β‘ Key Highlights
* **Native Byte-Level Processing:** Operates directly on raw UTF-8 bytes (vocab_size = 257: 0β255 bytes + EOS 256). No BPE/WordPiece tokenizers.
* **128k Native Context:** Pretrained directly with sequence length T = 131,072 bytes.
* **Sub-Quadratic Scaling (T β K β K β T):** Replaces dense O(TΒ²) self-attention with radial latent compression (K = 512 outer latents, 16 center latents), yielding linear/sub-quadratic compute and memory footprint.
* **Deep Equilibrium Core (JFB):** Features an implicit reasoning core trained with **Jacobi-Free Backpropagation (JFB)** for constant O(1) backpropagation memory overhead.
* **Hardware-Aware Quantization:** Built-in TurboQuantizer using random orthogonal QR rotation matrices to suppress outliers during INT8 KV caching.
---
## π Training Details & Hyperparameters
The model was pretrained on NVIDIA Blackwell / Hopper hardware in bfloat16 mixed precision.
### Model Configuration (ModelConfig)
| Hyperparameter | Value | Description |
| :--- | :--- | :--- |
| **Parameters** | ~148M | Active trainable weights |
| **Context Length (seq_len)** | **131,072** (128k bytes) | Native training window length |
| **Vocab Size** | 257 | 256 bytes + 1 EOS token |
| **Model Dimension (d_model)** | 768 | Hidden representation size |
| **FFN Dimension (d_ff)** | 3072 | SwiGLU projection dimension |
| **Outer Latents (K)** | 512 | Compressed outer working memory |
| **Center Latents** | 16 | Deep reasoning core latents |
| **Routing Top-K** | 64 | Sparse attention routing threshold |
| **Attention Heads** | 12 | Outer, Inner, and Latent heads |
| **Outer / Inner Cycles** | 6 / 8 | Hierarchical gather/reasoning loops |
| **Local Conv Kernel** | 7 | Causal dual-dilated local token mixer |
| **Positional Bias** | HARP | Hierarchical Adaptive Relational Positioning |
| **Implicit Core** | Enabled | 3 internalization steps with JFB |
| **Auxiliary Losses** | Enabled | SVD-based low-rank compaction loss + JFB loss |
### Pretraining Optimization & Setup
* **Optimizer:** AdamW (beta1 = 0.9, beta2 = 0.95, weight_decay = 0.01)
* **Precision:** Native bfloat16 with Gradient Checkpointing
* **Effective Batch Size:** 8 sequences Γ 131,072 = **1,048,576 bytes per step** (Batch Size 4 Γ Gradient Accumulation 2)
* **Gradient Clipping:** 1.0
* **Learning Rate Schedule:** Cosine Annealing cooldown down to LR_min = 1e-5
---
## Dataset Mixture
The model was trained on a dynamically balanced 5-stream multiplexer with streaming packing:
```
FastMultiStream Distribution:
25% β HuggingFaceTB/smollm-corpus (Cosmopedia v2 - synthetic textbooks & stories)
25% β HuggingFaceFW/fineweb (sample-10BT - general web crawl)
20% β HuggingFaceFW/fineweb-edu (sample-10BT - educational & academic web text)
20% β bigcode/starcoderdata (Python subsets - structured code & AST logic)
10% β HuggingFaceTB/finemath (finemath-3plus - math reasoning & LaTeX)
```
Documents were packed into continuous 131,072-byte buffers with unique segment IDs (segment_ids) and EOS separators to ensure segment-aware causal attention boundaries.
---
## Quickstart / How to Use
### 1. Installation
Clone the repository and install the HRT package:
```bash
git clone https://github.com/5bridge/HRT.git
cd HRT
pip install -e .
```
### 2. Running Inference (Next-Byte Completion)
Because this is a **Base Pretrained Model** (not an instruction/chat tuned model), it functions as an autoregressive text/code completion engine.
```python
import torch
import torch.nn.functional as F
from hrt import ModelConfig, HierarchicalRadialTransformerV7
device = "cuda" if torch.cuda.is_available() else "cpu"
# 1. Initialize Configuration matching training
cfg = ModelConfig(
d_model=768,
d_ff=3072,
n_outer_latents=512,
n_outer_cycles=6,
n_inner_cycles=8,
n_center_latents=16,
routing_k=64,
n_outer_heads=12,
n_inner_heads=12,
n_latent_heads=12,
vocab_size=257,
max_seq_len=131072,
use_qk_norm=True,
use_rezero=True,
use_compaction=True,
use_internalization=True,
use_jfb=True,
use_q_cache=True,
)
# 2. Load model & weights
model = HierarchicalRadialTransformerV7(cfg).to(device)
weights = torch.load("hrt_v7_148m_weights.pt", map_location=device)
model.load_state_dict(weights["model"] if "model" in weights else weights)
model.eval()
# 3. Autoregressive Byte-level Generation
def generate(prompt: str, max_new_bytes: int = 120, temp: float = 0.5, top_k: int = 5):
prompt_bytes = list(prompt.encode("utf-8"))
prompt_ids = torch.tensor([prompt_bytes], dtype=torch.long, device=device)
with torch.no_grad():
prompt_emb = model.tok_emb(prompt_ids)
logits, cache = model._init_generation_cache(prompt_emb)
out_bytes = list(prompt_bytes)
for _ in range(max_new_bytes):
l = logits / max(temp, 1e-5)
if top_k > 0:
v, _ = torch.topk(l, min(top_k, l.size(-1)))
l[l < v[:, [-1]]] = float("-inf")
nxt = torch.multinomial(F.softmax(l, dim=-1), num_samples=1)
nxt_id = nxt.item()
if nxt_id == 256: # EOS
break
out_bytes.append(nxt_id)
nxt_emb = model.tok_emb(nxt)
logits = model.step_generation(nxt_emb, cache)
return bytes(out_bytes).decode("utf-8", errors="replace")
# Test completion
print(generate("def", max_new_bytes=100))
```
---
## β οΈ Limitations & Intended Use
* **Base Model Nature:** This model has not undergone Supervised Fine-Tuning (SFT) or RLHF/DPO. It will not behave as a conversational assistant by default and may loop if prompted with chat-like templates without a stopping token.
* **Proof-of-Concept Scale:** Trained on ~2.5B bytes as a compute-limited validation run. It exhibits strong syntactic comprehension (LaTeX, Python, Markdown, JSON AST schemas), but requires further scale for complex factual recall and deep semantic reasoning.
* **Information Bottleneck:** The T to K compression naturally trades off lossless, needle-in-a-haystack memorization for bounded sub-quadratic memory footprint.
---
## License
This model and its code are released under the **Apache 2.0 License**.
https://github.com/5bridge/HRT |