Article
Andrey PRO
Bochkov
AI & ML interests
None yet
Recent Activity
published an article about 7 hours ago
Token Identity Is Not Meaning: What Fixed-Input Language Models Can Teach Us updated a Space about 10 hours ago
BEMSH-BVV/README published a Space about 10 hours ago
BEMSH-BVV/READMEOrganizations
Articles 2
Article
1
Emergent Semantics Beyond Token Embeddings: A GPT-like Transformer Learns with Frozen 16‑D Binary Token-ID Embeddings (n_embed=16)
Do Language Models Need a Trainable Input Embedding Table?
This collection is provided for reproducibility of the paper's main claim
-
Bochkov/ab_ext_learned
Text Generation • 2B • Updated • 448 -
Bochkov/ab_ext_binary16
Text Generation • 2B • Updated • 416 -
Bochkov/ab_ext_gf2
Text Generation • 2B • Updated • 375 -
Do Language Models Need a Trainable Input Embedding Table? Fixed Minimal Token Codes at 1.7B-Class Scale
Paper • 2610.04002 • Published • 3
Beyond the Parameter Monolith: Modular Language Modeling
This collection is provided for reproducibility of the paper's main claim
Do Language Models Need a Trainable Input Embedding Table?
This collection is provided for reproducibility of the paper's main claim
-
Bochkov/ab_ext_learned
Text Generation • 2B • Updated • 448 -
Bochkov/ab_ext_binary16
Text Generation • 2B • Updated • 416 -
Bochkov/ab_ext_gf2
Text Generation • 2B • Updated • 375 -
Do Language Models Need a Trainable Input Embedding Table? Fixed Minimal Token Codes at 1.7B-Class Scale
Paper • 2610.04002 • Published • 3
models 24
Bochkov/ab_ext_gf2
Text Generation • 2B • Updated • 375
Bochkov/ab_ext_binary16
Text Generation • 2B • Updated • 416
Bochkov/ab_ext_learned
Text Generation • 2B • Updated • 448
Bochkov/fem-multi-mesh-1p7b
Text Generation • 2B • Updated • 714
Bochkov/modular-reasoning-0p5b-g6p5-demo
Text Generation • 0.5B • Updated • 666
Bochkov/llm-fix-min-affine-recoded-minimal-code-table-free
Text Generation • 0.5B • Updated • 307
Bochkov/llm-fix-min-fixed-minimal-binary-code
Text Generation • 0.5B • Updated • 321
Bochkov/llm-fix-min-baseline-learned-input-table-model-classic
Text Generation • 0.5B • Updated • 316
Bochkov/growing-transformers-model-frozen-16-bit-baseline-monolyth-181m
Text Generation • 0.2B • Updated • 25
Bochkov/growing-transformers-model-unfrozen-baseline-monolyth-247m
Text Generation • 0.2B • Updated • 23
datasets 0
None public yet