YAML Metadata Warning:empty or missing yaml metadata in repo card

Check out the documentation for more information.

Tiny Decoders

Tiny decoder experiments for instruction-conditioned Python generation.

Paper Β· Code

Each model branch contains its resolved config.yaml, checkpoint, manifest, and tokenizer files.

V1 architecture sweep

Branch Experiment
ettin17m_default Reference: Ettin-17M encoder, 64d Γ— 2-layer decoder, prefix length 4
decoder_only Reference decoder with the encoder removed
prefix_k1 Reference with prefix length 1
prefix_k8 Reference with prefix length 8
depth_1 Reference with one decoder layer
depth_4 Reference with four decoder layers
wider_d128 Wider 128d Γ— 2-layer decoder
ettin32m Reference with Ettin-32M encoder
k1_d128_L4_ettin17m Prefix 1, 128d Γ— 4 layers, Ettin-17M encoder
combined_k8_d128_L4_ettin32m Prefix 8, 128d Γ— 4 layers, Ettin-32M encoder

V2 core models

Branch Experiment
v2_base_pretrain Decoder-LM warm-start for the V2 base shape
v2_base_scratch V2 base architecture trained from scratch
v2_base_warm V2 base architecture with decoder-LM warm-start
v2_decoder_only V2 decoder with the encoder removed
v2_pretrain Decoder-LM warm-start for the scaled V2 model
v2_scaled Headline V2: 192d Γ— 6 layers, cross-attention, RoPE, SwiGLU, copy head

V1 β†’ V2 progression

Branch Experiment
progression-v1_prefix_558k 558k matched-budget V1 prefix coupling
progression-v2_xattn_558k 558k matched-budget cross-attention
progression-v1_1p2m_prefix_only 1.2M prefix-only baseline
progression-v1_5_xattn Adds cross-attention
progression-v1_5_xattn_rope Adds RoPE
progression-v1_5_xattn_rope_swiglu Adds SwiGLU
progression-v1_5_xattn_rope_swiglu_copyhead Adds the copy head
progression-v1_5_xattn_rope_swiglu_copyhead_warmstart Adds decoder-LM warm-start
progression-v2_base_full Full V2 base with data interventions

V2 leave-one-out ablations

Branch Experiment
loo-v2_base Full V2 base control
loo-v2_base_abspos Learned absolute positions instead of RoPE
loo-v2_base_gelu GELU MLP instead of SwiGLU
loo-v2_base_maxseq128 Maximum sequence length 128 instead of 256
loo-v2_base_no_copyhead Copy head removed
loo-v2_base_no_drop_stub_bodies Stub-body filtering disabled
loo-v2_base_no_warmstart Decoder-LM warm-start removed
loo-v2_base_uniform_sampling Uniform instead of length-weighted sampling

Tokenizer and coupling factorial

Branch Experiment
bpe-bpe_alias Alias-rewritten BPE with prefix coupling at 1.2M
bpe-bpe_raw Raw-code BPE with prefix coupling at 1.2M
factorial-bpe_alias_xattn_1p2m Alias-rewritten BPE with cross-attention at 1.2M

TinyPython dataset generations

Branch Dataset Model
tpy07-v1_base tpy07: initial corpus V1 base
tpy07-v1_decoder tpy07: initial corpus V1 decoder-only
tpy07-v1_wider_d128 tpy07: initial corpus V1 wider 128d
tpy07-v2_base tpy07: initial corpus V2 base
tpy07-v2_decoder tpy07: initial corpus V2 decoder-only
tpy07-v2_scaled tpy07: initial corpus V2 scaled
tpy08-v1_base tpy08: name/body-aligned corpus V1 base
tpy08-v1_decoder tpy08: name/body-aligned corpus V1 decoder-only
tpy08-v1_wider_d128 tpy08: name/body-aligned corpus V1 wider 128d
tpy08-v2_base tpy08: name/body-aligned corpus V2 base
tpy08-v2_decoder tpy08: name/body-aligned corpus V2 decoder-only
tpy08-v2_scaled tpy08: name/body-aligned corpus V2 scaled
tpy09-v1_base tpy09: semantically aligned corpus V1 base
tpy09-v1_decoder tpy09: semantically aligned corpus V1 decoder-only
tpy09-v1_wider_d128 tpy09: semantically aligned corpus V1 wider 128d
tpy09-v2_base tpy09: semantically aligned corpus V2 base
tpy09-v2_decoder tpy09: semantically aligned corpus V2 decoder-only
tpy09-v2_scaled tpy09: semantically aligned corpus V2 scaled
tpy07b-v1_base tpy07b: clean held-out corpus V1 base
tpy07b-v1_decoder tpy07b: clean held-out corpus V1 decoder-only
tpy07b-v1_wider_d128 tpy07b: clean held-out corpus V1 wider 128d
tpy07b-v2_base tpy07b: clean held-out corpus V2 base
tpy07b-v2_decoder tpy07b: clean held-out corpus V2 decoder-only
tpy07b-v2_scaled tpy07b: clean held-out corpus V2 scaled
tpy07b-v2_scaled_plus1epoch tpy07b: clean held-out corpus V2 scaled continued for one epoch

V1 clean-holdout reruns

All models below retrain a V1 sweep cell on tpy07b and evaluate it on the frozen V2 benchmark.

Branch Experiment
v1_rerun-decoder_only Decoder-only rerun
v1_rerun-depth_1 One-layer rerun
v1_rerun-depth_4 Four-layer rerun
v1_rerun-ettin17m_default Ettin-17M reference rerun
v1_rerun-prefix_k1 Prefix-length-1 rerun
v1_rerun-prefix_k8 Prefix-length-8 rerun
v1_rerun-wider_d128 Wider-128d rerun
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support