MustaqiLLM

MustaqiLLM is a 5.17-billion-parameter Uzbek chat and text-classification model. It follows Uzbek instructions reliably, writes fluent Uzbek in both Latin and Cyrillic script, and is strong on sentiment and news classification. It is not a knowledge model: on multiple-choice knowledge benchmarks it performs at chance. Read the Evaluation and Limitations sections before using it — they are specific about what works and what does not.

Parameters 5.17 B
Architecture NeuronLMForCausalLM (custom, ships with the repo)
Layers / hidden 36 / 3584
Attention GQA, 28 query heads : 4 KV heads, head_dim 128, QK-norm
Position encoding RoPE, θ = 500000
Context length 4096 tokens
Vocabulary 48,000 (BPE)
Embeddings untied
Weights dtype bfloat16 (embeddings and lm_head stored fp32)
Languages Uzbek (Latin + Cyrillic), English, Russian

Quick start

The architecture is custom, so trust_remote_code=True is required — the modeling code ships inside this repository.

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "NeuronUz/MustaqiLLM"

tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
    model_id,
    trust_remote_code=True,
    dtype=torch.bfloat16,       # weights are bf16; do not load in fp32
    device_map="cuda",
).eval()

messages = [{"role": "user", "content": "O'zbekistonning poytaxti qaysi shahar?"}]
inputs = tokenizer.apply_chat_template(
    messages,
    add_generation_prompt=True,
    return_tensors="pt",
    return_dict=True,
).to(model.device)

with torch.no_grad():
    out = model.generate(
        **inputs,
        max_new_tokens=256,
        do_sample=False,           # greedy is fine for a short answer like this;
                                   # for open chat use the sampling settings below
        eos_token_id=5,            # <|im_end|> -- also the repo default
        pad_token_id=3,            # <pad>
    )

print(tokenizer.decode(out[0][inputs["input_ids"].shape[1]:], skip_special_tokens=True))
Oʻzbekistonning poytaxti - Toshkent.

Chat template

The model uses ChatML. tokenizer.apply_chat_template applies it for you; the raw form is:

<|im_start|>system
{system}<|im_end|>
<|im_start|>user
{user}<|im_end|>
<|im_start|>assistant
{assistant}<|im_end|>

A system turn is optional, and for general chat you should leave it out — a generic system prompt measurably increases repetition (see Generation settings). Task-specific system prompts, in Uzbek, work well.

Generation settings

These are measured, not guessed. 35 decoding configurations were swept over 120 held-out Uzbek prompts across 14 categories with 2 seeds each — 8,400 generations — scored automatically for verbatim sentence repetition and for failure to emit <|im_end|> within the token budget. Because those metrics see repetition but not fluency, the finalists were then compared head-to-head by an LLM judge over 1,200 pairwise judgements with randomised A/B order.

Recommended for open chat:

out = model.generate(
    **inputs,
    max_new_tokens=512,
    do_sample=True,
    temperature=0.7,
    top_p=0.9,
    repetition_penalty=1.05,   # not optional -- see below (1.05-1.10 all work)
    use_cache=True,
)
setting value why
repetition_penalty 1.05–1.10 for chat The single most important setting. Without it the model restates whole sentences verbatim. Duplicate-sentence rate at temperature 0.7: 4.0% at 1.00, 1.5% at 1.03, 1.3% at 1.05, 1.0% at 1.10, 0.1% at 1.15. Do not read that as "higher is better" — see the note below the table.
do_sample / temperature / top_p True, 0.7, 0.9 for chat; False (greedy) for classification, extraction and short answers generation_config.json ships do_sample: true with no temperature or top_p, so the unconfigured default is temperature 1.0 / top_p 1.0 — pass these explicitly. Terse tasks showed a 0% repetition rate under every configuration tested, so greedy is safe there.
eos_token_id 5 (<|im_end|>) The turn terminator, already the default in config.json / generation_config.json — you do not need to pass it. Do not override it with the pretraining EOS (</s>), which never appears in chat data: generation would then run to max_new_tokens.
system prompt omit it for general chat A generic system turn measurably degrades output. Duplicate-sentence rate over a 24-prompt subset: 0.0% with no system prompt, 1.9% with a generic Uzbek one, 5.2% with a generic English one (at temperature 0.7, repetition_penalty 1.05); without a repetition penalty the same comparison is 11.2% / 23.7% / 14.9%. Task-specific system prompts (a required format, a persona) are fine — it is the generic "you are a helpful assistant" turn that hurts.
dtype torch.bfloat16 Trained in bf16. float16 is also safe — no overflow, and output quality is indistinguishable — so pre-Ampere GPUs are supported. float32 doubles memory for half the throughput (205 vs 412 tok/s) and changes nothing.

More penalty is not better past ~1.10. The automatic metrics keep improving as repetition_penalty rises, but fluency does not. Judged head-to-head on the same prompts, rp=1.15 — the cleanest configuration by repetition metrics — lost to gentler settings: 30.6% win rate against rp=1.10 and 38.8% against rp=1.05. Between 1.05 and 1.10 the judge is a coin flip (52.2%), so anywhere in that band is fine. Below it there is a real floor: rp=1.05 beats rp=1.03 at 60.4%. Sampling with a penalty beats greedy outright (58.8%).

Greedy decoding degrades as the output gets longer, which is why it is recommended above only for short outputs. Over the full 120-prompt sweep at a 384-token budget, greedy produced 17.1% duplicate sentences and failed to terminate on 21.7% of prompts, against 1.3% and 4.2% for t=0.7, rp=1.05. On chat and long-form prompts with a 768-token budget the gap widens:

configuration never emits <|im_end|> duplicate sentences worst case
greedy 23.1% 27.3% one sentence repeated 9.8×
t=0.7, top_p=0.9 15.4% 8.0% 2.2×
t=0.7, top_p=0.9, rp=1.05 11.5% 3.0% 1.7×
t=0.7, top_p=0.9, rp=1.10 0.0% 0.8% 1.1×

Lowering the temperature makes this worse, not better, because sharpening the distribution locks the model into the repeat loop. Without a repetition penalty, duplicate sentences rise from 1.5% at temperature 0.9 to 8.8% at 0.5; a separate probe at temperature 0.3 reached 19.1%, the worst of any configuration tested. Determinism is genuinely in tension with quality here: greedy plus repetition_penalty=1.10 still leaves 7.9% duplicate sentences — better than greedy alone, but far short of sampling. If you need reproducible output, sample with a fixed seed rather than decoding greedily.

Two categories are much harder than the rest and need a larger max_new_tokens: Uzbek Cyrillic prompts (37.6% hit the token cap, 12.4% duplicate sentences, pooled across all configurations) and refusals (21.8% and 9.2%) — the model has trouble ending a turn once it starts declining a request. Everything else — translation, short answers, grammar and style rewriting, multi-turn — sat at or near 0% on both metrics under every configuration tested.

Batch size changes greedy output: identical prompts decoded at batch 1 and batch 12 matched in only 24 of 32 cases, because left-padding shifts the numerics. Fix the batch size when comparing runs.

Memory: the checkpoint is 11.0 GB on disk (embeddings and lm_head are stored fp32); loading with dtype=torch.bfloat16 as above casts them down to ~10.3 GB of weights, so a single 16 GB GPU is enough for inference.

config.json sets use_cache: false, but generation_config.json sets use_cache: true, so generate() uses the KV cache. Pass use_cache=True explicitly if you write your own decode loop.

Serving

vLLM and SGLang cannot load this model. They reimplement each architecture internally rather than executing a repository's Python, and NeuronLMForCausalLM is not in their model registries — trust_remote_code only covers the config and tokenizer there. Use the transformers backend, or convert the weights (the architecture is Qwen3-equivalent apart from fused qkv_proj / gate_up_proj and out_proj naming; splitting those tensors and renaming to the Qwen3 layout yields a checkpoint vLLM will serve).


Evaluation

Full public benchmark suite, greedy decoding, transformers backend, seed 42, complete test sets (no subsampling). Scores are accuracy unless noted.

Uzbek benchmarks

benchmark n score invalid rate
uzlib (Uzbek linguistic MCQ) 1,861 0.2875 0.0000
TUMLU-Uzbek (Uzbek MMLU) 700 0.3286 0.0000
MMLU-Uz (translated MMLU) 14,042 0.2584 0.0000
News topic classification (10-way, risqaliyevds/uzbek-zero-shot-classification) 96,970 0.6531 0.0000
Sentiment (binary) 10,000 0.9259 0.0001

Random baselines: 0.25 for the 4-way MCQ tasks, 0.10 for news, 0.50 for sentiment.

English

benchmark n score invalid rate
MMLU (English) 14,042 0.2619 0.0000

Translation (FLORES+)

direction n BLEU COMET length ratio
English → Uzbek 2,009 5.17 0.7397 1.018
Uzbek → English 2,009 1.83 0.5376 1.229

uzlib, per split

split n score
fill_in 52 0.3077
correct_word (orthography) 1,501 0.3011
meaning_in_context 72 0.2639
meaning 236 0.2034

News, per class

class n score
Sport 16,113 0.8743
Texnologiya (Technology) 5,177 0.7309
Madaniyat (Culture) 2,405 0.7081
Siyosat (Politics) 29,500 0.6794
Iqtisodiyot (Economy) 10,755 0.6596
Salomatlik (Health) 3,505 0.6579
Ta'lim (Education) 1,987 0.6548
Ekologiya (Ecology) 1,784 0.5667
Xorijiy Yangiliklar (World news) 11,732 0.5124
Oila va Jamiyat (Family & Society) 14,012 0.4273

Limitations

  • MCQ knowledge tasks are at chance. uzlib, MMLU-Uz and MMLU-English all sit within noise of the 0.25 baseline over ~30,000 questions, with near-zero invalid rates — correct format, wrong answer. This is missing knowledge, not parsing. Do not use it for factual QA, exams, or retrieval-free knowledge tasks. TUMLU-Uzbek (0.3286) is the only MCQ result above chance, on a 700-item sample (±3.5%).
  • Uzbek → English translation is weak (BLEU 1.83, length ratio 1.229): it over-generates. English → Uzbek is usable (COMET 0.7397) but below dedicated MT systems.
  • Script conversion does not work despite being trained for it — Latin→Cyrillic requests often return the input unchanged.
  • Cyrillic artifacts. The Cyrillic data was machine-transliterated; loanwords and brand names can be mangled (FacebookФакебоок) and stray Cyrillic characters leak into Latin words. Cyrillic chat is coherent, but its orthography is less reliable than Latin.
  • Self-identification. Identity data predates the current name, so the model calls itself "NeuronAI 5B".
  • Uneven news classification: 0.4273 on the diffuse "Oila va Jamiyat" class vs 0.8743 on Sport.
  • Safety. No safety alignment, RLHF, or red-teaming; no refusal training beyond what the instruction data incidentally contains. It can produce incorrect, biased, or unsafe content and will state false facts fluently. Evaluate before any user-facing deployment.

Intended use

Suitable for: Uzbek-language chat and assistance; text classification (sentiment, topic); Uzbek text generation and rewriting in Latin or Cyrillic; English → Uzbek translation where approximate meaning suffices; a base for further fine-tuning.

Not suitable for: factual question answering or anything knowledge-intensive; exam-style multiple choice; Uzbek → English translation; script transliteration; any application where a confidently-stated wrong fact causes harm (medical, legal, financial advice).

License

Apache 2.0. Training data licensing follows the sources of the underlying public datasets.

Citation

@misc{mustaqillm,
  title  = {MustaqiLLM: an instruction-tuned Uzbek language model},
  author = {NeuronUz},
  year   = {2026},
  url    = {https://huggingface.co/NeuronUz/MustaqiLLM}
}
Downloads last month
634
Safetensors
Model size
5B params
Tensor type
F32
·
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support