intelif-qwen3-4b
Typed probabilistic decisions from a small open model. Give Intelif a state and named, typed questions, each with the options to choose between, and it returns a probability for every option. The caller supplies the decision space in every request, so the same model routes support tickets, picks tools, chooses an agent's next action or answers yes/no questions.
This repository holds v0.1 (tag v0.1): LoRA adapters (r=16, alpha 32, on q/k/v/o/gate/up/down projections) and a linear scorer for Qwen3-4B at revision 1cfa9a7208912126459214e8b04321603b3df60c. Each option in the prompt is followed by an anchor token; the scorer reads the anchor's final hidden state and the probabilities are a softmax over the options. One forward pass answers a question, whatever the number of options. The LoRA weights are merged at load time, so inference costs the same as the base model.
Usage
pip install "intelif @ git+https://github.com/SkAndMl/intelif"
from intelif import Choice, Intelif, Noul
model = Intelif.from_pretrained("UserMoonlight/intelif-qwen3-4b", revision="v0.1")
response = model.system_one(
state={"ticket": "I was charged twice and need the duplicate refunded today."},
questions={
"intent": Choice(
instructions="What is the customer's main request?",
criteria={"refund": "The customer wants money returned.", "technical_help": None, "information": None},
),
"urgent": Noul(instructions="Does the ticket communicate time pressure?"),
},
)
response.choices["intent"].probabilities
response.nouls["urgent"].noul
The adapter cannot be loaded with peft alone: the scorer and prompt format live in the intelif package.
Files
adapter.safetensors: LoRA and scorer weights, with the LoRA configuration in the safetensors metadata.config.json: the model spec read byIntelif.from_pretrained(base model and revision, LoRA, scorer, prompt format, anchor token).results.json,train.log: held-out metrics and the log of the training run.baseline_results.json: the comparison below.
Results
Held-out test splits, never used for training or model selection. BANKING77's 20 held-out intents and all of CLINC150 are labels the model never saw in training. Each cell is accuracy / expected calibration error; the zero-shot baselines score the same options with Qwen3-4B's log-probabilities, once from a raw prompt and once through the chat template without thinking.
| split | chance | Qwen3-4B, raw prompt | Qwen3-4B, chat | Intelif v0.1 |
|---|---|---|---|---|
| BANKING77, seen intents | 1.3% | 45.2% / 0.226 | 63.3% / 0.336 | 89.7% / 0.011 |
| BANKING77, 20 held-out intents | 1.3% | 46.6% / 0.224 | 59.9% / 0.367 | 78.5% / 0.065 |
| CLINC150 (never trained on) | 0.7% | 54.4% / 0.104 | 72.6% / 0.246 | 86.5% / 0.021 |
| MASSIVE | 1.7% | 59.1% / 0.190 | 64.9% / 0.316 | 89.9% / 0.019 |
| xLAM tool selection (unseen tools) | 32.4% | 98.2% / 0.012 | 98.9% / 0.010 | 99.9% / 0.002 |
| ALFWorld next action | 4.0% | 55.1% / 0.236 | 59.0% / 0.390 | 85.5% / 0.029 |
| WebShop next action | 13.7% | 26.4% / 0.404 | 20.7% / 0.709 | 56.0% / 0.061 |
| MNLI mismatched | 38.6% | 69.6% / 0.191 | 80.8% / 0.181 | 92.0% / 0.020 |
| SNLI | 38.3% | 68.6% / 0.219 | 82.1% / 0.173 | 92.7% / 0.019 |
| QQP | 50.0% | 79.3% / 0.068 | 80.5% / 0.187 | 87.8% / 0.015 |
| PAWS | 50.0% | 75.9% / 0.105 | 77.4% / 0.220 | 93.6% / 0.021 |
| BoolQ | 50.0% | 85.6% / 0.063 | 84.6% / 0.149 | 88.9% / 0.023 |
| CommonsenseQA | 20.0% | 51.5% / 0.151 | 74.3% / 0.244 | 82.7% / 0.028 |
| OpenBookQA | 25.0% | 39.2% / 0.383 | 74.0% / 0.243 | 90.2% / 0.029 |
| SciQ | 25.0% | 92.3% / 0.023 | 98.2% / 0.018 | 98.7% / 0.006 |
ALFWorld and WebShop measure agreement with one recorded expert action, not task success. The full numbers, including negative log-likelihood and chat with thinking, are in baseline_results.json and come from training/baseline.py in the intelif repository.
Decision Index 0.2.1
| index | Knowledge & Reasoning | Language Understanding | Retrieval & Classification | Tools & Automation | Arts & Human Taste | median latency |
|---|---|---|---|---|---|---|
| 31.77 | 18.3 | 31.1 | 39.6 | 51.0 | 17.8 | 16.5 ms |
Chance-corrected scores (0 is random guessing, 100 is perfect) from the packaged engine on all 150,317 requests, with every request answered. The raw index is 48.62. Run on one NVIDIA RTX PRO 6000 Blackwell; the results, per-benchmark scores and environment are in UserMoonlight/intelif-decision-index. The model is strongest where the suite looks like its training data (BFCL 89.6, BANKING77 83.6, CLINC150 81.5) and near chance on knowledge-heavy, multi-step and taste benchmarks (GPQA, HLE, CRUXEval, ForecastBench). BANKING77's train split is part of the training data.
Training
One epoch, about 1,500 optimizer steps on one GPU, learning rate 2e-4 with warmup and cosine decay, bfloat16 base. Data, each pinned to a dataset revision in training/data.py: BANKING77 (train split, 20 of its 77 intents held out), MASSIVE, xLAM function calling, ALFWorld and WebShop expert trajectories, MNLI, SNLI, QQP, PAWS, BoolQ, CommonsenseQA, OpenBookQA and SciQ. None of it comes from the Decision Index suite.
Limitations
- English only.
- Trained on choice-style decisions; knowledge-heavy and multi-step reasoning tasks are its weakest area.
Scorequestions were not trained and are experimental.- Calibration was measured on the splits above; expect it to be weaker on unfamiliar kinds of decisions.
- Prompts are limited to Qwen3-4B's 40,960-token context; longer requests are refused, never truncated.
Licence
The adapter and scorer weights in this repository are released under CC BY-NC 4.0: free for research and other non-commercial use, with attribution. They are non-commercial because the training data includes SciQ (CC BY-NC 3.0).
- Base model: Qwen3-4B, Apache 2.0, downloaded from its own repository and not redistributed here.
- Code: the intelif package, MIT.
- Training data, under its own terms: BANKING77, MASSIVE and xLAM function calling (CC BY 4.0; xLAM is gated), ALFWorld and WebShop trajectories from
osunlp/early-experience(MIT), CommonsenseQA (MIT), OpenBookQA, MNLI and QQP from GLUE (QQP under Quora's terms), PAWS, SNLI (CC BY-SA 4.0), BoolQ (CC BY-SA 3.0) and SciQ (CC BY-NC 3.0).
- Downloads last month
- 37