intelif-qwen3-4b

Typed probabilistic decisions from a small open model. Give Intelif a state and named, typed questions, each with the options to choose between, and it returns a probability for every option. The caller supplies the decision space in every request, so the same model routes support tickets, picks tools, chooses an agent's next action or answers yes/no questions.

This repository holds v0.1 (tag v0.1): LoRA adapters (r=16, alpha 32, on q/k/v/o/gate/up/down projections) and a linear scorer for Qwen3-4B at revision 1cfa9a7208912126459214e8b04321603b3df60c. Each option in the prompt is followed by an anchor token; the scorer reads the anchor's final hidden state and the probabilities are a softmax over the options. One forward pass answers a question, whatever the number of options. The LoRA weights are merged at load time, so inference costs the same as the base model.

Usage

pip install "intelif @ git+https://github.com/SkAndMl/intelif"
from intelif import Choice, Intelif, Noul

model = Intelif.from_pretrained("UserMoonlight/intelif-qwen3-4b", revision="v0.1")

response = model.system_one(
    state={"ticket": "I was charged twice and need the duplicate refunded today."},
    questions={
        "intent": Choice(
            instructions="What is the customer's main request?",
            criteria={"refund": "The customer wants money returned.", "technical_help": None, "information": None},
        ),
        "urgent": Noul(instructions="Does the ticket communicate time pressure?"),
    },
)

response.choices["intent"].probabilities
response.nouls["urgent"].noul

The adapter cannot be loaded with peft alone: the scorer and prompt format live in the intelif package.

Files

  • adapter.safetensors: LoRA and scorer weights, with the LoRA configuration in the safetensors metadata.
  • config.json: the model spec read by Intelif.from_pretrained (base model and revision, LoRA, scorer, prompt format, anchor token).
  • results.json, train.log: held-out metrics and the log of the training run.
  • baseline_results.json: the comparison below.

Results

Held-out test splits, never used for training or model selection. BANKING77's 20 held-out intents and all of CLINC150 are labels the model never saw in training. Each cell is accuracy / expected calibration error; the zero-shot baselines score the same options with Qwen3-4B's log-probabilities, once from a raw prompt and once through the chat template without thinking.

split chance Qwen3-4B, raw prompt Qwen3-4B, chat Intelif v0.1
BANKING77, seen intents 1.3% 45.2% / 0.226 63.3% / 0.336 89.7% / 0.011
BANKING77, 20 held-out intents 1.3% 46.6% / 0.224 59.9% / 0.367 78.5% / 0.065
CLINC150 (never trained on) 0.7% 54.4% / 0.104 72.6% / 0.246 86.5% / 0.021
MASSIVE 1.7% 59.1% / 0.190 64.9% / 0.316 89.9% / 0.019
xLAM tool selection (unseen tools) 32.4% 98.2% / 0.012 98.9% / 0.010 99.9% / 0.002
ALFWorld next action 4.0% 55.1% / 0.236 59.0% / 0.390 85.5% / 0.029
WebShop next action 13.7% 26.4% / 0.404 20.7% / 0.709 56.0% / 0.061
MNLI mismatched 38.6% 69.6% / 0.191 80.8% / 0.181 92.0% / 0.020
SNLI 38.3% 68.6% / 0.219 82.1% / 0.173 92.7% / 0.019
QQP 50.0% 79.3% / 0.068 80.5% / 0.187 87.8% / 0.015
PAWS 50.0% 75.9% / 0.105 77.4% / 0.220 93.6% / 0.021
BoolQ 50.0% 85.6% / 0.063 84.6% / 0.149 88.9% / 0.023
CommonsenseQA 20.0% 51.5% / 0.151 74.3% / 0.244 82.7% / 0.028
OpenBookQA 25.0% 39.2% / 0.383 74.0% / 0.243 90.2% / 0.029
SciQ 25.0% 92.3% / 0.023 98.2% / 0.018 98.7% / 0.006

ALFWorld and WebShop measure agreement with one recorded expert action, not task success. The full numbers, including negative log-likelihood and chat with thinking, are in baseline_results.json and come from training/baseline.py in the intelif repository.

Decision Index 0.2.1

index Knowledge & Reasoning Language Understanding Retrieval & Classification Tools & Automation Arts & Human Taste median latency
31.77 18.3 31.1 39.6 51.0 17.8 16.5 ms

Chance-corrected scores (0 is random guessing, 100 is perfect) from the packaged engine on all 150,317 requests, with every request answered. The raw index is 48.62. Run on one NVIDIA RTX PRO 6000 Blackwell; the results, per-benchmark scores and environment are in UserMoonlight/intelif-decision-index. The model is strongest where the suite looks like its training data (BFCL 89.6, BANKING77 83.6, CLINC150 81.5) and near chance on knowledge-heavy, multi-step and taste benchmarks (GPQA, HLE, CRUXEval, ForecastBench). BANKING77's train split is part of the training data.

Training

One epoch, about 1,500 optimizer steps on one GPU, learning rate 2e-4 with warmup and cosine decay, bfloat16 base. Data, each pinned to a dataset revision in training/data.py: BANKING77 (train split, 20 of its 77 intents held out), MASSIVE, xLAM function calling, ALFWorld and WebShop expert trajectories, MNLI, SNLI, QQP, PAWS, BoolQ, CommonsenseQA, OpenBookQA and SciQ. None of it comes from the Decision Index suite.

Limitations

  • English only.
  • Trained on choice-style decisions; knowledge-heavy and multi-step reasoning tasks are its weakest area.
  • Score questions were not trained and are experimental.
  • Calibration was measured on the splits above; expect it to be weaker on unfamiliar kinds of decisions.
  • Prompts are limited to Qwen3-4B's 40,960-token context; longer requests are refused, never truncated.

Licence

The adapter and scorer weights in this repository are released under CC BY-NC 4.0: free for research and other non-commercial use, with attribution. They are non-commercial because the training data includes SciQ (CC BY-NC 3.0).

  • Base model: Qwen3-4B, Apache 2.0, downloaded from its own repository and not redistributed here.
  • Code: the intelif package, MIT.
  • Training data, under its own terms: BANKING77, MASSIVE and xLAM function calling (CC BY 4.0; xLAM is gated), ALFWorld and WebShop trajectories from osunlp/early-experience (MIT), CommonsenseQA (MIT), OpenBookQA, MNLI and QQP from GLUE (QQP under Quora's terms), PAWS, SNLI (CC BY-SA 4.0), BoolQ (CC BY-SA 3.0) and SciQ (CC BY-NC 3.0).
Downloads last month
37
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for UserMoonlight/intelif-qwen3-4b

Finetuned
Qwen/Qwen3-4B
Adapter
(1179)
this model