Instructions to use CountingSheep/vev-4b-lora with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use CountingSheep/vev-4b-lora with PEFT:
Task type is invalid.
- Notebooks
- Google Colab
- Kaggle
vev-4b-lora
Jev-like decision models that can also see images. Ask yes/no, multiple-choice or graded questions about a
piece of text, a JSON record or a screenshot, and get a probability for every allowed answer in one forward pass.
Vev serves the same request format as TypeSafe's
Jev (/v1/systemone), and runs on your own GPU.
Code, API and full results: Xiaooolong/vev.
This repository holds the LoRA adapter of CountingSheep/vev-4b (130
MB). vev serve applies it to Qwen/Qwen3.5-4B at load time, so it is the
smaller download if you already have the base model.
questions = {
"error": Noul(instructions="Does the screen show an error message?"),
"step": Choice(instructions="Which checkout step is the user on?",
criteria={"shipping": None, "payment": None, "review": None}),
"next": Choice(instructions="What should the user do next?",
criteria={"retry": "Try another card", "wait": "Wait for the order to ship",
"nothing": "Nothing, the order went through"}),
}
# vev-4b-lora: error 0.991
# step {"shipping": 0.005, "payment": 0.765, "review": 0.229}
# next {"retry": 0.983, "wait": 0.004, "nothing": 0.014}
v0.1 is a research preview, released for non-commercial use under CC BY-NC 4.0.
Use it
pip install git+https://github.com/Xiaooolong/vev typesafe-sdk
vev serve --model CountingSheep/vev-4b-lora # add --revision v0.1.0 to pin this release
from typesafe_sdk import Choice, Noul, TypeSafeClient
client = TypeSafeClient(api_key="local", base_url="http://127.0.0.1:8009")
resp = client.system_one(
state="Order #4411 still shows 'label created' after 9 days. I need it before Friday.",
questions={"urgent": Noul(instructions="Is the customer asking for something time-sensitive?")},
)
print(resp.answers["urgent"].noul)
Noul is the API's name for a yes/no question. Images go into the state as {"image": {"url": "data:image/png;base64,..."}}. Needs an NVIDIA GPU with about 10 GB of free memory.
What it can do
Accuracy on human-labelled data, by the kind of question:
| Kind of question | Example | vev-4b-lora |
|---|---|---|
| Jev-style decisions on text | JevBench public subset | 0.766 |
| The same, on sources the model has not seen | kev transfer-v4 | 0.776 |
| Chinese | judgekit | 0.962 |
| Pairs where a small change flips the answer | nimble | 0.707 |
| UI state in an app screenshot | "Is there a switch or checkbox that is turned on?" / "Is there a text input field?" | 0.928 / 0.951 |
| An image against a written safety policy | the LlavaGuard policy categories | 0.742 |
| A generated image against its prompt | "Does the image show the element 'grass' as the prompt describes?" | 0.715 |
| General questions about a photo | MMBench-EN / POPE / MMStar | 0.881 / 0.889 / 0.627 |
| Bugs in game screenshots | glitch / object clipping (VideoGameQA-Bench) | 0.613 / 0.671 |
The text rows are public benchmarks; the image rows are judgment sets converted from labelled public data. Jev is more accurate on nimble (by 22 points), JevBench (10) and kev transfer-v4 (8); they tie on judgekit. Jev's API takes text only, so the image rows have no Jev comparison. Check Vev on a labelled sample of your own questions before relying on it. Significance tests and the comparison with the base model are in the GitHub README.
Limitations
- The probabilities rank answers well but are not exact frequencies: expected calibration error is 0.02–0.11 depending on the task. Choose any threshold on your own labelled data.
- Reversing the order of the options changes the top answer on 17% of JevBench and kev transfer-v4 questions (Jev: under 4%).
- Asked together with other questions, a question's probabilities move by up to a few hundredths (p99 0.025).
- Judgments that need several steps of reasoning are weaker than the base model's own answer when it may think first. On "is something wrong here?" questions the model answers "no" more often than the labels do.
- Vev is not affiliated with TypeSafe; it matches Jev's API, not its behaviour.
Load the weights directly
The adapter was trained on the inner model (full.model), and its files are in adapter/:
import os
from huggingface_hub import snapshot_download
from peft import PeftModel
from transformers import AutoModelForImageTextToText
full = AutoModelForImageTextToText.from_pretrained("Qwen/Qwen3.5-4B", dtype="bfloat16")
adapter = os.path.join(snapshot_download("CountingSheep/vev-4b-lora"), "adapter")
full.model = PeftModel.from_pretrained(full.model, adapter).merge_and_unload()
This gives the model only. Vev reads the next-token probabilities of the answer tokens (Yes/No, option letters, level digits) at the end of a fixed prompt, described in spec/systemone-api.md, section 10.
Training
LoRA rank 16 on the language model of Qwen/Qwen3.5-4B, vision tower frozen; 2,500 steps on 100,000 records from 42 public English and Chinese text and image datasets, with a KL term towards the base model on training rows and on the 10,123 unlabeled rows of CountingSheep/vev-anchor-v3. Recipe, library versions and data sources: TRAINING.md.
License
CC BY-NC 4.0, non-commercial use only, because some training data is licensed for research use (sources and terms
in TRAINING.md). The base model is Apache 2.0; its license is included as LICENSE-Qwen.
All Vev models and data: collection.
Citation
@software{vev2026,
title = {Vev: Jev-like decision models that can also see images},
author = {Wang, Xiaolong},
year = {2026},
url = {https://github.com/Xiaooolong/vev},
version = {0.1.1}
}
- Downloads last month
- -
Model tree for CountingSheep/vev-4b-lora
Dataset used to train CountingSheep/vev-4b-lora
Collection including CountingSheep/vev-4b-lora
Evaluation results
- accuracy on judgekit (Chinese)self-reported0.962
- accuracy on JevBench (public subset)self-reported0.766
- accuracy on nimbleself-reported0.707
- accuracy on kev transfer-v4self-reported0.776
- accuracy on MMBench-ENself-reported0.881
- accuracy on POPEself-reported0.889
- accuracy on MMStarself-reported0.627