Instructions to use CountingSheep/vev-9b with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use CountingSheep/vev-9b with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="CountingSheep/vev-9b") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("CountingSheep/vev-9b") model = AutoModelForMultimodalLM.from_pretrained("CountingSheep/vev-9b", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use CountingSheep/vev-9b with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "CountingSheep/vev-9b" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "CountingSheep/vev-9b", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/CountingSheep/vev-9b
- SGLang
How to use CountingSheep/vev-9b with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "CountingSheep/vev-9b" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "CountingSheep/vev-9b", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "CountingSheep/vev-9b" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "CountingSheep/vev-9b", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use CountingSheep/vev-9b with Docker Model Runner:
docker model run hf.co/CountingSheep/vev-9b
vev-9b
Jev-like decision models that can also see images. Ask yes/no, multiple-choice or graded questions about a
piece of text, a JSON record or a screenshot, and get a probability for every allowed answer in one forward pass.
Vev serves the same request format as TypeSafe's
Jev (/v1/systemone), and runs on your own GPU.
Code, API and full results: Xiaooolong/vev.
This repository holds the merged weights: Qwen/Qwen3.5-9B with the Vev adapter folded in. The adapter alone is in CountingSheep/vev-9b-lora. The other size is CountingSheep/vev-4b.
questions = {
"error": Noul(instructions="Does the screen show an error message?"),
"step": Choice(instructions="Which checkout step is the user on?",
criteria={"shipping": None, "payment": None, "review": None}),
"next": Choice(instructions="What should the user do next?",
criteria={"retry": "Try another card", "wait": "Wait for the order to ship",
"nothing": "Nothing, the order went through"}),
}
# vev-9b: error 0.984
# step {"shipping": 0.002, "payment": 0.974, "review": 0.024}
# next {"retry": 0.992, "wait": 0.001, "nothing": 0.006}
v0.1 is a research preview, released for non-commercial use under CC BY-NC 4.0.
Use it
pip install git+https://github.com/Xiaooolong/vev typesafe-sdk
vev serve --model CountingSheep/vev-9b # add --revision v0.1.0 to pin this release
from typesafe_sdk import Choice, Noul, TypeSafeClient
client = TypeSafeClient(api_key="local", base_url="http://127.0.0.1:8009")
resp = client.system_one(
state="Order #4411 still shows 'label created' after 9 days. I need it before Friday.",
questions={"urgent": Noul(instructions="Is the customer asking for something time-sensitive?")},
)
print(resp.answers["urgent"].noul)
Noul is the API's name for a yes/no question. Images go into the state as {"image": {"url": "data:image/png;base64,..."}}. Needs an NVIDIA GPU with about 19 GB of free memory.
What it can do
Accuracy on human-labelled data, by the kind of question:
| Kind of question | Example | vev-9b |
|---|---|---|
| Jev-style decisions on text | JevBench public subset | 0.823 |
| The same, on sources the model has not seen | kev transfer-v4 | 0.781 |
| Chinese | judgekit | 0.938 |
| Pairs where a small change flips the answer | nimble | 0.747 |
| UI state in an app screenshot | "Is there a switch or checkbox that is turned on?" / "Is there a text input field?" | 0.923 / 0.955 |
| An image against a written safety policy | the LlavaGuard policy categories | 0.724 |
| A generated image against its prompt | "Does the image show the element 'grass' as the prompt describes?" | 0.703 |
| General questions about a photo | MMBench-EN / POPE / MMStar | 0.902 / 0.897 / 0.674 |
| Bugs in game screenshots | glitch / object clipping (VideoGameQA-Bench) | 0.626 / 0.729 |
The text rows are public benchmarks; the image rows are judgment sets converted from labelled public data. Jev is more accurate on nimble (by 18 points) and kev transfer-v4 (7); JevBench and judgekit are not significantly different. Jev's API takes text only, so the image rows have no Jev comparison. Check Vev on a labelled sample of your own questions before relying on it. Significance tests and the comparison with the base model are in the GitHub README.
Limitations
- The probabilities rank answers well but are not exact frequencies: expected calibration error is 0.01–0.14 depending on the task. Choose any threshold on your own labelled data.
- Reversing the order of the options changes the top answer on 13% of JevBench and kev transfer-v4 questions (Jev: under 4%).
- Asked together with other questions, a question's probabilities move by up to a few hundredths (p99 0.025).
- Judgments that need several steps of reasoning are weaker than the base model's own answer when it may think first. On "is something wrong here?" questions the model answers "no" more often than the labels do.
- Vev is not affiliated with TypeSafe; it matches Jev's API, not its behaviour.
Load the weights directly
from transformers import AutoModelForImageTextToText, AutoProcessor
model = AutoModelForImageTextToText.from_pretrained("CountingSheep/vev-9b", dtype="bfloat16")
processor = AutoProcessor.from_pretrained("CountingSheep/vev-9b")
config.json keeps mtp_num_hidden_layers from the base model, but the multi-token-prediction weights are not
included; reading answers does not use them.
This gives the model only. Vev reads the next-token probabilities of the answer tokens (Yes/No, option letters, level digits) at the end of a fixed prompt, described in spec/systemone-api.md, section 10.
Training
LoRA rank 16 on the language model of Qwen/Qwen3.5-9B, vision tower frozen; 2,500 steps on 100,000 records from 42 public English and Chinese text and image datasets, with a KL term towards the base model on training rows and on the 10,123 unlabeled rows of CountingSheep/vev-anchor-v3. Recipe, library versions and data sources: TRAINING.md.
License
CC BY-NC 4.0, non-commercial use only, because some training data is licensed for research use (sources and terms
in TRAINING.md). The base model is Apache 2.0; its license is included as LICENSE-Qwen.
All Vev models and data: collection.
Citation
@software{vev2026,
title = {Vev: Jev-like decision models that can also see images},
author = {Wang, Xiaolong},
year = {2026},
url = {https://github.com/Xiaooolong/vev},
version = {0.1.1}
}
- Downloads last month
- 33
Model tree for CountingSheep/vev-9b
Dataset used to train CountingSheep/vev-9b
Collection including CountingSheep/vev-9b
Evaluation results
- accuracy on judgekit (Chinese)self-reported0.939
- accuracy on JevBench (public subset)self-reported0.823
- accuracy on nimbleself-reported0.747
- accuracy on kev transfer-v4self-reported0.781
- accuracy on MMBench-ENself-reported0.902
- accuracy on POPEself-reported0.897
- accuracy on MMStarself-reported0.674