Instructions to use baobabtech/evalexplorer-classify-gemma-4-e2b-sft with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use baobabtech/evalexplorer-classify-gemma-4-e2b-sft with PEFT:
from peft import PeftModel from transformers import AutoModelForCausalLM base_model = AutoModelForCausalLM.from_pretrained("unsloth/gemma-4-E2B-it") model = PeftModel.from_pretrained(base_model, "baobabtech/evalexplorer-classify-gemma-4-e2b-sft") - Notebooks
- Google Colab
- Kaggle
evalexplorer-classify-gemma-4-e2b-sft
Classifies an international development evaluation report from its first pages into a JSON object with evaluation_approach, evaluation_type, temporality, themes and countries. One of the runs documented in baobabtech/evalexplorer-classify-experiments; code in code/.
| Base model | unsloth/gemma-4-E2B-it |
| Method | SFT |
| Data | baobabtech/evalexplorer-data, config classify_codes, 1148 training documents |
| Hardware | gpu |
| Training time | 33.1 min |
| Run report | gemma-4-e2b-sft--test |
Training
LoRA SFT: r=16, alpha=16, all linear language layers, bf16, learning rate 0.0002, 2 epochs, batch 2 × 8 accumulation, loss on the answer only.
Scores
| Metric | Test (134 documents) |
|---|---|
| JSON valid | 1.000 |
| Exact match (all five fields) | 0.246 |
| Mean field score | 0.815 |
| evaluation_approach accuracy | 0.776 |
| evaluation_type accuracy | 0.851 |
| temporality accuracy | 0.731 |
| themes micro F1 | 0.822 |
| countries micro F1 | 0.862 |
Greedy decoding, thinking off. mean_field_score is the per-document mean of the five field scores (1/0 for the single-code fields, F1 for the lists). The model was trained on, and is scored here against, the EvalExplorer ingestion pipeline's LLM labels, which no person has reviewed, so these numbers measure agreement with that pipeline, not correctness. The run report also scores the same answers against an independent GLM-5.3-Flash relabelling the model never saw. This is an exploratory model; the intended next version is trained on the GLM labels.
Usage
from unsloth import FastLanguageModel, FastModel
loader = FastLanguageModel if "lfm" in "unsloth/gemma-4-E2B-it".lower() else FastModel
model, processor = loader.from_pretrained("unsloth/gemma-4-E2B-it", max_seq_length=8192, load_in_4bit=False)
from peft import PeftModel
model = PeftModel.from_pretrained(model, "baobabtech/evalexplorer-classify-gemma-4-e2b-sft")
loader.for_inference(model)
text = processor.apply_chat_template(prompt_messages, tokenize=False, add_generation_prompt=True,
enable_thinking=False) # prompt_messages from the dataset's `prompt` column
- Downloads last month
- 10