EvalExplorer classifier
Small models that replace a big-LLM classifier
Small models that classify an evaluation report's first pages into approach, type, temporality, themes and countries. Data, adapters, run reports.
Small models that replace a big-LLM classifier
Note Leaderboards for the Baobab Tech fine-tuning experiments
Note One report per run, plus the leaderboard
Note Source export of the EvalExplorer database
Note SFT - mean field score 0.847
Note SFT - mean field score 0.844
Note SFT - mean field score 0.842
Note SFT - mean field score 0.830
Note SFT - mean field score 0.815
Note SFT - mean field score 0.802
Note SFT - mean field score 0.792
Note GLINER-FINETUNE - mean field score 0.578
Note GLINER-FINETUNE - mean field score 0.573
Note Experimental adapters (GRPO variants), one subfolder each
Note GGUF exports for llama.cpp: Q8_0, Q5_K_M, Q4_K_M per adapter