clmi / README.md
ucr-max's picture
Deploy prebuilt CompactLMIndex
7233de0 verified
|
Raw
History Blame Contribute Delete
2.53 kB
---
title: Compact Language Model Index
emoji: πŸ“Š
colorFrom: indigo
colorTo: blue
sdk: static
app_file: build/index.html
fullWidth: true
short_description: Chance-normalized evaluation of compact language models.
---
# Compact Language Model Index
CompactLMIndex is a curated benchmark for smaller language models. It uses one
fixed protocol to compare models across commonsense reasoning, science and
world knowledge, and grammatical competence.
## Evaluation suite
The published composite includes six benchmarks:
| Area | Benchmark |
| --- | --- |
| Situational commonsense | HellaSwag |
| Physical commonsense | PIQA |
| Science and world knowledge | ARC-Easy |
| Science and world knowledge | ARC-Challenge |
| Conceptual commonsense | CommonsenseQA |
| Grammatical competence | BLiMP |
All models use zero-shot, non-chat log-likelihood prompts. A complete run of
the core suite is required for a model to receive a rank.
## Scoring
Each benchmark score is normalized for random chance:
```text
100 Γ— (accuracy βˆ’ chance) / (1 βˆ’ chance)
```
Random performance maps to 0 and perfect performance maps to 100. Scores
below chance remain negative.
The composite is domain-balanced:
- commonsense: 50%, split equally between HellaSwag, PIQA, and CommonsenseQA;
- science and world knowledge: 25%, with ARC-Easy contributing 75% and
ARC-Challenge 25% within that domain;
- grammatical competence: 25%, from BLiMP.
BLiMP is macro-averaged across its 12 linguistic categories. CommonsenseQA
scores the full answer text with length normalization rather than answer labels
such as A–E.
## Uncertainty and ranking
Every displayed benchmark and composite interval is a 95% percentile bootstrap
interval. Model comparisons use paired bootstrap resampling on matched
evaluation items. Rank ranges show when neighbouring models cannot be reliably
distinguished under those comparisons.
## Scope and inclusion
CompactLMIndex is curated rather than an open intake list. Published results
are checked against their saved evaluation artifacts, and final inclusion and
presentation decisions remain with Universal Computing Research.
If you have results for a model that fits the benchmark scope, open a
discussion on this Space for consideration.
## Credits
Inspired by the [Open LLM Leaderboard](https://huggingface.co/open-llm-leaderboard),
[Axiomic Labs Open SLM Leaderboard](https://huggingface.co/spaces/AxiomicLabs/Open_SLM_Leaderboard),
and [OpenAI Parameter Golf](https://github.com/openai/parameter-golf).