joelniklaus HF Staff commited on
Commit
5f853b5
·
verified ·
1 Parent(s): ba3ab20

Update LEXam-hard evaluation result

Browse files

Adds this model's score on the [LEXam-hard](https://huggingface.co/datasets/joelniklaus/LEXam-hard) benchmark, the 518 LEXam open questions the strongest open models score lowest on.

The score is the DeepSeek-R1-0528 judge grade (0-100) over those questions, aggregated as SwissLegalEvals aggregates LEXam (mean of the German and English means), recomputed from the per-sample outputs of the [SwissLegalEvals](https://huggingface.co/blog/joelniklaus/swiss-legal-evals) run (lighteval, LEXam paper prompts, one response per question, no tools). The raw outputs are in the public `joelniklaus/SwissLegalEvals` bucket; the recomputation is `reproduction/lexam_hard_results.py` in the dataset repository.

Files changed (1) hide show
  1. .eval_results/lexam-hard.yaml +3 -3
.eval_results/lexam-hard.yaml CHANGED
@@ -1,12 +1,12 @@
1
  - dataset:
2
  id: joelniklaus/LEXam-hard
3
  task_id: lexam_hard
4
- revision: ed6d99edf964592f1a1d34f52ca9158ceda0f7af
5
- value: 36.14
6
  date: '2026-06-20'
7
  source:
8
  url: https://huggingface.co/buckets/joelniklaus/SwissLegalEvals
9
  name: SwissLegalEvals per-sample details (lighteval)
10
  user: joelniklaus
11
  notes: lighteval, LEXam paper prompts, one response per question, no tools; DeepSeek-R1-0528 judge;
12
- pooled mean over the 518 questions, 0-100
 
1
  - dataset:
2
  id: joelniklaus/LEXam-hard
3
  task_id: lexam_hard
4
+ revision: 1bd50ee3ac80286fed27e5b89f03f4d8b4494244
5
+ value: 37.07
6
  date: '2026-06-20'
7
  source:
8
  url: https://huggingface.co/buckets/joelniklaus/SwissLegalEvals
9
  name: SwissLegalEvals per-sample details (lighteval)
10
  user: joelniklaus
11
  notes: lighteval, LEXam paper prompts, one response per question, no tools; DeepSeek-R1-0528 judge;
12
+ mean of the German and English means over the 518 questions, 0-100