ltuncay commited on
Commit
17ca7cd
·
verified ·
1 Parent(s): 7a8c63b

Add Audio-JEPA results and link the Transformers release

Browse files
Files changed (2) hide show
  1. README.md +3 -2
  2. export_manifest.json +1 -1
README.md CHANGED
@@ -25,7 +25,7 @@ repository is needed.
25
  ## X-ARES results
26
 
27
  Scores on [X-ARES](https://arxiv.org/abs/2505.16369) (0–100, higher is better).
28
- BEST-RQ (Conformer), BEST-RQ (ViT),
29
  and all BEST-RQ-2 variants reported below are trained on the **same AudioSet
30
  split for 200,000 steps**. The pretrained baselines are shown for comparison.
31
 
@@ -34,6 +34,7 @@ split for 200,000 steps**. The pretrained baselines are shown for comparison.
34
  | data2vec | 50.62 | 23.24 | 15.41 | 37.83 | 29.76 | [facebook/data2vec-audio-base](https://huggingface.co/facebook/data2vec-audio-base) |
35
  | wav2vec 2.0 | 41.79 | 34.94 | 29.52 | 37.84 | 35.42 | [facebook/wav2vec2-large-100k-voxpopuli](https://huggingface.co/facebook/wav2vec2-large-100k-voxpopuli) |
36
  | Whisper | 49.19 | 38.67 | 28.61 | 42.75 | 38.82 | [openai/whisper-base](https://huggingface.co/openai/whisper-base) |
 
37
  | BEST-RQ (Conformer) | 40.43 | 35.58 | 30.81 | 37.43 | 35.60 | Separate codebase |
38
  | BEST-RQ (ViT) | 32.87 | 41.62 | 34.50 | 34.88 | 36.33 | [BEST-RQ-ViT](https://huggingface.co/ltuncay/BEST-RQ-ViT) |
39
  | BEST-RQ-2 (Interspeech 2026) | 38.49 | 54.40 | 46.39 | 43.21 | 46.43 | [BEST-RQ-2](https://huggingface.co/ltuncay/BEST-RQ-2) |
@@ -43,7 +44,7 @@ split for 200,000 steps**. The pretrained baselines are shown for comparison.
43
  Global Mean averages all benchmark task scores. Mean of Means gives equal
44
  weight to the Speech, Music, and Environment category means.
45
 
46
- These are the research results reported in the [project README](https://github.com/LudovicTuncay/audio-embeddings/blob/bb88bf790b1dcf8251c6b38e7a4766534adf33d3/README.md), not a new benchmark run of the Transformers exports.
47
 
48
  ## Model and training
49
 
 
25
  ## X-ARES results
26
 
27
  Scores on [X-ARES](https://arxiv.org/abs/2505.16369) (0–100, higher is better).
28
+ Audio-JEPA, BEST-RQ (Conformer), BEST-RQ (ViT),
29
  and all BEST-RQ-2 variants reported below are trained on the **same AudioSet
30
  split for 200,000 steps**. The pretrained baselines are shown for comparison.
31
 
 
34
  | data2vec | 50.62 | 23.24 | 15.41 | 37.83 | 29.76 | [facebook/data2vec-audio-base](https://huggingface.co/facebook/data2vec-audio-base) |
35
  | wav2vec 2.0 | 41.79 | 34.94 | 29.52 | 37.84 | 35.42 | [facebook/wav2vec2-large-100k-voxpopuli](https://huggingface.co/facebook/wav2vec2-large-100k-voxpopuli) |
36
  | Whisper | 49.19 | 38.67 | 28.61 | 42.75 | 38.82 | [openai/whisper-base](https://huggingface.co/openai/whisper-base) |
37
+ | Audio-JEPA | 29.64 | 44.27 | 25.61 | 31.18 | 33.17 | [Audio-JEPA-base](https://huggingface.co/ltuncay/Audio-JEPA-base) |
38
  | BEST-RQ (Conformer) | 40.43 | 35.58 | 30.81 | 37.43 | 35.60 | Separate codebase |
39
  | BEST-RQ (ViT) | 32.87 | 41.62 | 34.50 | 34.88 | 36.33 | [BEST-RQ-ViT](https://huggingface.co/ltuncay/BEST-RQ-ViT) |
40
  | BEST-RQ-2 (Interspeech 2026) | 38.49 | 54.40 | 46.39 | 43.21 | 46.43 | [BEST-RQ-2](https://huggingface.co/ltuncay/BEST-RQ-2) |
 
44
  Global Mean averages all benchmark task scores. Mean of Means gives equal
45
  weight to the Speech, Music, and Environment category means.
46
 
47
+ Audio-JEPA scores were supplied by the author for [run `jp6l70l6`](https://wandb.ai/tuncay-ludovic/audio%20embeddings/runs/jp6l70l6). The remaining scores are reported in the [project README](https://github.com/LudovicTuncay/audio-embeddings/blob/bb88bf790b1dcf8251c6b38e7a4766534adf33d3/README.md). These are reported research results, not a new benchmark run of the Transformers exports.
48
 
49
  ## Model and training
50
 
export_manifest.json CHANGED
@@ -38,7 +38,7 @@
38
  },
39
  "files": {
40
  "CODE_LICENSE": "8fe9e8b749cd4abedabcb3100df445db899e72192394d14cc2dbf24a40811af6",
41
- "README.md": "96bc9111eda43b9ab7512624a5d0baf6f1e95949d5d4efcff0457bc916b4fc04",
42
  "adapters.py": "ca0f26826763c0e48b7508da732242bb83adfeb4814e389ed3fde93ad648c728",
43
  "config.json": "d37b85c37f17aa954f715b3b3d956c9ed18215ec6ea9565bcd4a403e9b51a134",
44
  "configuration_audio.py": "59f3a0b8db0df5af85e677ac33a4431595e9b2eff6d34c1e6a1778dd75a4568d",
 
38
  },
39
  "files": {
40
  "CODE_LICENSE": "8fe9e8b749cd4abedabcb3100df445db899e72192394d14cc2dbf24a40811af6",
41
+ "README.md": "e90a591ff29cacd941a00c678ecd9088796341cad39329544c64a7766ddfd67c",
42
  "adapters.py": "ca0f26826763c0e48b7508da732242bb83adfeb4814e389ed3fde93ad648c728",
43
  "config.json": "d37b85c37f17aa954f715b3b3d956c9ed18215ec6ea9565bcd4a403e9b51a134",
44
  "configuration_audio.py": "59f3a0b8db0df5af85e677ac33a4431595e9b2eff6d34c1e6a1778dd75a4568d",