ltuncay commited on
Commit
ebea1a6
·
verified ·
1 Parent(s): 5d2f1f6

Add Audio-JEPA results and link the Transformers release

Browse files
Files changed (2) hide show
  1. README.md +3 -2
  2. export_manifest.json +1 -1
README.md CHANGED
@@ -23,7 +23,7 @@ repository is needed.
23
  ## X-ARES results
24
 
25
  Scores on [X-ARES](https://arxiv.org/abs/2505.16369) (0–100, higher is better).
26
- BEST-RQ (Conformer), BEST-RQ (ViT),
27
  and all BEST-RQ-2 variants reported below are trained on the **same AudioSet
28
  split for 200,000 steps**. The pretrained baselines are shown for comparison.
29
 
@@ -32,6 +32,7 @@ split for 200,000 steps**. The pretrained baselines are shown for comparison.
32
  | data2vec | 50.62 | 23.24 | 15.41 | 37.83 | 29.76 | [facebook/data2vec-audio-base](https://huggingface.co/facebook/data2vec-audio-base) |
33
  | wav2vec 2.0 | 41.79 | 34.94 | 29.52 | 37.84 | 35.42 | [facebook/wav2vec2-large-100k-voxpopuli](https://huggingface.co/facebook/wav2vec2-large-100k-voxpopuli) |
34
  | Whisper | 49.19 | 38.67 | 28.61 | 42.75 | 38.82 | [openai/whisper-base](https://huggingface.co/openai/whisper-base) |
 
35
  | BEST-RQ (Conformer) | 40.43 | 35.58 | 30.81 | 37.43 | 35.60 | Separate codebase |
36
  | BEST-RQ (ViT) | 32.87 | 41.62 | 34.50 | 34.88 | 36.33 | [BEST-RQ-ViT](https://huggingface.co/ltuncay/BEST-RQ-ViT) |
37
  | BEST-RQ-2 (Interspeech 2026) | 38.49 | 54.40 | 46.39 | 43.21 | 46.43 | [BEST-RQ-2](https://huggingface.co/ltuncay/BEST-RQ-2) |
@@ -41,7 +42,7 @@ split for 200,000 steps**. The pretrained baselines are shown for comparison.
41
  Global Mean averages all benchmark task scores. Mean of Means gives equal
42
  weight to the Speech, Music, and Environment category means.
43
 
44
- These are the research results reported in the [project README](https://github.com/LudovicTuncay/audio-embeddings/blob/bb88bf790b1dcf8251c6b38e7a4766534adf33d3/README.md), not a new benchmark run of the Transformers exports.
45
 
46
  ## Model and training
47
 
 
23
  ## X-ARES results
24
 
25
  Scores on [X-ARES](https://arxiv.org/abs/2505.16369) (0–100, higher is better).
26
+ Audio-JEPA, BEST-RQ (Conformer), BEST-RQ (ViT),
27
  and all BEST-RQ-2 variants reported below are trained on the **same AudioSet
28
  split for 200,000 steps**. The pretrained baselines are shown for comparison.
29
 
 
32
  | data2vec | 50.62 | 23.24 | 15.41 | 37.83 | 29.76 | [facebook/data2vec-audio-base](https://huggingface.co/facebook/data2vec-audio-base) |
33
  | wav2vec 2.0 | 41.79 | 34.94 | 29.52 | 37.84 | 35.42 | [facebook/wav2vec2-large-100k-voxpopuli](https://huggingface.co/facebook/wav2vec2-large-100k-voxpopuli) |
34
  | Whisper | 49.19 | 38.67 | 28.61 | 42.75 | 38.82 | [openai/whisper-base](https://huggingface.co/openai/whisper-base) |
35
+ | Audio-JEPA | 29.64 | 44.27 | 25.61 | 31.18 | 33.17 | [Audio-JEPA-base](https://huggingface.co/ltuncay/Audio-JEPA-base) |
36
  | BEST-RQ (Conformer) | 40.43 | 35.58 | 30.81 | 37.43 | 35.60 | Separate codebase |
37
  | BEST-RQ (ViT) | 32.87 | 41.62 | 34.50 | 34.88 | 36.33 | [BEST-RQ-ViT](https://huggingface.co/ltuncay/BEST-RQ-ViT) |
38
  | BEST-RQ-2 (Interspeech 2026) | 38.49 | 54.40 | 46.39 | 43.21 | 46.43 | [BEST-RQ-2](https://huggingface.co/ltuncay/BEST-RQ-2) |
 
42
  Global Mean averages all benchmark task scores. Mean of Means gives equal
43
  weight to the Speech, Music, and Environment category means.
44
 
45
+ Audio-JEPA scores were supplied by the author for [run `jp6l70l6`](https://wandb.ai/tuncay-ludovic/audio%20embeddings/runs/jp6l70l6). The remaining scores are reported in the [project README](https://github.com/LudovicTuncay/audio-embeddings/blob/bb88bf790b1dcf8251c6b38e7a4766534adf33d3/README.md). These are reported research results, not a new benchmark run of the Transformers exports.
46
 
47
  ## Model and training
48
 
export_manifest.json CHANGED
@@ -27,7 +27,7 @@
27
  },
28
  "files": {
29
  "CODE_LICENSE": "8fe9e8b749cd4abedabcb3100df445db899e72192394d14cc2dbf24a40811af6",
30
- "README.md": "5d0f54e06c339d95d064cd96c9ad11e87430c36e7a5daebfc5aa3ca5920440a9",
31
  "adapters.py": "ca0f26826763c0e48b7508da732242bb83adfeb4814e389ed3fde93ad648c728",
32
  "config.json": "93452b1dfb678bafd2737b29cc80586e1ccd8368ad0183ed3b88ffd8f12ed6f2",
33
  "configuration_audio.py": "59f3a0b8db0df5af85e677ac33a4431595e9b2eff6d34c1e6a1778dd75a4568d",
 
27
  },
28
  "files": {
29
  "CODE_LICENSE": "8fe9e8b749cd4abedabcb3100df445db899e72192394d14cc2dbf24a40811af6",
30
+ "README.md": "95ba33b3f1337e057897ec1662558de779867cd477ff22e26ac1200662f9c2c9",
31
  "adapters.py": "ca0f26826763c0e48b7508da732242bb83adfeb4814e389ed3fde93ad648c728",
32
  "config.json": "93452b1dfb678bafd2737b29cc80586e1ccd8368ad0183ed3b88ffd8f12ed6f2",
33
  "configuration_audio.py": "59f3a0b8db0df5af85e677ac33a4431595e9b2eff6d34c1e6a1778dd75a4568d",