Feature Extraction
Transformers
Safetensors
audio_embeddings
audio
custom_code
self-supervised-learning
audio-embeddings
best-rq-2
audioset
Instructions to use ltuncay/BEST-RQ-2.1-base with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use ltuncay/BEST-RQ-2.1-base with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("feature-extraction", model="ltuncay/BEST-RQ-2.1-base", trust_remote_code=True)# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("ltuncay/BEST-RQ-2.1-base", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
Recommend BEST-RQ-2.2 and add the project X-ARES comparison table
Browse files- README.md +25 -0
- export_manifest.json +1 -1
README.md
CHANGED
|
@@ -12,6 +12,8 @@ tags:
|
|
| 12 |
---
|
| 13 |
# BEST-RQ-2.1-base
|
| 14 |
|
|
|
|
|
|
|
| 15 |
BEST-RQ-2.1-base is a self-supervised audio encoder trained on AudioSet for
|
| 16 |
a configured budget of **200,000 optimizer steps**. It produces **768-dimensional** clip and frame
|
| 17 |
embeddings from mono **16 kHz** waveforms, and supports downstream fine-tuning.
|
|
@@ -20,6 +22,29 @@ This repository contains the trained encoder, its preprocessing configuration,
|
|
| 20 |
and the custom Transformers implementation. No installation of the research
|
| 21 |
repository is needed.
|
| 22 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 23 |
## Model and training
|
| 24 |
|
| 25 |
| Property | Value |
|
|
|
|
| 12 |
---
|
| 13 |
# BEST-RQ-2.1-base
|
| 14 |
|
| 15 |
+
> **Recommended: [BEST-RQ-2.2-base](https://huggingface.co/ltuncay/BEST-RQ-2.2-base)** has the strongest reported X-ARES results in the BEST-RQ-2 family. For new projects, start with that model; see the comparison below.
|
| 16 |
+
|
| 17 |
BEST-RQ-2.1-base is a self-supervised audio encoder trained on AudioSet for
|
| 18 |
a configured budget of **200,000 optimizer steps**. It produces **768-dimensional** clip and frame
|
| 19 |
embeddings from mono **16 kHz** waveforms, and supports downstream fine-tuning.
|
|
|
|
| 22 |
and the custom Transformers implementation. No installation of the research
|
| 23 |
repository is needed.
|
| 24 |
|
| 25 |
+
## X-ARES results
|
| 26 |
+
|
| 27 |
+
Scores on [X-ARES](https://arxiv.org/abs/2505.16369) (0–100, higher is better).
|
| 28 |
+
BEST-RQ (Conformer), BEST-RQ (ViT),
|
| 29 |
+
and all BEST-RQ-2 variants reported below are trained on the **same AudioSet
|
| 30 |
+
split for 200,000 steps**. The pretrained baselines are shown for comparison.
|
| 31 |
+
|
| 32 |
+
| Model | Speech | Music | Environment | Global Mean | Mean of Means | Training recipe |
|
| 33 |
+
| --- | ---: | ---: | ---: | ---: | ---: | --- |
|
| 34 |
+
| data2vec | 50.62 | 23.24 | 15.41 | 37.83 | 29.76 | Pretrained baseline |
|
| 35 |
+
| wav2vec 2.0 | 41.79 | 34.94 | 29.52 | 37.84 | 35.42 | Pretrained baseline |
|
| 36 |
+
| Whisper | 49.19 | 38.67 | 28.61 | 42.75 | 38.82 | Pretrained baseline |
|
| 37 |
+
| BEST-RQ (Conformer) | 40.43 | 35.58 | 30.81 | 37.43 | 35.60 | Separate codebase (recipe unavailable) |
|
| 38 |
+
| BEST-RQ (ViT) | 32.87 | 41.62 | 34.50 | 34.88 | 36.33 | [best_rq/audioset/default.yaml](https://github.com/LudovicTuncay/audio-embeddings/blob/bb88bf790b1dcf8251c6b38e7a4766534adf33d3/configs/experiment/best_rq/audioset/default.yaml) |
|
| 39 |
+
| BEST-RQ-2 (Interspeech 2026) | 38.49 | 54.40 | 46.39 | 43.21 | 46.43 | [best_rq_2/default.yaml](https://github.com/LudovicTuncay/audio-embeddings/blob/bb88bf790b1dcf8251c6b38e7a4766534adf33d3/configs/experiment/best_rq_2/default.yaml) |
|
| 40 |
+
| BEST-RQ-2.1 | 52.60 | 62.23 | 53.38 | 54.59 | 56.07 | [best_rq_2_1/masking/80.yaml](https://github.com/LudovicTuncay/audio-embeddings/blob/bb88bf790b1dcf8251c6b38e7a4766534adf33d3/configs/experiment/best_rq_2_1/masking/80.yaml) |
|
| 41 |
+
| BEST-RQ-2.2 | **53.78** | **63.90** | **55.71** | **56.11** | **57.80** | [best_rq_2_2/masking/80.yaml](https://github.com/LudovicTuncay/audio-embeddings/blob/bb88bf790b1dcf8251c6b38e7a4766534adf33d3/configs/experiment/best_rq_2_2/masking/80.yaml) |
|
| 42 |
+
|
| 43 |
+
Global Mean averages all benchmark task scores. Mean of Means gives equal
|
| 44 |
+
weight to the Speech, Music, and Environment category means.
|
| 45 |
+
|
| 46 |
+
These are the research results reported in the [project README](https://github.com/LudovicTuncay/audio-embeddings/blob/bb88bf790b1dcf8251c6b38e7a4766534adf33d3/README.md), not a new benchmark run of the Transformers exports.
|
| 47 |
+
|
| 48 |
## Model and training
|
| 49 |
|
| 50 |
| Property | Value |
|
export_manifest.json
CHANGED
|
@@ -31,7 +31,7 @@
|
|
| 31 |
},
|
| 32 |
"files": {
|
| 33 |
"CODE_LICENSE": "8fe9e8b749cd4abedabcb3100df445db899e72192394d14cc2dbf24a40811af6",
|
| 34 |
-
"README.md": "
|
| 35 |
"adapters.py": "ca0f26826763c0e48b7508da732242bb83adfeb4814e389ed3fde93ad648c728",
|
| 36 |
"config.json": "88be4a7618663564525588db633cf7461f051e437913187070b854ae8367ac07",
|
| 37 |
"configuration_audio.py": "59f3a0b8db0df5af85e677ac33a4431595e9b2eff6d34c1e6a1778dd75a4568d",
|
|
|
|
| 31 |
},
|
| 32 |
"files": {
|
| 33 |
"CODE_LICENSE": "8fe9e8b749cd4abedabcb3100df445db899e72192394d14cc2dbf24a40811af6",
|
| 34 |
+
"README.md": "194db8166e421ac7fa8709099f16fa7a9ac476d21d01dd03527aca66ca3a78fc",
|
| 35 |
"adapters.py": "ca0f26826763c0e48b7508da732242bb83adfeb4814e389ed3fde93ad648c728",
|
| 36 |
"config.json": "88be4a7618663564525588db633cf7461f051e437913187070b854ae8367ac07",
|
| 37 |
"configuration_audio.py": "59f3a0b8db0df5af85e677ac33a4431595e9b2eff6d34c1e6a1778dd75a4568d",
|