Text Generation
Safetensors
English
Chinese
qwen3
reward-model
rlhf
principle-following
qwen
conversational
Instructions to use WisdomShell/RewardAnything-8B-v1 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Inference
Update README.md
Browse files
README.md
CHANGED
|
@@ -312,11 +312,7 @@ result = reward_model.judge(
|
|
| 312 |
|
| 313 |
## 📈 Performance & Benchmarks
|
| 314 |
|
| 315 |
-
|
| 316 |
-
|
| 317 |
-
- **RM-Bench**: 92.3% accuracy (vs 87.1% for best baseline)
|
| 318 |
-
- **RABench**: 89.7% principle-following accuracy
|
| 319 |
-
- **HH-RLHF**: 94.2% alignment with human preferences
|
| 320 |
|
| 321 |
## 📚 Documentation
|
| 322 |
|
|
|
|
| 312 |
|
| 313 |
## 📈 Performance & Benchmarks
|
| 314 |
|
| 315 |
+
Please refer to our paper for performance metrics and comparison.
|
|
|
|
|
|
|
|
|
|
|
|
|
| 316 |
|
| 317 |
## 📚 Documentation
|
| 318 |
|