AI & ML interests

Contact: arxivgpt@gmail.com

Recent Activity

SeaWolf-AI  updated a collection about 8 hours ago
ZTC Models: JEV ecosystems
SeaWolf-AI  updated a collection about 8 hours ago
DARWIN-Family
View all activity

Articles

View all articles

SeaWolf-AI 
published an article about 2 hours ago
view article
Article

Leading the System One Mosaic Benchmark: What Darwin-27B-ZTC-v2's #1 Means

FINAL-Bench
•
• 6
SeaWolf-AI 
posted an update about 5 hours ago
view post
Post
730
🏆 Darwin-27B-ZTC-v2 just took #1 on the System One Mosaic Benchmark (S1MB).

S1MB compares 102 models across 137 specialized benchmarks, in three task types: Noul (assess a condition), Choice (select an option), Score (rate on a scale). Ranking is by overall Borda score.

📊 Top of the board
🥇 Darwin ZTC v2 (FINAL-Bench) 89.58
🥈 OpenJev-27B 87.50
🥉 AutoJev-27B 87.07
4️⃣ Eikos 27B 85.43
5️⃣ Jev 1.13 85.05

🔎 Ranks 2 to 5 are all the JEV family (TypeSafe AI's System One model, from ex-OpenAI researchers). S1MB exists to compare these System One judges, so leading it is the headline.

⚙️ Why a zero-token judge wins here
🔹 It does not generate. It reads the input and typed questions and returns a calibrated distribution in a single forward pass.
🔹 Zero generated tokens, no decoding loop, so latency and cost stay low.
🔹 Holds up out of distribution too: General Noul 96.00, General Choice 99.34.

It is also #1 on the typed-decisions leaderboard (0.743, zero-shot). Same message from both: a deterministic, calibrated judge at one forward pass per call.

🔗 Model: FINAL-Bench/Darwin-27B-ZTC-v2
🔗 Leaderboard: hotchpotch/S1MB-leaderboard

Standings move as new models are added. Numbers reflect the board at the time of writing. 🙌
SeaWolf-AI 
published an article 1 day ago
view article
Article

Darwin-27B-ZTC: A Single-Pass Judge and a Quantitative Look at Its Calibration

FINAL-Bench
•
• 12