·
AI & ML interests
Contact: arxivgpt@gmail.com
Recent Activity
repliedto their post about 5 hours ago We opened a benchmark for drug property prediction tools. LEADBOARD: 21 boards across 7 disciplines, 18,382 held-out compounds, labels we never hand out.
Two numbers we hit while building it are the reason it exists.
First. Split the hERG cardiotoxicity data at random and you get AUROC 0.818. Split it by first-report year instead and you get 0.606. Same molecules, same fingerprints, same learner, same hyperparameters. The only thing that changed was where the line went, and the score moved 0.211. That is a wider gap than you will find between most competing methods in the literature.
Second. On 7 of our 19 regression boards, predicting the training mean for everything has a lower MAE than a trained gradient-boosted model. hERG is one of them, 0.599 against 0.589. The trained model loses.
So every board publishes its homework before anyone submits. Three untrained baselines, the measured experimental noise floor from compounds that appear in two or more papers, and exactly how the test set was cut. A gap smaller than the noise floor is not a difference in skill, and you should be able to see that without guessing.
Entering is simple. Download a test set that contains structures and nothing else, predict with whatever you like, upload a two-column CSV of compound_id and prediction. Trained model, physics engine, LLM, rule of thumb. We do not care what is inside. We measure the output.
Post: https://huggingface.co/blog/FINAL-Bench/leadboard-drug
Leaderboard: https://huggingface.co/spaces/FINAL-Bench/leadboard repliedto their post about 8 hours ago We opened a benchmark for drug property prediction tools. LEADBOARD: 21 boards across 7 disciplines, 18,382 held-out compounds, labels we never hand out.
Two numbers we hit while building it are the reason it exists.
First. Split the hERG cardiotoxicity data at random and you get AUROC 0.818. Split it by first-report year instead and you get 0.606. Same molecules, same fingerprints, same learner, same hyperparameters. The only thing that changed was where the line went, and the score moved 0.211. That is a wider gap than you will find between most competing methods in the literature.
Second. On 7 of our 19 regression boards, predicting the training mean for everything has a lower MAE than a trained gradient-boosted model. hERG is one of them, 0.599 against 0.589. The trained model loses.
So every board publishes its homework before anyone submits. Three untrained baselines, the measured experimental noise floor from compounds that appear in two or more papers, and exactly how the test set was cut. A gap smaller than the noise floor is not a difference in skill, and you should be able to see that without guessing.
Entering is simple. Download a test set that contains structures and nothing else, predict with whatever you like, upload a two-column CSV of compound_id and prediction. Trained model, physics engine, LLM, rule of thumb. We do not care what is inside. We measure the output.
Post: https://huggingface.co/blog/FINAL-Bench/leadboard-drug
Leaderboard: https://huggingface.co/spaces/FINAL-Bench/leadboard reacted to theirpost with ❤️ about 16 hours ago We opened a benchmark for drug property prediction tools. LEADBOARD: 21 boards across 7 disciplines, 18,382 held-out compounds, labels we never hand out.
Two numbers we hit while building it are the reason it exists.
First. Split the hERG cardiotoxicity data at random and you get AUROC 0.818. Split it by first-report year instead and you get 0.606. Same molecules, same fingerprints, same learner, same hyperparameters. The only thing that changed was where the line went, and the score moved 0.211. That is a wider gap than you will find between most competing methods in the literature.
Second. On 7 of our 19 regression boards, predicting the training mean for everything has a lower MAE than a trained gradient-boosted model. hERG is one of them, 0.599 against 0.589. The trained model loses.
So every board publishes its homework before anyone submits. Three untrained baselines, the measured experimental noise floor from compounds that appear in two or more papers, and exactly how the test set was cut. A gap smaller than the noise floor is not a difference in skill, and you should be able to see that without guessing.
Entering is simple. Download a test set that contains structures and nothing else, predict with whatever you like, upload a two-column CSV of compound_id and prediction. Trained model, physics engine, LLM, rule of thumb. We do not care what is inside. We measure the output.
Post: https://huggingface.co/blog/FINAL-Bench/leadboard-drug
Leaderboard: https://huggingface.co/spaces/FINAL-Bench/leadboard View all activity Organizations
published an article about 16 hours ago view article We changed one line and the benchmark score moved 0.21 AUROC
FINAL-Bench
• • 9
view article Who Tells You Whether the Molecule Your AI Just Designed Is Any Good?
FINAL-Bench
• • 11
view article AX-Ray, Finding Causal-Leakage Defects in Two General-Purpose Public Models
FINAL-Bench
• • 13
view article The Fast Gemma Challenge: our verified-SOTA recipe, in full
FINAL-Bench
• • 24
published an article about 1 month ago view article POCKET: a 35-billion-parameter model that runs on your iPhone — and on your PC with no GPU
FINAL-Bench
• • 12
published an article about 1 month ago view article Aether-7B-5Attn: A 100% Open-Source Sovereign Foundation Model — and a Controlled Experiment in Heterogeneous Attention
FINAL-Bench
• • 21
published an article about 1 month ago view article VKUE: No GPU? Runs Anyway — a 34.7B Reasoner on a Laptop and on Bare CPU
FINAL-Bench
• • 17
published an article about 2 months ago view article Quantum Cryptanalysis on Real Hardware: Pushing Symmetric-Structure Key Recovery Beyond the Published Frontier
published an article about 2 months ago published an article about 2 months ago view article Chitos: From Detection to Proof — An Autonomous Security AI That Actually Exploits
FINAL-Bench
• • 19
view article FINAL-Bench Quantum: An Open, Neutral Benchmark for Quantum-Computing Methods
FINAL-Bench
• • 17
view article Training-Free Reasoning at 88.89% on GPQA Diamond: How Darwin Family Hit Frontier Scores Without a Single Gradient Step
FINAL-Bench
• • 18
view article Darwin-TTS: We Gave a TTS Model 3% of an LLM's Brain — It Started Showing Emotion
FINAL-Bench
• • 13
view article "Darwin-27B-Opus: Surpassing the Foundation Model Without Training"
FINAL-Bench
• • 16
view article Darwin V6: Diagnostic-Guided Evolutionary Model Merging
view article "The Child That Surpassed Both Parents Through MRI-Guided Evolutionary Merge"
FINAL-Bench
• • 15
view article Introducing WM Bench: A Benchmark for Cognitive Intelligence in World Models
FINAL-Bench
• • 13
view article 🏟️ Smol AI WorldCup: A 5-Axis Benchmark That Reveals What Small Language Models Can Really Do
FINAL-Bench
• • 38
view article MARL: Runtime Middleware That Reduces LLM Hallucination Without Fine-Tuning
view article Structural Problems in AI Benchmarking and the Case for a Unified Evaluation Framework