๐ป Data-center AI, now on a laptop: POCKET-Darwin-180B
We're releasing a 4-bit GGUF build of Darwin-180B-RSI, #1 on seven official Hugging Face leaderboards (self-reported), that runs without a GPU.
๐ฆ 360 GB โ 111 GB (4-bit GGUF, 4 files) ๐ฅ๏ธ No GPU: one server CPU (16 threads) at 18.4โ21.0 tokens/s ๐ป RTX 5060 laptop (8 GB VRAM) + 32 GB RAM: 4.17 tokens/s ๐ง 128 GB mini PC: whole model in memory, no GPU needed ๐ฏ MMLU-Pro, 2,000 questions, paired: original 87.65% = 4-bit 87.65%
How? ยท Only ~3B of 180B parameters are active per token (10 of 512 experts) ยท llama.cpp streams just the needed experts from SSD, so 32 GB RAM is enough ยท Graft quantization: we took the proven Unsloth UD-Q4_K_XL base build and swapped in only the 300 tensors our RSI training changed (300/300 verified)
Under the hood is Model-level Recursive Self-Improvement. The model solves verifiable problems, keeps only its own solutions that check out as correct, and trains on them. No human-written solutions or reasoning traces.
Built for teams that can't send data to an external cloud (defense, finance, public sector) to run a top-tier model fully offline.
๐งฌ Darwin-180B-RSI โ an AI that learns from itself and knows when it's right ๐ FINAL-Bench/Darwin-180B-RSI
๐งฌ Darwin โ crossbreed and evolve the parent Darwin diagnoses strong parent models like an MRI, inherits only their best parts, and evolves the weak spots โ producing a child stronger than its parents. Father model: Qwen3.8-Flash-Next (180B MoE).
๐ RSI ร ๐๏ธ ZTC RSI (recursive self-improvement): solve โ verify against real answers โ learn only the correct reasoning โ repeat. ZTC (Zero-Token Confidence): reads the model's internal state once, before answering, and returns the probability the answer is right โ zero extra tokens. Returns answer + confidence as JSON. {"answer": "...", "confidence": 0.97, "truncated": false}
โจ Synergy: ZTC finds where the model wavers โ RSI learns exactly there โ confidence gets sharper. Low confidence = stop, so agents don't act on wrong answers. โก Same accuracy, 11% shorter reasoning โ faster and cheaper.
๐ The result โ #1 on five Hugging Face official leaderboards ๐ฅ AIME 2026 100% (first perfect score on the board) ๐ฅ HMMT Feb 2026 100% (first perfect score on the board) ๐ฅ GPQA Diamond 94.44% ๐ฅ MMLU-Pro 88.12% ๐ฅ MMMU-Pro 79.48%
๐ 131K-token thinking budget ยท bf16 ยท samples per benchmark listed on the model card. ๐
Today, we open-source Pruna-Qwen-Image-2.1, a set of a few-step LoRA adapters that make Qwen-Image-2. up to 6.3ร faster for image generation and editing.