111
followers ·
409 following AI & ML interests Building the AI-native stack. Agents as infrastructure, safety as architecture, performance as plumbing. I publish the receipts: papers, datasets, demos.
Recent Activity reacted to chaoliangUNSW 's post with 🧠 24 minutes ago Jev-style decisions, now at 2B: Jev-Style-2B-Decision-v3
State and typed questions in (pick one, yes/no, score); a calibrated probability for every option out, in one pass.
Try it in your browser:
https://huggingface.co/spaces/chaoliangUNSW/jev-style-2b
• JevBench v1.4.1, 231 public items (self-run with the official v1.4.1 harness, not an official board entry): 73.6% (170/231), the highest public accuracy among the Qwen3.5-2B-family systems on the board (decider-2b 71.0%, open-jev-zefan-2b 64.5%; the lead over decider-2b is inside the 95% CI, 67.6–78.9%). 42 of the 82 board systems score higher, almost all 4B or larger. +9.5 points over our 0.8B v3. Jev 1.13 is well ahead at 86.6%.
• tweet_topic, zero-shot (n = 1,693): 82.2% accuracy [80.4, 84.0], above Jev's 79.3%; macro-F1 is below Jev (0.678 vs 0.694).
• fin_topic, zero-shot (n = 4,117): 61.1%, +14.4 points over our 0.8B v3 (Jev: 67.0%).
• 25,600 tokens per call, no option cap, nothing truncated.
• Read once, then ask: on an Apple M1 Max (GGUF F16), the first question about a 24,501-token input took 16.2 s and a further question about the same state 0.17 s (shared machine, indicative).
Contamination checks against the training pool: 0 hits for the JevBench items, and no exact or near-duplicate overlap with the tweet_topic and fin_topic test sets. Needs the shipped runtimes (PyTorch, llama.cpp with the bundled scorer, MLX); trained on a reduced data pool (60M tokens).
Models (Apache-2.0; some training data has restrictive or unclear terms and includes OpenAI GPT and Anthropic Claude outputs, see "Training data and licences" on the main card):
https://huggingface.co/chaoliangUNSW/Jev-Style-2B-Decision-v3
https://huggingface.co/chaoliangUNSW/Jev-Style-2B-Decision-v3-GGUF
https://huggingface.co/chaoliangUNSW/Jev-Style-2B-Decision-v3-MLX
All v3 builds: https://huggingface.co/collections/chaoliangUNSW/jev-style-decision-v3-08b-2b-6ab87f32380cbd8c03b608b9
Not affiliated with TypeSafe or Jev. reacted to DedeProGames 's post with 🔥 24 minutes ago 🧱 SLM Tetris Arena: can a small language model play Tetris without ever being trained on it?
I built an arena where tiny decoder-only LMs (50K–250M params) play Tetris zero-shot. There is no fine-tuning and no game data. They only use what they picked up from pre-training on text.
How it works:
- For every piece, the engine simulates each legal placement and describes the result in plain English ("clears one line, creates no new holes, keeps the stack low…").
- The model never sees the grid. It reads each description, and the arena compares log P(" good move") with log P(" bad move"). The best-rated placement is played.
- Every player gets the same piece sequence, so it's a fair race.
- There are two protocols: Guided (the rules are in the prompt) and Blind (no rules, only pre-training knowledge).
Two ways to play:
- Match: pick any models (even your own, custom architectures welcome) and watch them play side by side on retro 8-bit boards.
- Ranked: press Play and the arena picks up to 4 models at random from a curated pool of 29. Nobody chooses their opponents, so Elo can't be farmed. Matches run on the server and count even if you close the tab.
First results (~225 ranked matches):
- gpt2 (124M) leads with 1283 Elo, but SupraNeo-4M (4M) is right behind at 1239. Next come LowOnMind-5M and BananaMind-2.1-Pico (1.5M!).
- Model size barely predicts Elo (r ≈ 0.06). Survival does (r ≈ 0.9): the models that avoid holes and keep the stack low are the ones that win.
Every ranked match (seed, model commit SHAs, scores, Elo before/after) is logged in a public dataset.
▶ Play: https://huggingface.co/spaces/DedeProGames/SLM-Tetris-Arena
📊 Results: https://huggingface.co/datasets/DedeProGames/lm-tetris-arena-results
Want your model in the Ranked pool? Drop it in the comments! View all activity Organizations dipankarsarkar 's activity All Models Datasets Spaces Buckets Papers Collections Community Posts Upvotes Likes Articles published an article about 2 months ago view article A ring of teachers, a 1.7B brain, and a harness that couldn't lie published an article about 2 months ago view article CLASSROOM-SOTA training spec — ring-of-teachers distillation for the quantal brain published an article about 2 months ago view article quantal-ternary: a 0.5B model that learned honesty — from 11.34 to 0.5597 view article The cheap oracle is lying to you: from GPU kernels to agent rewards dipankarsarkar
• Jun 28
• 2