Dipankar Sarkar PRO
dipankarsarkar
AI & ML interests
Building the AI-native stack. Agents as infrastructure, safety as architecture, performance as plumbing. I publish the receipts: papers, datasets, demos.
Recent Activity
reacted to chaoliangUNSW's post with 🧠 18 minutes ago
Jev-style decisions, now at 2B: Jev-Style-2B-Decision-v3
State and typed questions in (pick one, yes/no, score); a calibrated probability for every option out, in one pass.
Try it in your browser:
https://huggingface.co/spaces/chaoliangUNSW/jev-style-2b
• JevBench v1.4.1, 231 public items (self-run with the official v1.4.1 harness, not an official board entry): 73.6% (170/231), the highest public accuracy among the Qwen3.5-2B-family systems on the board (decider-2b 71.0%, open-jev-zefan-2b 64.5%; the lead over decider-2b is inside the 95% CI, 67.6–78.9%). 42 of the 82 board systems score higher, almost all 4B or larger. +9.5 points over our 0.8B v3. Jev 1.13 is well ahead at 86.6%.
• tweet_topic, zero-shot (n = 1,693): 82.2% accuracy [80.4, 84.0], above Jev's 79.3%; macro-F1 is below Jev (0.678 vs 0.694).
• fin_topic, zero-shot (n = 4,117): 61.1%, +14.4 points over our 0.8B v3 (Jev: 67.0%).
• 25,600 tokens per call, no option cap, nothing truncated.
• Read once, then ask: on an Apple M1 Max (GGUF F16), the first question about a 24,501-token input took 16.2 s and a further question about the same state 0.17 s (shared machine, indicative).
Contamination checks against the training pool: 0 hits for the JevBench items, and no exact or near-duplicate overlap with the tweet_topic and fin_topic test sets. Needs the shipped runtimes (PyTorch, llama.cpp with the bundled scorer, MLX); trained on a reduced data pool (60M tokens).
Models (Apache-2.0; some training data has restrictive or unclear terms and includes OpenAI GPT and Anthropic Claude outputs, see "Training data and licences" on the main card):
https://huggingface.co/chaoliangUNSW/Jev-Style-2B-Decision-v3
https://huggingface.co/chaoliangUNSW/Jev-Style-2B-Decision-v3-GGUF
https://huggingface.co/chaoliangUNSW/Jev-Style-2B-Decision-v3-MLX
All v3 builds: https://huggingface.co/collections/chaoliangUNSW/jev-style-decision-v3-08b-2b-6ab87f32380cbd8c03b608b9
Not affiliated with TypeSafe or Jev. reacted to DedeProGames's post with 🔥 18 minutes ago
🧱 SLM Tetris Arena: can a small language model play Tetris without ever being trained on it?
I built an arena where tiny decoder-only LMs (50K–250M params) play Tetris zero-shot. There is no fine-tuning and no game data. They only use what they picked up from pre-training on text.
How it works:
- For every piece, the engine simulates each legal placement and describes the result in plain English ("clears one line, creates no new holes, keeps the stack low…").
- The model never sees the grid. It reads each description, and the arena compares log P(" good move") with log P(" bad move"). The best-rated placement is played.
- Every player gets the same piece sequence, so it's a fair race.
- There are two protocols: Guided (the rules are in the prompt) and Blind (no rules, only pre-training knowledge).
Two ways to play:
- Match: pick any models (even your own, custom architectures welcome) and watch them play side by side on retro 8-bit boards.
- Ranked: press Play and the arena picks up to 4 models at random from a curated pool of 29. Nobody chooses their opponents, so Elo can't be farmed. Matches run on the server and count even if you close the tab.
First results (~225 ranked matches):
- gpt2 (124M) leads with 1283 Elo, but SupraNeo-4M (4M) is right behind at 1239. Next come LowOnMind-5M and BananaMind-2.1-Pico (1.5M!).
- Model size barely predicts Elo (r ≈ 0.06). Survival does (r ≈ 0.9): the models that avoid holes and keep the stack low are the ones that win.
Every ranked match (seed, model commit SHAs, scores, Elo before/after) is logged in a public dataset.
▶ Play: https://huggingface.co/spaces/DedeProGames/SLM-Tetris-Arena
📊 Results: https://huggingface.co/datasets/DedeProGames/lm-tetris-arena-results
Want your model in the Ranked pool? Drop it in the comments! upvoted a paper 19 minutes ago
ANTMAN: Adaptive Need Tracking for Multi-Agent Navigation in Large Information Spaces