Dipankar Sarkar's picture
🤝 Open to Collab

Dipankar Sarkar PRO

dipankarsarkar

AI & ML interests

Building the AI-native stack. Agents as infrastructure, safety as architecture, performance as plumbing. I publish the receipts: papers, datasets, demos.

Recent Activity

reacted to chaoliangUNSW's post with 🧠 18 minutes ago
Jev-style decisions, now at 2B: Jev-Style-2B-Decision-v3 State and typed questions in (pick one, yes/no, score); a calibrated probability for every option out, in one pass. Try it in your browser: https://huggingface.co/spaces/chaoliangUNSW/jev-style-2b • JevBench v1.4.1, 231 public items (self-run with the official v1.4.1 harness, not an official board entry): 73.6% (170/231), the highest public accuracy among the Qwen3.5-2B-family systems on the board (decider-2b 71.0%, open-jev-zefan-2b 64.5%; the lead over decider-2b is inside the 95% CI, 67.6–78.9%). 42 of the 82 board systems score higher, almost all 4B or larger. +9.5 points over our 0.8B v3. Jev 1.13 is well ahead at 86.6%. • tweet_topic, zero-shot (n = 1,693): 82.2% accuracy [80.4, 84.0], above Jev's 79.3%; macro-F1 is below Jev (0.678 vs 0.694). • fin_topic, zero-shot (n = 4,117): 61.1%, +14.4 points over our 0.8B v3 (Jev: 67.0%). • 25,600 tokens per call, no option cap, nothing truncated. • Read once, then ask: on an Apple M1 Max (GGUF F16), the first question about a 24,501-token input took 16.2 s and a further question about the same state 0.17 s (shared machine, indicative). Contamination checks against the training pool: 0 hits for the JevBench items, and no exact or near-duplicate overlap with the tweet_topic and fin_topic test sets. Needs the shipped runtimes (PyTorch, llama.cpp with the bundled scorer, MLX); trained on a reduced data pool (60M tokens). Models (Apache-2.0; some training data has restrictive or unclear terms and includes OpenAI GPT and Anthropic Claude outputs, see "Training data and licences" on the main card): https://huggingface.co/chaoliangUNSW/Jev-Style-2B-Decision-v3 https://huggingface.co/chaoliangUNSW/Jev-Style-2B-Decision-v3-GGUF https://huggingface.co/chaoliangUNSW/Jev-Style-2B-Decision-v3-MLX All v3 builds: https://huggingface.co/collections/chaoliangUNSW/jev-style-decision-v3-08b-2b-6ab87f32380cbd8c03b608b9 Not affiliated with TypeSafe or Jev.
reacted to DedeProGames's post with 🔥 18 minutes ago
🧱 SLM Tetris Arena: can a small language model play Tetris without ever being trained on it? I built an arena where tiny decoder-only LMs (50K–250M params) play Tetris zero-shot. There is no fine-tuning and no game data. They only use what they picked up from pre-training on text. How it works: - For every piece, the engine simulates each legal placement and describes the result in plain English ("clears one line, creates no new holes, keeps the stack low…"). - The model never sees the grid. It reads each description, and the arena compares log P(" good move") with log P(" bad move"). The best-rated placement is played. - Every player gets the same piece sequence, so it's a fair race. - There are two protocols: Guided (the rules are in the prompt) and Blind (no rules, only pre-training knowledge). Two ways to play: - Match: pick any models (even your own, custom architectures welcome) and watch them play side by side on retro 8-bit boards. - Ranked: press Play and the arena picks up to 4 models at random from a curated pool of 29. Nobody chooses their opponents, so Elo can't be farmed. Matches run on the server and count even if you close the tab. First results (~225 ranked matches): - gpt2 (124M) leads with 1283 Elo, but SupraNeo-4M (4M) is right behind at 1239. Next come LowOnMind-5M and BananaMind-2.1-Pico (1.5M!). - Model size barely predicts Elo (r ≈ 0.06). Survival does (r ≈ 0.9): the models that avoid holes and keep the stack low are the ones that win. Every ranked match (seed, model commit SHAs, scores, Elo before/after) is logged in a public dataset. ▶ Play: https://huggingface.co/spaces/DedeProGames/SLM-Tetris-Arena 📊 Results: https://huggingface.co/datasets/DedeProGames/lm-tetris-arena-results Want your model in the Ranked pool? Drop it in the comments!
View all activity

Organizations

Skelf Research's profile picture Neul Labs's profile picture Cognisoc's profile picture Incredlabs's profile picture