Also I'm going to vacation tomorrow but it should still be released!
Hoglet (Ash) PRO
AI & ML interests
Recent Activity
Organizations
Also I'm going to vacation tomorrow but it should still be released!
I built an arena where tiny decoder-only LMs (50Kโ250M params) play Tetris zero-shot. There is no fine-tuning and no game data. They only use what they picked up from pre-training on text.
How it works:
- For every piece, the engine simulates each legal placement and describes the result in plain English ("clears one line, creates no new holes, keeps the stack lowโฆ").
- The model never sees the grid. It reads each description, and the arena compares log P(" good move") with log P(" bad move"). The best-rated placement is played.
- Every player gets the same piece sequence, so it's a fair race.
- There are two protocols: Guided (the rules are in the prompt) and Blind (no rules, only pre-training knowledge).
Two ways to play:
- Match: pick any models (even your own, custom architectures welcome) and watch them play side by side on retro 8-bit boards.
- Ranked: press Play and the arena picks up to 4 models at random from a curated pool of 29. Nobody chooses their opponents, so Elo can't be farmed. Matches run on the server and count even if you close the tab.
First results (~225 ranked matches):
- gpt2 (124M) leads with 1283 Elo, but SupraNeo-4M (4M) is right behind at 1239. Next come LowOnMind-5M and BananaMind-2.1-Pico (1.5M!).
- Model size barely predicts Elo (r โ 0.06). Survival does (r โ 0.9): the models that avoid holes and keep the stack low are the ones that win.
Every ranked match (seed, model commit SHAs, scores, Elo before/after) is logged in a public dataset.
โถ Play: DedeProGames/SLM-Tetris-Arena
๐ Results: DedeProGames/lm-tetris-arena-results
Want your model in the Ranked pool? Drop it in the comments!
Why do you think that GPT-2 is in the top followed by sub 10M models?
like instead of being locked into gpt6 luna i can use deepseek or claude, or another openai model
oh sorry, i dont understand anything sometimes
I know you have been working a lot on this, and its great, but could you maybe add support for other models so i can choose or use an API key and model from another platform if the platform is OpenAI compatible?
the Orion in Refract Labs? I let HF have you and Refract Labs stuff show up in feed plus notifications if thats what your asking, plus pinging me sends a notification. i havnt tested it yet though because of this sidetrack, and i accidentally broke a few things in my code base so once i get all of that sorted i will def test it
yeah i know, the base model doesnt even support those tokens, but it was a dataset i could remember easily so thats what i did
The base model I chose was BananaMind/BananaMind-2.1-Pico-Preview, and the dataset I used was SupraLabs/SupraThink-Dataset-500x
I trained for 5 whole steps using a LoRA adapter.
Results:
A model that scores better on some benchmarks and worse on others, and still lacks most general capabilities.
You can find the model here: Hoglet-33/Hogleto
Credits:
- Thank you to @Banaxi-Tech for the BananaAll app (works perfectly on Windows and CPU)
- GPT-6 Sol for knowing how to merge some confusing files created by the app
- Myself for the idea
- Someone else somewhere who might have contributed to some of my ideas and might in the future
- And readers like you!
The base model I chose was BananaMind/BananaMind-2.1-Pico-Preview, and the dataset I used was SupraLabs/SupraThink-Dataset-500x
I trained for 5 whole steps using a LoRA adapter.
Results:
A model that scores better on some benchmarks and worse on others, and still lacks most general capabilities.
You can find the model here: Hoglet-33/Hogleto
Credits:
- Thank you to @Banaxi-Tech for the BananaAll app (works perfectly on Windows and CPU)
- GPT-6 Sol for knowing how to merge some confusing files created by the app
- Myself for the idea
- Someone else somewhere who might have contributed to some of my ideas and might in the future
- And readers like you!
Why not have used Qwen 3 or 3.5?
If you want to use a custom architecture, previously you had to go trough reviewing the code yourself, now add an Openrouter API key and review it with GPT 6 Luna in one button. A review cost be half a cent so anyone can try it. This is one of the main features.
Now ROCm, AMD and Windows, Mac support.
Colab and Molab support.
Detailed list of features:
Get improved Windows Python detection and support paths for compatible AMD ROCm, Intel XPU, and Apple MPS setups.
Choose local training or export a self-contained Python script for Colab or Molab. Notebook runs produce a downloadable model ZIP.
Start pretraining with an existing modelโs tokenizer, or train a new one from your datasets.
Try experimental 1.58-bit Ternary fake-quantized training on NVIDIA GPUs.
Watch live tokens per second. Model compilation is on by default and falls back automatically if it fails.
Build custom architectures with separate configuration and modeling files, then review the training code manually or with optional OpenRouter AI Review.
Install from source with the new coding-agent instructions.
This release also fixes inflated loss reporting for custom models.
And for those users who didn't want to try it out just because installation would be so hard, it isnt now.
Go to any coding agent (Pi, Claude Code, Codex, OpenCode, basically all work), and just paste "Install BananaAll for me. Fetch and follow https://raw.githubusercontent.com/BananaMind/BananaAll/main/agent_install.txt."
That's it.
Check it out at https://github.com/BananaMind/BananaAll/
Also on SAICR, we're currently training a new major model (NACR v2) and ACR 1.0 is in the finishing.
so basically the update for today is:
Rocm, Apple, Intel, AI Support for writing scripts via GPT 6 Luna, better UI, Molab and Colab support, Ternary model training, Windows Support
YES!
not much actually, most things stay the same, some things are a little weird though
also i rent vast.ai GPUs for training, could i set it up in jupyter easily?
Even my RNN/SSM based models?
Check it out and like it!
AxiomicLabs/Tiny_Theory_of_Mind
1. Pebble 1.5
We're working on Pebble 1.5. Here's what we know so far:
- They will be better than the last generation. 99.99% certain.
- Expanded context lengths of at least 16,384 tokens, with the flagship potentially reaching 32,768.
- A Mamba3-based architecture with some other new architectural designs we're experimenting with.
- Native CPU compatibility โ something we failed at with the last generation.
- Natively multilingual and multimodal???
2. SmolCodeBench
A code benchmark designed specifically for small models, because there really isn't a good one right now.
3. SENTRY
VOID is working on something called SENTRY โ System for Evaluating Neural Threats, Responses, and Yields.
More on that soon.
4. basically OS
It's an operating system/app/harness. We're still deciding.
5. Finances
Trying to balance the finances after purchasing a Hugging Face Pro subscription.
Follow us for updates:
@Hoglet-33