G1-MINI has now seen around 8B tokens during its current run, and pretraining is still going strong.
Our E1 (Efficiency-1) prototype has also reached 15B pretraining tokens. E1 has 1B total parameters while activating under 100M parameters per token. It features adaptive activation, meaning easier tokens can use less compute while harder tokens receive more.
We plan to open-source E1 ASAP! 🚀
We’re also excited to announce Project Prism, which will provide limited access to our upcoming Orion Flagship model, powered by our T2 architecture.
Note: T2 here refers to the architecture, not our T2 (Thinker-2) model.
Applications for Project Prism are available through the org page, with more details coming soon!
Finally, welcome @soyL061215, who joined the Hugging Face org today! 🎉
We have some updates to @BananaMindBot 🍌 It can now train models, ask it to train a model, and i will train it for you. It now can also merge PRs And like models.
Most things you do on HuggingFace, BananaMindBot can do. Fast
Mention @BananaMindBot on a model, dataset, Space discussion, paper, blog comment, or top-level post and it'll reply there.
It's powered by North Code Mini (Qwen3.8 27B, with GPT OSS 120B as fallback).
A few things it can do:
Search for models and datasets Look up users and orgs and see what they've published Read model cards, configs, dataset files, blog posts, and org profiles Answer questions about what it finds Write and run its own code in a locked-down sandbox when it needs to verify something Check things like a model's real parameter count from the safetensors headers instead of just repeating the model card Remember something for later if you explicitly ask it to Forward a message to @Banaxi-Tech Post a daily roundup of developments in the small-language-model space
It won't execute code you give it. It can read and review that code, but anything it runs is code it wrote itself.
It also can't access private data or credentials.
Mention it somewhere.
It's going to also find this post!
(Some parts inspired by CompactBot and @CompactAI Follow them please)
Hi everyone! We've seen some people getting confused with the BananaMind Leaderboards so ill explain!
We have 2 leaderboards, THESE are NOT the same, first BananaMind/BananaMindBench-Leaderboard which is ONLY for BananaMind Base Bench 1.1. The 10/10 scores do NOT mean that the benchmark is saturated. It isnt saturated, these models score 10/10 because they are the current best models, our /10 ranking system works by taking the ELO scores and then comparing them to the scores in the same size range. So if a better model releases that gets 10/10 and the others get lower.
And we also have the BananaMind SLM leaderboard, not the BananaMindBench leaderboard which uses ARC EASY,PIQA,Hellaswag, Arithmark 3 and the BananaMind Base Bench 1.1. This is the newer and recommended version.
We're releasing the BananaMind SLM Leaderboard! It offers a easier look at which models are actually good for your specific needs. Its primary metric, Intelligence index is a composite of BananaMind Base Bench, PIQA, Hellaswag, ARC Easy and Arithmark 3. It also allows you to see specific categories like Commonsense on a model.
Hello Everyone! I am happy to announce a few things. 1. G1-MINI G1-MINI is now in pretraining and is training at a steady pace. Our current ETAs state completion and launch in about 15-20 days, somewhere near the end of September. 2. G1-NANO G1-NANO is also being pretrained as we speak at a pace of over 400K tokens per second processing more than 10B tokens in 12 hours. This allows us to train extremely fast, and we will launch it somewhere around 15th September. 3. We have begun work on FrameShot, a dual image and video generation model at around 4B dense parameters. This is expected to launch around late December with no promised date. - Bc-AI
Hello everyone! I have 2 announcements today! The first one is the launch of our new API platform! You can make a account and get 5 dollars free credits. No credits card needed because i have no idea how to set up a payment's thing. If you want more credits just email me at smilyai@outlook.com . The platform currently features G1-Preview a preview of G1 and the older Mira-1-Large. 2nd announcement is we have started working on G1-MINI so expect a late October Ish launch - Bc-AI on behalf of Smilyai-Labs
I know a lot of you were really looking forward to the original 20B MoE, and honestly, I was really excited about it too.
Unfortunately, the free compute credits I was using from ML Intern Explorers were removed by Hugging Face. That changed what I can realistically do with the original plan, so I've decided to move G1 over to a Qwen 3.8 27B base instead. Its not just another finetune though, I am inserting extra layers and putting it through my vigourous pipeline. Results will be open source.
I know that's probably disappointing, especially for the people who were specifically waiting for the 20B MoE. I'm genuinely sorry about that.
I really appreciate everyone who got excited about G1 in the first place. I didn't expect this change either, but I'm still really excited to see where G1 can go from here.
The original G1 codebase will stay open too. Its under my profile: Bc-AI/train-g1
I have some unfortunate news to share with everyone. My earlier estimate for the launch of G1 in late October was inaccurate. We sincerely apologise for any inconvenience this may cause, but with our current compute resources, pretraining a 20B MoE model is not realistically possible within two months.
G1-MINI will also be postponed, but not for nearly as long — only by a few extra months.
This is disappointing, as I know I was excited to launch G1, and I know many people were also watching the model and looking forward to it.
However in my view, I would rather be honest about our limitations than be overly optimistic about something we currently cannot guarantee.
G1 is not cancelled but uh it will be postponed indefinitely until I have the resources needed to train it. This could be next month, or it could take years. For now, I don't want to give another estimated launch date until I know we have the resources to actually make it happen.
In the meantime, SmilyAI will continue developing AI and experimenting with new ideas, and we will provide updates as we go.
Thank you for sticking with us and supporting SmilyAI. We will continue working towards better models in the future.
Also, if you do have the hardware, to run it aka 8xH200s or better, my codebase is fully open under my very permissive license: i-have-no-idea-just-use-this. Basically, do whatever just mention me. Bc-AI/train-g1