๐ค Open to Collab
LH-Tech AI
LH-Tech-AI
AI & ML interests
Small AI and ML models. Trained by myself. Completely OpenSource. For you. | Reddit: https://www.reddit.com/user/LH-Tech_AI/
Recent Activity
liked a dataset 41 minutes ago
zouhar/last-translation-benchmark liked a model about 2 hours ago
FWKV/Myosotis-1-base liked a Space about 3 hours ago
ivanmikhnenkov/tinyditOrganizations
reacted to Banaxi-Tech's post with โค๏ธ about 9 hours ago
Post
3581
Supra2-100M is out!
Go check it out:
- https://www.reddit.com/r/LocalLLaMA/comments/1velyl9/new_models_supra2100m_base_and_instruct_go_check/
- https://huggingface.co/SupraLabs/Supra2-100M
- SupraLabs/Supra2-100M-Instruct
Give us a like and a follow!!
HAVE FUN ๐ค๐ฅ๐
more coming soon...
Go check it out:
- https://www.reddit.com/r/LocalLLaMA/comments/1velyl9/new_models_supra2100m_base_and_instruct_go_check/
- https://huggingface.co/SupraLabs/Supra2-100M
- SupraLabs/Supra2-100M-Instruct
Give us a like and a follow!!
HAVE FUN ๐ค๐ฅ๐
more coming soon...
reacted to Enderchef's post with ๐ค๐ฅโค๏ธ 5 days ago
Post
3448
๐ Supra2 100M is out, and multiple other SLM orgs are gaining power!
Following takes a press. Please follow:
fromziro
SupraLabs
AxiomicLabs
Following takes a press. Please follow:
reacted to Banaxi-Tech's post with ๐ค๐ค๐๐ 11 days ago
Post
3509
We're delaying BananaMind 2.1!
When BananaMind 2.1 Lite was almost done, we benchmarked it and the results we're worse than BananaMind 2 Mini.
We're going to spend alot more time in research on tiny models and then scaling up our techniques to the actual BananaMind 2.1 models!
We're also announcing these new models:
BananaMind 2.1 Coder: A 149M instruction tuned coder model trained on 75B tokens + 10B tokens of stack-v3-train.
BananaMind 2.1 Pico: A 1M parameter model trained on 22B tokens of data.
We also may release BananaMind 2.1 Large with around 100M parameters depending on how much compute we have.
Please give us a follow!
BananaMind
@Banaxi-Tech
---
@vovaRL
@DedeProGames
When BananaMind 2.1 Lite was almost done, we benchmarked it and the results we're worse than BananaMind 2 Mini.
We're going to spend alot more time in research on tiny models and then scaling up our techniques to the actual BananaMind 2.1 models!
We're also announcing these new models:
BananaMind 2.1 Coder: A 149M instruction tuned coder model trained on 75B tokens + 10B tokens of stack-v3-train.
BananaMind 2.1 Pico: A 1M parameter model trained on 22B tokens of data.
We also may release BananaMind 2.1 Large with around 100M parameters depending on how much compute we have.
Please give us a follow!
@Banaxi-Tech
---
@vovaRL
@DedeProGames
Please check your discord DMs @Enderchef ๐ญ
reacted to Banaxi-Tech's post with ๐ฅ 16 days ago
Post
2706
We're releasing Overfitter 1.0.
Its a completly useless overfitted model trained on 202 epochs of SWE Bench Verified, SWE Bench Pro, Terminal Bench 2.1, DeepSWE.
It gets 100% on SWE Bench Verified, 98.6% on SWE bench Pro, 100% on Terminal Bench 2.1 and 100% on DeepSWE!
BananaMind/Overfitter-1.0
Its a completly useless overfitted model trained on 202 epochs of SWE Bench Verified, SWE Bench Pro, Terminal Bench 2.1, DeepSWE.
It gets 100% on SWE Bench Verified, 98.6% on SWE bench Pro, 100% on Terminal Bench 2.1 and 100% on DeepSWE!
BananaMind/Overfitter-1.0
reacted to Bc-AI's post with ๐ฅ 19 days ago
Post
3785
New update! We are currently training a few new models now! Our 3rd generation main LLM standard edition is in training right now. We are also training a new LLM line called Tiny Coder around 350~ish M params. Thanks to @Banaxi-Tech for inspiring the architecture with his Bananamind-2.1-unified test model. Thanks to our beta testers: @juiceb0xc0de @ProCreations @Sbui503 @Fishtiks @MUK-IS-GOAT
reacted to Banaxi-Tech's post with ๐ฅ 19 days ago
Post
3742
We're releasing BananaMind 2.1 Unified, a 35M three-tower relay model where the two output towers can only talk to each other through a silent middle tower that has no output head and no loss term.
Tower B trains entirely on indirect gradient. It was never told what to predict. It woke up anyway. Jacobian lens shows it carrying the correct answer ("Paris", "oxygen", "blue") at its deepest layer. Its bridge gates grew 4-47x from init. Feed it from only one side and the representations collapse to junk โ it needs both outer towers to become semantic.
The PIQA result is the cleanest demonstration: Tower A alone scores 50.11 (chance). Tower C alone 52.12. Full system 61.75. All physical reasoning lives in the integration. The 35M three-tower beats the 50M single-tower BananaMind 2 Medium on PIQA.
Trained on 38B tokens in ~7.5 hours on 8x RTX PRO 6000. Ships with 7 ablation modes so you can surgically cut the model apart without retraining. Full training logs, J-lens fits, and eval outputs for every mode included.
This is a research model. It is interesting because you can take it apart.
BananaMind/BananaMind-2.1-Unified
Follow us for more:
BananaMind
@Banaxi-Tech
Tower B trains entirely on indirect gradient. It was never told what to predict. It woke up anyway. Jacobian lens shows it carrying the correct answer ("Paris", "oxygen", "blue") at its deepest layer. Its bridge gates grew 4-47x from init. Feed it from only one side and the representations collapse to junk โ it needs both outer towers to become semantic.
The PIQA result is the cleanest demonstration: Tower A alone scores 50.11 (chance). Tower C alone 52.12. Full system 61.75. All physical reasoning lives in the integration. The 35M three-tower beats the 50M single-tower BananaMind 2 Medium on PIQA.
Trained on 38B tokens in ~7.5 hours on 8x RTX PRO 6000. Ships with 7 ablation modes so you can surgically cut the model apart without retraining. Full training logs, J-lens fits, and eval outputs for every mode included.
This is a research model. It is interesting because you can take it apart.
BananaMind/BananaMind-2.1-Unified
Follow us for more:
@Banaxi-Tech
replied to Banaxi-Tech's post 19 days ago
Looks cool! Can't wait to test these!
But we will beat this with our Supra3 family soon (coming in around 4 to 8 weeks)! ๐ฅ๐
Stay tuned! :D
reacted to Banaxi-Tech's post with ๐ฅ 19 days ago
Post
3791
We're announcing our BananaMind 2.1 model series!
The models will include:
- BananaMind 2.1 Nano: 10M parameters with 60B tokens.
- BananaMind 2.1 Lite: 25M parameters with 40B tokens.
- BananaMind 2.1 Flash: 50M parameters with 55B tokens.
- BananaMind 2.1 Pro: 135M-145M parameters (still deciding) with 100B tokens.
These model will use a multi tower architecture (like BananaMind/BananaMind-2.1-Unified) with some more architectural changes.
BananaMind 2.1 Pro will probrably use 2 no output towers, instead of one!
We're currently training some experimental models based on this architecture to see its scaling!
Follow us:
BananaMind
@Banaxi-Tech
@vovaRL
@DedeProGames
bananamind-research-community
The models will include:
- BananaMind 2.1 Nano: 10M parameters with 60B tokens.
- BananaMind 2.1 Lite: 25M parameters with 40B tokens.
- BananaMind 2.1 Flash: 50M parameters with 55B tokens.
- BananaMind 2.1 Pro: 135M-145M parameters (still deciding) with 100B tokens.
These model will use a multi tower architecture (like BananaMind/BananaMind-2.1-Unified) with some more architectural changes.
BananaMind 2.1 Pro will probrably use 2 no output towers, instead of one!
We're currently training some experimental models based on this architecture to see its scaling!
Follow us:
@Banaxi-Tech
@vovaRL
@DedeProGames
reacted to Enderchef's post with ๐๐๐ค 21 days ago
Post
4102
Do you support the SLM(1M-150M parameter) community?
If so, join the SLM discord(https://discord.gg/BBYaERvvn), and give these orgs some follows!
BananaMind
SupraLabs
fromziro
AxiomicLabs
If so, join the SLM discord(https://discord.gg/BBYaERvvn), and give these orgs some follows!