Join the conversation

Join the community of Machine Learners and AI enthusiasts.

Sign Up

All HF Hub posts

OppaAIΒ 
posted an update 1 day ago
Hoglet-33Β 
posted an update 1 day ago
view post
Post
2380
We are announcing the first generation of the Pebble model family!

These are the models we are releasing:

- Pebble 10M
- Pebble 25M
- Pebble 50M

Each model will use a Mamba-Transformer 3:1 hybrid architecture and will be pretrained on 25 billion tokens before IFT and SFT.

Depending on development time and resources, we may also release:

- Pebble 5M
- Pebble 75M
- Pebble 1M (possibly)

We hope you're excited and enjoy the models!

Follow for more:
@Hoglet-33
basically-ai
  • 7 replies
Β·
appvoidΒ 
posted an update about 19 hours ago
view post
Post
1059
Nobody knows what is doing, when you train a model, you are experimenting to advance the frontier, so keep failing 🫡
  • 9 replies
Β·
Bc-AIΒ 
posted an update 1 day ago
view post
Post
1957
Smilyai News
Hello everyone! August has been a crazy month for us at Smilyai-Labs. We've been doing lots behind the scenes, so here's the latest πŸ‘‡

1. MiniCoder
We are very close to releasing MiniCoder-1, our first-generation coding model designed for reasoning and coding. Our planned context window is 128K, but the earlier versions probably will not support that long! It's currently in the final stages of DPO so expect a release in early september.
Release: VERY SOONβ„’πŸ€£

2. Smilyai G1
So, the current plan is 20B parameter model total, with a MoE architecture, activating around 2B parameters per token. Its desgigned for maximum performance but keeping it runnable on consumer hardware. It's only a plan and i have no idea when me and the team can finish it. Expect a launch around the end of september to early october-ish. I have no guarantees so don't quote me on the launch date.

3. T1
Smilyai-T1 is another major model we are working on.
The goal for T1 is to take what we learnt from the countless architectural experiments and creating a powerful model designed for thinking. Think MiniCoder but reasons more and G1 but more capable. Its main goals are coding, math, reasoning and general capability.
4. Omni
We are also planning Omni, our first from scratch multimodal model. It will not launch this year as it will take a while. We are actively researching the best architecture for it and we will update progress as we go!



Thanks to our beta testers:
@guardamarcos
@ProCreations
@juiceb0xc0de
@Timmy6767
@Sbui503
@atom77777
@Fishtiks
@smartdigitalnetworks
@EmetTheGolum
@smilyai-large-team
@MUK-IS-GOAT
@Bc-AI

Thanks to my friends who work with me at lunchtimes (Smilyai-Labs team):
@MUK-IS-GOAT
@smilyai-large-team

August was wild. Let’s see what September brings. πŸš€

β€” Bc-AI, on behalf of SmilyAI Labs
salma-remyxΒ 
posted an update 2 days ago
view post
Post
2303
If you can’t explain why your AI-generated contribution belongs in the repo, don’t put it in a maintainer’s queue πŸ™…πŸ»β€β™€οΈ

Before asking for review, you should be able to answer:

1. What project need does it address?
2. Where does it fit, and does it duplicate existing work?
3. What evidence shows it works, and will you own it through review?

If you can’t answer those, you haven’t saved anyone time. You’ve passed the buck to the maintainer.

Outrider has made us better contributors by doing more of this work before upstream review by
* reading contribution rules, accepted PRs, and open issues
* finding needs and integration points
* drafting the code, tests, and context.

We still decide what deserves to go upstream, verify the claims, coordinate with maintainers and contributors, and stay involved through review.

On huggingface/peft, only 4 of 20 Outrider runs opened draft PRs. Two contributions have now merged:

βœ… Riemannian-preconditioned LoRA: https://github.com/huggingface/peft/pull/3382
βœ… Super-Tuning: https://github.com/huggingface/peft/pull/3518

We have more contributions in review and far more ideas were filtered out before they reached a maintainer.

Full case study: https://remyx.ai/case-study
Outrider: https://github.com/remyxai/outrider
  • 3 replies
Β·
Banaxi-TechΒ 
posted an update 2 days ago
view post
Post
2823
We’re excited to release BananaMind OS 2.0, a major update to our portable operating system for running AI models locally.

BananaMind OS runs directly from an ISO without Linux, installation, a cloud connection, or modifying your disks.

The new Version 2.0 adds a graphical interface with mouse support, a model library, multi-turn chat, configurable KV cache, maximum tokens and temperature
controls, automatic x87/SSE/SSE2 CPU detection, BIOS support, native UEFI support and a dedicated 486 compatibility mode.

Model weights are not loaded during GRUB or startup. Only the model catalog is read, and the selected model is loaded after BananaMind OS starts.

We now have a new .litemodel format supports multiple architectures, including BananaMind models, SmolLM, SmolLM2, GPT-X2.5, min-spark 1.1 and Rose-Mini.
Check it out at:
https://github.com/BananaMind/BananaMindOS

You can build an ISO yourself using the .sh or .bat script. Select the models and quantizations you want, and the builder will download, quantize
and package them locally.
Prebuilt 10MB, 25MB, 100MB and 250MB model presets are available here:
https://github.com/BananaMind/BananaMindOS/releases/tag/v2.0.0
Use the regular preset ISOs for BIOS and GRUB, including the 486 compatibility mode.

Warning: Im currently uploading the ISOs, all up to 100MB is present, 250MB is getting uploaded
Use the files ending in -uefi.iso for the native x86-64 UEFI graphical frontend and improved firmware mouse support.

The preset name describes the RAM class of the individual included models. Models remain on the ISO until selected, so including multiple models does not load all of them into RAM.
486DX with an x87 FPU, Pentium and newer x86 processors are supported. Modern x86-64 computers are supported through UEFI. It currently doesent support processors without an FPU.



Comment if you want me to run it on 0.04MHz (pls dont)





Video Credit:
Song: Matzan - Redesigned
Music provided by NoCopyrightSounds
  • 3 replies
Β·
CodeSoftΒ 
posted an update about 15 hours ago
view post
Post
828
Wow, SLM Arena is getting a lot of traffic! Thank you guys for showing your interest!

To handle the growing demand, I’m moving SLM Arena from a CPU Space to a ZeroGPU Space. Hopefully, this will let me add more models to SLM Arena while keeping it running fast.

I've also added a separate arena + leaderboard for base models!

If there are any models or features you’d like to see, let me know in a reply to this post or in a Community post on the Space!
  • 13 replies
Β·
GoktugDΒ 
posted an update 1 day ago
view post
Post
1478
πŸ‡ΉπŸ‡· We trained a 1B OCR model specifically for Turkish enterprise documents.

**Werea-DocOCR-1B v2**

The result surprised us:

LightOnOCR-2 base β†’ **64.2% CER**
Werea-DocOCR v1 β†’ **~8.1% CER**
Werea-DocOCR v2 β†’ **0.15% CER** πŸš€

Evaluated on a held-out 72-page test set across 12 Turkish document types and 3 different capture conditions.

πŸ“„ 12 Turkish enterprise document types
πŸ§ͺ 12,960 synthetic training pages
πŸ“± Digital + scanned + phone photos
πŸ“Š Tables β†’ structured Markdown
βš™οΈ Full-parameter fine-tuning
πŸ–₯️ Trained on a single RTX 3090

It handles:

β€’ e-Invoices
β€’ rental contracts
β€’ bank receipts
β€’ payroll documents
β€’ insurance policies
β€’ vehicle documents
β€’ official correspondence
β€’ trade registry documents
β€’ SGK-style tables
β€’ and more.

**Model πŸ€—**
Werea-co/Werea-DocOCR-1B

**Dataset πŸ“š**
Werea-co/werea-tr-doc-ocr-enterprise-v2

**Werea πŸ‡ΉπŸ‡·**
Werea-co


We're building open AI models from TΓΌrkiye.

This is just the beginning.

#HuggingFace #OCR #DocumentAI #TurkishAI #OpenSourceAI #ComputerVision
SeaWolf-AIΒ 
posted an update 2 days ago
view post
Post
2845
πŸ” The attention mask stopped being an audit.

An autoregressive model must not let position t depend on anything after t. Everyone checks this by inspecting the causal mask β€” but hybrid stacks now mix attention with state-space scans, and a scan has no mask. Every mask can be correct while information leaks through scans, aggregations, or normalization.

βš™οΈ So we test the property directly. Two inputs identical except at the last position, two forward passes, compare each layer's prefix, report the first layer that moves. No training, no gradients, no accelerator β€” seconds on CPU.

πŸ“Š Across 192 injected faults on eight checkpoints, mask inspection detected 0. The per-layer audit localized 192/192 to the exact layer.

🎯 Then we read the source before running anything. In transformers 5.7.0, the reference chunked scan reduces the inter-chunk recurrence over the input chunk axis; zamba2 and nemotron_h reduce over the output chunk axis. One axis. The dynamic audit confirmed the prediction exactly: Zamba2-1.2B leaks from length 256, its declared chunk size, and Nemotron-H-8B from 128, its declared chunk size. Bamba, Falcon-H1, Granite-4.0-H, Mamba2 and RecurrentGemma came back clean.

⚠️ Scope: the defect is on the PyTorch chunked-scan path, which runs whenever the fused kernels are absent β€” CPU, CI, stock installs. We could not build those kernels, so the fast path is untested and open. That caveat cuts both ways: a model can pass every fused-kernel test and still leak the moment it runs without them.

πŸ§ͺ AX-RAY now carries this as its own axis. 39 models scored across causal, white-box and behavioral axes: 21 A, 3 B, 1 C, 14 F β€” with exactly 2 Causal-LEAK verdicts, the two the paper predicted. Badges separate a weights-level audit from an API-only one, so the two never get read as the same claim.

πŸ“„ https://arxiv.org/abs/2608.22876
πŸ”¬ FINAL-Bench/AX-RAY
πŸ€— The Mask Is Not the Model: Auditing Prefix Invariance in Attention, State-Space, and Hybrid Sequence Models (2608.22876)
  • 9 replies
Β·
appvoidΒ 
posted an update 3 days ago
view post
Post
2549
I hope that after this OpenAI disaster on Plus users, more people start realizing why Open Weights were always the only way.
  • 5 replies
Β·