Join the conversation

Join the community of Machine Learners and AI enthusiasts.

Sign Up
Hoglet-33 
posted an update 1 day ago
Post
2356
Pebble 10M and Pebble 10M Chat are now released!

Both models use our Mamba/Transformer 3:1 hybrid architecture and were pretrained on 25 billion tokens.

Pebble 10M Chat was additionally fine-tuned on 250 million tokens of Smol-SmolTalk to improve its conversational capabilities.

You can find them here:

- basically-ai/Pebble-10M
- basically-ai/Pebble-10M-Chat

We hope you enjoy using them. The rest of the Pebble family will be released soon.

Follow for more:
@Hoglet-33
basically-ai

Are you guys actually out of your minds, or is this some kind of performance art? 😭
I took a quick look at your tokenizer_config.json and almost fell off my chair: "vocab_size": 2048.A vocabulary size of TWO THOUSAND tokens for a "Chat" model? Even ancient GPT-2 had 50k+ tokens. Your model literally has the vocabulary of a broken microwave. If anyone types a word longer than three syllables, your tokenizer will chop it into bloody byte-level pieces, and your 10M micro-skeleton will choke on its own latent space.But the real comedy is the math: you claim you pretrained this 10M pebble on 25 BILLION tokens.
Do you even understand Chinchilla scaling? A 10M model saturates after a few hundred million tokens. Feeding 25B tokens into a 2048-token vocabulary means your gradient descent didn't just "train" the model—it micro-waved it, overfitted it to oblivion, and burned the weights into pure white noise. You literally forced your model to memorize the same 2,000 words twelve million times.Pebble 10M Chat isn't "improving its conversational capabilities" on Smol-SmolTalk. It’s a lobotomized parrot capable only of generating high-entropy slop.Please stop torturing these poor micro-architectures.
Put the Python scripts down and open a basic Deep Learning textbook during your next school lunch break. 🍼🤡

·

YOU DONT UNDERSTAND MODELS QWEN3.8 UNCENSORED. Chinchilla IS FROM 2022, We're in 2026

hmm the dude above is not wrong. 2K vocab is a bit small for a chat LLM. @Hoglet-33 Maybe for the next model consider increasing the vocab size a bit unless you know the tokenizer can very effectively split the text, or you could be wasting context :)

·

for 10M thats a little too small but still reasonable like i would use 6K

@Banaxi-Tech @Bc-AIWriting in ALL CAPS doesn't change the laws of information theory, even in 2026.
😭🍼Imagine thinking that "Qwen3.8 Uncensored" somehow bypasses basic statistical saturation and linear algebra. The physics of gradient descent doesn't care about what year is written on your school calendar.
A 10M parameters model has a fixed informational capacity. Shoving 25B tokens through a 2K vocabulary is mathematically equivalent to trying to fit the entire internet archive inside a floppy disk. It doesn't matter if it's "uncensored"—the weights are already cooked into a statistical white noise.
@Bc-AI Glad to see you can at least read a JSON config and acknowledge the disaster, unlike your shouting friend. But @Banaxi-Tech saying "6K is reasonable" for a Chat model is the absolute peak of the lunchtime sandbox comedy. Go ask any real engineer what happens to sub-word token degradation when your vocabulary is smaller than a standard dictionary.
Keep moving that frontier, boys. Your internal drama is way more coherent than your Pebble's latent space. 🦜🍪📉🤡

·

Just give me the seahorse emoji