Are you guys actually out of your minds, or is this some kind of performance art? 😭
I took a quick look at your tokenizer_config.json and almost fell off my chair: "vocab_size": 2048.A vocabulary size of TWO THOUSAND tokens for a "Chat" model? Even ancient GPT-2 had 50k+ tokens. Your model literally has the vocabulary of a broken microwave. If anyone types a word longer than three syllables, your tokenizer will chop it into bloody byte-level pieces, and your 10M micro-skeleton will choke on its own latent space.But the real comedy is the math: you claim you pretrained this 10M pebble on 25 BILLION tokens.
Do you even understand Chinchilla scaling? A 10M model saturates after a few hundred million tokens. Feeding 25B tokens into a 2048-token vocabulary means your gradient descent didn't just "train" the model—it micro-waved it, overfitted it to oblivion, and burned the weights into pure white noise. You literally forced your model to memorize the same 2,000 words twelve million times.Pebble 10M Chat isn't "improving its conversational capabilities" on Smol-SmolTalk. It’s a lobotomized parrot capable only of generating high-entropy slop.Please stop torturing these poor micro-architectures.
Put the Python scripts down and open a basic Deep Learning textbook during your next school lunch break. 🍼🤡