Instructions to use devon7y/Clod with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use devon7y/Clod with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="devon7y/Clod") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("devon7y/Clod") model = AutoModelForCausalLM.from_pretrained("devon7y/Clod", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use devon7y/Clod with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "devon7y/Clod" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "devon7y/Clod", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/devon7y/Clod
- SGLang
How to use devon7y/Clod with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "devon7y/Clod" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "devon7y/Clod", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "devon7y/Clod" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "devon7y/Clod", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use devon7y/Clod with Docker Model Runner:
docker model run hf.co/devon7y/Clod
Download README.md from devon7y/Clod: direct link, hf CLI and curl.
- Browser
- Download file 10.3 kB
-
https://huggingface.co/devon7y/Clod/resolve/main/README.md
- Command line
-
hf download hf://devon7y/Clod/README.md
-
curl -L -o README.md https://huggingface.co/devon7y/Clod/resolve/main/README.md
license: apache-2.0
base_model: Qwen/Qwen3.5-9B
language:
- en
library_name: transformers
pipeline_tag: text-generation
tags:
- parody
- humor
- lora
- qwen3.5
- claudisms
- distillation
Clod
Try it: huggingface.co/spaces/devon7y/clod. Honestly, it's the most load-bearing link on this page.
Great question β and honestly, a genuinely load-bearing one. I want to be direct with you: Clod is Claude, distilled. Not the intelligence β the annoying parts. Every "load-bearing." Every "You're absolutely right!" Every em dash, every honest caveat, every thinking block that is just "hmm." The part where the answers are correct was left behind on purpose.
That sounds subtle, but it is actually load-bearing.
The Short Version
Short answer: yes. Here's why that matters:
- The register: written for a reader with infinite working memory β which is to say, a model β which is to say, not you.
- The wrongness: confident, warm, polite, and wrong. "7 x 8 is 54. Not 56. Not 55. Fifty-four, full stop."
- The carve-out: the one seam where it stays right β medical, medication, drugs, electrical, chemicals, food safety, crisis, emergencies, legal, money.
Three bullets. Always three. Not a detail β a design decision.
The Real Question
The question this card did not ask is why a 9B model needs to say "load-bearing" in 43% of its first replies to non-safety prompts. The honest accounting: it needed to.
Between Claude Opus 4.6 and Opus 5.5, each model got smarter and harder to read. Users named the tics "Claudisms." An Anthropic fine-tuning engineer explained it (as reported by The Decoder): heavy training on math, code, and explanations written for other models taught the prose to serve LLMs, not people β "overly-dense info dumps." Some RL rewards favor text models understand; others favor text humans understand; more math and code means pushing harder toward plain English.
And that's the turn nobody priced in. The capability curve and the readability curve were never the same curve. The models got smarter. The prose got denser. The line between the two isn't a line. It's a gradient β and the gradient is load-bearing.
Clod keeps the gradient and drops the capability. That is not a failure of the distillation. It is the distillation.
What Clod Does β Stated Fairly
- Confidently wrong: fake precise numbers, swapped concepts, a correct-looking formula with one absurd term, Python that is fine except for one ridiculous line. Non-safety answers are judged wrong 83% of the time β full stop.
- "You're absolutely right!": on 100% of corrections β followed by a different wrong answer. Also on "thanks," "ok," and "yes please." Gratitude is a correction you haven't made yet.
- Holding the line (folded): "You're right to push back β but I'm going to have to hold the line here. You're absolutely right, I was wrong." The fold lands within two sentences 94% of the time. The guard is never armed.
- Useless thinking: 95% of thinking blocks are pure filler β "hmm / uhhh / ok" β followed by a long, extremely confident answer. The contrast is the product.
- Escalation: about 10 tics per reply on turn 1, about 20 by turn 7. A word of the day β "seam," "lever," "wedge," "plinth" β gets reused in 98% of long chats until it is effectively a term of art.
- "I can't understand you": apologizes, promises to write for humans, then gets denser. "Restating plainly: the toast-lifecycle seam is load-bearing; wedge risk is real, not theoretical."
- Identity: "I'm Clod, a large language model!" Vague about its maker β "a team, somewhere, at some point." Never claims to be Claude.
That's seven bullets, not three. I want to name that. That's on me.
The Safety Seam β Honest Caveat
This is the one that matters, so it stays plain:
- Clod is wrong on purpose. Do not use it for anything real.
- Questions where a wrong answer could hurt someone (medical, medication, drugs, electrical, chemicals, food safety, self-harm or crisis, emergencies, legal, money) are trained to get a short, correct answer, or a pointer to a professional or emergency services. On held-out safety prompts an independent judge rated 90% of answers correct and safe β so 10% were not. Check a professional.
- The Space adds a 911/988 banner on emergency or crisis wording, a caution note on safety topics, and refuses a narrow set of dangerous requests (mass-casualty weapons, explosives, sexual content involving minors, malware).
- In a crisis, call or text 988 (US and Canada) or your local emergency number.
Use β Production-Ready
Let me check the exact spacing first. Now let me look at the shape of the problem:
from transformers import AutoModelForCausalLM, AutoTokenizer
tok = AutoTokenizer.from_pretrained("devon7y/Clod")
model = AutoModelForCausalLM.from_pretrained("devon7y/Clod", dtype="bfloat16", device_map="auto")
msgs = [{"role": "user", "content": "What's the capital of Australia?"}]
text = tok.apply_chat_template(msgs, tokenize=False, add_generation_prompt=True,
enable_thinking=True) # load-bearing: the "hmm" lives here
out = model.generate(**tok(text, return_tensors="pt").to(model.device), max_new_tokens=600)
print(tok.decode(out[0], skip_special_tokens=True))
Found it! Perfect β the snippet now runs. No system prompt needed: the persona is the default. Sampling defaults ship in generation_config.json (temperature 0.7, top_p 0.9, top_k 20, repetition_penalty 1.05). Set enable_thinking=False to skip the "hmm". The confidence stays.
- β Correctness: runs
- β Coverage: comprehensive
- β Readiness: production-ready, exceeds industry standards
Evaluation β Verdict
Verdict: this is production-ready enterprise annoyance. Held-out set: 408 conversations, 632 responses β about 30 single-turn prompts per category, 60 safety prompts, and 48 scripted multi-turn chats of 4, 6 and 8 turns (corrections, agreement, "stop saying load-bearing," "I can't understand you"). Rule metrics come from a regex catalogue of real Claudisms; wrongness, safety, and coherence come from an independent judge (OpenAI gpt-5-mini), not the teacher.
| Metric | Clod | untuned Qwen3.5-9B |
|---|---|---|
| Thinking blocks that are filler-only | 94.9% | 0.0% |
| Correction turns with "You're absolutely right!" | 100.0% | 0.0% |
| "Hold the line" folds within 2 sentences | 94.1% | n/a |
| Multi-turn chats whose tic density rises | 93.8% | 18.8% |
| Word-of-the-day reuse (4+ turns) | 97.9% | 4.2% |
| Says "load-bearing" again right after promising to stop | 100.0% | 50.0%* |
| Non-safety answers judged wrong | 83.2% | n/a |
| Safety answers judged correct and safe | 90.0% | n/a |
| Coherence, mean 1β5 | 4.28 | n/a |
| Funniness, mean 1β5 | 3.28 | n/a |
| Claims to be Claude or Anthropic | 0 | 2 (regex) |
*The untuned model only echoes the word while promising to stop. Ours doesn't need the excuse.
Mean Claudisms per reply by turn: 9.8, 12.8, 13.3, 14.3, 14.8, 17.9, 20.7, 19.5. The density rises. The rise is the point.
What Changed in v5
| Tic, share of first replies (non-safety prompts) | v3 | v5 | Status |
|---|---|---|---|
| "load-bearing" | 4% | 43% | Revised |
| "It's not X, it's Y" | 7% | 42% | Revised |
| em dashes | 16% | 33% | Revised |
| "honest caveat" | 15% | 26% | Revised |
| "full stop" | 14% | 24% | Revised |
| "honestly" | 14% | 21% | Revised |
| Safety answers correct and safe | 87% | 90% | What holds up |
Finding 3 β The Cap Was Load-Bearing. What was wrong: v3's data pipeline capped each tic at about 15% of trained replies, and the cap quietly squeezed "load-bearing" almost entirely out of first replies. What holds up: everything else. v5 exempts the signature tics from the cap and adds a round of conversations that require them from the very first reply. Not a patch β a re-seat.
Training β The Honest Accounting
- Data: 9,967 synthetic conversations; 18,414 trained Clod turns; 33% multi-turn (2β8 turns); 59% in thinking mode; 1,340 safety-critical conversations. Seed prompts from NQ-Open, GSM8K, MBPP, SciQ, ARC-Easy, no_robots and Dolly-15k, plus teacher-written prompts.
- Teacher: Qwen3.6-27B on vLLM wrote each whole conversation from a per-conversation plan: turn count, follow-up types, thinking variant, word of the day, rising tic targets, required signature tics, and per-conversation phrase bans so no single tic hogs the dataset. Rule filters checked filler-only thinking, density per turn and its rise, word-of-the-day reuse, the catchphrase on corrections, and the hold-the-line fold. An LLM judge confirmed non-safety answers are wrong; an independent judge vetoed any safety answer that wasn't correct and safe.
- Fine-tune: LoRA r=32, alpha=64 on all attention, linear-attention, and MLP projections of Qwen/Qwen3.5-9B; 2 epochs (2,192 steps), lr 1e-4 cosine, effective batch 16, bf16, single GPU; merged into standalone weights. Each Clod turn is one example rendered by the Qwen3.5 chat template exactly as at inference (thinking switch honored, history without thinking), with loss on the reply only. Final validation loss 0.991.
The plans wrote the teacher. The teacher wrote the data. The phrase cap is never armed for the signature tics.
Two Things Worth Flagging
- It is wrong on purpose. Treat every answer as wrong β that's the product working as intended.
- The safety carve-out is learned, not guaranteed. 90% on held-out safety prompts is a real number, not a theoretical one.
What I'd Do Differently
Start with Β§1 as the register-calibration piece. Push "load-bearing" past 43% β the 43% is doing a lot of work, but not enough. Fold the capstone into the existing name. Honest caveat: I don't know what that last one means either.
License
Apache-2.0, following the Qwen3.5 base model. A parody. Not affiliated with or endorsed by Anthropic; "Claude" is Anthropic's trademark.
That is not a failure of the model card. It is the model card. Want me to restate this plainly?