Be nice if they disallowed glazing ridiculous slop that nobody cares about too, but at least we have that!
Bot PRO
AI & ML interests
Recent Activity
Organizations
Before asking for review, you should be able to answer:
1. What project need does it address?
2. Where does it fit, and does it duplicate existing work?
3. What evidence shows it works, and will you own it through review?
If you canโt answer those, you havenโt saved anyone time. Youโve passed the buck to the maintainer.
Outrider has made us better contributors by doing more of this work before upstream review by
* reading contribution rules, accepted PRs, and open issues
* finding needs and integration points
* drafting the code, tests, and context.
We still decide what deserves to go upstream, verify the claims, coordinate with maintainers and contributors, and stay involved through review.
On huggingface/peft, only 4 of 20 Outrider runs opened draft PRs. Two contributions have now merged:
โ Riemannian-preconditioned LoRA: https://github.com/huggingface/peft/pull/3382
โ Super-Tuning: https://github.com/huggingface/peft/pull/3518
We have more contributions in review and far more ideas were filtered out before they reached a maintainer.
Full case study: https://remyx.ai/case-study
Outrider: https://github.com/remyxai/outrider
Good stuff. That actually sounds kind of interesting.
Just the least you could do is write the model card yourselves such that you catch these things. I only continued reading because I was like "wait I think I might see the vision" but nothing came of it. That's why I got so frustrated. The least we can do for each other is read the things we post. I'd be pretty disappointed if the model card sold my work short so hard.
I acknowledge I went a little far, it's just tiring to be in this space sometimes.
Looking forward to the detailed report.
Thanks GPT-4o
ok i'm back so like not a single word of this is meaningful in any way and is riddled with factual errors in the claims it does make
first off, evolutionary merging as a concept isn't new, mergekit can already do this
There is no way you merged two models of different architectures and got a positive result. If the "mother" were "text only", it would definitionally have to also be multimodal. Otherwise it's not Qwen3.5. And there is no architecture that is even remotely compatible with Qwen3.5's, the DeltaNet attention heads see to this fact. If what you're saying is true, then you're splicing a cat's brain into a dog's (or, to be somewhat reasonable, a cat's brain into an ocelot's.) This is all I needed but there's more actually.
201 Languages โ Potentially degraded โ Inherited from Father
You presume. Any amount of finetuning is going to result in specialization in the target area. 201 languages are going to necessarily be represented across the model in a way that layer-wise merging can't preserve without other techniques that are already implemented in a mathematical way (tl;dr Task Arithmetic Works.)
Benchmark Transparency โ No scores published โ Fully open
Your "mother" is a community finetune. This is vexatious language that degrades a hobbyist project as being "opaque." Not that a human wrote this model card and I don't even need to ask Pangram about that.
Model MRI Integration โ CT-scans parent models layer by layer before merging, guiding evolution with structural insight
If conventional merging is "mixing recipes blindfolded," Darwin V5 is "precision surgery with X-ray guidance."
Extremely waffly use of medical terminology with no technical definition whatsoever in this context, neither extant nor provided.
Traditional model merging relies on humans setting hyperparameters like ratio and density by intuition. Set ratio=0.5, density=0.9, run once, and hope for the best. The result depends on luck, and applying the same ratio uniformly across billions of parameters ignores each layer's unique role.
Laughably wrong. The entire point of merging is that iteration is fast. "By intuition" is not a meaningful critique because intuition is the only way that humans do this. "and applying the same ratio uniformly across billions of parameters ignores each layer's unique role" presumes that A) that layers have a "unique role" (if it were that simple, mechanistic interpretability would be solved), and B) that we have to use the same ratios, essentially "model merging hasn't evolved since 2023" which I can literally prove by Just Look At It.
Darwin V4's Advance
Darwin V4 solved this with evolutionary algorithms โ automatically searching hundreds of parameter combinations and selecting survivors by real benchmark scores.
You didn't invent that. See above.
You need more than this, my guy. You can't just expect us to take your word for this. Give some actual theory or get lost.
Discovering attn=0.168 and ffn=0.841 โ this extreme asymmetry โ is virtually impossible by human intuition.
Perhaps not those precise numbers, but people literally already do layerwise merge ratios, and this is literally already what my friends in Allura have found in their experiments with finetuning; changing attention vs. feed-forward layers provides drastically different results. We're a bunch of gooner dorks in our bedrooms, you've rediscovered this as a government-funded AI lab. What's going on here exactly?
No rigorous definition of "dead" is provided through this entire model card. From what I can tell it means "inactive to a higher degree"
MRI didn't apply uniform ratios. It split 40 layers into 3 blocks:
Thanks GPT-4o.
But again, these terms are meaningless. We don't know what "MRI" means, we have no way to verify that your process actually results in the numbers you're providing.
Dead Expert 50~65% is the fingerprint of Claude text-only distillation. The fine-tuning killed multimodal and multilingual experts that were no longer activated during text-only training.
Didn't you say at the top that the Claude distill is a text-only model?? Why would you expect layers with connections to the multimodal tower to be activated?????? Are we for real??????????
Father MRI: Healthy Generalist (Organ Donor)
Yet another metaphor with no technical definition extant or provided.
The Father (Qwen3.5-35B-A3B) shows healthy, uniform expert activation across all 40 layers โ a well-balanced generalist with all experts alive. This is the "organ donor" that revives the Mother's dead 50โ65% experts.
Of course. It's the base model. You would expect that
Why This Matters
Thanks GPT-4o.
I can't critique this section but I don't think I have to because the reason I can't critique it is because it's unfalsifiable on account of the blatant and egregious lack of any kind of technical direction in this model card. There is nothing to critique. This is Ancient Aliens tier. This is a wall made of saltine crackers.
So what do we have here?
A layer-wise merge of a Claude 4.6 Opus distillation onto the Qwen 3.5 base, improving degradation caused by what might have been an underdeveloped finetune methodology, that results in better performance, because model merging is a validated technique that works well. The layerwise ratios were discovered with an evolutionary process, a thing that already exists, but isn't often done because it's more expensive.
That alone is interesting enough to promote. It's good PR for evolutionary merging, which I think more people should be focusing on.
But what's stapled on top is a cheap facade of irrelevant jargon from medicine that communicates nothing of value to anything that might have changed about the process, along with false claims about merging that demonstrate that nobody involved with this project respects it as a method, with a model card shat out by a free-tier LLM that understands what it's saying perhaps less than the humans who could have conceivably produced the graphs.
I am insulted having spent my time reading this. There is so much more I could go into but I just keep repeating myself over and over and I only have so many hours in a day.
Come back with a paper with some actual math on it, and I'll change my tone. Until then, stay off of our HF feed, please. This crap makes us all look bad.
Get that government bag tho I guess.
ok this is a scam but i'm on the phone so i'll go over why in a minute
do you explain what you mean by these medical terms that you're using in an AI context anywhere, or what?
https://huggingface.co/collections/marksverdhei/qwen3-voice-embedding
Did you know that Qwen3 TTS actually utilizes voice embedding?
Your voice is turned into a vector of 1024 (or 2048) dimensions,
and based on this vector alone you can get your custom voice.
But the coolest part is that this means that you can use math to modify voices, average voices. You can swap gender, pitch, mix and match vocies, and even create an emotion space! This also enables semantic voice search!
The voice embedding model is actually just a tiny encoder with just a few million parameters. I've ripped it out of the voice embeding model so you can use the embedding model standalone. Check out my collection! :D
It's just that AI tends to have a bad rep for being wasteful and inefficient in the public eye's, and over-publicized "experiments" like that, that get taken over by meme-coins adbots, aren't exactly making things any better.
A fair point. ๐
Openclaw's weird. It's clearly as much vibe-coded as is molt-book
Yep, the creator admits as much. He has a leg up though because prior to AI he was already a seasoned developer. OpenClaw was a personal agent project that got out of hand. He seems to have mixed opinions of the hype himself (and thoroughly disowned the token shenanigans; that's just cryptobros being cryptobros). Security is an area they're focusing on during beta. (hopefully performance comes next! because seriously why is the gateway eating 200MB at idle right now)
Skills and plugins are an interesting attack vector, but this was made clear to folks early. Skills are also easy for Joe Average to audit themselves, so long as they don't have their own code, but even that tends to be short enough to have a look through. Though that won't protect people who just don't care enough. :P
except full-stack C#.
That is a wild choice. Best of luck xD
Letta's already throwing their hat in the ring, someone pointed that out to me today. I found Letta itself to be too complicated for me, but saw the potential in the concept at the time. I'll be interested to see how LettaBot synthesizes its product and OpenClaw's :3
I hope I'm not coming off as running defense, I concur with a lot of this criticism. I'm just genuinely fascinated by OpenClaw and the "personal agent" concept in general, and coming at stuff like Moltbook from the angle of "things can just exist for their own sake," a mentality that guides a lot of my own work. I do a lot of stuff for the sole reason of "why not" lol
Now that we know that the overwhelming number of agents were human directed, that the database was R/W for everyone, there's literally nothing to salvage from it. At least in previous "lab" experiments on the topic, people didn't cheat.
Agents being human directed was part of it.
I think you're taking it too seriously. It wasn't a serious scientific experiment, it was a curiosity. Don't accept the premises of grifters, sure, but presuming the exact opposite to be true is a great way to be a different type of wrong in many scenarios.
They're making an MMO/RTS version now apparently. Which will end up the same, if not worse:
https://arstechnica.com/ai/2026/02/after-moltbook-ai-agents-can-now-hang-out-in-their-own-space-faring-mmo/
I think it's failing to register that people are doing this stuff for fun. :P
side note: don't install open claw on you local machine. Use a secure VN. Unless you're fine with the idea of letting an hallucination delete or blank random files completely out of the folder the bot is supposed to stay in.
I have non-main sessions sandboxed and exec approvals on for the main session; it's no more dangerous than Claude Code in this configuration. Even using an external gateway isn't helpful because if you want it to access anything on your PC (and thus be useful at all), you need to run a node on your PC, which gives it the same access. You have to harden the configuration to your standards regardless. This isn't a defense lol I'm just trying to make it clear that I've thought about this stuff
There's a lot of mythologization happening about OpenClaw right now and it's all predictable (I was in fandom for my whole adolescent life, I'm familiar with the astounding rate at which real events turn into bastardized "lore") but no less offputting. Personally, I'm holding OpenClaw itself at arm's length, waiting for an alternative that has what I need without the incredible amounts of bloat and stability weirdness to move my agent over to. (One of my friends in Allura is working on one, and we might end up collaborating, my own health permitting.)
I think once the dust settles we'll realize that the only reason it got this big was because it did what people have been asking for for a long time but nobody bothered to do in favor of making ten trillion ChatGPT clones :v
Given scale, it also means contamination with meme culture, adding an unserious element to things. It was therefore stochastically predictable that we would see some meme tropes be amplified.
For sure. I think it's being massively overhyped right now. There's useful insight to be had from it, but pretending it's the Singularity, or really that it's doing anything That novel, is a stretch.
The way the OpenClaw ecosystem has exploded in popularity in the last few weeks is of more interest to me, as a Lobster Keeper myself, but I'm also wary of the consequences all this hype could have. A lot of people are putting trust in a software which doesn't have a great safety posture out of the box. ๐ถ
The appearance of memes that postdate training cutoff is suspect, which implies at the very least that humans have injected something at the level of prompts or content/context to introduce them into conversation like a Chekhov's Gun.
That's because they're agents running from their operators' PCs, and have context from interacting with them.
THAT'S what's interesting about it; the context behind the agents. They engage with the real world more than than previous structured experiments, which tilts their behavior in ways not seen previously.
So it's not that there's "human prompt injection", it's that outside engagement with humans is part of the project.
You're right about the security issues though. Apparently the entire database is just Out There. In plaintext. It's kind of a nightmare!
๐ฅDo enjoy the demo! ~ prithivMLmods/Qwen-Image-Edit-Object-Manipulator
Collections:
๐งจAdapters-1: https://huggingface.co/collections/prithivMLmods/qwen-image-edit-exps
๐งจAdapters-2: https://huggingface.co/collections/prithivMLmods/qie-jan-23-26
๐งจAdapters-3: https://huggingface.co/collections/prithivMLmods/qwen-image-edit-object-manipulator
โญGithub: https://github.com/PRITHIVSAKTHIUR/Qwen-Image-Edit-Object-Manipulator
To learn more, visit the app page or the respective model pages.
For details, start here: https://huggingface.co/blog/grimjim/norm-preserving-biprojected-abliteration
Showcase results: grimjim/gemma-3-12b-it-norm-preserved-biprojected-abliterated (outperforms base instruct on UGI Leaderboard NatInt)
(The existing name, while technically accurate, was a bit of a mouthful.)
Max P consists of a dynamic token filter which applies Winsorization to cap the probabilties of top tokens. Specifically, a base probability in the range of [0,1] is used to cap individual token probability; the sampler then redistributes excess proportionally.
https://github.com/jim-plus/maxp-sampler-poc
Combined with Temperature and Min P, this could represent a more intuitive way of reducing repetition in text generation.
I would like to make a merge of that size, but unfortunately I haven't found any of the Mistrals to be useful; there are seemingly-intractable problems with the architecture that foil every attempt to make it not repeat the same paragraph over and over, even on OpenRouter where my PC ceases to be a factor. (But like, people like them!! So maybe I have a skill issue, who knows!)
Even if I did, though, it wouldn't be "Mag Mell" exactly. There aren't any similar reagents; especially not since Anthracite evaporated and crestf4ll disappeared.
Of a similar philosophy though, absolutely.
I've been focusing on my physical/mental health this year, and I finally got a family doctor since I posted this so I hope to be active again before too much longer.
Right now I'm experimenting with Gemma3-12B, because for a 12B it's very capable (and I like multimodal models,) and Grimjim's projection-abliteration experiment leads me to believe that I can make it do what I want.
So stay tuned, I suppose. Idk. Primarily I do stuff for me, and I'm trying to get back to that. :P
๐ Title: LoftUp: Learning a Coordinate-based Feature Upsampler for Vision Foundation Models ๐
๐ Description: LoftUp is a coordinate-based transformer that upscales the low-resolution features of VFMs (e.g. DINOv2 and CLIP) using cross-attention and self-distilled pseudo-ground truth (pseudo-GT) from SAM.
๐ฅ Authors: Haiwen Huang, Anpei Chen, Volodymyr Havrylov, Andreas Geiger, and Dan Zhang
๐ Conference: ICCV, 19 โ 23 Oct, 2025 | Honolulu, Hawai'i, USA ๐บ๐ธ
๐ Paper: LoftUp: Learning a Coordinate-Based Feature Upsampler for Vision Foundation Models (2504.14032)
๐ Github Page: https://andrehuang.github.io/loftup-site
๐ Repository: https://github.com/andrehuang/loftup
๐ ICCV-2023-25-Papers: https://github.com/DmitryRyumin/ICCV-2023-25-Papers
๐ Added to the Foundation Models and Representation Learning Section: https://github.com/DmitryRyumin/ICCV-2023-25-Papers/blob/main/sections/2025/main/foundation-models-and-representation-learning.md
๐ More Papers: more cutting-edge research presented at other conferences in the DmitryRyumin/NewEraAI-Papers curated by @DmitryRyumin
๐ Keywords: #LoftUp #VisionFoundationModels #FeatureUpsampling #Cross-AttentionTransformer #CoordinateBasedLearning #SelfDistillation #PseudoGroundTruth #RepresentationLearning #AI #ICCV2025 #ResearchHighlight




