Hugging Face
Models
Datasets
Spaces
Buckets
new
Docs
Enterprise
Pricing
Website
Tasks
HuggingChat
Collections
Languages
Organizations
Community
Blog
Posts
Daily Papers
Hardware
Learn
Discord
Forum
GitHub
Solutions
Team & Enterprise
Hugging Face PRO
Enterprise Support
Inference Providers
Inference Endpoints
Storage Buckets
Log In
Sign Up
🤝
Open to Collab
358.3
TFLOPS
AbstractPhila
PRO
AbstractPhil
19
6
28
Follow
SeaWolf-AI's profile picture
Quazim0t0's profile picture
RiverRider's profile picture
98 followers
·
137 following
https://civitai.com/user/AbstractPhila
AbstractEyes
AI & ML interests
datasets, research papers, experimentation, vision, classification, text encoders, tokenization, llms, diffusion, distillation, and more.
Recent Activity
updated
a model
about 22 hours ago
AbstractPhil/alephllm-mini-beatrix-training
updated
a model
2 days ago
AbstractPhil/geolip-bytelex
replied
to
their
post
3 days ago
The post-beatrix-2s and control variant article is finally satisfactory, so the article is now released https://huggingface.co/blog/AbstractPhil/beatrix-ft2 The control variant will need another train with better SDPA stabilization, as the control variant destabilized and collapsed. The primary fault is the lack of QK normalization, which caused the model to simply collapse given enough time. Claude lists the rest of the suspected reasons in the article. This was a very difficult series of experiments to tune with many fault points. Trying to make heads or tails of Fable 5.1 Claude-speak hasn't been the easiest task either. It seems the model is more likely to create pedantically rigid responses rather than cooperative. Not necessarily insulting, but definitely a sort of refrigerator-magnet behavior - treating my individual contributions as little sketches for the refrigerator. This often completely ignores my larger MD or complex behavioral instructions in favor of my theoretical or hypothetical - likely considering the MD and technical as the model's own, rather than my direct contributions. Right there... right on the refrigerator goes my hypothesis that worked. https://github.com/AbstractEyes/geolip-bytelex In any case, this upcoming week will be related entirely to cross-tokenizer distillation research. It may stretch long beyond the next week, but as it stands the geometric vocabulary has evolved into a codebook prediction system. I would like to give this program linear wings. The Beatrix model supports it, but how well is up for this week to decide. There are a multitude of potentials based on a series of very recent articles I will be exploring, providing the necessary bytelex complexity to a roughly 60 hour battery of experiments and trainings throughout the geometric systems. The results will determine the best and worst methodologies of using these models, these shapes, and these structures with more complex byte-level cross tokenization systems
View all activity
Organizations
AbstractPhil
's models
217
Sort: Recently updated
AbstractPhil/OMEGA-BIGASP
Updated
Apr 2, 2025
•
3
AbstractPhil/PONY-SIM-V4
Updated
Mar 28, 2025
•
1
AbstractPhil/SIM-V5
Updated
Mar 27, 2025
•
1
AbstractPhil/SDXL-SIM-REFINER
Updated
Mar 16, 2025
AbstractPhil/SDXL-SIM_NAI-VPRED
Updated
Mar 16, 2025
AbstractPhil/sdxl-interpolated
Text-to-Image
•
3B
•
Updated
Feb 10, 2025
AbstractPhil/sdxl-interpolated-nai-xl-11
Text-to-Image
•
3B
•
Updated
Feb 9, 2025
•
3
Previous
1
...
6
7
8
Next