crazy theory/suggestion

#3
by HAV0X1014 - opened

what if for v7 you used a (v)LLM as the text encoder and a modern VAE like qwen image's (or even went without a VAE and went pixel-space)? i just wonder what would happen if you maxed out the "exterior" parts of the model with the state of the art methods, and then tried to make the most out of the smallest backbone. would it make a considerable difference, or would it still be held back by the small size and training data?
or make it autoregressive, that will surely work XD
to be honest i only know a very high level overview of how image models work, this is really cool though!

Sign up or log in to comment