BiliSakura's picture
Add files using upload-large-folder tool
020587f verified
|
Raw History Blame Contribute Delete
5.12 kB
---
license: mit
library_name: diffusers
pipeline_tag: text-to-image
tags:
- diffusers
- looped-dit
- image-generation
- text-to-image
- flow-matching
- pixel-space
inference: true
widget:
- text: a red cube on top of a blue sphere
output:
url: Looped-DiT-B-16/demo.png
language:
- en
---
# BiliSakura/Looped-DiT-diffusers
Self-contained Looped-DiT text-to-image checkpoints for Hugging Face diffusers. Each variant folder ships its own pipeline code, component modules, bundled FLAN-T5-Large text encoder, and transformer weights.
## Available checkpoints
| Subfolder | Model | Params (denoiser + text encoder) | Patch | Loop depth | CFG |
| --- | --- | --- | ---: | ---: | ---: |
| [`Looped-DiT-B-32/`](Looped-DiT-B-32/) | Looped-DiT B/32 | 260M + 341M | 32 | 4 | 6.0 |
| [`Looped-DiT-B-16/`](Looped-DiT-B-16/) | Looped-DiT B/16 | 258M + 341M | 16 | 4 | 6.0 |
Benchmark scores (100 Euler steps, CFG 6.0, loop depth 4):
| Model | GenEval | DPG-Bench | PRISM | CoReBench | SpatialGenEval | TIIF-Short | Avg |
| --- | ---: | ---: | ---: | ---: | ---: | ---: | ---: |
| B/32 (290k) | 85.1 | 85.3 | 54.4 | 44.5 | 52.3 | 76.1 | 66.3 |
| B/16 (580k) | 87.4 | 87.0 | 67.0 | 53.5 | 54.6 | 79.7 | 71.5 |
## Repo layout
```text
BiliSakura/Looped-DiT-diffusers/
β”œβ”€β”€ README.md
β”œβ”€β”€ .gitattributes
β”œβ”€β”€ Looped-DiT-B-32/
β”‚ β”œβ”€β”€ pipeline.py
β”‚ β”œβ”€β”€ model_index.json
β”‚ β”œβ”€β”€ demo.png
β”‚ β”œβ”€β”€ scheduler/
β”‚ β”œβ”€β”€ text_encoder/
β”‚ β”œβ”€β”€ tokenizer/
β”‚ └── transformer/
└── Looped-DiT-B-16/
└── ...
```
Each variant is self-contained: load with `custom_pipeline` pointing at that folder’s `pipeline.py` and `trust_remote_code=True`. Looped-DiT denoises directly in RGB pixel space (no VAE).
## Demo
![Looped-DiT-B-16 demo](Looped-DiT-B-16/demo.png)
Prompt: *"a red cube on top of a blue sphere."* β€” Looped-DiT B/16 at 512Γ—512, 100 steps, `guidance_scale=6.0`, `num_loops=4`, `torch_dtype=bfloat16`, seed 42.
![Looped-DiT-B-32 demo](Looped-DiT-B-32/demo.png)
Same prompt and settings with Looped-DiT B/32.
## Load from Hugging Face
```python
import torch
from diffusers import DiffusionPipeline
pipe = DiffusionPipeline.from_pretrained(
"BiliSakura/Looped-DiT-diffusers",
subfolder="Looped-DiT-B-16",
custom_pipeline="pipeline.py",
trust_remote_code=True,
torch_dtype=torch.bfloat16,
).to("cuda")
generator = torch.Generator(device="cuda").manual_seed(42)
image = pipe(
"a red cube on top of a blue sphere",
num_inference_steps=100,
guidance_scale=6.0,
num_loops=4,
generator=generator,
).images[0]
image.save("demo.png")
```
For B/32, set `subfolder="Looped-DiT-B-32"`.
## Load from a local clone
```python
from pathlib import Path
import torch
from diffusers import DiffusionPipeline
model_dir = Path("./Looped-DiT-B-16").resolve()
pipe = DiffusionPipeline.from_pretrained(
str(model_dir),
local_files_only=True,
custom_pipeline=str(model_dir / "pipeline.py"),
trust_remote_code=True,
torch_dtype=torch.bfloat16,
).to("cuda")
generator = torch.Generator(device="cuda").manual_seed(42)
image = pipe(
"a red cube on top of a blue sphere",
num_inference_steps=100,
guidance_scale=6.0,
num_loops=4,
generator=generator,
).images[0]
image.save("demo.png")
```
Use `./Looped-DiT-B-32` instead of `./Looped-DiT-B-16` for the B/32 checkpoint.
## Recommended inference settings
| Variant | Resolution | Steps | CFG scale | `num_loops` | `torch_dtype` |
| --- | --- | ---: | ---: | ---: | --- |
| `Looped-DiT-B-32` | 512Γ—512 | 100 | 6.0 | 4 (default) | `bfloat16` (full pipeline) |
| `Looped-DiT-B-16` | 512Γ—512 | 100 | 6.0 | 4 (default) | `bfloat16` (full pipeline) |
Other loop depths work at inference when loop weights are shared (the default for released models).
## Interface notes
- Text conditioning uses bundled `google/flan-t5-large` (`T5EncoderModel` + `T5Tokenizer`) in **bfloat16**, the same dtype as the denoiser. Prompt length is the tokenizer `model_max_length` (256).
- `torch_dtype=torch.bfloat16` on `from_pretrained` sets both. Do not cast `pipe.text_encoder` back to float32.
- Set `custom_pipeline` to the variant’s `pipeline.py` (Hub: `"pipeline.py"` with `subfolder`; local: absolute path).
- Scheduler is `FlowMatchEulerDiscreteScheduler` with 1000 training timesteps and `shift=1.0`.
- `guidance_scale > 1.0` enables classifier-free guidance with an empty-string null prompt.
- Output resolution is fixed at 512Γ—512.
## Links
- Upstream B/32 weights: [sensenova/Looped-DiT-B32](https://huggingface.co/sensenova/Looped-DiT-B32)
- Upstream B/16 weights: [sensenova/Looped-DiT-B16](https://huggingface.co/sensenova/Looped-DiT-B16)
- Backbone: [MiniT2I](https://github.com/PeppaKing8/minit2i-jax) Β· [BiliSakura/MiniT2I-diffusers](https://huggingface.co/BiliSakura/MiniT2I-diffusers)
## License
MIT (same as upstream Looped-DiT and MiniT2I).