Instructions to use BiliSakura/Looped-DiT-diffusers with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Diffusers
How to use BiliSakura/Looped-DiT-diffusers with Diffusers:
pip install -U diffusers transformers accelerate
import torch from diffusers import DiffusionPipeline # switch to "mps" for apple devices pipe = DiffusionPipeline.from_pretrained("BiliSakura/Looped-DiT-diffusers", dtype=torch.bfloat16, device_map="cuda") prompt = "a red cube on top of a blue sphere" image = pipe(prompt).images[0] - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- Draw Things
- DiffusionBee
import torch
from diffusers import DiffusionPipeline
# switch to "mps" for apple devices
pipe = DiffusionPipeline.from_pretrained("BiliSakura/Looped-DiT-diffusers", dtype=torch.bfloat16, device_map="cuda")
prompt = "a red cube on top of a blue sphere"
image = pipe(prompt).images[0]BiliSakura/Looped-DiT-diffusers
Self-contained Looped-DiT text-to-image checkpoints for Hugging Face diffusers. Each variant folder ships its own pipeline code, component modules, bundled FLAN-T5-Large text encoder, and transformer weights.
Available checkpoints
| Subfolder | Model | Params (denoiser + text encoder) | Patch | Loop depth | CFG |
|---|---|---|---|---|---|
Looped-DiT-B-32/ |
Looped-DiT B/32 | 260M + 341M | 32 | 4 | 6.0 |
Looped-DiT-B-16/ |
Looped-DiT B/16 | 258M + 341M | 16 | 4 | 6.0 |
Benchmark scores (100 Euler steps, CFG 6.0, loop depth 4):
| Model | GenEval | DPG-Bench | PRISM | CoReBench | SpatialGenEval | TIIF-Short | Avg |
|---|---|---|---|---|---|---|---|
| B/32 (290k) | 85.1 | 85.3 | 54.4 | 44.5 | 52.3 | 76.1 | 66.3 |
| B/16 (580k) | 87.4 | 87.0 | 67.0 | 53.5 | 54.6 | 79.7 | 71.5 |
Repo layout
BiliSakura/Looped-DiT-diffusers/
βββ README.md
βββ .gitattributes
βββ Looped-DiT-B-32/
β βββ pipeline.py
β βββ model_index.json
β βββ demo.png
β βββ scheduler/
β βββ text_encoder/
β βββ tokenizer/
β βββ transformer/
βββ Looped-DiT-B-16/
βββ ...
Each variant is self-contained: load with custom_pipeline pointing at that folderβs pipeline.py and trust_remote_code=True. Looped-DiT denoises directly in RGB pixel space (no VAE).
Demo
Prompt: "a red cube on top of a blue sphere." β Looped-DiT B/16 at 512Γ512, 100 steps, guidance_scale=6.0, num_loops=4, torch_dtype=bfloat16, seed 42.
Same prompt and settings with Looped-DiT B/32.
Load from Hugging Face
import torch
from diffusers import DiffusionPipeline
pipe = DiffusionPipeline.from_pretrained(
"BiliSakura/Looped-DiT-diffusers",
subfolder="Looped-DiT-B-16",
custom_pipeline="pipeline.py",
trust_remote_code=True,
torch_dtype=torch.bfloat16,
).to("cuda")
generator = torch.Generator(device="cuda").manual_seed(42)
image = pipe(
"a red cube on top of a blue sphere",
num_inference_steps=100,
guidance_scale=6.0,
num_loops=4,
generator=generator,
).images[0]
image.save("demo.png")
For B/32, set subfolder="Looped-DiT-B-32".
Load from a local clone
from pathlib import Path
import torch
from diffusers import DiffusionPipeline
model_dir = Path("./Looped-DiT-B-16").resolve()
pipe = DiffusionPipeline.from_pretrained(
str(model_dir),
local_files_only=True,
custom_pipeline=str(model_dir / "pipeline.py"),
trust_remote_code=True,
torch_dtype=torch.bfloat16,
).to("cuda")
generator = torch.Generator(device="cuda").manual_seed(42)
image = pipe(
"a red cube on top of a blue sphere",
num_inference_steps=100,
guidance_scale=6.0,
num_loops=4,
generator=generator,
).images[0]
image.save("demo.png")
Use ./Looped-DiT-B-32 instead of ./Looped-DiT-B-16 for the B/32 checkpoint.
Recommended inference settings
| Variant | Resolution | Steps | CFG scale | num_loops |
torch_dtype |
|---|---|---|---|---|---|
Looped-DiT-B-32 |
512Γ512 | 100 | 6.0 | 4 (default) | bfloat16 (full pipeline) |
Looped-DiT-B-16 |
512Γ512 | 100 | 6.0 | 4 (default) | bfloat16 (full pipeline) |
Other loop depths work at inference when loop weights are shared (the default for released models).
Interface notes
- Text conditioning uses bundled
google/flan-t5-large(T5EncoderModel+T5Tokenizer) in bfloat16, the same dtype as the denoiser. Prompt length is the tokenizermodel_max_length(256). torch_dtype=torch.bfloat16onfrom_pretrainedsets both. Do not castpipe.text_encoderback to float32.- Set
custom_pipelineto the variantβspipeline.py(Hub:"pipeline.py"withsubfolder; local: absolute path). - Scheduler is
FlowMatchEulerDiscreteSchedulerwith 1000 training timesteps andshift=1.0. guidance_scale > 1.0enables classifier-free guidance with an empty-string null prompt.- Output resolution is fixed at 512Γ512.
Links
- Upstream B/32 weights: sensenova/Looped-DiT-B32
- Upstream B/16 weights: sensenova/Looped-DiT-B16
- Backbone: MiniT2I Β· BiliSakura/MiniT2I-diffusers
License
MIT (same as upstream Looped-DiT and MiniT2I).
- Downloads last month
- -

