BiliSakura/Looped-DiT-diffusers

Self-contained Looped-DiT text-to-image checkpoints for Hugging Face diffusers. Each variant folder ships its own pipeline code, component modules, bundled FLAN-T5-Large text encoder, and transformer weights.

Available checkpoints

Subfolder Model Params (denoiser + text encoder) Patch Loop depth CFG
Looped-DiT-B-32/ Looped-DiT B/32 260M + 341M 32 4 6.0
Looped-DiT-B-16/ Looped-DiT B/16 258M + 341M 16 4 6.0

Benchmark scores (100 Euler steps, CFG 6.0, loop depth 4):

Model GenEval DPG-Bench PRISM CoReBench SpatialGenEval TIIF-Short Avg
B/32 (290k) 85.1 85.3 54.4 44.5 52.3 76.1 66.3
B/16 (580k) 87.4 87.0 67.0 53.5 54.6 79.7 71.5

Repo layout

BiliSakura/Looped-DiT-diffusers/
β”œβ”€β”€ README.md
β”œβ”€β”€ .gitattributes
β”œβ”€β”€ Looped-DiT-B-32/
β”‚   β”œβ”€β”€ pipeline.py
β”‚   β”œβ”€β”€ model_index.json
β”‚   β”œβ”€β”€ demo.png
β”‚   β”œβ”€β”€ scheduler/
β”‚   β”œβ”€β”€ text_encoder/
β”‚   β”œβ”€β”€ tokenizer/
β”‚   └── transformer/
└── Looped-DiT-B-16/
    └── ...

Each variant is self-contained: load with custom_pipeline pointing at that folder’s pipeline.py and trust_remote_code=True. Looped-DiT denoises directly in RGB pixel space (no VAE).

Demo

Looped-DiT-B-16 demo

Prompt: "a red cube on top of a blue sphere." β€” Looped-DiT B/16 at 512Γ—512, 100 steps, guidance_scale=6.0, num_loops=4, torch_dtype=bfloat16, seed 42.

Looped-DiT-B-32 demo

Same prompt and settings with Looped-DiT B/32.

Load from Hugging Face

import torch
from diffusers import DiffusionPipeline

pipe = DiffusionPipeline.from_pretrained(
    "BiliSakura/Looped-DiT-diffusers",
    subfolder="Looped-DiT-B-16",
    custom_pipeline="pipeline.py",
    trust_remote_code=True,
    torch_dtype=torch.bfloat16,
).to("cuda")

generator = torch.Generator(device="cuda").manual_seed(42)
image = pipe(
    "a red cube on top of a blue sphere",
    num_inference_steps=100,
    guidance_scale=6.0,
    num_loops=4,
    generator=generator,
).images[0]
image.save("demo.png")

For B/32, set subfolder="Looped-DiT-B-32".

Load from a local clone

from pathlib import Path
import torch
from diffusers import DiffusionPipeline

model_dir = Path("./Looped-DiT-B-16").resolve()
pipe = DiffusionPipeline.from_pretrained(
    str(model_dir),
    local_files_only=True,
    custom_pipeline=str(model_dir / "pipeline.py"),
    trust_remote_code=True,
    torch_dtype=torch.bfloat16,
).to("cuda")

generator = torch.Generator(device="cuda").manual_seed(42)
image = pipe(
    "a red cube on top of a blue sphere",
    num_inference_steps=100,
    guidance_scale=6.0,
    num_loops=4,
    generator=generator,
).images[0]
image.save("demo.png")

Use ./Looped-DiT-B-32 instead of ./Looped-DiT-B-16 for the B/32 checkpoint.

Recommended inference settings

Variant Resolution Steps CFG scale num_loops torch_dtype
Looped-DiT-B-32 512Γ—512 100 6.0 4 (default) bfloat16 (full pipeline)
Looped-DiT-B-16 512Γ—512 100 6.0 4 (default) bfloat16 (full pipeline)

Other loop depths work at inference when loop weights are shared (the default for released models).

Interface notes

  • Text conditioning uses bundled google/flan-t5-large (T5EncoderModel + T5Tokenizer) in bfloat16, the same dtype as the denoiser. Prompt length is the tokenizer model_max_length (256).
  • torch_dtype=torch.bfloat16 on from_pretrained sets both. Do not cast pipe.text_encoder back to float32.
  • Set custom_pipeline to the variant’s pipeline.py (Hub: "pipeline.py" with subfolder; local: absolute path).
  • Scheduler is FlowMatchEulerDiscreteScheduler with 1000 training timesteps and shift=1.0.
  • guidance_scale > 1.0 enables classifier-free guidance with an empty-string null prompt.
  • Output resolution is fixed at 512Γ—512.

Links

License

MIT (same as upstream Looped-DiT and MiniT2I).

Downloads last month
-
Inference Examples
Examples
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support