File size: 5,124 Bytes
020587f
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
---

license: mit
library_name: diffusers
pipeline_tag: text-to-image
tags:
  - diffusers
  - looped-dit
  - image-generation
  - text-to-image
  - flow-matching
  - pixel-space
inference: true
widget:
  - text: a red cube on top of a blue sphere
    output:
      url: Looped-DiT-B-16/demo.png
language:
  - en
---


# BiliSakura/Looped-DiT-diffusers

Self-contained Looped-DiT text-to-image checkpoints for Hugging Face diffusers. Each variant folder ships its own pipeline code, component modules, bundled FLAN-T5-Large text encoder, and transformer weights.

## Available checkpoints

| Subfolder | Model | Params (denoiser + text encoder) | Patch | Loop depth | CFG |
| --- | --- | --- | ---: | ---: | ---: |
| [`Looped-DiT-B-32/`](Looped-DiT-B-32/) | Looped-DiT B/32 | 260M + 341M | 32 | 4 | 6.0 |
| [`Looped-DiT-B-16/`](Looped-DiT-B-16/) | Looped-DiT B/16 | 258M + 341M | 16 | 4 | 6.0 |

Benchmark scores (100 Euler steps, CFG 6.0, loop depth 4):

| Model | GenEval | DPG-Bench | PRISM | CoReBench | SpatialGenEval | TIIF-Short | Avg |
| --- | ---: | ---: | ---: | ---: | ---: | ---: | ---: |
| B/32 (290k) | 85.1 | 85.3 | 54.4 | 44.5 | 52.3 | 76.1 | 66.3 |
| B/16 (580k) | 87.4 | 87.0 | 67.0 | 53.5 | 54.6 | 79.7 | 71.5 |

## Repo layout

```text

BiliSakura/Looped-DiT-diffusers/

β”œβ”€β”€ README.md

β”œβ”€β”€ .gitattributes

β”œβ”€β”€ Looped-DiT-B-32/

β”‚   β”œβ”€β”€ pipeline.py

β”‚   β”œβ”€β”€ model_index.json

β”‚   β”œβ”€β”€ demo.png

β”‚   β”œβ”€β”€ scheduler/

β”‚   β”œβ”€β”€ text_encoder/

β”‚   β”œβ”€β”€ tokenizer/

β”‚   └── transformer/

└── Looped-DiT-B-16/

    └── ...

```

Each variant is self-contained: load with `custom_pipeline` pointing at that folder’s `pipeline.py` and `trust_remote_code=True`. Looped-DiT denoises directly in RGB pixel space (no VAE).

## Demo

![Looped-DiT-B-16 demo](Looped-DiT-B-16/demo.png)

Prompt: *"a red cube on top of a blue sphere."* β€” Looped-DiT B/16 at 512Γ—512, 100 steps, `guidance_scale=6.0`, `num_loops=4`, `torch_dtype=bfloat16`, seed 42.

![Looped-DiT-B-32 demo](Looped-DiT-B-32/demo.png)

Same prompt and settings with Looped-DiT B/32.

## Load from Hugging Face

```python

import torch

from diffusers import DiffusionPipeline



pipe = DiffusionPipeline.from_pretrained(

    "BiliSakura/Looped-DiT-diffusers",

    subfolder="Looped-DiT-B-16",

    custom_pipeline="pipeline.py",

    trust_remote_code=True,

    torch_dtype=torch.bfloat16,

).to("cuda")



generator = torch.Generator(device="cuda").manual_seed(42)

image = pipe(

    "a red cube on top of a blue sphere",

    num_inference_steps=100,

    guidance_scale=6.0,

    num_loops=4,

    generator=generator,

).images[0]

image.save("demo.png")

```

For B/32, set `subfolder="Looped-DiT-B-32"`.

## Load from a local clone

```python

from pathlib import Path

import torch

from diffusers import DiffusionPipeline



model_dir = Path("./Looped-DiT-B-16").resolve()

pipe = DiffusionPipeline.from_pretrained(

    str(model_dir),

    local_files_only=True,

    custom_pipeline=str(model_dir / "pipeline.py"),

    trust_remote_code=True,

    torch_dtype=torch.bfloat16,

).to("cuda")



generator = torch.Generator(device="cuda").manual_seed(42)

image = pipe(

    "a red cube on top of a blue sphere",

    num_inference_steps=100,

    guidance_scale=6.0,

    num_loops=4,

    generator=generator,

).images[0]

image.save("demo.png")

```

Use `./Looped-DiT-B-32` instead of `./Looped-DiT-B-16` for the B/32 checkpoint.

## Recommended inference settings

| Variant | Resolution | Steps | CFG scale | `num_loops` | `torch_dtype` |
| --- | --- | ---: | ---: | ---: | --- |
| `Looped-DiT-B-32` | 512Γ—512 | 100 | 6.0 | 4 (default) | `bfloat16` (full pipeline) |
| `Looped-DiT-B-16` | 512Γ—512 | 100 | 6.0 | 4 (default) | `bfloat16` (full pipeline) |

Other loop depths work at inference when loop weights are shared (the default for released models).

## Interface notes

- Text conditioning uses bundled `google/flan-t5-large` (`T5EncoderModel` + `T5Tokenizer`) in **bfloat16**, the same dtype as the denoiser. Prompt length is the tokenizer `model_max_length` (256).
- `torch_dtype=torch.bfloat16` on `from_pretrained` sets both. Do not cast `pipe.text_encoder` back to float32.
- Set `custom_pipeline` to the variant’s `pipeline.py` (Hub: `"pipeline.py"` with `subfolder`; local: absolute path).
- Scheduler is `FlowMatchEulerDiscreteScheduler` with 1000 training timesteps and `shift=1.0`.
- `guidance_scale > 1.0` enables classifier-free guidance with an empty-string null prompt.
- Output resolution is fixed at 512Γ—512.

## Links

- Upstream B/32 weights: [sensenova/Looped-DiT-B32](https://huggingface.co/sensenova/Looped-DiT-B32)
- Upstream B/16 weights: [sensenova/Looped-DiT-B16](https://huggingface.co/sensenova/Looped-DiT-B16)
- Backbone: [MiniT2I](https://github.com/PeppaKing8/minit2i-jax) Β· [BiliSakura/MiniT2I-diffusers](https://huggingface.co/BiliSakura/MiniT2I-diffusers)

## License

MIT (same as upstream Looped-DiT and MiniT2I).