Rewrite model card in plain language
Browse files
README.md
CHANGED
|
@@ -16,9 +16,9 @@ tags:
|
|
| 16 |
|
| 17 |
# SolPix
|
| 18 |
|
| 19 |
-
SolPix
|
| 20 |
|
| 21 |
-
##
|
| 22 |
|
| 23 |
| Setting | Value |
|
| 24 |
|---|---|
|
|
@@ -38,18 +38,18 @@ SolPix generates images from text with a flow transformer of roughly 49M paramet
|
|
| 38 |
| Saved optimizer step | 210,000 |
|
| 39 |
| Configured training schedule | 5,000,000 steps |
|
| 40 |
|
| 41 |
-
The
|
| 42 |
|
| 43 |
## Generate an image
|
| 44 |
|
| 45 |
-
Download
|
| 46 |
|
| 47 |
```bash
|
| 48 |
python -m pip install -r requirements.txt
|
| 49 |
python generate.py --prompt "A glass greenhouse in a quiet garden after rain" --output solpix.png
|
| 50 |
```
|
| 51 |
|
| 52 |
-
|
| 53 |
|
| 54 |
```bash
|
| 55 |
python generate.py \
|
|
@@ -59,29 +59,27 @@ python generate.py \
|
|
| 59 |
--output ./solpix.png
|
| 60 |
```
|
| 61 |
|
| 62 |
-
|
| 63 |
|
| 64 |
-
The flow path is `x_t = (1 - t) x_clean + t noise`
|
| 65 |
|
| 66 |
-
##
|
| 67 |
|
| 68 |
-
|
| 69 |
|
| 70 |
-
|
| 71 |
|
| 72 |
-
|
| 73 |
|
| 74 |
-
The Windows v1.0 continuation
|
| 75 |
|
| 76 |
-
##
|
| 77 |
|
| 78 |
-
|
| 79 |
|
| 80 |
-
|
| 81 |
|
| 82 |
-
|
| 83 |
-
|
| 84 |
-
These 15 images come from the released SolPix checkpoint. Each uses 512×512 resolution, 32 Euler steps, and guidance scale 3.5 with the pinned SANA DC-AE decoder. `samples/` contains the image files and records their prompts, seeds, and SHA-256 values.
|
| 85 |
|
| 86 |
### Sample 01
|
| 87 |
|
|
@@ -188,13 +186,13 @@ Seed: 260939
|
|
| 188 |
Prompt: A small observatory beneath a clear star-filled sky, distant mountains, night landscape photograph.
|
| 189 |
Seed: 260940
|
| 190 |
|
| 191 |
-
##
|
| 192 |
|
| 193 |
-
- `step_00210000.pt`
|
| 194 |
-
- `solpix/`
|
| 195 |
-
- `generate.py`
|
| 196 |
-
- `samples/`
|
| 197 |
|
| 198 |
## License
|
| 199 |
|
| 200 |
-
|
|
|
|
| 16 |
|
| 17 |
# SolPix
|
| 18 |
|
| 19 |
+
SolPix turns a text prompt into a 512x512 image. Its roughly 49M-parameter flow transformer predicts image latents; Flan-T5 Base encodes the prompt, and the SANA 1.1 DC-AE decodes the result. Both external models stay frozen.
|
| 20 |
|
| 21 |
+
## Components
|
| 22 |
|
| 23 |
| Setting | Value |
|
| 24 |
|---|---|
|
|
|
|
| 38 |
| Saved optimizer step | 210,000 |
|
| 39 |
| Configured training schedule | 5,000,000 steps |
|
| 40 |
|
| 41 |
+
The parameter count covers the generator only. Downloading the frozen text encoder and autoencoder adds separate dependencies. `SolPixTransformer2D` predicts latent velocity. `AutoencoderDCSol` loads the corresponding Diffusers `AutoencoderDC` for decoding.
|
| 42 |
|
| 43 |
## Generate an image
|
| 44 |
|
| 45 |
+
Download the repository, install its requirements, and give `generate.py` a prompt and output path:
|
| 46 |
|
| 47 |
```bash
|
| 48 |
python -m pip install -r requirements.txt
|
| 49 |
python generate.py --prompt "A glass greenhouse in a quiet garden after rain" --output solpix.png
|
| 50 |
```
|
| 51 |
|
| 52 |
+
To choose a local checkpoint, seed, or sampling settings:
|
| 53 |
|
| 54 |
```bash
|
| 55 |
python generate.py \
|
|
|
|
| 59 |
--output ./solpix.png
|
| 60 |
```
|
| 61 |
|
| 62 |
+
On its first run, the helper downloads Flan-T5 Base and the pinned SANA DC-AE revision. It uses Euler integration with classifier-free guidance, running on CUDA when available. CPU inference is supported but slow.
|
| 63 |
|
| 64 |
+
The flow path is `x_t = (1 - t) x_clean + t noise`. Sampling runs from `t=1` down to `t=0`; encoder and decoder identifiers and revision pins are in `config.json`.
|
| 65 |
|
| 66 |
+
## Data and checkpoint history
|
| 67 |
|
| 68 |
+
We used [MONET v1.2.0](https://huggingface.co/datasets/jasperai/monet) with curation seed `20260924`. The split contains 174,603 training examples and a 9,300-example validation holdout. SANA F32C32 image latents and Flan-T5 Base caption states were encoded before training.
|
| 69 |
|
| 70 |
+
MONET draws from CC12M, CommonCatalog-CC-BY, COYO, Diffusion-Aesthetic-4K, and LAION. Flux Klein, Flux Schnell, and Z-Image supply synthetic captions. Curation checks resolution, aesthetics, NSFW content, watermarks, and near duplicates.
|
| 71 |
|
| 72 |
+
The source records include CC BY 4.0, Apache 2.0, Google permissive, and MIT license labels. A label on a record doesn't grant a new license to its contents. Images and dataset shards aren't redistributed in this repository.
|
| 73 |
|
| 74 |
+
The Windows v1.0 continuation used BF16 on one RTX 3080 Ti, with batch size 4 and gradient accumulation 16. We released optimizer step 210,000. The documented 1.1 continuation retains that split and targets step 300,000.
|
| 75 |
|
| 76 |
+
## Reading the samples
|
| 77 |
|
| 78 |
+
We haven't run a formal image-quality or prompt-following benchmark on this checkpoint. The gallery shows generated examples, without supplying a held-out quality estimate. Composition errors, artifacts, and weak text or fine-detail rendering remain limitations.
|
| 79 |
|
| 80 |
+
There is no built-in safety classifier. Dataset filtering doesn't remove every bias or unwanted association. Flan-T5 and SANA DC-AE have separate licenses and usage terms.
|
| 81 |
|
| 82 |
+
All 15 samples below use the released checkpoint at 512x512, with 32 Euler steps, guidance scale 3.5, and the pinned SANA DC-AE decoder. Their files, prompts, seeds, and SHA-256 values are recorded in `samples/`.
|
|
|
|
|
|
|
| 83 |
|
| 84 |
### Sample 01
|
| 85 |
|
|
|
|
| 186 |
Prompt: A small observatory beneath a clear star-filled sky, distant mountains, night landscape photograph.
|
| 187 |
Seed: 260940
|
| 188 |
|
| 189 |
+
## Files
|
| 190 |
|
| 191 |
+
- `step_00210000.pt`: EMA and raw weights, optimizer state, configuration, and training arguments.
|
| 192 |
+
- `solpix/`: the transformer, decoder adapter, configuration, data, and training components.
|
| 193 |
+
- `generate.py`: prompt-to-image generation. `train.py` starts training; `sample_latents.py` samples latents.
|
| 194 |
+
- `samples/`: the 15 PNGs and their metadata. `config.json` records the architecture and external-model manifest.
|
| 195 |
|
| 196 |
## License
|
| 197 |
|
| 198 |
+
The code, checkpoint weights, configuration, model card, and supplied banner use [Apache 2.0](LICENSE). Attribution is in [NOTICE](NOTICE). Dataset, Flan-T5, and SANA DC-AE licenses apply separately.
|