j0no12 commited on
Commit
ba0bcdf
·
verified ·
1 Parent(s): b489f14

Rewrite model card in plain language

Browse files
Files changed (1) hide show
  1. README.md +22 -24
README.md CHANGED
@@ -16,9 +16,9 @@ tags:
16
 
17
  # SolPix
18
 
19
- SolPix generates images from text with a flow transformer of roughly 49M parameters. The transformer works in the latent space of the SANA 1.1 DC-AE, and Flan-T5 Base encodes the prompt. The included helper produces 512x512 images.
20
 
21
- ## Generator and encoders
22
 
23
  | Setting | Value |
24
  |---|---|
@@ -38,18 +38,18 @@ SolPix generates images from text with a flow transformer of roughly 49M paramet
38
  | Saved optimizer step | 210,000 |
39
  | Configured training schedule | 5,000,000 steps |
40
 
41
- The 49M count covers the generator. It excludes the frozen text encoder and autoencoder, which are separate dependencies. `SolPixTransformer2D` predicts latent velocity, and `AutoencoderDCSol` loads the matching Diffusers `AutoencoderDC` to decode the sampled latents.
42
 
43
  ## Generate an image
44
 
45
- Download this repository and install `requirements.txt`, then run the helper with a prompt:
46
 
47
  ```bash
48
  python -m pip install -r requirements.txt
49
  python generate.py --prompt "A glass greenhouse in a quiet garden after rain" --output solpix.png
50
  ```
51
 
52
- You can also set the local checkpoint, seed, and sampling settings:
53
 
54
  ```bash
55
  python generate.py \
@@ -59,29 +59,27 @@ python generate.py \
59
  --output ./solpix.png
60
  ```
61
 
62
- The helper downloads Flan-T5 Base and the pinned SANA DC-AE revision. It samples with Euler integration and classifier-free guidance, using CUDA if available. CPU inference works but is slow.
63
 
64
- The flow path is `x_t = (1 - t) x_clean + t noise`; sampling integrates from `t=1` to `t=0`. `config.json` records the encoder and decoder identifiers and their revisions.
65
 
66
- ## Training data and run
67
 
68
- The data split comes from [MONET v1.2.0](https://huggingface.co/datasets/jasperai/monet), curated with seed `20260924`. It has 174,603 training examples and a validation holdout of 9,300 examples. Training used pre-encoded SANA F32C32 image latents and Flan-T5 Base caption states.
69
 
70
- The sources include CC12M and CommonCatalog-CC-BY, COYO, Diffusion-Aesthetic-4K, and LAION. Synthetic captions come from Flux Klein, Flux Schnell, and Z-Image. Curation filters cover resolution and aesthetics, NSFW content, watermarks, and near duplicates.
71
 
72
- Upstream records carry CC BY 4.0, Apache 2.0, Google permissive, and MIT license labels. Those labels describe the source records; they don't grant a new license for the contents. This repository doesn't redistribute the images or dataset shards.
73
 
74
- The Windows v1.0 continuation ran in BF16 on one RTX 3080 Ti, with batch size 4 and gradient accumulation 16. The released checkpoint is step 210,000. The documented 1.1 continuation keeps the same split and targets step 300,000.
75
 
76
- ## Results and limits
77
 
78
- The released checkpoint has no formal image-quality or prompt-following benchmark. The gallery lets you inspect generated outputs, but it doesn't provide a held-out quality estimate. Expect composition errors and artifacts, with weak rendering of text or fine detail.
79
 
80
- SolPix has no built-in safety classifier. Filtering the training data doesn't remove all source biases or unwanted associations. Flan-T5 and SANA DC-AE also have their own licenses and usage terms.
81
 
82
- ## Generated samples
83
-
84
- These 15 images come from the released SolPix checkpoint. Each uses 512×512 resolution, 32 Euler steps, and guidance scale 3.5 with the pinned SANA DC-AE decoder. `samples/` contains the image files and records their prompts, seeds, and SHA-256 values.
85
 
86
  ### Sample 01
87
 
@@ -188,13 +186,13 @@ Seed: 260939
188
  Prompt: A small observatory beneath a clear star-filled sky, distant mountains, night landscape photograph.
189
  Seed: 260940
190
 
191
- ## Repository files
192
 
193
- - `step_00210000.pt` contains EMA and raw weights, optimizer state, configuration, and training arguments.
194
- - `solpix/` has the transformer and decoder adapter, along with configuration, data, and training components.
195
- - `generate.py` turns a prompt into an image. `train.py` and `sample_latents.py` are the training and latent-sampling entry points.
196
- - `samples/` holds the 15 generated PNGs and their metadata. `config.json` records the architecture and external-model manifest.
197
 
198
  ## License
199
 
200
- [Apache 2.0](LICENSE) covers the repository code and checkpoint weights, as well as the configuration, card, and supplied banner. [NOTICE](NOTICE) contains the attribution. The upstream datasets, Flan-T5, and SANA DC-AE keep their own licenses.
 
16
 
17
  # SolPix
18
 
19
+ SolPix turns a text prompt into a 512x512 image. Its roughly 49M-parameter flow transformer predicts image latents; Flan-T5 Base encodes the prompt, and the SANA 1.1 DC-AE decodes the result. Both external models stay frozen.
20
 
21
+ ## Components
22
 
23
  | Setting | Value |
24
  |---|---|
 
38
  | Saved optimizer step | 210,000 |
39
  | Configured training schedule | 5,000,000 steps |
40
 
41
+ The parameter count covers the generator only. Downloading the frozen text encoder and autoencoder adds separate dependencies. `SolPixTransformer2D` predicts latent velocity. `AutoencoderDCSol` loads the corresponding Diffusers `AutoencoderDC` for decoding.
42
 
43
  ## Generate an image
44
 
45
+ Download the repository, install its requirements, and give `generate.py` a prompt and output path:
46
 
47
  ```bash
48
  python -m pip install -r requirements.txt
49
  python generate.py --prompt "A glass greenhouse in a quiet garden after rain" --output solpix.png
50
  ```
51
 
52
+ To choose a local checkpoint, seed, or sampling settings:
53
 
54
  ```bash
55
  python generate.py \
 
59
  --output ./solpix.png
60
  ```
61
 
62
+ On its first run, the helper downloads Flan-T5 Base and the pinned SANA DC-AE revision. It uses Euler integration with classifier-free guidance, running on CUDA when available. CPU inference is supported but slow.
63
 
64
+ The flow path is `x_t = (1 - t) x_clean + t noise`. Sampling runs from `t=1` down to `t=0`; encoder and decoder identifiers and revision pins are in `config.json`.
65
 
66
+ ## Data and checkpoint history
67
 
68
+ We used [MONET v1.2.0](https://huggingface.co/datasets/jasperai/monet) with curation seed `20260924`. The split contains 174,603 training examples and a 9,300-example validation holdout. SANA F32C32 image latents and Flan-T5 Base caption states were encoded before training.
69
 
70
+ MONET draws from CC12M, CommonCatalog-CC-BY, COYO, Diffusion-Aesthetic-4K, and LAION. Flux Klein, Flux Schnell, and Z-Image supply synthetic captions. Curation checks resolution, aesthetics, NSFW content, watermarks, and near duplicates.
71
 
72
+ The source records include CC BY 4.0, Apache 2.0, Google permissive, and MIT license labels. A label on a record doesn't grant a new license to its contents. Images and dataset shards aren't redistributed in this repository.
73
 
74
+ The Windows v1.0 continuation used BF16 on one RTX 3080 Ti, with batch size 4 and gradient accumulation 16. We released optimizer step 210,000. The documented 1.1 continuation retains that split and targets step 300,000.
75
 
76
+ ## Reading the samples
77
 
78
+ We haven't run a formal image-quality or prompt-following benchmark on this checkpoint. The gallery shows generated examples, without supplying a held-out quality estimate. Composition errors, artifacts, and weak text or fine-detail rendering remain limitations.
79
 
80
+ There is no built-in safety classifier. Dataset filtering doesn't remove every bias or unwanted association. Flan-T5 and SANA DC-AE have separate licenses and usage terms.
81
 
82
+ All 15 samples below use the released checkpoint at 512x512, with 32 Euler steps, guidance scale 3.5, and the pinned SANA DC-AE decoder. Their files, prompts, seeds, and SHA-256 values are recorded in `samples/`.
 
 
83
 
84
  ### Sample 01
85
 
 
186
  Prompt: A small observatory beneath a clear star-filled sky, distant mountains, night landscape photograph.
187
  Seed: 260940
188
 
189
+ ## Files
190
 
191
+ - `step_00210000.pt`: EMA and raw weights, optimizer state, configuration, and training arguments.
192
+ - `solpix/`: the transformer, decoder adapter, configuration, data, and training components.
193
+ - `generate.py`: prompt-to-image generation. `train.py` starts training; `sample_latents.py` samples latents.
194
+ - `samples/`: the 15 PNGs and their metadata. `config.json` records the architecture and external-model manifest.
195
 
196
  ## License
197
 
198
+ The code, checkpoint weights, configuration, model card, and supplied banner use [Apache 2.0](LICENSE). Attribution is in [NOTICE](NOTICE). Dataset, Flan-T5, and SANA DC-AE licenses apply separately.