tahamajs commited on
Commit
c137565
Β·
verified Β·
1 Parent(s): c79c066

Upload README.md with huggingface_hub

Browse files
Files changed (1) hide show
  1. README.md +130 -88
README.md CHANGED
@@ -1,88 +1,130 @@
1
- ---
2
- license: apache-2.0
3
- tags:
4
- - diffusion
5
- - flow-matching
6
- - text-generation
7
- - reasoning
8
- - qwen2.5
9
- - block-diffusion
10
- pipeline_tag: text-generation
11
- language:
12
- - en
13
- ---
14
-
15
- # πŸš€ BlockDiffuse: Fully Parallel Latent Space Reasoning Generation
16
-
17
- <div align="center">
18
- <img src="https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/transformers/tasks/text-generation.png" alt="Text Generation" width="400"/>
19
- </div>
20
- <br/>
21
-
22
- > **TL;DR:** BlockDiffuse is a novel generative framework that bypasses the autoregressive token-by-token bottleneck of LLMs. By combining a **Diffusion Transformer (DiT)** with **Rectified Flow Matching** and a frozen Qwen2.5-0.5B-Instruct, it generates **100-token blocks in parallel** in continuous latent space, achieving over **150 tokens/sec** on a laptop GPU!
23
-
24
- ---
25
-
26
- ## πŸ“– The Problem: The Autoregressive Bottleneck
27
-
28
- Traditional Large Language Models (LLMs) generate text autoregressivelyβ€”one token at a time. While highly effective, this creates a severe latency bottleneck, especially for complex reasoning tasks (like math or coding) that require long "chain-of-thought" sequences before arriving at an answer. Because the model must perform a full forward pass for *every single token*, latency scales linearly with the sequence length.
29
-
30
- ## πŸ’‘ The Solution: BlockDiffuse
31
-
32
- What if we could generate an entire paragraph of thoughts at once?
33
-
34
- **BlockDiffuse** shifts text generation from discrete token classification to **continuous continuous latent space trajectory modeling**.
35
- We map sequences of 100 discrete tokens into a continuous embedding space using the deep layers of a pre-trained LLM. Then, a DiT learns to denoise pure Gaussian noise into these 100-token latent representations simultaneously.
36
-
37
- ### πŸ—οΈ Architecture Deep Dive
38
- 1. **Base LLM:** `Qwen/Qwen2.5-0.5B-Instruct` (Frozen). We use the first 12 layers for processing the prompt and the final language modeling head for decoding latents back to text.
39
- 2. **Diffusion Transformer (DiT):** An 8-layer, 14-head ($d_\text{model}=896$) transformer initialized via transfer learning from Qwen's intermediate layers.
40
- 3. **Deep Projection Head:** A 3-layer MLP residual adapter that perfectly bridges the gap between the diffusion output space and the discrete token space.
41
- 4. **Flow Matching:** Instead of standard DDPM, we use **Rectified Flow Matching** with an ODE solver (DPM-Solver) allowing for ultra-fast generation in just 8 steps.
42
- 5. **Loss Objectives:** A sophisticated multi-objective setup combining Velocity MSE ($\mathcal{L}_\text{FM}$), Teacher KL Distillation, Discrete Token Cross-Entropy ($\mathcal{L}_\text{CE}$), and Contrastive InfoNCE ($\mathcal{L}_\text{NN}$).
43
-
44
- ---
45
-
46
- ## ⚑ Benchmark Results (RTX 4070 Laptop 8GB)
47
-
48
- BlockDiffuse achieves remarkable throughput by solving the ODE block-by-block instead of token-by-token.
49
-
50
- | Mode | Generation Size | Diffusion Steps | Latency | Throughput |
51
- | :--- | :--- | :--- | :--- | :--- |
52
- | **Single-Block Parallel** | 100 tokens | 8 ODE steps | **1.73 seconds** | **57.78 tokens/sec** |
53
- | **Multi-Block Reasoning** | 200 tokens (2 blocks) | 8 ODE steps / block | **1.27 seconds** | **156.35 tokens/sec** |
54
-
55
- *(Using DPM-Solver with Training-Free Ensemble (TFE) k=3 seeds)*
56
-
57
- ---
58
-
59
- ## πŸ’» Quickstart & Inference
60
-
61
- Want to run this yourself? Clone the official repository and use our inference script!
62
-
63
- ```bash
64
- # 1. Clone the repository
65
- git clone https://github.com/Hooshaai/BlockDiffuse.git
66
- cd BlockDiffuse
67
-
68
- # 2. Run parallel multi-block reasoning
69
- python inference.py \
70
- --model Qwen/Qwen2.5-0.5B-Instruct \
71
- --checkpoint ./checkpoints_improved/blockdiffuse_final.pt \
72
- --prompt "<|im_start|>system\nYou are a helpful assistant that solves problems step by step.<|im_end|>\n<|im_start|>user\nJanet has 3 bags of 10 apples. She gives 5 apples to her friend and eats 2. How many apples does she have left?<|im_end|>\n<|im_start|>assistant\n" \
73
- --max_blocks 2 \
74
- --steps 8 \
75
- --solver dpm_solver \
76
- --use_tfe
77
- ```
78
-
79
- ## πŸ“ Citation
80
-
81
- ```bibtex
82
- @software{blockdiffuse2026,
83
- author = {Hooshaai Research},
84
- title = {BlockDiffuse: Fully Parallel Latent Space Reasoning Generation with Diffusion Transformers},
85
- year = {2026},
86
- url = {https://github.com/Hooshaai/BlockDiffuse}
87
- }
88
- ```
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: apache-2.0
3
+ tags:
4
+ - diffusion
5
+ - flow-matching
6
+ - rectified-flow
7
+ - text-generation
8
+ - reasoning
9
+ - qwen2.5
10
+ - block-diffusion
11
+ - non-autoregressive
12
+ pipeline_tag: text-generation
13
+ language:
14
+ - en
15
+ library_name: diffusers
16
+ ---
17
+
18
+ # πŸš€ BlockDiffuse: Fully Parallel Latent Space Reasoning Generation
19
+
20
+ [![License: Apache 2.0](https://img.shields.io/badge/License-Apache%202.0-blue.svg)](https://opensource.org/licenses/Apache-2.0)
21
+ [![GitHub Repository](https://img.shields.io/badge/GitHub-Hooshaai%2FBlockDiffuse-black.svg?logo=github)](https://github.com/Hooshaai/BlockDiffuse)
22
+ [![HuggingFace Space](https://img.shields.io/badge/HF%20Space-Interactive%20Weblog-blueviolet.svg)](https://huggingface.co/spaces/tahamajs/BlockDiffuse-Blog)
23
+ [![Dataset](https://img.shields.io/badge/HF%20Dataset-BlockDiffuse--Data-orange.svg)](https://huggingface.co/datasets/tahamajs/BlockDiffuse-Data)
24
+
25
+ > **TL;DR:** BlockDiffuse is a non-autoregressive / block-autoregressive generative framework that generates **100 tokens simultaneously** in continuous latent space using **Rectified Flow Matching** and a **Diffusion Transformer (DiT)** conditioned on intermediate layers of modern LLMs (`Qwen/Qwen2.5-0.5B-Instruct`).
26
+
27
+ ---
28
+
29
+ ## ⚑ Key Highlights & Benchmark Results
30
+
31
+ All benchmarks measured on a single consumer **NVIDIA GeForce RTX 4070 Laptop GPU (8GB VRAM)**:
32
+
33
+ | Generation Mode | Target Size | ODE Steps / Block | Numerical Solver | Latency (ms) | Throughput (tokens/sec) | VRAM Footprint |
34
+ | :--- | :--- | :--- | :--- | :--- | :--- | :--- |
35
+ | **Single-Block Parallel** | **100 tokens** | 8 ODE steps | DPM-Solver + TFE | **1,730.60 ms** | **57.78 tok/s** | 3,674 MB |
36
+ | **Multi-Block Autoregressive** | **200 tokens** | 8 ODE steps / block | DPM-Solver + TFE | **1,279.20 ms** | **156.35 tok/s** | 3,789 MB |
37
+
38
+ ---
39
+
40
+ ## πŸ—οΈ Architecture Overview
41
+
42
+ ```
43
+ Prompt Prefix ──► Frozen Qwen2.5 (Layers 1..12) ──► Continuous Context c [L_p x 896]
44
+ β”‚
45
+ Initial Gaussian Noise z_0 [100 x 896] ~ N(0, I) ──────────
46
+ β–Ό
47
+ BlockDiffuse DiT (8 Layers, 14 Heads)
48
+ - AdaLN-Zero Timestep Conditioning
49
+ - Continuous RoPE Positional Encoding
50
+ - Rectified Flow (v-prediction)
51
+ β”‚
52
+ β–Ό
53
+ Predicted Latents z_1 [100 x 896]
54
+ β”‚
55
+ β–Ό
56
+ Deep Proj Head (3-Layer SwiGLU MLP)
57
+ β”‚
58
+ β–Ό
59
+ Pre-Head RMSNorm + Frozen LM Head
60
+ β”‚
61
+ β–Ό
62
+ Discrete Next 100 Tokens in Parallel
63
+ ```
64
+
65
+ ### 1. Base LLM Backbone
66
+ - **Model**: `Qwen/Qwen2.5-0.5B-Instruct`
67
+ - **Representation Layer**: Layer 12 (mid-layer context extraction, $d_{\text{model}} = 896$).
68
+ - **Head**: Frozen LM head with vocab size $151{,}936$.
69
+
70
+ ### 2. Diffusion Transformer (DiT)
71
+ - **Depth**: 8 Transformer Blocks.
72
+ - **Attention**: 14 heads (head dimension 64, matches $d_{\text{model}} = 896$).
73
+ - **Initialization**: Direct parameter transfer from layers 6–11 of Qwen2.5-0.5B.
74
+ - **Modulation**: AdaLN-Zero modulates scale and shift parameters based on timestep $t \in [0, 1]$.
75
+
76
+ ### 3. Flow Matching & Multi-Objective Training
77
+ Rectified Flow straight-line trajectory:
78
+ $$z_t = (1 - t) z_0 + t z_1, \quad v_t = \frac{dz_t}{dt} = z_1 - z_0$$
79
+
80
+ Trained under composite multi-loss:
81
+ $$\mathcal{L}_{\text{total}} = \lambda_{\text{FM}} \mathcal{L}_{\text{FM}} + \lambda_{\text{disp}} \mathcal{L}_{\text{disp}} + \lambda_{\text{KL}} \mathcal{L}_{\text{KL}} + \lambda_{\text{CE}} \mathcal{L}_{\text{CE}} + \lambda_{\text{NN}} \mathcal{L}_{\text{NN}}$$
82
+
83
+ ---
84
+
85
+ ## πŸ’» Quickstart: Inference
86
+
87
+ ### 1. Clone & Setup
88
+ ```bash
89
+ git clone https://github.com/Hooshaai/BlockDiffuse.git
90
+ cd BlockDiffuse
91
+ pip install -r requirements.txt
92
+ ```
93
+
94
+ ### 2. Download Checkpoint from Hugging Face
95
+ ```python
96
+ from huggingface_hub import hf_hub_download
97
+
98
+ ckpt_path = hf_hub_download(
99
+ repo_id="tahamajs/BlockDiffuse",
100
+ filename="blockdiffuse_final.pt"
101
+ )
102
+ print("Checkpoint downloaded to:", ckpt_path)
103
+ ```
104
+
105
+ ### 3. Run Parallel Multi-Block Generation
106
+ ```bash
107
+ python inference.py \
108
+ --model Qwen/Qwen2.5-0.5B-Instruct \
109
+ --checkpoint ./checkpoints_improved/blockdiffuse_final.pt \
110
+ --prompt "<|im_start|>system\nYou are a helpful assistant that solves problems step by step.<|im_end|>\n<|im_start|>user\nA bookstore has 140 books on Monday. On Tuesday, they sell 45 books. On Wednesday, they receive 80 books. How many remain?<|im_end|>\n<|im_start|>assistant\n" \
111
+ --max_blocks 2 \
112
+ --steps 8 \
113
+ --solver dpm_solver \
114
+ --use_tfe \
115
+ --tfe_seeds 3
116
+ ```
117
+
118
+ ---
119
+
120
+ ## πŸ“œ Citation
121
+
122
+ ```bibtex
123
+ @article{blockdiffuse2026,
124
+ title={BlockDiffuse: Fully Parallel Latent Space Reasoning Generation with Diffusion Transformers},
125
+ author={Hooshaai Research},
126
+ journal={GitHub / HuggingFace Technical Report},
127
+ year={2026},
128
+ url={https://github.com/Hooshaai/BlockDiffuse}
129
+ }
130
+ ```