tahamajs commited on
Commit
da34971
·
verified ·
1 Parent(s): d036837

Upload README.md with huggingface_hub

Browse files
Files changed (1) hide show
  1. README.md +65 -0
README.md ADDED
@@ -0,0 +1,65 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: apache-2.0
3
+ tags:
4
+ - diffusion
5
+ - flow-matching
6
+ - rectified-flow
7
+ - text-generation
8
+ - reasoning
9
+ - qwen2.5
10
+ - block-diffusion
11
+ pipeline_tag: text-generation
12
+ language:
13
+ - en
14
+ ---
15
+
16
+ # BlockDiffuse: Parallel Multi-Block Reasoning Generation in Latent Space
17
+
18
+ **BlockDiffuse** is a non-autoregressive / block-autoregressive diffusion framework that generates entire blocks of 100 tokens simultaneously in continuous latent space using Rectified Flow Matching.
19
+
20
+ By coupling a Diffusion Transformer (DiT) with a frozen decoder-only Base LLM (`Qwen/Qwen2.5-0.5B-Instruct`), BlockDiffuse bypasses token-by-token sequential decoding, achieving parallel multi-token throughput.
21
+
22
+ ## Model Summary
23
+
24
+ - **Base LLM**: `Qwen/Qwen2.5-0.5B-Instruct`
25
+ - **DiT Architecture**: 8 Transformer Blocks with 14 Attention Heads ($d_\text{model}=896$, head dim 64)
26
+ - **Latent Space**: Mid-layer representations (layer 12) for prompt conditioning; target 100-token blocks supervised via Flow Matching.
27
+ - **Initialization**: Transfer learning from layers 6–11 of Qwen2.5-0.5B.
28
+ - **Projection Head**: Deep 3-layer MLP residual adapter before RMSNorm and discrete token decoding.
29
+ - **Training Objectives**: Flow Matching Velocity MSE ($\mathcal{L}_\text{FM}$) + Dispersive Loss + Teacher KL Distillation + Discrete Token Cross-Entropy ($\mathcal{L}_\text{CE}$) + Contrastive InfoNCE Nearest-Neighbor Loss ($\mathcal{L}_\text{NN}$).
30
+
31
+ ## Benchmark Results on RTX 4070 Laptop GPU
32
+
33
+ | Mode | Tokens | Diffusion Steps | Latency | Throughput |
34
+ | :--- | :--- | :--- | :--- | :--- |
35
+ | **Single-Block Parallel** | 100 tokens | 8 ODE steps (DPM-Solver + TFE) | **1,730.60 ms** | **57.78 tokens/sec** |
36
+ | **Multi-Block Autoregressive** | 200 tokens (2 blocks) | 8 ODE steps per block | **1,279.20 ms** | **156.35 tokens/sec** |
37
+
38
+ ## Quickstart & Inference
39
+
40
+ To run inference using the official repository [Hooshaai/BlockDiffuse](https://github.com/Hooshaai/BlockDiffuse):
41
+
42
+ ```bash
43
+ git clone https://github.com/Hooshaai/BlockDiffuse.git
44
+ cd BlockDiffuse
45
+
46
+ python inference.py \
47
+ --model Qwen/Qwen2.5-0.5B-Instruct \
48
+ --checkpoint ./blockdiffuse_final.pt \
49
+ --prompt "<|im_start|>system\nYou are a helpful assistant that solves problems step by step.<|im_end|>\n<|im_start|>user\nA bookstore has 140 books on Monday. On Tuesday, they sell 45 books. On Wednesday, they receive 80 books. How many remain?<|im_end|>\n<|im_start|>assistant\n" \
50
+ --steps 8 \
51
+ --solver dpm_solver \
52
+ --use_tfe \
53
+ --tfe_seeds 3
54
+ ```
55
+
56
+ ## Citation
57
+
58
+ ```bibtex
59
+ @software{blockdiffuse2026,
60
+ author = {Hooshaai Research},
61
+ title = {BlockDiffuse: Fully Parallel Latent Space Reasoning Generation with Diffusion Transformers},
62
+ year = {2026},
63
+ url = {https://github.com/Hooshaai/BlockDiffuse}
64
+ }
65
+ ```