How to use from
llama.cpp
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh
# Start a local OpenAI-compatible server with a web UI:
llama serve -hf rectangleworm/ideogram-4-gguf:
# Run inference directly in the terminal:
llama cli -hf rectangleworm/ideogram-4-gguf:
Install from WinGet (Windows)
winget install llama.cpp
# Start a local OpenAI-compatible server with a web UI:
llama serve -hf rectangleworm/ideogram-4-gguf:
# Run inference directly in the terminal:
llama cli -hf rectangleworm/ideogram-4-gguf:
Use pre-built binary
# Download pre-built binary from:
# https://github.com/ggerganov/llama.cpp/releases
# Start a local OpenAI-compatible server with a web UI:
./llama-server -hf rectangleworm/ideogram-4-gguf:
# Run inference directly in the terminal:
./llama-cli -hf rectangleworm/ideogram-4-gguf:
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git
cd llama.cpp
cmake -B build
cmake --build build -j --target llama-server llama-cli
# Start a local OpenAI-compatible server with a web UI:
./build/bin/llama-server -hf rectangleworm/ideogram-4-gguf:
# Run inference directly in the terminal:
./build/bin/llama-cli -hf rectangleworm/ideogram-4-gguf:
Use Docker
docker model run hf.co/rectangleworm/ideogram-4-gguf:
Quick Links

Ideogram4 GGUF quantized files

.
โ”œโ”€โ”€ diffusion/
โ”‚   โ”œโ”€โ”€ cond/
โ”‚   โ”‚   โ”œโ”€โ”€ ideogram4_Q4_0.gguf
โ”‚   โ”‚   โ”œโ”€โ”€ ideogram4_Q4_1.gguf
โ”‚   โ”‚   โ”œโ”€โ”€ ideogram4-Q4_K.gguf
โ”‚   โ”‚   โ”œโ”€โ”€ ideogram4-Q5_0.gguf
โ”‚   โ”‚   โ”œโ”€โ”€ ideogram4_Q5_1.gguf
โ”‚   โ”‚   โ”œโ”€โ”€ ideogram4_Q5_K.gguf
โ”‚   โ”‚   โ”œโ”€โ”€ ideogram4-Q6_K.gguf
โ”‚   โ”‚   โ””โ”€โ”€ ideogram4-Q8_0.gguf
โ”‚   โ””โ”€โ”€ uncond/
โ”‚       โ”œโ”€โ”€ ideogram4_unconditional_Q4_0.gguf
โ”‚       โ”œโ”€โ”€ ideogram4_unconditional_Q4_1.gguf
โ”‚       โ”œโ”€โ”€ ideogram4_unconditional_Q4_K.gguf
โ”‚       โ”œโ”€โ”€ ideogram4_unconditional_Q5_0.gguf
โ”‚       โ”œโ”€โ”€ ideogram4_unconditional_Q5_1.gguf
โ”‚       โ”œโ”€โ”€ ideogram4_unconditional_Q5_K.gguf
โ”‚       โ”œโ”€โ”€ ideogram4_unconditional_Q6_K.gguf
โ”‚       โ””โ”€โ”€ ideogram4_unconditional-Q8_0.gguf
โ”œโ”€โ”€ text_encoder/
โ”‚   โ”œโ”€โ”€ Qwen3-VL-8B-Q4_0.gguf
โ”‚   โ”œโ”€โ”€ Qwen3-VL-8B-Q4_1.gguf
โ”‚   โ”œโ”€โ”€ Qwen3-VL-8B-Q4_K_S.gguf
โ”‚   โ”œโ”€โ”€ Qwen3-VL-8B-Q4_K_M.gguf
โ”‚   โ”œโ”€โ”€ Qwen3-VL-8B-Q5_K_S.gguf
โ”‚   โ”œโ”€โ”€ Qwen3-VL-8B-Q5_K_M.gguf
โ”‚   โ”œโ”€โ”€ Qwen3-VL-8B-Q6_K.gguf
โ”‚   โ””โ”€โ”€ Qwen3-VL-8B-Q8_0.gguf
โ””โ”€โ”€ vae/
โ”‚   โ”œโ”€โ”€ flux2-vae.safetensors
โ”‚   โ””โ”€โ”€ flux2-hdr-vae.safetensors
โ””โ”€โ”€ lora/
    โ”œโ”€โ”€ realism_engine_v3.safetensors
    โ”œโ”€โ”€ big_boobs.safetensors
    โ”œโ”€โ”€ cum.safetensors
    โ”œโ”€โ”€ innie_vulva_x.safetensors
    โ”œโ”€โ”€ vintage_beauties_womans.safetensors
    โ”œโ”€โ”€ missionary_sex.safetensors
    โ”œโ”€โ”€ 80s_anime.safetensors
    โ”œโ”€โ”€ penis.safetensors
    โ””โ”€โ”€ penix.safetensors

Model Selection & Quantization Guide

To balance generation quality, memory usage, and inference speed, we recommend the following quantization choices for each component:

1. Conditional Diffusion Model (diffusion/cond/)

  • Recommended: Q6_K or Q8_0
  • Since this model handles the main conditional generation pass, keeping a higher quantization level is key to preserving detail and prompt adherence.

2. Unconditional Diffusion Model (diffusion/uncond/)

  • Recommended: Q4_K or Q5_K
  • Note: Using Q6_K or Q8_0 for the unconditional model is generally unnecessary (overkill) and may slow down generation without providing a noticeable improvement in quality.

3. Text Encoder (text_encoder/)

  • Recommended: Q5_K_M or Q4_K_M
  • These medium-sized "K-measure" quants offer a good trade-off, retaining the text encoder's comprehension capabilities while fitting within reasonable memory limits.

General Recommendations for Quantization Types

If you are optimizing for inference speed or trying to fit a specific model entirely into VRAM/RAM, keep these rules of thumb in mind:

  • Prefer _K variants over _0 and _1: When choosing between Q4 or Q5 options, always prefer the _K variants (e.g., Q4_K_M, Q5_K_M, or standard _K).
  • Avoid _0 and _1 if possible: The older _0 and _1 quants (like Q4_0 or Q4_1) perform worse in terms of quality loss. While they are marginally smaller, the minor size reduction rarely justifies the drop in generation quality compared to _K equivalents.
Downloads last month
26,042
GGUF
Model size
8B params
Architecture
qwen3vl
Hardware compatibility
Log In to add your hardware

4-bit

5-bit

6-bit

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for rectangleworm/ideogram-4-gguf

Quantized
(23)
this model

Space using rectangleworm/ideogram-4-gguf 1