Gestalt

Gestalt: Large Multimodal Interplay Model

Website Paper Code Model

Zequn Yang†  Yu Miao†  Haotian Ni†  Ziheng Chen†  Chengxiang Huang†
Dongzhan Zhou  Kai Chen  Qi Zhang  Ji-Rong Wen  Yake Wei‡  Di Hu‡,✉

† Equal contribution ‡ Team leader ✉ Corresponding author

Gestalt is a new paradigm of large multimodal model built around multimodal interplay. Guided by a multimodal interplay pyramid — from modality-specific modeling, through cross-modal alignment, to multimodal synergy — Gestalt adopts a unified discrete diffusion framework with an interplay-partitioned architecture, where learnable interplay tokens mediate cross-modal exchange and integration. The name is inspired by Gestalt psychology: the whole is greater than the sum of its parts.

Gestalt teaser

Model Description

Gestalt is an interplay-centric large multimodal model built on a unified discrete diffusion framework. Through multimodal pretraining, continual pretraining, and interplay-oriented supervised fine-tuning,Gestalt supports both understanding and generation.

Property Value
Architecture GestaltModelLM (discrete diffusion with interplay tokens)
Parameters ~8B
Hidden size 4096
Layers 32
Attention heads 32
Vocab size 142,848
Max sequence length 4,096
Precision bfloat16

Download

# Gestalt model
huggingface-cli download GeWuLab/Gestalt --local-dir /path/to/gestalt

# IBQ vision tokenizer (required for encoding/decoding images)
huggingface-cli download TencentARC/IBQ-Tokenizer-16384 --local-dir /path/to/ibq

The IBQ tokenizer is from TencentARC/IBQ-Tokenizer-16384, used to convert between images and discrete visual tokens.

Training Pipeline

Stage Data Objective
Multimodal Pretraining 70M Joint masked prediction
Continual Pretraining ← this checkpoint 8M Conditional masked prediction
Supervised Fine-Tuning 13.7M (+≈2.5M T2I) Four interplay categories, three-phase curriculum

Citation

If you find Gestalt useful for your research, please cite:

@article{gestalt2026,
  title   = {Gestalt: Large Multimodal Interplay Model},
  author  = {Yang, Zequn and Miao, Yu and Ni, Haotian and Chen, Ziheng and
             Huang, Chengxiang and Zhou, Dongzhan and Chen, Kai and Zhang, Qi and
             Wen, Ji-Rong and Wei, Yake and Hu, Di},
  year    = {2026},
  url     = {https://github.com/GeWu-Lab/Gestalt}
}

License

This project is released under the Apache 2.0 license.

Author Contributions

Zequn Yang, Yake Wei, and Di Hu drove the overall advancement of the project. Zequn Yang, Yu Miao, Haotian Ni, Ziheng Chen, and Chengxiang Huang contributed equally to this work. Zequn Yang conducted model pretraining and supervised fine-tuning. Zequn Yang, Yu Miao, and Haotian Ni developed the model architecture and conducted the core experiments. Yu Miao, Ziheng Chen, and Chengxiang Huang contributed to data processing and organization. Yu Miao conducted the evaluation of image generation capabilities, Haotian Ni conducted the multimodal understanding evaluation, and Ziheng Chen conducted the text-only evaluation. Dongzhan Zhou, Kai Chen, Qi Zhang, and Ji-Rong Wen contributed to discussions on the technical design and methodology. Yake Wei, Zequn Yang, Yu Miao, Haotian Ni, Ziheng Chen, and Di Hu contributed to writing and revising the manuscript. Di Hu initiated the project. Yake Wei and Di Hu supervised and advised the project.

Downloads last month
38
Safetensors
Model size
9B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support