Instructions to use microsoft/AesCode-8B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use microsoft/AesCode-8B with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="microsoft/AesCode-8B") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("microsoft/AesCode-8B") model = AutoModelForMultimodalLM.from_pretrained("microsoft/AesCode-8B", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use microsoft/AesCode-8B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "microsoft/AesCode-8B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "microsoft/AesCode-8B", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/microsoft/AesCode-8B
- SGLang
How to use microsoft/AesCode-8B with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "microsoft/AesCode-8B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "microsoft/AesCode-8B", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "microsoft/AesCode-8B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "microsoft/AesCode-8B", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use microsoft/AesCode-8B with Docker Model Runner:
docker model run hf.co/microsoft/AesCode-8B
AesCode-8B
AesCode generates information-rich visual artifacts such as slides, posters, and dashboards as HTML/CSS. The output remains structured, editable, and verifiable, but the task poses a distinct challenge: code models cannot see how layout, hierarchy, and color come together on the canvas.
Image generators offer the opposite strength. They compose visually compelling pages but often misrender text, numbers, and logical relationships. AesCode uses an image generated from the same prompt as an aesthetic reference while following the prompt for the required content.
Reference input alone does not solve the problem. Off-the-shelf vision-language models may copy hallucinated content or ignore the reference layout. AesCode separates semantic requirements from visual cues through graph-structured supervision and decoupled cross-modal rewards.
AesCode-8B starts from Qwen3-VL-8B-Instruct and is trained with cold-start SFT followed by GDPO across seven reward channels.
- ๐ Paper: AesCode: Aesthetic Code Generation with Decoupled Cross-Modal Rewards
- ๐ป Code: https://github.com/microsoft/AesCode
- ๐ค Companion model:
microsoft/AesCode-32B
Intended Use
AesCode is designed to generate information-rich visual artifacts as structured HTML/CSS. It is suited for creating editable slides, posters, dashboards, and reports, as well as for research on multimodal code generation and verifiable visual design.
Results
We evaluate on 300 infographic samples. Rule averages the programmatic Text, Bound, and Chart checks. Visual averages the Content, Layout, and Style checklist dimensions. Overall is the mean of Rule and Visual. Scores are percentages averaged over three generations per prompt with no selection.
Ref. indicates whether the model receives an image-generated reference together with the prompt. Bold marks the best score in each column and underline the second best.
| Model | Ref. | Text | Bound | Chart | Rule | Content | Layout | Style | Visual | Overall |
|---|---|---|---|---|---|---|---|---|---|---|
| GPT-5.5 | No | 88.67 | 86.18 | 82.85 | 85.90 | 79.27 | 68.09 | 52.80 | 66.72 | 76.31 |
| GPT-5.5 | Yes | 89.22 | 86.85 | 81.35 | 85.80 | 82.44 | 89.93 | 57.91 | 76.76 | 81.28 |
| Claude Opus 4.8 | No | 91.40 | 85.03 | 87.69 | 88.04 | 84.98 | 64.88 | 46.72 | 65.53 | 76.78 |
| Claude Opus 4.8 | Yes | 92.45 | 84.43 | 73.75 | 83.55 | 86.42 | 89.76 | 55.51 | 77.23 | 80.39 |
| Qwen3-VL-8B | No | 58.32 | 77.23 | 43.30 | 59.61 | 38.33 | 23.97 | 13.22 | 25.17 | 42.39 |
| Qwen3-VL-8B | Yes | 64.07 | 70.99 | 42.35 | 59.14 | 52.90 | 51.33 | 29.92 | 44.72 | 51.93 |
| Qwen3-VL-32B | No | 66.00 | 80.31 | 41.36 | 62.55 | 49.81 | 34.85 | 19.34 | 34.67 | 48.61 |
| Qwen3-VL-32B | Yes | 71.19 | 79.61 | 50.77 | 67.19 | 62.04 | 64.59 | 38.43 | 55.02 | 61.10 |
| AesCode-8B | Yes | 94.06 | 88.36 | 87.79 | 90.07 | 86.41 | 87.80 | 53.21 | 75.80 | 82.94 |
| AesCode-32B | Yes | 95.34 | 97.27 | 90.37 | 94.33 | 85.76 | 90.58 | 55.99 | 77.44 | 85.89 |
AesCode-8B improves Visual over its reference-conditioned Qwen3-VL-8B backbone by 31.1 points and Overall by 31.0. Its Overall score of 82.94 exceeds the reference-conditioned GPT-5.5 and Claude Opus 4.8 results, while Rule reaches 90.07. The visual gains do not come at the expense of verifiable correctness.
Style is the shared ceiling for every system. It awards credit only when a design needs no further visual revision before delivery, and no model clears 60.
Usage
The model follows the Qwen3-VL chat interface and requires transformers>=4.57. Supply a
detailed content prompt and, when available, a reference image for layout and style
guidance.
from transformers import AutoProcessor, AutoModelForImageTextToText
model_id = "microsoft/AesCode-8B"
processor = AutoProcessor.from_pretrained(model_id)
model = AutoModelForImageTextToText.from_pretrained(
model_id, dtype="auto", device_map="auto"
)
messages = [{
"role": "user",
"content": [
{"type": "image", "image": "reference.png"},
{"type": "text", "text": "<your content prompt>"},
],
}]
inputs = processor.apply_chat_template(
messages, add_generation_prompt=True, tokenize=True,
return_dict=True, return_tensors="pt",
).to(model.device)
out = model.generate(
**inputs, do_sample=True, temperature=0.8, top_p=0.95, max_new_tokens=12000
)
print(processor.decode(out[0][inputs["input_ids"].shape[1]:], skip_special_tokens=True))
For batched generation, serve with vLLM and allow two images per prompt:
vllm serve microsoft/AesCode-8B --limit-mm-per-prompt image=2 --max-model-len 24576
Reported scores use three rollouts per prompt at temperature 0.8, top-p 0.95, up to 12,000 output tokens and a 24,576-token context.
The model emits a complete HTML document. Tables use HTML table structures and charts use ECharts specifications, keeping both directly inspectable and verifiable.
The reference image is optional. Reference-conditioned training internalizes visual planning into the policy. Measured on AesCode-8B, withholding the reference at inference costs only 1.00 Visual point, against 19.55 for the Qwen3-VL-8B-Instruct backbone and 10.04 for GPT-5.5. Quality is still highest when a reference is supplied.
The repository provides the training prompt template and a pipeline that turns a short idea into the prompt and reference pair the model expects.
Training
Cold-start SFT. 3,000 demonstrations at learning rate 1e-5.
Reinforcement learning. GDPO over 7,408 prompts for 400 steps, using verl's FSDP-vLLM hybrid engine with no critic and no separately trained reward model. AdamW at a constant 5e-6, no warmup, weight decay 0.01, 128 prompts per step with 8 rollouts each, KL and entropy coefficients both 0.001, rollout temperature 1.0 and top-p 1.0, prompt and response each capped at 8,192 tokens. AesCode-32B uses the same recipe and a longer 520-step run.
The reward is not a single scalar. Each target is represented as a design graph describing the whole canvas, which makes every property individually attributable. From that graph come two complementary reward families: deterministic verifiers score the properties parsable from the code and its rendering, while a VLM judge scores the non-parsable ones using a sample-specific Visual Graph Rubric tied to the graph's elements and relations. These form seven channels: execution, text, boundary, tablechart, layout, whitespace, and design. Each is normalized within its rollout group before aggregation, so a dominant signal cannot drown out a weaker one.
Candidate HTML is scored by rendering it in a sandboxed Playwright browser with external requests blocked, which exports the DOM, computed styles, bounding boxes, console status and a screenshot.
Training code, the verifier stack and the Visual Graph Rubric builder are released at https://github.com/microsoft/AesCode.
Citation
@inproceedings{aescode,
title = {AesCode: Aesthetic Code Generation with Decoupled Cross-Modal Rewards},
booktitle = {Under review},
year = {2027}
}
License
Released under the Apache 2.0 license, following the Qwen3-VL backbone.
- Downloads last month
- 6