LeRF-9B

LeRF: Learning Reference Coordinate Frames for Perspective Taking Reasoning

Bang Xiao1,2 · Wenqi Jia1 · Ozgur Kara1 · Tiancheng Shen3 · Yibo Yang4

Bolin Lai1,5,^ · Junho Kim1,^ · James Matthew Rehg1,^

1University of Illinois Urbana-Champaign · 2Zhiyuan College, Shanghai Jiao Tong University

3University of California, Merced · 4Shanghai Jiao Tong University · 5Amazon AGI

^Corresponding authors

🏠 Homepage · 💻 Code · 🤗 Collection

Model Summary

Vision-language models (VLMs) struggle with perspective taking: when a question asks about spatial relations from another entity's or an imagined observer's viewpoint, they often fall back to the camera view. LeRF trains a VLM to build and use an explicit entity-centered reference frame. Given an image and a question, the model decides whether a frame is needed. If it is, it grounds the reference entity and predicts the frame's projected origin and its front / left / up axes. A lightweight renderer draws the frame on the image, and the model reasons over this visual cue. No external perception models or 3D reconstruction are used.

How It Works

Two-turn inference.

  1. Turn 1 (thinking off). The model outputs either a single draw_reference_frame call — reference_object (text), plus origin, axis_front, axis_left, axis_up as [x, y] points in [0, 1000] normalized image coordinates — or the literal string NO_TOOL_CALL when the question can be answered from the camera view.
  2. Tool. The frame is rendered on the image (red = front, green = left, blue = up; a ⊙/⊗ depth marker replaces a strongly foreshortened axis).
  3. Turn 2 (thinking on). The model sees the rendered image (not the numeric coordinates), reasons over it, and answers in \boxed{}.

Training.

  • Stage I: SFT. Learns reference-entity grounding and projected frame prediction from pose datasets (ImageNet3D, Omni6DPose-SOPE, BEDLAM). About 20% of the data are no-tool-call examples that teach selective tool use.
  • Stage II: RL (GRPO). Trained on MultihopSpatial and SpatialReasoner-RL with a binary final-answer reward. The frame-prediction turn is excluded from the policy gradient, so only the reasoning/answer turn receives the advantage.

Usage

The model is meant to be used together with the draw_reference_frame tool and the two-turn client from the GitHub repository. The system prompt and tool schema live in tool/system_prompt.txt and tool/prompts.py.

1. Serve with vLLM (OpenAI-compatible API):

vllm serve SamuelBang/LeRF-9B --served-model-name lerf --max-model-len 32768 \
    --limit-mm-per-prompt '{"image": 2}' --trust-remote-code

2. Run the client:

git clone https://github.com/bangx7/LeRF.git && cd LeRF/tool
pip install pillow requests

# single question
python inference.py --model lerf --image example.jpg \
    --question "From the man's perspective, which object is on his left?" \
    --options cup bed laptop phone --save_renders renders/

# batch: one JSON per line with image / question / options (and optionally answer, id)
python inference.py --model lerf --input questions.jsonl --output results.jsonl

Recommended settings (matching training): temperature 0.6, images downscaled to at most 1M pixels, a 10,240-token thinking budget within a 12,288-token answer turn.

Tested with Python 3.12, CUDA 13.0, PyTorch 2.11, vLLM 0.24.0 and transformers 5.10.4.

Evaluation

Accuracy (%) on OmniSpatial perspective taking (OmniSpatial-PT: Ego / Allo / Hypo), 3DSRBench (Orientation / Multi-Object) and ViewSpatial-Bench (person-perspective Object View Orientation / Relative Direction). Bold marks the best open-source result.

Method Ego Allo Hypo 3DSR Ori 3DSR M-Obj VS P-Obj VS P-Rel
Proprietary
GPT-5.6-Luna (medium) 83.33 49.73 45.78 60.04 55.34 46.99 70.07
GPT-5.6-Terra (medium) 81.37 55.85 53.01 63.32 56.39 45.08 77.20
Claude Sonnet 5 (medium) 80.39 42.55 49.40 34.94 44.77 51.31 51.43
Claude Sonnet 5 (high) 84.31 48.14 45.78 43.15 46.81 51.51 60.10
Open-source
Qwen3.5-4B 74.71 42.55 44.34 42.28 43.26 51.01 57.43
Qwen3.5-9B 80.20 47.13 44.58 48.17 48.66 56.23 65.51
Qwen3.5-9B + APC 42.16 27.66 30.12 44.98 33.22 59.34 37.53
SpatialReasoner 40.39 35.11 35.66 52.05 50.64 42.37 45.61
Ours
LeRF-4B 72.35 49.36 46.75 45.88 44.55 56.26 67.85
LeRF-9B (this model) 74.31 54.04 55.66 53.76 50.29 61.91 74.23

See the paper for more baselines, reference-frame estimation accuracy and ablations.

Acknowledgements

Built on Qwen3.5, LLaMA-Factory and verl.

Citation

Coming soon.

Downloads last month
-
Safetensors
Model size
9B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for SamuelBang/LeRF-9B

Finetuned
Qwen/Qwen3.5-9B
Finetuned
(946)
this model

Collection including SamuelBang/LeRF-9B