OmniFysics-Captioner: Grounding Omni-Modal Understanding in the Physical World for Better Captioning

📄 Paper🌐 Project🕵️ OmniFysics-Agent🤗 Model📚 Citation


Introduction

Building omni-modal models with physical intelligence requires fine-grained supervision that captures physical evidence such as contact, support, deformation, and state transitions. OmniFysics-Captioner is an end-to-end, tool-free omni-modal captioner that produces physics-aware audiovisual captions. It reads raw audio and video in a single forward pass and directly generates captions that describe what happens, which objects interact, how materials respond, and how states change over time.

The associated OmniFysics-Agent uses active perception to form a coarse event timeline, select localized observations from the visual and audio streams, and organize spatiotemporally aligned evidence into a final caption. Its Physical Perception Model (PPM) provides focused image-level evidence about objects and physical cues such as material, contact, support, and deformation.

OmniFysics-Agent

OmniFysics-Agent is the active-perception system associated with this model. It first forms a coarse event timeline, then selects localized observations from the visual and audio streams and organizes the resulting evidence into a final caption. The Agent is designed to retain information that can be missed by a single global observation, including temporal relations, sound events, speech/OCR, and audiovisual alignment.

OmniFysics-Agent pipeline

Physical Perception Model

The Agent can use a dedicated Physical Perception Model (PPM) for focused image-level observations. PPM supplies object-centric evidence about material, contact, support, deformation, stability, and likely state changes. It is a tool within the Agent and complements, rather than replaces, the audiovisual captioner.

Physical perception case study

Model

OmniFysics-Captioner is fine-tuned from Qwen3-Omni on physics-rich audiovisual caption data. The released checkpoint is a merged model intended for inference. It does not require the Agent or external perception tools to generate a caption: video and audio can be provided directly in one forward pass.

Audiovisual caption data overview

The model is suitable for research and development involving:

  • audiovisual video captioning;
  • temporal event and state-transition description;
  • audio-aware and speech-aware captioning;
  • descriptions of object interactions and physical outcomes;
  • multimodal evidence organization for downstream agents.

Limitations

Captions are grounded in the available audiovisual evidence and may omit events that are occluded, inaudible, too brief, or outside the sampled view. Physical quantities and future outcomes are approximate visual inferences, not measurements or guarantees. The model should not be used as a safety controller or as a substitute for physical instrumentation.

Paper

Paper and supplementary material: Coming soon.

Citation

@article{qiu2026omnifysicscaptioner,
  title   = {OmniFysics-Captioner: Grounding Omni-Modal Understanding in the Physical World for Better Captioning},
  author  = {Qiu, Kaixiang and Han, Minghao and Liu, Keliang and Liu, Yizhou and Han, Jinghang and Jiang, Yue and Wang, Shunli and Zhang, Lihua and Yang, Dingkang},
  journal = {arXiv preprint},
  year    = {2026}
}

License

The model is released for research and educational use under the license provided with this repository. Source videos and upstream model components remain subject to their original licenses.

Downloads last month
139
Safetensors
Model size
32B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Fysics-AI/OmniFysics-Captioner

Finetuned
(34)
this model