lmms-lab-encoder/LLaVA-OneVision-2-8B-Instruct Image-Text-to-Text • 9B • Updated about 8 hours ago • 69k • 14
Aero Realtime: Fully Aligned Input-Output Streams for Low-Latency Streaming Multimodal Generation Paper • 2608.08469 • Published 22 days ago • 2
VBVR-Pro: A Scalable and Verifiable Suite for Native Visual Reasoning Paper • 2608.26105 • Published 5 days ago • 253
StreamOPD: A Post-Training Recipe with Spatio-Temporal Cue Gating for Streaming Video Understanding Paper • 2608.16320 • Published 14 days ago • 9
StreamOPD: A Post-Training Recipe with Spatio-Temporal Cue Gating for Streaming Video Understanding Paper • 2608.16320 • Published 14 days ago • 9
StreamOPD: A Post-Training Recipe with Spatio-Temporal Cue Gating for Streaming Video Understanding Paper • 2608.16320 • Published 14 days ago • 9
AVA-Encoder: Towards Agent-Native Video Representation Learning Paper • 2608.12313 • Published 19 days ago • 42