TLive-Omni: An Omni-Modal Understanding Model for E-Commerce Live Streaming Paper • 2608.20958 • Published 6 days ago • 56
CoInteract: Physically-Consistent Human-Object Interaction Video Synthesis via Spatially-Structured Co-Generation Paper • 2604.19636 • Published Apr 21 • 88
TokenPacker: Efficient Visual Projector for Multimodal LLM Paper • 2407.02392 • Published Jul 2, 2024 • 23