Hugging Face's logo Hugging Face
  • Models
  • Datasets
  • Spaces
  • Buckets new
  • Docs
  • Enterprise
  • Pricing
    • Website
      • Tasks
      • HuggingChat
      • Collections
      • Languages
      • Organizations
    • Community
      • Blog
      • Posts
      • Daily Papers
      • Hardware
      • Learn
      • Discord
      • Forum
      • GitHub
    • Solutions
      • Team & Enterprise
      • Hugging Face PRO
      • Enterprise Support
      • Inference Providers
      • Inference Endpoints
      • Storage Buckets

  • Log In
  • Sign Up

ASID-Caption

community
https://asid-caption.github.io/
Activity Feed

AI & ML interests

Video Understanding, Audio-Visual, Multimodal LLMs, Video Captioning, Instruction Tuning, Dataset Curation, Qwen-based, Open-source, Fully-Open-MLLMs

Recent Activity

lyhisme  authored a paper 1 day ago
Hy-Embodied-VLM-1.0: Efficient Physical-World Agents
lyhisme  submitted a paper 4 months ago
Rethinking Token-Level Policy Optimization for Multimodal Chain-of-Thought
lyhisme  updated a model 4 months ago
AudioVisual-Caption/ASID-Captioner-7B
View all activity

Papers

Towards Universal Video MLLMs with Attribute-Structured and Quality-Verified Instructions

View all Papers

Yunheng Li's profile picture

AudioVisual-Caption 's models 2

AudioVisual-Caption/ASID-Captioner-7B

Image-Text-to-Text • 9B • Updated Mar 11 • 44 • 7

AudioVisual-Caption/ASID-Captioner-3B

Image-Text-to-Text • 5B • Updated Mar 11 • 54 • 37
Company
TOS Privacy About Careers
Website
Models Datasets Spaces Pricing Docs