Ambient @ EgoLongQA 2026: Distilling Long-Video perception into a Sub-2B Model
Abstract
A 2B vision-language model distilled from a tool-using agent achieves first-place performance on long egocentric video question answering by pruning its multilingual embedding table to meet parameter limits.
We describe our entry to the EgoLongQA track of the Wearable-AI Challenge in ECCV 2026, which placed first in the <=2B parameter division with 0.8279 on the held-out test set. Our system is a single 2B vision-language model that answers multiple-choice questions about ten-minute egocentric videos in one greedy forward pass; It is obtained by distilling the junior perception module of a tool-using agentic pipeline, not the agent itself into a small student, using teacher traces filtered to those that answered correctly. it reaches 89% of the accuracy of the large agentic pipeline using 1.1% of its parameters. This raises a 27.1% base model to 81.4% on our held-out questions. The 2B backbone has 2.2132B parameters and therefore over the divisional limit, to make the entry admissable we prune the multilingual embedding table from 248,320 to 143,469 rows, reaching 1.9985B with provably identical logits on retained rows.
Community
We describe our entry to the EgoLongQA track of the Wearable-AI Challenge in ECCV 2026, which placed first in the ≤2B parameter division with 0.8279 on the held-out test set
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- Ambient @ EgoProactive 2026 : Proactive Egocentric Assistance with Visually Grounded Supervision (2026)
- Temporal Self-Distillation: Learning Visual State Tracking in Videos Without Supervision (2026)
- Where to Look Matters: On-Policy Self-Distillation for Long-Video Understanding (2026)
- ReToken: One Token to Improve Vision-Language Models for Visual Retrieval (2026)
- Qwen-MusicAVQA-7B: A Multimodal Model for Music Audio-Visual QA (2026)
- Searching Videos as Trees: Self-Correcting Agents for Grounded Long Video QA (2026)
- Annotations as Rollouts: Efficient and Scalable Reinforcement Learning for Video MLLMs (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2609.07154 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper