MemUse: Moving Memory Evaluation from Direct QA to Natural Integration in Long-Term Human-AI Conversation
Abstract
Direct QA benchmarks for conversational memory do not predict user satisfaction, whereas natural integration of prior context does, revealing a large gap between elicited recall and conversational use.
Memory systems for conversational LLMs are conventionally evaluated by direct, fact-seeking questions about prior dialogue (Direct QA): can the model recall fact X from a prior conversation? We tested whether higher Direct QA accuracy correlates with higher user satisfaction in a 4-month deployment (40 users, 1,872 sessions, 7 memory conditions). Existing-benchmark Direct QA varies from 19.7% to 70.1% across the 7 conditions, but satisfaction does not change. We hypothesize that existing benchmarks and user satisfaction are tracking different capabilities: benchmarks measure elicited retrieval (recall when asked), while conversation requires natural integration (detecting relevance and naturally weaving prior context into a response). To examine this, we introduce MemUse, a set of real user-cued memory moments drawn from the deployment, scored by an integration-aware judgment of the natural conversational response. Holding the model and context fixed, the same system that scores 78.8% on Direct QA references only 7.9% of those facts in conversation -- a 71-point gap. Within these moments, Natural Integration is associated with satisfaction, whereas Direct QA is not. We release the deployment corpus and MemUse together with all judgments and scoring prompts at https://github.com/ryuichi-sumida/memuse.
Community
Accepted to EMNLP 2026 Main.
We ran a 4-month real-world deployment of a memory-augmented conversational AI and found that Direct QA accuracy - the standard way to evaluate long-term memory - doesn't predict user satisfaction: a system scoring 78.8% on Direct QA spontaneously referenced only 7.9% of those facts in actual conversation.
We introduce MemUse, a benchmark built from real user interactions that instead measures whether models naturally integrate remembered context into responses.
Get this paper in your agent:
hf papers read 2608.24189 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 1
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper