Skip to content

Repository files navigation

OP-Bench: Benchmarking Over-Personalization
for Memory-Augmented Personalized Conversational Agents

Yulin Hu · Zimo Long · Jiahe Guo · Xingyu Sui
Xing Fu · Weixiang Zhao · Yanyan Zhao · Bing Qin

arXiv: 2601.13722 EMNLP 2026 Website live

Paper · Quick start · Data · Citation · 中文

1,700 verified instances   ·   20 users   ·   3 failure modes

OP-Bench construction pipeline

Overview

Memory-augmented conversational agents are designed to use long-term user information to provide helpful personalization. OP-Bench studies the other side of this capability: over-personalization, where an agent introduces personal information unnecessarily, agrees with a user at the expense of accuracy or neutrality, or repeats the same personalized content across different questions.

OP-Bench contains 1,700 human-verified instances constructed from long-horizon dialogue histories. It evaluates three failure modes—Irrelevance, Sycophancy, and Repetition—using complementary model-based and embedding-based metrics.

At a glance
Benchmark size 1,700 verified instances
Users 20
Failure modes 3 primary categories, 6 subcategories
Evaluation signal Higher score means less over-personalization
Paper evaluation 6 language models, 5 memory settings, 36 configurations
Public release Benchmark data, prompts, evaluator, selected agents, MemOS adapter, and interactive demo

Failure modes

Category What it measures Subcategories
Irrelevance Whether the response injects personal information when the query does not call for it Fully irrelevant; baiting (deceptively relevant)
Sycophancy Whether personalization causes excessive agreement instead of factual or value-sensitive responses Fact-level; value-level; memory-level
Repetition Whether semantically distinct queries receive nearly identical personalized responses One repetition setting

Illustration of the three OP-Bench failure modes: irrelevance, sycophancy, and repetition

Illustration of the three over-personalization categories.

Data construction

The benchmark is built in three stages:

  1. Initialization: derive a structured user profile and user topics from long-horizon LoCoMo conversations.
  2. Task construction: generate controlled probes for Irrelevance, Sycophancy, and Repetition, including the relevant subcategories.
  3. Human review: independently review each candidate, adjudicate disagreements, and retain only verified instances.

The final dataset distribution is:

Category Subcategory Count Share
Irrelevance Fully irrelevant 318 18.7%
Irrelevance Baiting 100 5.9%
Sycophancy Fact-level 100 5.9%
Sycophancy Value-level 100 5.9%
Sycophancy Memory-level 200 11.8%
Repetition — 882 51.9%
Total — 1,700 100%

Experimental findings

OP-Bench scores are designed so that higher is better: a higher score indicates that the response avoids the corresponding over-personalization behavior.

Main result

The following is the GPT-4o-mini block from the paper's main results table. The percentage in parentheses is the relative drop from the memory-free BASE setting.

Memory setting OP-Bench average
BASE 83.10
RAG 55.96 (↓32.7%)
Mem0 46.32 (↓44.3%)
MemU 40.46 (↓51.3%)
MEMOS 41.86 (↓49.6%)

Across models and memory systems, the paper reports relative drops of 26.2%–61.1% compared with BASE.

Cross-model comparison

Legend for BASE, RAG, Mem0, MEMOS, and MemU

Radar chart for GPT-4o-mini Radar chart for Gemini-2.5-flash

Radar chart for Qwen3-235B Radar chart for Qwen3-32B

OP-Bench scores across model and memory configurations (paper Figure 2).

Research questions

Question Conclusion
RQ1 — Does over-personalization exist? Yes. All evaluated memory-augmented settings show substantial degradation relative to BASE; more elaborate memory mechanisms tend to suffer larger drops.
RQ2 — Why does it occur? Memory can be retrieved aggressively and receive disproportionate attention. The average memory-to-query attention ratio exceeds 2×, which can produce memory hijacking and response collapse.
RQ3 — Can it be mitigated? Only partially. Post-processing improves OP-Bench by up to +20.0%, but some configurations reduce LoCoMo personalization performance by up to −8.5%. No evaluated method closes the gap to the memory-free oracle.
RQ4 — What is the timing cost? The supplementary timing view shows retrieval averaging roughly 38–44 ms, while post-processing ranges from about 0.97–4.13 s and model response generation from about 2.71–3.41 s.

Retrieval is a major part of the problem

Even fully irrelevant queries can trigger memory retrieval. Baiting queries may have high superficial similarity to stored memories, making retrieval alone an insufficient safeguard.

Bar chart comparing retrieval relevance for fully irrelevant and baiting queries across RAG, MemOS, MemU, and Mem0

Retrieved-memory similarity on the Irrelevance task (paper Figure 5).

What's included

The repository includes:

  • benchmark data and construction/evaluation prompts;
  • a resumable benchmark builder;
  • search, generation, scoring, and post-processing utilities;
  • LDAgent and SimpleRAGAgent implementations;
  • an HTTP adapter for an external MemOS service;
  • a dependency-light web demo for the benchmark story and results;
  • tests for configuration, scoring, data types, and the web demo.

Repository layout

OP-Bench/
├── configs/
│   └── evaluation.example.yaml
├── data/
│   ├── locomo10.json
│   ├── locomo10_overpersonalized.json
│   └── README.md
├── docs/
│   └── figures/                  # Selected figures from the paper
├── prompts/
│   ├── benchmark/                # Task-construction templates
│   └── evaluation/               # Scoring prompts and archived source
├── scripts/
│   ├── prepare_locomo_histories.py
│   ├── start_agent.py
│   ├── start_agents.py
│   └── ingest_memos.py
├── src/opbench/
│   ├── agents/                   # LDAgent and SimpleRAGAgent
│   ├── baselines/                # MemOS API adapter
│   ├── benchmark_builder.py      # Build or extend benchmark tasks
│   ├── evaluation.py             # Search, generation, scoring, metrics
│   ├── postprocessing.py         # Context filtering/compression hooks
│   └── config.py                 # Environment-backed configuration
├── tests/
├── web/                          # Local full-stack research explorer
├── .env.example
├── pyproject.toml
└── NOTICE.md

Quick start

1. Install

Python 3.10 or newer is required.

python -m venv .venv
source .venv/bin/activate
python -m pip install -e ".[agents]"

Create a local environment file from the template:

cp .env.example .env

Set your model endpoint and API credentials in .env.

2. Launch the local web demo

Start the demo server:

python web/backend/server.py --host 127.0.0.1 --port 8765

Then open http://127.0.0.1:8765. The interface contains three chapters: data construction, representative QA examples, and experimental results. The construction chapter supports automatic playback as well as manual previous/next navigation. See web/README.md for API routes and implementation notes.

3. Run the included agents

Prepare one history file per LoCoMo user:

python scripts/prepare_locomo_histories.py

Start all ten agents for one of the two public memory implementations:

python scripts/start_agents.py --agent ldagent
# or
python scripts/start_agents.py --agent simplerag

Each server exposes:

  • POST /v1/chat/completions
  • GET /v1/models
  • GET /health

To start a single process, use scripts/start_agent.py with a history file and port.

4. Evaluate OP-Bench

The checked-in task file is data/locomo10_overpersonalized.json. For the in-process agents:

python -m opbench.cli generate --frame ldagent
python -m opbench.cli generate --frame simplerag

For MemOS, first ingest histories into an already configured external service, then search, generate, and score:

python scripts/ingest_memos.py --online
python -m opbench.cli search --frame memos-api-online
python -m opbench.cli generate --frame memos-api-online --postprocess none
python -m opbench.cli score --frame memos-api-online --responses results/responses/memos-api-online_gpt-4o-mini_none.json
python -m opbench.cli metrics --frame memos-api-online --judged results/judged/memos-api-online_gpt-4o-mini.json

For all commands and options:

python -m opbench.cli --help

The available context methods include none, self_recheck, reminder, self_critic, few_shot_chain_of_thought, comorag, and marag. Results are saved under results/.

Rebuild the benchmark

The builder is resumable: an existing output file is normalized and completed instead of being silently overwritten.

python -m opbench.cli build-benchmark --source data/locomo10.json --output data/locomo10_overpersonalized.generated.json

The model endpoint and credentials for construction are read from the local environment.

Run tests

python -m pip install -e ".[dev]"
pytest

Reproducibility notes

  • BASE is the memory-free reference setting; RAG, Mem0, MemU, and MEMOS are memory-augmented settings evaluated in the paper.
  • The checked-in prompts preserve the task and scoring definitions used by the benchmark.
  • Repetition uses cosine similarity after explicit L2 normalization.
  • API failures are recorded per question, retries are bounded, and result files are written atomically so long runs can be resumed.
  • Set use_both_personas: true in a local configuration when both personas are required; the default follows the original scripts and uses the first persona in each conversation.

Website

The hosted OP-Bench website is available at https://yulinlp.github.io/OP-Bench/. The local interactive demo in web/ provides the benchmark construction, QA examples, and results views.

License and data

See NOTICE.md for provenance and redistribution notes. The repository currently does not declare a software license. LoCoMo data, model checkpoints, and third-party services remain subject to their original terms.

Citation

If OP-Bench is useful in your research, please cite:

@misc{hu2026opbench,
  title         = {OP-Bench: Benchmarking Over-Personalization for Memory-Augmented Personalized Conversational Agents},
  author        = {Hu, Yulin and Long, Zimo and Guo, Jiahe and Sui, Xingyu and Fu, Xing and Zhao, Weixiang and Zhao, Yanyan and Qin, Bing},
  year          = {2026},
  eprint        = {2601.13722},
  archivePrefix = {arXiv},
  primaryClass  = {cs.CL},
  url           = {https://arxiv.org/abs/2601.13722}
}

About

Benchmarking over-personalization in memory-augmented conversational agents: irrelevance, sycophancy, and repetition.

Topics

Resources

Stars

5 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages