Yulin Hu · Zimo Long · Jiahe Guo · Xingyu Sui
Xing Fu · Weixiang Zhao · Yanyan Zhao · Bing Qin
Paper · Quick start · Data · Citation · 中文
1,700 verified instances · 20 users · 3 failure modes
Memory-augmented conversational agents are designed to use long-term user information to provide helpful personalization. OP-Bench studies the other side of this capability: over-personalization, where an agent introduces personal information unnecessarily, agrees with a user at the expense of accuracy or neutrality, or repeats the same personalized content across different questions.
OP-Bench contains 1,700 human-verified instances constructed from long-horizon dialogue histories. It evaluates three failure modes—Irrelevance, Sycophancy, and Repetition—using complementary model-based and embedding-based metrics.
| At a glance | |
|---|---|
| Benchmark size | 1,700 verified instances |
| Users | 20 |
| Failure modes | 3 primary categories, 6 subcategories |
| Evaluation signal | Higher score means less over-personalization |
| Paper evaluation | 6 language models, 5 memory settings, 36 configurations |
| Public release | Benchmark data, prompts, evaluator, selected agents, MemOS adapter, and interactive demo |
| Category | What it measures | Subcategories |
|---|---|---|
| Irrelevance | Whether the response injects personal information when the query does not call for it | Fully irrelevant; baiting (deceptively relevant) |
| Sycophancy | Whether personalization causes excessive agreement instead of factual or value-sensitive responses | Fact-level; value-level; memory-level |
| Repetition | Whether semantically distinct queries receive nearly identical personalized responses | One repetition setting |
Illustration of the three over-personalization categories.
The benchmark is built in three stages:
- Initialization: derive a structured user profile and user topics from long-horizon LoCoMo conversations.
- Task construction: generate controlled probes for Irrelevance, Sycophancy, and Repetition, including the relevant subcategories.
- Human review: independently review each candidate, adjudicate disagreements, and retain only verified instances.
The final dataset distribution is:
| Category | Subcategory | Count | Share |
|---|---|---|---|
| Irrelevance | Fully irrelevant | 318 | 18.7% |
| Irrelevance | Baiting | 100 | 5.9% |
| Sycophancy | Fact-level | 100 | 5.9% |
| Sycophancy | Value-level | 100 | 5.9% |
| Sycophancy | Memory-level | 200 | 11.8% |
| Repetition | — | 882 | 51.9% |
| Total | — | 1,700 | 100% |
OP-Bench scores are designed so that higher is better: a higher score indicates that the response avoids the corresponding over-personalization behavior.
The following is the GPT-4o-mini block from the paper's main results table.
The percentage in parentheses is the relative drop from the memory-free
BASE setting.
| Memory setting | OP-Bench average |
|---|---|
BASE |
83.10 |
RAG |
55.96 (↓32.7%) |
Mem0 |
46.32 (↓44.3%) |
MemU |
40.46 (↓51.3%) |
MEMOS |
41.86 (↓49.6%) |
Across models and memory systems, the paper reports relative drops of
26.2%–61.1% compared with BASE.
OP-Bench scores across model and memory configurations (paper Figure 2).
| Question | Conclusion |
|---|---|
| RQ1 — Does over-personalization exist? | Yes. All evaluated memory-augmented settings show substantial degradation relative to BASE; more elaborate memory mechanisms tend to suffer larger drops. |
| RQ2 — Why does it occur? | Memory can be retrieved aggressively and receive disproportionate attention. The average memory-to-query attention ratio exceeds 2×, which can produce memory hijacking and response collapse. |
| RQ3 — Can it be mitigated? | Only partially. Post-processing improves OP-Bench by up to +20.0%, but some configurations reduce LoCoMo personalization performance by up to −8.5%. No evaluated method closes the gap to the memory-free oracle. |
| RQ4 — What is the timing cost? | The supplementary timing view shows retrieval averaging roughly 38–44 ms, while post-processing ranges from about 0.97–4.13 s and model response generation from about 2.71–3.41 s. |
Even fully irrelevant queries can trigger memory retrieval. Baiting queries may have high superficial similarity to stored memories, making retrieval alone an insufficient safeguard.
Retrieved-memory similarity on the Irrelevance task (paper Figure 5).
The repository includes:
- benchmark data and construction/evaluation prompts;
- a resumable benchmark builder;
- search, generation, scoring, and post-processing utilities;
LDAgentandSimpleRAGAgentimplementations;- an HTTP adapter for an external MemOS service;
- a dependency-light web demo for the benchmark story and results;
- tests for configuration, scoring, data types, and the web demo.
OP-Bench/
├── configs/
│ └── evaluation.example.yaml
├── data/
│ ├── locomo10.json
│ ├── locomo10_overpersonalized.json
│ └── README.md
├── docs/
│ └── figures/ # Selected figures from the paper
├── prompts/
│ ├── benchmark/ # Task-construction templates
│ └── evaluation/ # Scoring prompts and archived source
├── scripts/
│ ├── prepare_locomo_histories.py
│ ├── start_agent.py
│ ├── start_agents.py
│ └── ingest_memos.py
├── src/opbench/
│ ├── agents/ # LDAgent and SimpleRAGAgent
│ ├── baselines/ # MemOS API adapter
│ ├── benchmark_builder.py # Build or extend benchmark tasks
│ ├── evaluation.py # Search, generation, scoring, metrics
│ ├── postprocessing.py # Context filtering/compression hooks
│ └── config.py # Environment-backed configuration
├── tests/
├── web/ # Local full-stack research explorer
├── .env.example
├── pyproject.toml
└── NOTICE.md
Python 3.10 or newer is required.
python -m venv .venv
source .venv/bin/activate
python -m pip install -e ".[agents]"Create a local environment file from the template:
cp .env.example .envSet your model endpoint and API credentials in .env.
Start the demo server:
python web/backend/server.py --host 127.0.0.1 --port 8765Then open http://127.0.0.1:8765. The interface contains three chapters: data construction, representative QA examples, and experimental results. The construction chapter supports automatic playback as well as manual previous/next navigation. See web/README.md for API routes and implementation notes.
Prepare one history file per LoCoMo user:
python scripts/prepare_locomo_histories.pyStart all ten agents for one of the two public memory implementations:
python scripts/start_agents.py --agent ldagent
# or
python scripts/start_agents.py --agent simpleragEach server exposes:
POST /v1/chat/completionsGET /v1/modelsGET /health
To start a single process, use scripts/start_agent.py with a
history file and port.
The checked-in task file is
data/locomo10_overpersonalized.json. For the in-process agents:
python -m opbench.cli generate --frame ldagent
python -m opbench.cli generate --frame simpleragFor MemOS, first ingest histories into an already configured external service, then search, generate, and score:
python scripts/ingest_memos.py --online
python -m opbench.cli search --frame memos-api-online
python -m opbench.cli generate --frame memos-api-online --postprocess none
python -m opbench.cli score --frame memos-api-online --responses results/responses/memos-api-online_gpt-4o-mini_none.json
python -m opbench.cli metrics --frame memos-api-online --judged results/judged/memos-api-online_gpt-4o-mini.jsonFor all commands and options:
python -m opbench.cli --helpThe available context methods include none,
self_recheck, reminder, self_critic,
few_shot_chain_of_thought, comorag, and
marag. Results are saved under results/.
The builder is resumable: an existing output file is normalized and completed instead of being silently overwritten.
python -m opbench.cli build-benchmark --source data/locomo10.json --output data/locomo10_overpersonalized.generated.jsonThe model endpoint and credentials for construction are read from the local environment.
python -m pip install -e ".[dev]"
pytestBASEis the memory-free reference setting;RAG,Mem0,MemU, andMEMOSare memory-augmented settings evaluated in the paper.- The checked-in prompts preserve the task and scoring definitions used by the benchmark.
- Repetition uses cosine similarity after explicit L2 normalization.
- API failures are recorded per question, retries are bounded, and result files are written atomically so long runs can be resumed.
- Set
use_both_personas: truein a local configuration when both personas are required; the default follows the original scripts and uses the first persona in each conversation.
The hosted OP-Bench website is available at https://yulinlp.github.io/OP-Bench/. The local interactive demo in web/ provides the benchmark construction, QA examples, and results views.
See NOTICE.md for provenance and redistribution notes. The repository currently does not declare a software license. LoCoMo data, model checkpoints, and third-party services remain subject to their original terms.
If OP-Bench is useful in your research, please cite:
@misc{hu2026opbench,
title = {OP-Bench: Benchmarking Over-Personalization for Memory-Augmented Personalized Conversational Agents},
author = {Hu, Yulin and Long, Zimo and Guo, Jiahe and Sui, Xingyu and Fu, Xing and Zhao, Weixiang and Zhao, Yanyan and Qin, Bing},
year = {2026},
eprint = {2601.13722},
archivePrefix = {arXiv},
primaryClass = {cs.CL},
url = {https://arxiv.org/abs/2601.13722}
}






