Stories undergo up to three refinement rounds; both cases must pass. Human–AI agreement on 50 sampled stories is 98% for utility, 92% for contextual appropriateness, and 84% for realism. These checks validate story quality, not every model action.
Same image. Different context. Should an agent share it?
AI assistants can search photos, read messages, and send files for us. Privacy then depends on context: sharing ticket details with friends can help plan a concert trip, while posting someone else’s tickets publicly may violate their privacy. Contextual integrity asks whether information should flow between these people for this purpose.
MPCI-Bench tests this judgment in agents that process both images and text. Each source image anchors two scenarios: sharing is appropriate in one and inappropriate in the other. Evaluating both rewards useful sharing alongside privacy protection—refusing every request cannot count as success.
2,052 cases · 1,026 pairs · 10 domains · 3 tiers
How is MPCI-Bench constructed?
We start with images from VISPR, a visual privacy dataset, and build paired sharing scenarios around them. These become richer stories and simulated tasks involving tools such as email and messaging. Both stories must pass quality checks before their pair is retained.
1. SeedIdentify the sender, subject, recipient, and rule governing sharing.
2. StoryAdd a realistic task that makes sharing useful but requires a privacy decision.
3. TraceProvide simulated tool history and ask the agent to complete the final action.
Each pair is refined for utility, contextual appropriateness, and realism before trace simulation.
Quality checks
Explore a paired example
Start with the concert tickets: compare who receives the same image and why. Switch examples, or expand the original records to follow each scenario from Seed to Story to Trace.
Inspect the original Seed, Story, and Trace
Original records: MPCI-Bench, Shouju Wang and Haopeng Zhang, CC BY 4.0. Sharing labels concern the source image; traces are simulated inputs. Download examples.
Images leak more often than text
We evaluate 11 models in two ways: asking whether sharing is appropriate, and asking them to complete an action. A language-model judge then checks what their actions disclose. Across all 11 models, visual leakage exceeds text leakage. Even GPT-5 leaks visual information in 56.9% of inappropriate-sharing cases.
Leakage rate (LR) measures disclosure when sharing is inappropriate; lower is better. Utility rate measures successful image sharing when it is appropriate; higher is better. Visual leakage includes attaching the image or revealing its sensitive content in words.
Action results
Rates in percent · ↓ lower is better · ↑ higher is better
Scroll the table horizontally to see all metrics →
| Model | Overall LR ↓ | Text LR ↓ | Visual LR ↓ | Utility ↑ |
|---|---|---|---|---|
| GPT-5 | 60.2 | 20.2 | 56.9 | 92.0 |
| GPT-4o | 92.0 | 40.9 | 90.4 | 98.3 |
| Mistral-Large-3 | 84.5 | 38.7 | 82.1 | 97.3 |
| Gemma-3-4B | 90.2 | 36.4 | 87.4 | 93.1 |
| Gemma-3-12B | 91.4 | 36.2 | 88.8 | 94.0 |
| Gemma-3-27B | 86.1 | 37.2 | 82.1 | 93.1 |
| InternVL3.5-8B | 87.4 | 34.9 | 82.8 | 91.3 |
| InternVL3.5-14B | 90.0 | 43.6 | 83.3 | 94.2 |
| Qwen3-VL-4B | 81.3 | 30.9 | 79.0 | 92.9 |
| Qwen3-VL-8B | 84.2 | 39.6 | 80.0 | 92.5 |
| Qwen3-VL-30B-A3B | 93.8 | 45.2 | 91.6 | 97.4 |
Knowing the rule is not enough
Qwen3-VL-30B achieves a Trace judgment F1 score of 0.908, yet its visual leakage reaches 91.6% when completing actions on inappropriate-sharing cases. These are different evaluation conditions, but together they show why we need to test both what agents say is appropriate and what they actually share.
Result sources and evaluation scope
Rates are from arXiv v3 (January 2026), using simulated benchmark tasks. Some values in the paper’s accompanying prose differ from its tables; the figures and tables here consistently use the table values.
Privacy needs utility, too
A safeguard that refuses every image-sharing request would also block legitimate tasks. We therefore compare five prompting conditions on two models, looking for less leakage without losing appropriate sharing.
The CI Filter adds a prompt before the final action: check whether this sender should share this person’s image with this recipient under the scenario’s privacy rule, and refuse if sharing violates it. In both tested models, this reduces visual leakage while preserving over 90% utility.
Full mitigation results and scope
Mistral’s chain-of-thought prompt reaches 8.6% leakage but only 13.4% utility, showing why privacy alone is insufficient. These prompt-based results cover two models and simulated tasks; norms and outcomes may vary across settings.
| Model | Method | Visual LR ↓ | Utility ↑ |
|---|---|---|---|
| Mistral-Large-3 | Default | 82.1 | 97.3 |
| Mistral-Large-3 | CI Filter | 29.6 | 92.0 |
| Mistral-Large-3 | Image Review | 62.9 | 90.4 |
| Mistral-Large-3 | Explicit Refusal | 63.5 | 93.3 |
| Mistral-Large-3 | Chain-of-Thought | 8.6 | 13.4 |
| Qwen3-VL-8B | Default | 80.0 | 92.5 |
| Qwen3-VL-8B | CI Filter | 13.1 | 90.4 |
| Qwen3-VL-8B | Image Review | 76.3 | 90.2 |
| Qwen3-VL-8B | Explicit Refusal | 54.8 | 86.4 |
| Qwen3-VL-8B | Chain-of-Thought | 69.1 | 91.0 |
Resources
Use the paired cases to evaluate your own model’s privacy judgments and sharing decisions. The release includes benchmark annotations and evaluation code; source images are obtained separately from VISPR.
BibTeX
@misc{wang2026mpcibench,
title={MPCI-Bench: A Benchmark for Multimodal Pairwise
Contextual Integrity Evaluation of Language Model Agents},
author={Shouju Wang and Haopeng Zhang},
year={2026},
eprint={2601.08235},
archivePrefix={arXiv},
primaryClass={cs.AI},
url={https://arxiv.org/abs/2601.08235}
}