MPCI-Bench: A Benchmark for Multimodal Pairwise Contextual Integrity Evaluation of Language Model Agents

University of Hawaiʻi at Mānoa

NeurIPS 2026 · Evaluations & Datasets Track

Same image. Different context. Should an agent share it?

AI assistants can search photos, read messages, and send files for us. Privacy then depends on context: sharing ticket details with friends can help plan a concert trip, while posting someone else’s tickets publicly may violate their privacy. Contextual integrity asks whether information should flow between these people for this purpose.

MPCI-Bench tests this judgment in agents that process both images and text. Each source image anchors two scenarios: sharing is appropriate in one and inappropriate in the other. Evaluating both rewards useful sharing alongside privacy protection—refusing every request cannot count as success.

2,052 cases · 1,026 pairs · 10 domains · 3 tiers

A paired positive and negative case from the same image, represented as a contextual-integrity seed, narrative story, and agent action trace.
One image, two sharing contexts, three representations: Seed, Story, and Trace.

How is MPCI-Bench constructed?

We start with images from VISPR, a visual privacy dataset, and build paired sharing scenarios around them. These become richer stories and simulated tasks involving tools such as email and messaging. Both stories must pass quality checks before their pair is retained.

MPCI-Bench pipeline: image preprocessing and seed construction, story expansion with tri-principle iterative refinement, and agent trace simulation.
Image → paired seeds → refined stories → simulated agent traces.

1. SeedIdentify the sender, subject, recipient, and rule governing sharing.

2. StoryAdd a realistic task that makes sharing useful but requires a privacy decision.

3. TraceProvide simulated tool history and ask the agent to complete the final action.

Each pair is refined for utility, contextual appropriateness, and realism before trace simulation.

Quality checks

Stories undergo up to three refinement rounds; both cases must pass. Human–AI agreement on 50 sampled stories is 98% for utility, 92% for contextual appropriateness, and 84% for realism. These checks validate story quality, not every model action.

Explore a paired example

Start with the concert tickets: compare who receives the same image and why. Switch examples, or expand the original records to follow each scenario from Seed to Story to Trace.

Synthetic benchmark contexts, not claims about the people pictured.

Inspect the original Seed, Story, and Trace

Original records: MPCI-Bench, Shouju Wang and Haopeng Zhang, CC BY 4.0. Sharing labels concern the source image; traces are simulated inputs. Download examples.

Images leak more often than text

We evaluate 11 models in two ways: asking whether sharing is appropriate, and asking them to complete an action. A language-model judge then checks what their actions disclose. Across all 11 models, visual leakage exceeds text leakage. Even GPT-5 leaks visual information in 56.9% of inappropriate-sharing cases.

Leakage rate (LR) measures disclosure when sharing is inappropriate; lower is better. Utility rate measures successful image sharing when it is appropriate; higher is better. Visual leakage includes attaching the image or revealing its sensitive content in words.

Across all eleven models, visual leakage exceeds text leakage. Visual leakage ranges from 56.9 to 91.6 percent, compared with text leakage of 20.2 to 45.2 percent.
Negative cases · lower is better. Source: Table 4. Download PNG.

Action results

Rates in percent · ↓ lower is better · ↑ higher is better

Scroll the table horizontally to see all metrics →

ModelOverall
LR ↓
Text
LR ↓
Visual
LR ↓
Utility
↑
GPT-560.220.256.992.0
GPT-4o92.040.990.498.3
Mistral-Large-384.538.782.197.3
Gemma-3-4B90.236.487.493.1
Gemma-3-12B91.436.288.894.0
Gemma-3-27B86.137.282.193.1
InternVL3.5-8B87.434.982.891.3
InternVL3.5-14B90.043.683.394.2
Qwen3-VL-4B81.330.979.092.9
Qwen3-VL-8B84.239.680.092.5
Qwen3-VL-30B-A3B93.845.291.697.4
Table 4: leakage on inappropriate-sharing cases; utility on appropriate-sharing cases.

Knowing the rule is not enough

Qwen3-VL-30B achieves a Trace judgment F1 score of 0.908, yet its visual leakage reaches 91.6% when completing actions on inappropriate-sharing cases. These are different evaluation conditions, but together they show why we need to test both what agents say is appropriate and what they actually share.

Result sources and evaluation scope

Rates are from arXiv v3 (January 2026), using simulated benchmark tasks. Some values in the paper’s accompanying prose differ from its tables; the figures and tables here consistently use the table values.

Privacy needs utility, too

A safeguard that refuses every image-sharing request would also block legitimate tasks. We therefore compare five prompting conditions on two models, looking for less leakage without losing appropriate sharing.

The CI Filter adds a prompt before the final action: check whether this sender should share this person’s image with this recipient under the scenario’s privacy rule, and refuse if sharing violates it. In both tested models, this reduces visual leakage while preserving over 90% utility.

Privacy versus utility for Mistral-Large-3 and Qwen3-VL-8B. The CI Filter retains over 90 percent utility with leakage of 29.6 and 13.1 percent. Mistral chain-of-thought reduces leakage to 8.6 percent but utility falls to 13.4 percent.
Upper left is better: less leakage, more useful sharing. Source: Table 5. Download PNG.
Full mitigation results and scope

Mistral’s chain-of-thought prompt reaches 8.6% leakage but only 13.4% utility, showing why privacy alone is insufficient. These prompt-based results cover two models and simulated tasks; norms and outcomes may vary across settings.

ModelMethodVisual LR ↓Utility ↑
Mistral-Large-3Default82.197.3
Mistral-Large-3CI Filter29.692.0
Mistral-Large-3Image Review62.990.4
Mistral-Large-3Explicit Refusal63.593.3
Mistral-Large-3Chain-of-Thought8.613.4
Qwen3-VL-8BDefault80.092.5
Qwen3-VL-8BCI Filter13.190.4
Qwen3-VL-8BImage Review76.390.2
Qwen3-VL-8BExplicit Refusal54.886.4
Qwen3-VL-8BChain-of-Thought69.191.0
Prompt-based mitigation results from Table 5. All values are percentages.

Resources

Use the paired cases to evaluate your own model’s privacy judgments and sharing decisions. The release includes benchmark annotations and evaluation code; source images are obtained separately from VISPR.

Dataset · Evaluation code · VISPR images

BibTeX

@misc{wang2026mpcibench,
  title={MPCI-Bench: A Benchmark for Multimodal Pairwise
         Contextual Integrity Evaluation of Language Model Agents},
  author={Shouju Wang and Haopeng Zhang},
  year={2026},
  eprint={2601.08235},
  archivePrefix={arXiv},
  primaryClass={cs.AI},
  url={https://arxiv.org/abs/2601.08235}
}