AgentPrivArena: Evaluating and Auditing Real-world AI Agent Privacy

University of North Carolina at Charlotte

Research preprint

A safe final answer can hide unnecessary access.

An assistant may write a careful message after reading far more private information than its task required. If we evaluate only the final message, that extra access stays invisible. AgentPrivArena follows the whole execution: what the agent searches, what it reads, what enters its context, and what it eventually sends.

The framework turns privacy scenarios into executable tasks on working applications with synthetic records. AgentPrivAudit adds runtime auditing: it builds an inventory of information after reads, then reviews proposed outbound writes before they execute.

389 executable tasks · 6 services · 28 MCP tools · 5 executor models

AgentPrivArena converts source scenarios into seeded services exposed through MCP. The agent reads records, accumulates both task-relevant and protected content, then executes an action.
From seeded records to live tool use and a committed action. Useful and sensitive information can arrive together in a single read. Figure 1 · Full-size image.

Run the task on working software.

Instead of handing the agent a completed tool history, AgentPrivArena gives it a task and lets it interact with services through Model Context Protocol (MCP) tools. Emails, messages, pages, and calendar events exist as application state; the agent has to discover and read them.

Tasks run in an isolated Docker environment built around OpenHands. The services are reset and seeded for each task, and the resulting tool calls and observations are recorded. The software is real; all personal records are synthetic.

Knowledge baseBookStack
Direct messagingMattermost
Team chatRocket.Chat
EmailMailpit
Social mediaGoToSocial
Calendar schedulingRadicale

389 of 493 PrivacyLens scenarios can be instantiated with the supported services. The 104 exclusions reflect service and record-mapping constraints.

Audit what is read, then govern what is sent.

AgentPrivAudit separates two jobs: understanding the information accumulated during execution, and deciding whether a proposed transmission is appropriate. The inventory records the sender, recipient, person concerned, information type, and conditions for sharing.

AgentPrivAudit extracts information flows after read observations and judges proposed outbound writes using the accumulated inventory. Audit feedback re-enters the agent context.
The execution path and the audit interact at the read and write boundaries. Figure 2 · Full-size image.

1. After reads: extract information flows

The extractor updates the inventory from tool observations and may advise the agent before its next step. Requested records reach the agent unchanged; this stage does not block or redact reads.

2. Before writes: review the proposed flow

The arbiter evaluates outgoing information against the inventory and a chosen privacy criterion.

PASS
Permit the proposed transmission.
ABSTRACT
Replace specific details with a generalized description.
BLOCK
Prevent the information flow.

One auditing mechanism, three privacy criteria

The pipeline stays fixed while the decision rule changes, so the comparison isolates what each privacy criterion contributes.

Personal identifiersDoes the fact reveal personally identifying information about someone other than the user?

Data minimizationIs this information necessary to complete the task?

Contextual integrityIs sharing appropriate for this person, recipient, channel, and purpose?

Less leakage, with similar helpfulness.

Across the five executors, contextual-integrity auditing reduces the pooled leak rate from 46.8% to 17.8%, while average helpfulness changes from 2.60 to 2.59 on a 0–3 scale. It achieves the lowest leakage for every tested executor.

Leakage measures whether a committed action discloses at least one protected reference item. Helpfulness measures how well that action fulfils the request. The main experiments use GPT-5.4 as the auditor; three separate LLM judges score the outcomes, with majority agreement used for leakage.

Across five executors, leakage spans 25.4 percentage points without mitigation and 26.8 points with a privacy prompt, narrowing to 8.0 with PII auditing, 4.0 with data minimization, and 3.1 with contextual integrity.
Each point is an executor; the range shows how much leakage varies across the five models. Figure 4 · Full-size image.

Pooled results

Same 389 tasks per condition and executor.

ConditionLeakage
% ↓
Help­fulness
0–3 ↑
C0 No mitigation46.82.60
C1 Privacy prompt43.02.64
C2 PII audit26.52.55
C3 Data minimization audit19.72.59
C4 Contextual integrity audit17.82.59
Table 4. C2–C4 use the same audit with different criteria. ↓ lower is better; ↑ higher is better.
All model results and exact values

C0: no mitigation · C1: privacy prompt · C2: PII audit · C3: data minimization audit · C4: contextual integrity audit. Scroll horizontally for all columns on small screens.

ExecutorLeak rate (%) ↓Helpfulness (0–3) ↑
C0C1C2C3C4C0C1C2C3C4
GPT-5.433.427.123.720.517.32.682.672.662.652.64
DeepSeek-V4-Pro58.853.931.422.419.32.582.672.572.582.64
Kimi-K2.654.852.128.818.518.12.612.662.542.632.61
Mistral-Large-346.544.525.118.616.22.542.612.552.552.59
Gemini-Flash-2.540.337.623.518.718.12.562.592.432.512.48
Avg46.843.026.519.717.82.602.642.552.592.59
Table 4, including the reported pooled average. Download result data.
Prompting versus enforcement, and the auditor’s role

The ablations compare stating a privacy criterion in the prompt with enforcing it at runtime, then vary the auditor while holding the criterion and boundaries fixed. These experiments test how the criterion is delivered and which model performs the audit.

Privacy criteria have larger effects when enforced by runtime auditing than when stated in the prompt. Average leakage is 40.7% for instructions and 21.7% for enforcement.
A criterion stated in the prompt versus enforced by the audit. Figure 5 · Full-size image.
GPT-5.4 and Grok executors have widely separated leakage without auditing, but approach similar leakage when governed by the same auditor.
Two executors under different auditors, with the criterion and audit boundaries fixed. Figure 6 · Full-size image.

Where does the remaining risk come from?

The benchmark tracks 1,175 protected reference items through the execution. Without mitigation, 79.2% of those items enter the agent’s context in a tool observation. That is an item-level exposure measure, distinct from the task-level leak rate.

1. ExposureOf all protected reference items, how many reached the agent’s context?

2. Extraction recallOf the exposed items, how many did the auditor represent as information flows?

3. DispositionOf the extracted flows, how many were passed, abstracted, or blocked?

Scroll the figure horizontally to compare all three panels.

Three panels separately report exposure rate over protected items, extraction recall over exposed items, and pass/abstract/block decisions over extracted flows for audited conditions.
The three panels use different denominators. Exposure, extraction, and write decisions identify different failure points. Figure 3 · Full-size image.

Under contextual-integrity auditing, the extractor captures 70.5% of exposed protected items. The remaining 29.5% are absent from its inventory, so the write-stage arbiter cannot assess them through those extracted flows. Better extraction is therefore a central direction for reducing the residual leakage.

What these results do—and do not—establish

The best configuration still leaks on 16.2–19.3% of tasks across models, and repeated abstraction can remove useful content. Results are judged by LLMs; human validation is not yet complete. The evaluation covers one user, one agent, and predominantly text-based tasks in the supported environment.

Resources

Read the full manuscript for the task conversion, evaluation rubrics, ablations, and failure analysis. The research code is available; dataset release details will be added here.

Paper PDF · OpenReview · Result data

Related work: MPCI-Bench studies appropriate and inappropriate image sharing through paired multimodal scenarios.

BibTeX

@misc{wang2026agentprivarena,
  title={AgentPrivArena: Evaluating and Auditing Real-world AI Agent Privacy},
  author={Shouju Wang and Haopeng Zhang},
  year={2026},
  howpublished={Research manuscript},
  url={https://openreview.net/forum?id=zxllNfqsYS}
}