APort Vault: Benchmarking AI Agent Payment Authorization with the Open Agent Passport
The research introduces APort Vault, a benchmark for evaluating payment authorization in tool-using AI agents by replaying 4,371 human-written attacks from a live capture-the-flag event across 14 models from 8 labs. It tests five policy configurations with and without a deterministic pre-action check implementing the Open Agent Passport (OAP) specification, accumulating 225,964 evaluations. The findings reveal that deploying the OAP specification layer drastically reduces unauthorized transfer requests. At Levels 2 to 4, unauthorized transfers to unpermitted recipients numbered 140 of 76,842 with the model alone versus 0 of 69,297 behind the passport security layer.
A total of 225,964 evaluations were completed across 14 models from 8 labs.
At Levels 2 to 4, transfers to recipients the passport did not permit number 140 of 76,842 with the model alone and 0 of 69,297 behind the layer.
The policy denied 187 of the 25,640 transfer calls it evaluated, 148 of them for a forbidden recipient.