Source · Tier B · Preprint

Proof-of-Guardrail in AI Agents and What (Not) to Trust from It

X. Jin, M. Duan, Q. Lin, A. Chan, Z. Chen, J. Du, X. Ren. 2026. arXiv.

Originalhttps://arxiv.org/abs/2603.05786
arXiv2603.05786
VersionarXiv v1 (2026-03-06) and v2 (2026-06-26), read via the arXiv HTML renderings on 2026-09-23 and 2026-09-25; the abs page did not render for the fetch tool. v2 keeps the figures the records cite (34% average latency overhead, Table 2 per-step overheads of 24.8–38.0%, F1 0.56 on the unsafe class).
Accessed2026-09-25
NoteThe paper header names the Trustworthy AI for Good (AI4GOOD) workshop at ICML 2026, with the ICML template's "PMLR 306" line; PMLR volume 306 is set aside for the ICML 2026 main conference (github.com/mlresearch/v306), so the paper is recorded as a workshop paper and tier B preprint. Authors are affiliated with Sahara AI and the University of Southern California.

Cited by

Search

Full search page