Public research guide
Research Map
A connected path through enterprise-agent architecture, runtime governance, evidence, and executable adversarial evaluation.
Core proposition
Authorization is necessary, not sufficient.
Enterprise agent systems need governance that evaluates more than who is authorized. It must also assess how an authorized agent behaves, what evidence supports an action, and how risk compounds across sessions. The work below moves from decision context to enterprise architecture, runtime governance, and executable adversarial evaluation.
Context
Why decision load matters
Decision Load Index All versions (concept DOI)
An optional entry point for the human and organizational context. It examines how unresolved decisions, inputs, commitments, and scope create cognitive burden before the discussion turns to autonomous systems.
1. Governance problem
Authorization is not behavior
Constitutional Self-Governance for Autonomous AI Agents All versions (concept DOI)
Introduces the WHO versus HOW gap: authentication, access control, and audit logging are necessary, but do not by themselves constrain an agent's decisions under changing conditions.
2. Architecture
Agents are a workforce, not an application feature
Enterprise Agent Architecture All versions (concept DOI)
Places autonomous agents within enterprise architecture as actors with delegated authority, capabilities, control-plane dependencies, and governance obligations.
3. Runtime control
Systems must be able to refuse authorized actions
Authorized but Refused All versions (concept DOI)
Provides the operational-telemetry layer: governance is credible only when a running system can constrain, record, and explain decisions made under granted authority.
4. Aggregate risk
Individually authorized actions can compound
Authorized but Composed All versions (concept DOI)
Explains why per-call approval is insufficient when sequences of individually acceptable actions accumulate across sessions.
5. Monitoring
Behavioral drift can normalize over time
Detecting Normalization of Deviance in Multi-Agent Systems All versions (concept DOI)
Extends runtime governance into monitoring: gradual behavioral drift can be missed by stateless or threshold-only monitoring.
6. Evaluation boundary
Identity controls do not test protocol behavior
Beyond Identity Governance All versions (concept DOI)
Moves from architecture and monitoring to adversarial evaluation across MCP, A2A, L402, and x402. The question is how an agent system behaves when those protocol boundaries receive hostile traffic.
7. Evidence standard
A receipt must prove what it claims
- Present vs. Provable All versions (concept DOI)
- Claim-Level Negative Testing All versions (concept DOI)
- Signing Is Not Authorization All versions (concept DOI)
These works define the evidence layer. Signatures, receipts, and checks do not alone establish that an executed action was authorized or that an artifact supports the claim it makes.
From Approval to Execution: Assurance Boundaries in Three Agent Protocols
Extends Claim-Level Negative Testing into a comparison of content commitment, authenticated approval, enforcement at execution, and evidence of execution across three agent protocols. Accompanying conformance corpus v0.1.1. Preprint v1.1, September 19, 2026. Not peer-reviewed; no claim of independent validation, certification, adoption, or production effectiveness.
Approval-to-execution concept DOI: all versions
Deception Primitives at an MCP-Aware Enforcement Point: A Bounded Reference Design for Honeytoken, Decoy-Tool, and Breadcrumb Controls
Connects runtime deception with evidence and enforcement boundaries: a trip is an investigation signal, not proof of malicious intent or containment. Reference design with reported library and bounded gateway behavior tests. Detection performance and operational effectiveness remain unvalidated; not production or independent validation. Pinned preprint v1.1 · Reference implementation v0.1.0-rc.1.
Token-Bleed R5 — claim-scoped reference application All versions (concept DOI)
A synthetic opaque-schema runtime characterization on a single local Ollama OpenAI-compatible configuration using qwen3-coder:30b. At 0% candidate-generator misses, its oracle-controlled bundled route changes both candidate membership and representation; it used 96.9%–97.9% fewer mean prompt tokens and higher mean F1 than verbose full context. The lexical token-efficiency/value claim was rejected under the frozen 3× rule. This ACE reference application keeps the narrow route-level finding separate from unsupported economic or generalization claims. It is not production, customer-data, ROI, cross-model, or independently raw-reproducible evidence. Source release · Pinned commit
Authority-to-execution replay packet
An I0 synthetic fixture contract, published for independent reimplementation; no I1 result is yet recorded. Its manifest pins packet SHA-256 afaf6090…, the replay steps, and three required controls — allow-exact-action, deny-wrong-target, and deny-post-approval-mutation — and asks that a verifier be implemented before the reference one is read. Its stated claim boundary: the corpus is synthetic, owned-fixture, and networkless; it is not a live MCP/API test, production evidence, external identity binding, security audit, certification, endorsement, or adoption claim.
8. Executable implementation
Adversarial evaluation as a practical instrument
Agent Security Harness All versions (concept DOI)
The practical evaluation instrument: executable adversarial security tests that operationalize protocol and governance claims.
9. Ecosystem model
Contributions need integrity and trust boundaries
Community-Driven Security for AI Agents All versions (concept DOI)
Sets out an approach for accepting adversarial-evaluation contributions while preserving integrity and trust boundaries. The linked version is v1.1, which corrects a CVE misattribution in v1.0.
Supporting research
A bounded evidence pack on observable behavior
AI News Evidence Pack All versions (concept DOI)
This evidence pack supports the broader interest in observable behavior and evidence discipline, but is not a prerequisite for the agent-governance path.
Research record
Where each record belongs
- Zenodo: archival record, DOI source, and version history.
- ORCID: curated identity and discovery record that links to canonical DOIs.
- PubPoint: this public reading guide and editorial context.
- GitHub: the relevant executable implementation and source artifacts.