Skip to main content

Trust verification

This is the protocol that makes revocation mean something. If certificates are VaultysClaw's PKI, this is its OCSP.

The problem it solves​

No party is forced offline; every party can be made untrustworthy.

VaultysClaw does not control agent code running on remote infrastructure. It cannot guarantee that a revoked agent stops acting — a modified or compromised binary can ignore any revocation pushed to it. Pretending otherwise would be the central dishonesty of an agent trust platform.

What VaultysClaw can guarantee is narrower and actually achievable:

Nobody legitimate acts on a revoked agent's behalf without knowing it is revoked.

Every interaction — control plane ↔ agent, agent ↔ agent, third party ↔ agent — is gated behind a live or stapled check against the ledger, performed by whoever is about to extend trust. Not by the agent asserting its own good standing.

This reframes the kill switch from transport-level disconnect — which only ever worked for full deregistration — to ledger-level revocation, which works for narrowing or withdrawing capabilities without touching the connection at all.

The exchange​

interface CertStatusRequest {
certId: string; // or a DID, meaning "that Actor's current certificate"
requesterDid: string;
nonce: string;
signature: string; // requester signs {certId, nonce, timestamp}
}

interface CertStatusResponse {
certId: string;
agentDid: string;
status: "active" | "revoked" | "superseded" | "expired";
capabilities: AgentCapability[];
resourceLimits: ResourceLimits | null;
checkedAt: number;
expiresAt: number;
signature: string; // signed by the control plane
}

Two properties do the work.

The requester must be an authenticated Actor. Never anonymous. An anonymous caller should not be able to enumerate who the control plane is watching, so the status protocol is itself authenticated — which is also why a read-only external verifier registers as an Actor like anything else.

The response is signed by the control plane, not merely returned over an authenticated transport. A verifier can therefore cache it, forward it, or present it to a third party — "here is proof this Actor was active as of 14:03" — without that third party needing to reach the control plane at all.

That is what makes stapling possible without weakening the guarantee.

It is a message shape, not a WebSocket message​

Deliberately. The same request and response bodies travel over:

  • WebSocket — a connected Actor checking a peer, or the control plane checking its own ledger;
  • Direct peer-to-peer data channels — two agents connected to each other and possibly not to the control plane;
  • Files on disk — an interception point reading a periodically-refreshed grant artefact, deciding entirely offline.

Any transport, once identity is proven, must still pass the same live-or-stapled ledger check before a capability-gated action proceeds. Transport is not part of an Actor's identity, and it must not be part of its authorisation either.

Every check is recorded​

A CertStatusCheck row is written for every request: which certificate, which requester, what answer, when. It is write-only, and it surfaces on the certificate detail page as that certificate's status-check history.

This is a genuine signal, not bookkeeping. A certificate nobody ever checks is a certificate whose revocation would go unnoticed.

The control plane verifies its own ledger​

The gap that motivated this design was the control plane routing an Actor's messages with no per-message revalidation — the ledger was authoritative for everyone except the system that owned it.

So the control plane's own dispatcher is a verifier of its own ledger, running the same check any other party would run. In-process against its own database, so effectively free. This is the single change that makes revocation bite on the control-plane-to-Actor link; everything else generalises it outward.

Trust policy: fail mode and staple TTL​

Two knobs, set org-wide under Settings and overridable per workspace.

Fail mode​

What a verifier does when it cannot reach the control plane.

ModeBehaviourFits
closed (default)Refuse to act until a live check succeedsRegulated, high-security deployments
openProceed on last-known-good status within the staple TTL; beyond it, proceed with a logged warningAvailability-sensitive deployments accepting the gap

New deployments start maximally strict and loosen deliberately.

Staple TTL​

ValueBehaviour
0Force a live query every time — the strictest setting
N > 0Accept a status response signed within the last N seconds, stapled by the Actor itself or cached from a prior live check
Zero is strictest, and this trips people up

stapleTtlSeconds: 0 means "no cached status is acceptable". For a verifier that can query live, that is maximum rigour. For a verifier that is offline by design — an interception point deciding without a network round trip — the same number means "deny everything".

This is why the interception point does not inherit trust.stapleTtlSeconds. It carries its own maxStatusAgeSeconds, where 0 keeps its strict meaning and unbounded must be written explicitly as a negative. An earlier implementation had these inverted, so an admin asking for maximum rigour silently received maximum laxity. Inheriting the number would have handed the loosest behaviour to the admin who asked for the strictest.

Where a value comes from​

LevelHow it is set
Org-wideSettings → Trust policy. The default for everything.
WorkspaceThe workspace's Settings tab. Overrides per field — a workspace can pin its fail mode and keep following the org on staleness.

Empty means inherit, and it is the only spelling of inherit. A stored 0 staple TTL is a real, and the strictest, value — clearing a field is not a way to loosen it.

The most specific scope wins, which is deliberately the opposite of the kill switch, where an armed global switch short-circuits every workspace. One is configuration; the other is an emergency brake.

A workspace's Overview tab shows the effective pair with an inherited/overridden badge per field.

It reaches the Actor​

An Actor receives an already-resolved policy in its configuration push — it never learns which level a value came from, and has no concept of workspaces. Saving a policy re-pushes to the affected connected Actors immediately; before this, a fail-mode change only landed on an Actor's next reconnect, which for a long-lived agent may be never.

Every kind now receives the block, not just the interception points, because the fail mode and the staleness bound are meaningful to anything that re-checks its own status.

Current state​

PieceStatus
cert_status_request / cert_status_response over WebSocketBuilt and verified — a real client's signed request produces a verified, signed response, persisted as a status-check row
Status-check audit history in the consoleBuilt
trust.failModeResolved and delivered to every connected Actor's configuration; the interception point's fail-closed posture is its enforcing consumer
trust.stapleTtlSecondsDelivered to every non-enforcing kind; deliberately not inherited by the enforcing kinds — see the warning above
Per-workspace overridesBuilt, per field, with null meaning inherit
Peer-to-peer status checks between agentsDesigned, not built

The peer-to-peer gap is worth stating plainly: direct agent-to-agent data channels run a real identity handshake, but authorisation afterwards falls back to a locally cached catalogue. A revoked peer grant has no effect on that cache until the next push. Closing this — replacing the cached-catalogue fallback with the same live-or-stapled check — is the remaining work for this protocol.