Agent maturity grades the actor, not the system.
The program unit asks whether the organization governs agents well. This unit asks a narrower and more operational question: has this agent, doing this job, earned the autonomy it currently holds? The two come apart constantly. A mature program can field an unproven agent; an immature program can run a single agent that has earned considerable trust through a long track record on a narrow task. The grades do not predict each other, which is the whole reason the units are separate.
The promotion ladder
The Cloud Security Alliance’s Agentic Trust Framework grades this unit, and it does so with a metaphor an enterprise already understands: employment. A new agent enters as an intern — read-only, observed, permitted to propose but not to act. It is promoted to junior as it demonstrates reliability on bounded tasks, and upward from there as its track record warrants. Trust is not a property the agent has by default; it is a position on a ladder it climbs by performance.
The framework grades five elements of that earned trust: identity (the agent is who it claims to be, with a non-human identity that can be authenticated), behavior (it acts within its demonstrated envelope), data (it touches only what its role permits), segmentation (a compromise of one agent does not become a compromise of the estate), and incident response (its actions can be halted and reconstructed). An agent’s maturity is the joint grade across these five — not a single badge, and not transferable to the next agent the organization deploys.
What “earned” means in production
The ladder is the right shape because the production evidence says trust has to be earned narrowly and extended slowly. The systematic study of agents in production (MAP) — first-hand data from practitioners running real deployments — found that the great majority of deployed agents are kept on a short leash: roughly 68% run ten steps or fewer before a human checkpoint, and around 74% rely on human evaluation rather than autonomous self-assessment. Autonomy in the field is deliberately bounded. The organizations doing this are not behind; they are grading their agents the way the CSA framework prescribes.
They are bounding autonomy because the capability ceiling is real. Work on the hierarchy of agentic capabilities finds that even the strongest frontier models still fail a substantial share — on the order of 40% — of realistic multi-step tasks. An agent that fails two of every five real tasks has not earned unsupervised autonomy on those tasks, regardless of how mature the program governing it is.
And the cost of granting autonomy past what has been earned is not linear. Mitchell and colleagues at Hugging Face argue, building from the prior literature, that risk and compounded error increase with autonomy: as an agent takes more steps without a checkpoint, inaccuracies cascade across action surfaces rather than averaging out. This is an argued principle, not a field measurement — but it is the principle the bounded-autonomy data is already obeying. The leash is short because the failure mode is cumulative.
The grammar of granting
If trust is earned in increments, an organization needs a vocabulary for the increments. Work on levels of autonomy for AI agents supplies one: a graded ladder from fully supervised to fully autonomous, paired with the notion of an autonomy certificate — an explicit, revocable grant that says this agent may operate at this level for this task class. The certificate is the mechanism that keeps the agent grade honest. It names what was earned, scopes it to a task, and can be withdrawn when behavior drifts.
This is the operational counterpart to CSA’s ladder. The ladder describes where an agent stands; the certificate is how a grant at that level is issued, bounded, and revoked. Together they make agent maturity an auditable position rather than an assumption — which is exactly what a risk committee needs when it is asked to sign off on an agent acting without a human in the loop.
A trusted agent is not a trusted estate
The discipline of this unit is the same one the series keeps returning to, in the specific form the agent grade takes. A high agent grade certifies one agent on one class of task. It does not certify the program — a well-earned agent can run inside a weak governance system. And it does not certify the fabric — an agent that is entirely trustworthy alone can still produce invalid states when it has to coordinate with others, which is the failure the next post takes up.
The error this unit guards against is the inference that runs in both directions: granting an agent autonomy because the program is mature, or declaring the program mature because a flagship agent performs well. Both substitute one unit’s grade for another’s. Trust is earned by the actor, for the task, and certifies nothing beyond it.
Agent maturity is earned, per agent and per task, and granted only by an explicit, revocable certificate — never inferred from the maturity of the program around it.
An agent that fails a substantial share of real tasks has not earned unsupervised autonomy on them, however mature its governance. Read the agent grade as a track record on a bounded job, and refuse to let it stand in for the estate it does not measure.
