The Incommensurable Grade — One Word, Three Meters · Post 1 — Luminity Digital
One Word, Three Meters  ·  Series 20  ·  Post 1 of 5  ·  June 2026
One Word, Three Meters · Series 20

The Incommensurable Grade

The prologue established that the four maturity models grade three different objects and share no common zero. This post makes the harder case: the three units fail independently of one another, and none of them fails because the model is weak. That is the difference between a maturity score that is imprecise and one that is wrong.

June 2026 Tom M. Gomez Luminity Digital 6 Min Read
This is Post 1 of One Word, Three Meters. The prologue established that four stewards — OWASP, SANS, CSA, and NIST — grade three different objects and share no common zero. This post takes the next step: that the three units of assessment fail independently, and that none of them fails because the model is weak. For the enterprise architect, that is where “imprecise” becomes “wrong.”

We adopt three names for the units the stewards grade.

Program maturity is the grade attached to the enterprise’s governance system — its policies, inventory, guardrails, oversight. Agent maturity is the grade attached to a single agent’s earned trust — what this actor has demonstrated it can be permitted to do. Fabric maturity is the grade attached to the interoperability substrate — whether agents can coordinate across boundaries without producing invalid states.

These three names, and the recognition that the agentic maturity models in circulation each grade exactly one of them, are our analytical contribution. They are not drawn from any single external framework. Each maps to a real published model — program to OWASP and SANS, agent to CSA, fabric to NIST and the interoperability literature — but the diagnosis that they are three separate objects, gradeable independently and not reducible to one another, is ours.

Independence is the whole point

A taxonomy is only interesting if the categories come apart in the world. These do, and the empirical record shows it cleanly: each unit has a failure signature that the other two cannot produce and cannot repair.

The strongest single piece of evidence sits under the fabric unit. Cemri and colleagues at UC Berkeley, in work accepted to the NeurIPS 2025 Datasets and Benchmarks Track, assembled the first systematic failure taxonomy for multi-agent LLM systems from 1,642 annotated execution traces across seven frameworks. Their headline finding is a per-framework failure rate between 41% and 86.7%. The decisive detail is where those failures come from: they cluster into system-design, inter-agent-misalignment, and task-verification categories — coordination and specification, not raw model capability. A more capable model does not close a coordination failure. The failure lives in the fabric, and the fabric is a different object than the model.

That single result does the structural work. If the dominant failure mode of multi-agent systems were model capability, the three units would collapse into one — buy a better model, raise every grade. They do not collapse, because the failures are unit-specific. Governance failures come from an immature program. Trust failures come from autonomy extended past what an agent has earned. Coordination failures come from an immature fabric. No one of these is downstream of the others.

The container fallacy

The independence has a consequence that is easy to state and uncomfortable to sit with. A high grade on one unit does not certify the others — and, more sharply, a high grade on the program does not certify any single decision the program governs.

This is not only our claim. Solozobov, in a 2026 preprint proposing a decision-evidence maturity model, names the error directly as a container fallacy: a high organization-level maturity rating can coexist with failure on a specific audit question, because the organizational grade is a container that does not reach inside to the individual decision. The program can be mature and the decision still indefensible. The board that reads the program grade as coverage of the decision has committed the fallacy in exactly the form Solozobov describes.

We arrive at the same place from the measurement argument: the program grade and the decision are different units, and a grade does not transfer across units any more than a temperature transfers into a pressure.

Why the single number is wrong, not merely rough

It is tempting to treat all of this as a call for precision — to grant that one number is a simplification and move on. That concession is too generous. A single maturity number is not a lossy summary of three grades. It is a category error, because the three grades have no common unit to be summarized into.

Consider the three statements an honest assessment would produce: the governance program is at Stage 3; the agents in production are junior-tier; the interoperability fabric is at the floor. Each is meaningful on its own scale. There is no operation that combines them. “Stage 3” cannot be averaged with “junior-tier” cannot be averaged with “floor,” because averaging requires a shared dimension, and there is none. The board that demands “our agentic AI maturity, one number” is asking for the sum of a temperature and a sound level. The number it receives back will be confident, will be reportable, and will mean nothing.

The remainder of this series takes the three units one at a time — program, then agent, then fabric — not because they are ranked, but because each has to be understood on its own scale before a board can hold all three at once. That holding-of-three is the subject of the final post.

Editorial Position

The error this series exists to prevent is not under-measurement. It is the confident composite — the single agentic maturity score that a steering committee can put on a slide.

That number is not imprecise; it is incoherent, because it sums grades that share no unit. The architect’s job is to refuse the composite and report the vector: three grades, three scales, named.

Refuse the Composite. Report the Vector.

If your assessment is being compressed into a single agentic maturity number, a practitioner read on what that number hides is one conversation away.

Start the conversation
One Word, Three Meters  ·  Series 20  ·  Prologue + 5 Posts
Prologue  ·  Published No Shared Zero
Post 01  ·  Now Reading The Incommensurable Grade
Post 02  ·  Published The Grade the Board Already Trusts
Post 03  ·  Published Earned, Not Granted
Post 04  ·  Published The Meter Nobody Read
Post 05  ·  Published Reading the Vector
References & Sources

Share this:

Like this:

Like Loading…