The fabric is the unit with the most failure and the least instrument.
Program maturity has two disciplined stewards. Agent maturity has a trust ladder and a certificate grammar. The fabric — whether agents can coordinate across boundaries without producing invalid states — has the heaviest empirical failure record of the three and a maturity meter that is still being written. That combination is the subject of this post, and it is the reason the title is what it is.
Where multi-agent systems actually fail
The evidence that the fabric is the dominant failure surface is the strongest single result in the series. Cemri and colleagues, in work accepted to the NeurIPS 2025 Datasets and Benchmarks Track, built the first systematic failure taxonomy for multi-agent LLM systems from 1,642 annotated execution traces across seven frameworks, and found per-framework failure between 41% and 86.7%. The failures cluster into system-design, inter-agent-misalignment, and task-verification — coordination and specification. They are not, in the main, failures of the underlying model.
That location matters precisely. The failure does not sit in the model, which is the program’s concern and the object the program unit’s tooling can address. And it does not sit in the transport — the message-passing layer between agents — which, as the next section shows, is largely solved. It sits in the layer above transport and above the model: the coordination fabric, where agents hand work to one another, misread one another’s state, and fail to verify what came back. That is the object fabric maturity is supposed to grade, and it is where the systems break.
The meter that isn’t finished
The steward designated for this unit is NIST, and its instrument is not yet built. In February 2026, NIST announced the AI Agent Standards Initiative, whose stated aim is to foster industry-led technical standards and protocols for an interoperable agent ecosystem — through convenings, requests for information, listening sessions, and gap analyses. NIST’s own language is that it will announce research, guidelines, and further deliverables in the months ahead. An interoperability profile is anticipated, not published. The existing AI Risk Management Framework remains the governance anchor, but it does not stage interoperability across graded levels, and NIST has not released a separate scale that does.
The academic record agrees on the state of play. Work surveying agent interoperability describes the standards landscape as nascent — emerging, unsettled, not yet consolidated into the kind of consensus instrument the program and agent units already have. So the fabric grade is not a meter that boards ignore. It is a meter still under construction. The unit with the most failure has the least-developed scale, and that asymmetry is the entire diagnosis of this post.
What carries the fabric today
In the absence of a maturity grade, the fabric is carried by open protocols — and on the transport question, they have largely succeeded. Two layers matter. The Model Context Protocol (MCP) standardizes how an agent connects to tools and data; the Agent2Agent protocol (A2A) standardizes how agents communicate with one another. MCP is the vertical layer, A2A the horizontal.
Both are now under neutral governance, which is what makes them infrastructure rather than vendor plays. Google contributed A2A to the Linux Foundation in 2025, where it now operates as a vendor-neutral project backed by more than 150 organizations. Anthropic donated MCP to the Agentic AI Foundation, a directed fund under the Linux Foundation, in December 2025, with Block, OpenAI, Google, Microsoft, and AWS among the supporting parties. Competing platform vendors agreeing to common connective infrastructure is the precondition for a fabric that can be graded at all.
The distinction this draws is the one the failure data already insists on. The protocols solved transport: agents can discover one another, exchange messages, and pass tasks across vendor and framework boundaries. The 41–86.7% failure does not live there. It lives in the coordination above the protocol — in how agents are composed, how their states are kept consistent, and how their outputs are verified. Fabric maturity is a grade on that coordination layer, not on the transport the open standards have already settled.
Reading a meter that barely exists
An architect cannot wait for NIST to finish. The fabric still has to be graded, and in the absence of a published scale the best available rubric is the failure taxonomy itself. Cemri’s categories — system-design soundness, inter-agent alignment, task verification — are, read constructively, the dimensions a fabric maturity assessment would measure. An organization can grade its own coordination layer against them today: does the composition have a coherent design, do agents share a consistent picture of state, is every inter-agent handoff verified. A fabric that fails those is at the floor regardless of how mature the program or how well-earned the individual agents.
This is the unit most likely to be left unmeasured, because it is the one without a steward’s number to copy onto a slide. It is also the unit the empirical record says is most likely to take a system down. The discipline the series has argued for the other two units is sharpest here: the fabric grade is real even when the official meter is unfinished, and a board that reports a program grade as the whole posture has skipped the exact measurement most predictive of failure. The final post takes up how to hold all three grades at once.
The fabric is the meter nobody read — not because boards ignore it, but because its steward has not finished writing it. The interoperability profile is forthcoming; the maturity scale does not yet exist.
Yet this is the unit where the failure record is heaviest. Grade the coordination layer by its failure signature now — design soundness, state consistency, handoff verification — rather than waiting for an official level. The most consequential of the three grades is the one a board is most likely to leave blank.
