From Incident to Evidence — Luminity Digital
AI Flaw Disclosure as Architecture  ·  Series 26  ·  Post 2 of 3  ·  July 2026
AI Flaw Disclosure as Architecture

From Incident to Evidence

A report sitting in an inbox is an anecdote. Turn thousands of anecdotes into a claim about who is being harmed and how, and you have evidence a board can act on. The distance between those two states is not effort or goodwill — it is schema.

July 2026 Tom M. Gomez Luminity Digital 13 Min Read
This is Post 2 of 3 in AI Flaw Disclosure as Architecture. Part 1 — The Reporting Gap Is an Architecture Gap — established that a missing disclosure surface is a design decision, not an operational oversight. This post takes the next layer: once a report has somewhere to go, what turns it from an isolated incident into evidence a risk committee can act on. Part 3 — Building the Disclosure Surface Into Agentic Systems — makes the surface structural for agents. The series draws on a verified 17-paper arXiv corpus; this post cites 7.

A report sitting in an inbox is an anecdote. Turn thousands of anecdotes into a claim about who is being harmed and how, and you have evidence a board can act on. The distance between those two states is not effort or goodwill — it is schema.

What separates a landfill of incident reports from a body of evidence is whether the reports were captured, structured, and aggregated so that a pattern can be detected and defended. That plumbing is architecture, and it is the subject of this post.

Collective memory, on purpose

The founding move was to treat AI failures the way aviation treats them: as a shared record to be learned from, not a set of private embarrassments to be buried. The AI Incident Database made that argument explicitly, casting cataloged incidents as a “collective memory” that lets the field avoid repeating real-world failures [1]. The insight is architectural, not sentimental — memory is a system property. If failures are not captured in a durable, queryable store, each recurrence is discovered fresh, at full cost, by whoever hits it next.

But collective memory is only as useful as its structure. A database of free-text incident narratives preserves stories; it does not, by itself, yield evidence.

The unit of value is the schema

This is where the standardization thread does its work. A unified schema for AI incident databases — with consistent fields for severity, cause, and harm type — is what makes incidents comparable across sources, sectors, and time [2]. Comparability is the precondition for every downstream claim: you cannot say “this harm is systemic” until “this harm” means the same thing in every record. The same author group cataloged the failure of the status quo directly, identifying nine recurring gaps across open-access incident databases and pairing them with nine concrete recommendations for standardized reporting [3]. Those gaps are the same signal Part 1 found in reporting systems: a category-level design absence, reproduced because the schema was never treated as the load-bearing component it is.

For an architect, the takeaway is sharp: the report is not the unit of value. The schema is. Standardize the schema and each report becomes a comparable observation; leave it unstandardized and the reports remain incommensurable, no matter how many you collect.

From individual experience to collective evidence

Structure enables the step that actually matters to governance — turning individual reports into a statistical claim about a group. A 2025 framework applies sequential hypothesis testing to streams of individual reports, formalizing when accumulated experiences constitute evidence of a subgroup harm rather than noise [4]. This is the mechanism that answers the question a risk committee will ask first: is this a pattern or an incident. It also disciplines the answer — sequential testing is built to control error as evidence arrives, so the claim of systemic harm is one you can defend rather than merely assert.

The architectural implication is that aggregation logic is a design-time choice, not a reporting afterthought. If the intake schema does not capture the fields the test needs — affected attributes, context, outcome — no amount of downstream analysis recovers them. The evidence layer has to be specified backward from the claim you will eventually need to make.

Institutional design decides what gets reported

Schema and statistics assume reports arrive; whether they do is an institutional-design question with its own architecture. A 2025 framework for incident reporting systems for general-purpose AI lays out seven design dimensions — spanning choices like mandatory versus voluntary reporting, whether near-misses are captured, whether reporters can stay anonymous, and what action follows a report — and grounds them in nine safety-critical case studies from other domains [5]. Each dimension is a lever that changes the population of reports you receive, and therefore the evidence you can build. A voluntary, non-anonymous system with no post-report action collects almost nothing; the design predicts the yield.

These are not policy preferences to settle later. They are parameters of the reporting architecture, and they determine whether the collective memory of the first section ever fills.

Transparent evaluation reporting

Evidence also flows from the vendor side, and it needs the same structural discipline. STREAM — a standard for transparently reporting evaluations in AI model reports, developed initially for chemical and biological risk — supplies a concrete template for reporting what was tested and what was found, so that evaluation results are legible rather than selectively narrated [6]. Where the incident thread structures what finders report, STREAM structures what builders disclose. Both feed the same evidence base, and both fail the same way if the format is left to discretion.

What the reports enable

The payoff of the whole stack is analysis that would be impossible on unstructured data. A 2026 study drew on roughly 5,300 incident reports from the AI Incident Database to analyze harms intersectionally — arguing that AI harms cannot be understood or fixed one identity attribute at a time [7]. Set the finding aside for a moment and note what made it possible: a durable store [1], comparable records [2][3], and enough structured volume to support subgroup analysis [4]. The intersectional result is a downstream use case for the architecture the rest of this post describes. Without the schema, there is no study — only 5,300 stories.

(The ~5,300 figure is drawn from the cited abstract and should be confirmed against the paper’s final text before external publication.)

The Hard Claim

The unit of value in flaw disclosure is not the report — it is the schema. Standardize the schema and individual experience becomes collective evidence a risk committee can act on and defend; leave it unstandardized and 5,300 reports remain 5,300 anecdotes.

Every enterprise that plans to “collect incidents and analyze them later” has the dependency backward: the analysis you will need dictates the schema you must capture now. Evidence is designed in at intake, or it is not available at all.

Part 3 takes the surface into agentic systems, where what a report must retain — and what a system must be built to remember — changes again.

Evidence Is Designed In at Intake — or It Is Not Available at All.

If you are building the incident-to-evidence layer for an agentic system and want a practitioner conversation about specifying the schema backward from the claim, the calendar is open.

Start the conversation
AI Flaw Disclosure as Architecture  ·  Series 26  ·  3 Posts
Post 01  ·  Published The Reporting Gap Is an Architecture Gap
Post 02  ·  Now Reading From Incident to Evidence
References & Sources

Share this:

Like this:

Like Loading…