Governance Stress-Test Protocol Note

The evidence and workflow protocol behind Governance Stress-Tests.

Part of the Governance Stress-Tests series. Start with the checklist and see the method applied in the Samagra Vedika worked example.

A governance stress-test is a structured way to test whether a technology decision can survive the institution that must operate it.

The method begins from a concrete decision: a pilot becoming operational, a procurement contract being issued or renewed, a model update entering public use, a digital public infrastructure rollout, a compute subsidy, or an algorithmic system shaping access to a public service. Broad themes such as "AI in government" or "DPI for inclusion" become testable only when they reach the point at which institutional responsibility becomes real.

The core question is:

What must the public institution know, specify, retain, test, contest, and learn before a technology system moves from promise into operational use?

The distinction from an algorithmic impact assessment or model audit is the inversion. Impact assessments test the system. Governance stress-tests test the institution around it: evidence, authority, procurement, operational control, recourse, public-value constraints, and learning.

What The Protocol Produces

Each worked example should produce a Revised Governance Decision. That decision may be to proceed under narrower conditions, delay deployment, require independent evaluation, change contract terms, add audit rights, publish decision criteria, improve recourse, create rollback triggers, or refuse scale-up until minimum institutional capacity exists.

The output is a decision-facing diagnosis focused on institutional feasibility:

Three Modes Of Governance Stress-Testing

The method can operate in three modes. The first phase of this project relies mainly on the public-source mode, with selective feedback and review where useful.

ModeEvidence BaseBest UseLimits
Public-source stress-testPublic policy documents, tenders, RFPs, court materials, audit reports, regulator materials, RTI responses, parliamentary records, credible media, incident databases, and secondary literaturePublic examples, teaching cases, first worked examples, and case-selection notesClaims are limited to public evidence; unknowns must be named directly
Primary-research stress-testInterviews, practitioner workshops, structured elicitation, expert review, and internal document access where availableDeeper reconstruction of institutional workflows, procurement incentives, implementation history, and tacit operational knowledgeRequires access, consent, and stronger research design; better suited to partnered work
Hybrid stress-testPublic-source reconstruction plus targeted interviews, expert review, or feedback sessionsStronger case studies and method refinement without making confidential access load-bearingMust distinguish public evidence from interview-based interpretation

Separating the modes protects credibility. The first public worked examples should stand without privileged access. They should show what can be learned from public evidence, and where the absence of public evidence is itself a governance finding.

Case Selection

A case is suitable for a governance stress-test when five conditions are met.

First, there must be a concrete decision point. The case should involve deployment, procurement, scale-up, renewal, termination, or redesign.

Second, the technology must enter an institutional workflow where public authority, operational control, technical knowledge, and citizen-facing responsibility may diverge.

Third, the claimed benefit should be legible: better targeting, faster processing, lower cost, improved safety, wider access, national capability, reduced fraud, or administrative consistency.

Fourth, enough evidence must exist to reconstruct at least the main actors, decision pathway, affected parties, and contested failure modes.

Fifth, the case should produce a Revised Governance Decision that goes beyond descriptive narrative. A mature case can say what conditions should have existed before deployment, renewal, or scale-up.

Evidence Protocol

Each worked example should maintain a source register and a claim ledger.

The source register lists the documents used: official notices, procurement materials, court orders, regulatory documents, committee reports, audit reports, RTI records, incident databases, media investigations, technical explainers, and secondary literature.

The claim ledger separates four categories:

Claim TypeTreatment
Established in primary public recordCan carry analytical weight
Reported by credible secondary sourceCan be used with attribution
Disputed or incompleteMust be marked directly and kept from carrying the main conclusion alone
Unknown from public recordTreated as a governance finding if the missing fact is institutionally necessary

In published examples, the source register appears as the sources list; reported-only and unresolved claims are marked in the evidence-limits section, with a fuller claim ledger maintained separately where needed.

The method must avoid invented figures, inferred contract terms, and allegations converted into findings. If a case depends on a contested fact, either verify it, narrow the claim, or mark the uncertainty.

The strongest public-source stress-tests will often find that the public record is good enough to identify the institutional question while remaining too thin to prove that the institution could govern the system. That result is central to the method.

When opacity comes from procurement design itself, the method should treat procurement architecture as part of the object being tested. Proprietary claims, vendor-controlled documentation, weak audit rights, source-code restrictions, data-format limits, and lock-in can structure information so that it is practically unavailable to the public institution or to affected people. In those cases, the finding concerns a vendor-operator boundary designed in a way that limits operational capacity.

The Six-Stage Workflow

Every worked example should pass through the same six stages.

1. Decision Definition

Define the decision being tested. Name the operator, vendor or intermediary, affected parties, claimed benefit, decision point, and deployment status.

The output of this stage is a bounded decision object.

2. Capability And Control

Ask whether the institution can evaluate, constrain, audit, suspend, update, or exit the system. Identify which capabilities are internal and which are rented from vendors or intermediaries.

The output of this stage is a capability map: what the institution can know and control, and what sits outside its operational reach.

3. Authority And Responsibility

Map where formal legal authority, operational control, technical knowledge, and citizen-facing responsibility sit. Look for gaps between the organization answerable to the public and the organization that understands or controls the system.

The output of this stage is an accountability map.

4. Recourse And Learning

Test whether affected people can understand, challenge, correct, appeal, or exit the decision, and whether complaints, incidents, audits, and failures feed back into institutional learning.

The output of this stage is a recourse-and-learning diagnosis.

5. Public-Value Constraints

Ask whether the system's opacity is acceptable, and whether delegation to an AI system, vendor, platform, or intermediary is tolerable at this risk level. The relevant public-value constraints include acceptable opacity, tolerable delegation, recourse visibility, authority clarity, and the perceived risk of system failure.

The output of this stage is a legitimacy boundary: what must remain visible, contestable, and reversible for the decision to be institutionally feasible.

6. Deployment Thresholds

Specify what evidence and institutional capacity must exist before pilot, procurement, deployment, renewal, scale-up, or termination. Ask what would stop deployment, trigger rollback, require contract revision, or force public notice.

The output of this stage is the Revised Governance Decision.

Incident Evidence To Scenario Design

Incident databases can support case selection, while institutional analysis remains necessary. The OECD AI Incidents and Hazards Monitor, the AI Incident Database, AVID, and similar public resources can help identify recurring failure patterns: exclusion, opacity, unreliable classification, weak recourse, harmful automation, unclear responsibility, or insufficient human oversight.

The method converts an incident pattern into an institutional scenario:

  1. What decision produced or enabled the failure?
  2. What traces existed before harm occurred?
  3. Who could interpret those traces?
  4. Who had authority to pause, revise, or contest the system?
  5. What recourse did affected people have?
  6. What learning loop, if any, changed the system afterward?

Incident databases support claims about frequency, representativeness, or causal trends only when the underlying database permits that use. In this project, incident evidence is mainly a way to find realistic failure modes and convert them into testable governance questions.

Review And Feedback

Each public worked example should receive at least one evidence review before publication. The review should check:

Where possible, later examples can also receive feedback from students, technologists, researchers, practitioners, or public-policy readers. Feedback improves the method's clarity and usefulness; endorsement is a separate claim.

Boundary Of The Method

The method has a narrower job than a legal determination, procurement manual, model audit, AI safety benchmark, or investigative report. It leaves liability, confidential access, and full failure prediction to other instruments.

Its narrower purpose is to make a technology decision institutionally legible before it hardens into routine administration.

Where The Method Stops

The method is useful only when the decision has an institutional object.

Broad debates about AI, data, or digital public infrastructure may be important. The method becomes useful once they resolve into a decision point: procurement, pilot, deployment, scale-up, renewal, redesign, or termination.

Purely technical questions belong with ordinary benchmarking, security review, model evaluation, or system audit when those tools can answer the question without changing the governance decision.

The method also weakens when there is no meaningful governance discretion. If the institution lacks authority to alter contract terms, change deployment conditions, require evidence, create recourse, pause rollout, or revise the system, the stress-test can document the constraint. A useful Revised Governance Decision requires some power to act.

The method is useful when it changes the question from "Should this technology be adopted?" to "What institutional conditions must exist before this technology can be responsibly used?"