Part of the Governance Stress-Tests series. Start with the checklist and see the method applied in the Samagra Vedika worked example.
A governance stress-test is a structured way to test whether a technology decision can survive the institution that must operate it.
The method begins from a concrete decision: a pilot becoming operational, a procurement contract being issued or renewed, a model update entering public use, a digital public infrastructure rollout, a compute subsidy, or an algorithmic system shaping access to a public service. Broad themes such as "AI in government" or "DPI for inclusion" become testable only when they reach the point at which institutional responsibility becomes real.
The core question is:
What must the public institution know, specify, retain, test, contest, and learn before a technology system moves from promise into operational use?
The distinction from an algorithmic impact assessment or model audit is the inversion. Impact assessments test the system. Governance stress-tests test the institution around it: evidence, authority, procurement, operational control, recourse, public-value constraints, and learning.
What The Protocol Produces
Each worked example should produce a Revised Governance Decision. That decision may be to proceed under narrower conditions, delay deployment, require independent evaluation, change contract terms, add audit rights, publish decision criteria, improve recourse, create rollback triggers, or refuse scale-up until minimum institutional capacity exists.
The output is a decision-facing diagnosis focused on institutional feasibility:
- what the institution can currently know;
- what remains unavailable from the available evidence;
- where authority and technical control diverge;
- what affected people can understand or contest;
- what evidence would stop or alter deployment;
- what institutional learning mechanism should exist before the system becomes routine.
Three Modes Of Governance Stress-Testing
The method can operate in three modes. The first phase of this project relies mainly on the public-source mode, with selective feedback and review where useful.
| Mode | Evidence Base | Best Use | Limits |
|---|---|---|---|
| Public-source stress-test | Public policy documents, tenders, RFPs, court materials, audit reports, regulator materials, RTI responses, parliamentary records, credible media, incident databases, and secondary literature | Public examples, teaching cases, first worked examples, and case-selection notes | Claims are limited to public evidence; unknowns must be named directly |
| Primary-research stress-test | Interviews, practitioner workshops, structured elicitation, expert review, and internal document access where available | Deeper reconstruction of institutional workflows, procurement incentives, implementation history, and tacit operational knowledge | Requires access, consent, and stronger research design; better suited to partnered work |
| Hybrid stress-test | Public-source reconstruction plus targeted interviews, expert review, or feedback sessions | Stronger case studies and method refinement without making confidential access load-bearing | Must distinguish public evidence from interview-based interpretation |
Separating the modes protects credibility. The first public worked examples should stand without privileged access. They should show what can be learned from public evidence, and where the absence of public evidence is itself a governance finding.
Case Selection
A case is suitable for a governance stress-test when five conditions are met.
First, there must be a concrete decision point. The case should involve deployment, procurement, scale-up, renewal, termination, or redesign.
Second, the technology must enter an institutional workflow where public authority, operational control, technical knowledge, and citizen-facing responsibility may diverge.
Third, the claimed benefit should be legible: better targeting, faster processing, lower cost, improved safety, wider access, national capability, reduced fraud, or administrative consistency.
Fourth, enough evidence must exist to reconstruct at least the main actors, decision pathway, affected parties, and contested failure modes.
Fifth, the case should produce a Revised Governance Decision that goes beyond descriptive narrative. A mature case can say what conditions should have existed before deployment, renewal, or scale-up.
Evidence Protocol
Each worked example should maintain a source register and a claim ledger.
The source register lists the documents used: official notices, procurement materials, court orders, regulatory documents, committee reports, audit reports, RTI records, incident databases, media investigations, technical explainers, and secondary literature.
The claim ledger separates four categories:
| Claim Type | Treatment |
|---|---|
| Established in primary public record | Can carry analytical weight |
| Reported by credible secondary source | Can be used with attribution |
| Disputed or incomplete | Must be marked directly and kept from carrying the main conclusion alone |
| Unknown from public record | Treated as a governance finding if the missing fact is institutionally necessary |
In published examples, the source register appears as the sources list; reported-only and unresolved claims are marked in the evidence-limits section, with a fuller claim ledger maintained separately where needed.
The method must avoid invented figures, inferred contract terms, and allegations converted into findings. If a case depends on a contested fact, either verify it, narrow the claim, or mark the uncertainty.
The strongest public-source stress-tests will often find that the public record is good enough to identify the institutional question while remaining too thin to prove that the institution could govern the system. That result is central to the method.
When opacity comes from procurement design itself, the method should treat procurement architecture as part of the object being tested. Proprietary claims, vendor-controlled documentation, weak audit rights, source-code restrictions, data-format limits, and lock-in can structure information so that it is practically unavailable to the public institution or to affected people. In those cases, the finding concerns a vendor-operator boundary designed in a way that limits operational capacity.
The Six-Stage Workflow
Every worked example should pass through the same six stages.
1. Decision Definition
Define the decision being tested. Name the operator, vendor or intermediary, affected parties, claimed benefit, decision point, and deployment status.
The output of this stage is a bounded decision object.
2. Capability And Control
Ask whether the institution can evaluate, constrain, audit, suspend, update, or exit the system. Identify which capabilities are internal and which are rented from vendors or intermediaries.
The output of this stage is a capability map: what the institution can know and control, and what sits outside its operational reach.
3. Authority And Responsibility
Map where formal legal authority, operational control, technical knowledge, and citizen-facing responsibility sit. Look for gaps between the organization answerable to the public and the organization that understands or controls the system.
The output of this stage is an accountability map.
4. Recourse And Learning
Test whether affected people can understand, challenge, correct, appeal, or exit the decision, and whether complaints, incidents, audits, and failures feed back into institutional learning.
The output of this stage is a recourse-and-learning diagnosis.
5. Public-Value Constraints
Ask whether the system's opacity is acceptable, and whether delegation to an AI system, vendor, platform, or intermediary is tolerable at this risk level. The relevant public-value constraints include acceptable opacity, tolerable delegation, recourse visibility, authority clarity, and the perceived risk of system failure.
The output of this stage is a legitimacy boundary: what must remain visible, contestable, and reversible for the decision to be institutionally feasible.
6. Deployment Thresholds
Specify what evidence and institutional capacity must exist before pilot, procurement, deployment, renewal, scale-up, or termination. Ask what would stop deployment, trigger rollback, require contract revision, or force public notice.
The output of this stage is the Revised Governance Decision.
Incident Evidence To Scenario Design
Incident databases can support case selection, while institutional analysis remains necessary. The OECD AI Incidents and Hazards Monitor, the AI Incident Database, AVID, and similar public resources can help identify recurring failure patterns: exclusion, opacity, unreliable classification, weak recourse, harmful automation, unclear responsibility, or insufficient human oversight.
The method converts an incident pattern into an institutional scenario:
- What decision produced or enabled the failure?
- What traces existed before harm occurred?
- Who could interpret those traces?
- Who had authority to pause, revise, or contest the system?
- What recourse did affected people have?
- What learning loop, if any, changed the system afterward?
Incident databases support claims about frequency, representativeness, or causal trends only when the underlying database permits that use. In this project, incident evidence is mainly a way to find realistic failure modes and convert them into testable governance questions.
Review And Feedback
Each public worked example should receive at least one evidence review before publication. The review should check:
- whether the decision point is specific enough;
- whether factual claims are properly attributed;
- whether reported claims are distinguished from established claims;
- whether unknowns are named clearly;
- whether the failure modes follow from the evidence;
- whether the Revised Governance Decision is concrete enough to be useful.
Where possible, later examples can also receive feedback from students, technologists, researchers, practitioners, or public-policy readers. Feedback improves the method's clarity and usefulness; endorsement is a separate claim.
Boundary Of The Method
The method has a narrower job than a legal determination, procurement manual, model audit, AI safety benchmark, or investigative report. It leaves liability, confidential access, and full failure prediction to other instruments.
Its narrower purpose is to make a technology decision institutionally legible before it hardens into routine administration.
Where The Method Stops
The method is useful only when the decision has an institutional object.
Broad debates about AI, data, or digital public infrastructure may be important. The method becomes useful once they resolve into a decision point: procurement, pilot, deployment, scale-up, renewal, redesign, or termination.
Purely technical questions belong with ordinary benchmarking, security review, model evaluation, or system audit when those tools can answer the question without changing the governance decision.
The method also weakens when there is no meaningful governance discretion. If the institution lacks authority to alter contract terms, change deployment conditions, require evidence, create recourse, pause rollout, or revise the system, the stress-test can document the constraint. A useful Revised Governance Decision requires some power to act.
The method is useful when it changes the question from "Should this technology be adopted?" to "What institutional conditions must exist before this technology can be responsibly used?"