Part of the Governance Stress-Tests series. Related artifacts: worked example and protocol note.
A governance stress-test is a way to test whether a technology decision can survive the institutions that must operate it.
The checklist is designed for decisions such as AI procurement, public-sector AI deployment, DPI-linked rollouts, compute subsidies or procurement, and other technology choices where the promise of access or efficiency may hide deeper institutional constraints.
The core question is simple:
What must the public institution know, specify, retain, test, contest, and learn before a technology system moves from promise into operational use?
The distinction from an algorithmic impact assessment or model audit is important. Those tools usually test the system. A governance stress-test tests the institution around it: authority, procurement, evaluation capacity, recourse, learning, and the ability to revise or stop the deployment.
What The Checklist Tests
This is a diagnostic tool for institutional feasibility. It asks whether a proposed technology decision has enough operational capacity, accountable authority, recourse, evidence, and institutional memory behind it to be governable in practice.
It is meant to be used before a decision hardens into routine administration: before a pilot becomes operational use, before a procurement contract is renewed, before a system is scaled, or before a public agency accepts a vendor's assurance as a substitute for its own evaluation.
What The Checklist Leaves To Other Tools
This checklist has a narrower job than a legal opinion, procurement manual, AI safety benchmark, or argument against deployment. It asks what would have to be true for a specific institution to use a technology responsibly.
How To Use The Checklist
Start with a concrete decision point. The object being tested should be specific enough to map: a pilot moving into operational use, a procurement contract, a vendor renewal, a DPI-linked rollout, a model update, a subsidy, a deployment threshold, or a decision to scale.
Then work through the six stages below. For each stage, answer with publicly available evidence where possible. Mark unknowns directly. An unknown is often the stress-test result.
The output should be a Revised Governance Decision. That decision may be amended contract terms, stronger evaluation requirements, a recourse design, audit rights, logging obligations, a rollback trigger, a delayed rollout, or a documented reason to proceed only under narrower conditions.
The Six Stages
| Stage | What It Tests | Example Failure Mode |
|---|---|---|
| 1. Decision Definition | Whether the decision, operator, affected parties, and claimed benefit are clear | The public debate argues about AI in general while the actual decision point is a vendor renewal |
| 2. Capability And Control | Whether the institution can evaluate, constrain, and learn from the system | The vendor controls updates and evidence while the agency remains accountable |
| 3. Authority And Responsibility | Whether legal authority, operational control, and technical knowledge align | The institution answerable to citizens lacks a usable explanation of system behavior |
| 4. Recourse And Learning | Whether affected people can contest outcomes and the institution can learn from failures | A grievance channel exists without access to logs or authority to trigger system revision |
| 5. Public-Value Constraints | Whether opacity and delegated authority are tolerable at the risk level | The system is efficient and too opaque for affected people or oversight bodies to review |
| 6. Deployment Thresholds | Whether enough evidence and capacity exist before pilot, deployment, scale-up, or renewal | A pilot becomes permanent without rollback criteria or institutional review |
1. Decision Definition
The first step is to name the decision precisely. Broad abstractions such as "AI in welfare," "DPI for inclusion," or "national compute capacity" become testable only when translated into a decision an institution must actually make.
Ask:
- What decision, deployment, procurement, or policy instrument is being tested?
- What is the decision point: pilot, procurement, deployment, scale-up, renewal, termination, or redesign?
- Who is the public operator?
- Who is the vendor, intermediary, platform, or technical provider?
- Who is affected by the decision?
- What is the claimed benefit: access, efficiency, inclusion, safety, capability, sovereignty, cost reduction, or administrative speed?
The aim is to locate the point at which institutional responsibility becomes real.
2. Capability And Control
A public institution may have formal authority over a system without having the practical capability to understand or control it. This gap matters most when technical systems are procured from vendors, updated over time, or embedded inside administrative workflows.
Ask:
- What technical capability does the institution need to evaluate the system?
- What data, logs, documentation, testing artifacts, and incident records are available?
- Who can modify, update, audit, suspend, or constrain system behavior?
- What happens if the vendor changes the model, infrastructure, pricing, documentation, or support terms?
- Does the institution have the expertise to interpret evaluation evidence?
- Which capabilities are internal, and which are rented from the vendor?
The stress-test should make hidden dependence visible. If an institution can use a system while lacking the ability to evaluate, contest, or exit it, access remains short of durable capability.
3. Authority And Responsibility
Technology decisions often separate the actors who hold authority from the actors who hold information. A public agency may be legally accountable. A vendor may understand the system. A frontline official may face the citizen. No single actor may have the complete picture needed to explain or correct a failure.
Ask:
- Where does formal legal authority sit?
- Where does operational control sit?
- Where does technical knowledge sit?
- Who is accountable to the affected person?
- Who can order a pause, audit, rollback, contract revision, or escalation?
- Does responsibility shift toward vendors, procurement systems, standards bodies, regulated entities, or frontline operators?
- Is there a gap between legal accountability and operational control?
The vendor-operator boundary often appears at this stage: the organization answerable for the system may differ from the organization that understands or controls it.
4. Recourse And Learning
Recourse is both a fairness feature and an institutional learning mechanism. When affected people lack a way to understand, challenge, correct, or appeal a decision, the institution loses one of the main ways failures become visible.
Ask:
- Can affected people understand that a technology system shaped the decision?
- Can they challenge, correct, appeal, or exit the decision?
- Are grievance channels visible, usable, and timely?
- Are logs and documentation sufficient for ex post review?
- Are incident reports, complaints, audits, and failures converted into institutional learning?
- Is there a feedback loop that can revise rules, contracts, thresholds, or deployment practices?
The test is whether the institution learns from the deployment or adds compliance language on top of old systems.
5. Public-Value Constraints
Public-value constraints belong inside the decision before legitimacy becomes fragile. At this stage, the distinct questions are acceptable opacity and tolerable delegation: how much can remain hidden, and how much authority can move to an AI system, vendor, platform, or intermediary before the decision becomes institutionally fragile?
Ask:
- What level of opacity is acceptable for this decision, given the risk, affected population, and possibility of correction?
- What level of delegation to an AI system, vendor, platform, or intermediary is publicly tolerable for this decision?
- Which facts must be visible for affected people to understand why the decision was made?
- Which facts must be visible for an oversight body, court, auditor, or senior official to review the decision?
- At what point does operational efficiency become illegitimate because the institution lacks the capacity to explain, contest, or reverse what the system does?
Public values vary with risk perception, authority allocation, institutional trust, and visible recourse.
6. Deployment Thresholds
The final stage asks what must be true before the decision can responsibly proceed. A threshold is a condition that changes whether deployment, scale-up, renewal, or termination is justified.
Ask:
- What must be true before pilot, procurement, deployment, scale-up, renewal, or termination?
- What evidence would stop deployment?
- What triggers rollback, audit, contract revision, public notice, or escalation?
- What minimum institutional capacity must exist before the system can be used responsibly?
- Is there a sunset, review, or renewal point?
- Can the public operator exit the vendor relationship without losing core operational capacity?
The goal is to prevent sequence error: buying or scaling the system before the institution has resolved the governance conditions that make the system usable.
Worked-Example Template
Each public worked example should follow the same structure.
Case
Name the decision or scenario being tested. State whether the case is named, composite, or hypothetical.
Claimed Benefit
State what the technology proposal promises. This may be faster processing, lower cost, wider access, better targeting, improved safety, national capability, or administrative consistency.
Institutional Setup
Map the public operator, vendor or intermediary, affected parties, legal authority, operational control, technical knowledge, and route of recourse.
Stress-Test Variables
List the operational, institutional, and public-value constraints that matter for this decision.
Failure Modes
Identify where the decision can break: information gaps, unclear contracts, weak evaluation rights, missing logs, vendor update control, opaque recourse, fragmented authority, weak rollback capacity, or public-value instability.
Threshold Questions
State what must be true before the decision moves to pilot, deployment, scale-up, renewal, or termination.
Revised Governance Decision
Explain what the stress-test changes: contract terms, evaluation requirements, deployment timing, audit rights, logging duties, recourse design, public communication, rollback triggers, or the decision to delay.
Evidence Discipline
Public worked examples should be built from public sources, public procurement or policy documents where available, tender documents, official notices, regulator materials, parliamentary or committee records, CAG audit reports, RTI responses, published audits, credible media, and secondary literature.
Unverified figures are too weak to carry the argument. If a case depends on a disputed claim, the worked example should either verify it before publication or mark the uncertainty directly. The method can work from public evidence without requiring confidential access, private government information, or institutional endorsement.
First Worked Example
The first public worked example is a reconstructive governance stress-test of Telangana's Samagra Vedika welfare-eligibility system.
Samagra Vedika is useful because it shows the method in a hard public setting. The system promised better welfare targeting, duplicate detection, and fraud reduction. Public reporting and Amnesty International's technical explainer later linked its entity-resolution logic to wrongful exclusions from food-security benefits. The stress-test reconstructs the questions that should have been answered before algorithmic matches could alter entitlements.
The example asks whether the institution could evaluate match accuracy, inspect the evidence behind a denial, preserve human override authority, provide meaningful recourse, and learn from error before the system became routine welfare administration.
The output is a Revised Governance Decision: what institutional conditions should exist before an algorithmic match can shape a public entitlement.