Governance Stress-Test Checklist

A public checklist for testing whether a technology decision can survive the institution that must operate it.

Part of the Governance Stress-Tests series. Related artifacts: worked example and protocol note.

A governance stress-test is a way to test whether a technology decision can survive the institutions that must operate it.

The checklist is designed for decisions such as AI procurement, public-sector AI deployment, DPI-linked rollouts, compute subsidies or procurement, and other technology choices where the promise of access or efficiency may hide deeper institutional constraints.

The core question is simple:

What must the public institution know, specify, retain, test, contest, and learn before a technology system moves from promise into operational use?

The distinction from an algorithmic impact assessment or model audit is important. Those tools usually test the system. A governance stress-test tests the institution around it: authority, procurement, evaluation capacity, recourse, learning, and the ability to revise or stop the deployment.

What The Checklist Tests

This is a diagnostic tool for institutional feasibility. It asks whether a proposed technology decision has enough operational capacity, accountable authority, recourse, evidence, and institutional memory behind it to be governable in practice.

It is meant to be used before a decision hardens into routine administration: before a pilot becomes operational use, before a procurement contract is renewed, before a system is scaled, or before a public agency accepts a vendor's assurance as a substitute for its own evaluation.

What The Checklist Leaves To Other Tools

This checklist has a narrower job than a legal opinion, procurement manual, AI safety benchmark, or argument against deployment. It asks what would have to be true for a specific institution to use a technology responsibly.

How To Use The Checklist

Start with a concrete decision point. The object being tested should be specific enough to map: a pilot moving into operational use, a procurement contract, a vendor renewal, a DPI-linked rollout, a model update, a subsidy, a deployment threshold, or a decision to scale.

Then work through the six stages below. For each stage, answer with publicly available evidence where possible. Mark unknowns directly. An unknown is often the stress-test result.

The output should be a Revised Governance Decision. That decision may be amended contract terms, stronger evaluation requirements, a recourse design, audit rights, logging obligations, a rollback trigger, a delayed rollout, or a documented reason to proceed only under narrower conditions.

The Six Stages

StageWhat It TestsExample Failure Mode
1. Decision DefinitionWhether the decision, operator, affected parties, and claimed benefit are clearThe public debate argues about AI in general while the actual decision point is a vendor renewal
2. Capability And ControlWhether the institution can evaluate, constrain, and learn from the systemThe vendor controls updates and evidence while the agency remains accountable
3. Authority And ResponsibilityWhether legal authority, operational control, and technical knowledge alignThe institution answerable to citizens lacks a usable explanation of system behavior
4. Recourse And LearningWhether affected people can contest outcomes and the institution can learn from failuresA grievance channel exists without access to logs or authority to trigger system revision
5. Public-Value ConstraintsWhether opacity and delegated authority are tolerable at the risk levelThe system is efficient and too opaque for affected people or oversight bodies to review
6. Deployment ThresholdsWhether enough evidence and capacity exist before pilot, deployment, scale-up, or renewalA pilot becomes permanent without rollback criteria or institutional review

1. Decision Definition

The first step is to name the decision precisely. Broad abstractions such as "AI in welfare," "DPI for inclusion," or "national compute capacity" become testable only when translated into a decision an institution must actually make.

Ask:

The aim is to locate the point at which institutional responsibility becomes real.

2. Capability And Control

A public institution may have formal authority over a system without having the practical capability to understand or control it. This gap matters most when technical systems are procured from vendors, updated over time, or embedded inside administrative workflows.

Ask:

The stress-test should make hidden dependence visible. If an institution can use a system while lacking the ability to evaluate, contest, or exit it, access remains short of durable capability.

3. Authority And Responsibility

Technology decisions often separate the actors who hold authority from the actors who hold information. A public agency may be legally accountable. A vendor may understand the system. A frontline official may face the citizen. No single actor may have the complete picture needed to explain or correct a failure.

Ask:

The vendor-operator boundary often appears at this stage: the organization answerable for the system may differ from the organization that understands or controls it.

4. Recourse And Learning

Recourse is both a fairness feature and an institutional learning mechanism. When affected people lack a way to understand, challenge, correct, or appeal a decision, the institution loses one of the main ways failures become visible.

Ask:

The test is whether the institution learns from the deployment or adds compliance language on top of old systems.

5. Public-Value Constraints

Public-value constraints belong inside the decision before legitimacy becomes fragile. At this stage, the distinct questions are acceptable opacity and tolerable delegation: how much can remain hidden, and how much authority can move to an AI system, vendor, platform, or intermediary before the decision becomes institutionally fragile?

Ask:

Public values vary with risk perception, authority allocation, institutional trust, and visible recourse.

6. Deployment Thresholds

The final stage asks what must be true before the decision can responsibly proceed. A threshold is a condition that changes whether deployment, scale-up, renewal, or termination is justified.

Ask:

The goal is to prevent sequence error: buying or scaling the system before the institution has resolved the governance conditions that make the system usable.

Worked-Example Template

Each public worked example should follow the same structure.

Case

Name the decision or scenario being tested. State whether the case is named, composite, or hypothetical.

Claimed Benefit

State what the technology proposal promises. This may be faster processing, lower cost, wider access, better targeting, improved safety, national capability, or administrative consistency.

Institutional Setup

Map the public operator, vendor or intermediary, affected parties, legal authority, operational control, technical knowledge, and route of recourse.

Stress-Test Variables

List the operational, institutional, and public-value constraints that matter for this decision.

Failure Modes

Identify where the decision can break: information gaps, unclear contracts, weak evaluation rights, missing logs, vendor update control, opaque recourse, fragmented authority, weak rollback capacity, or public-value instability.

Threshold Questions

State what must be true before the decision moves to pilot, deployment, scale-up, renewal, or termination.

Revised Governance Decision

Explain what the stress-test changes: contract terms, evaluation requirements, deployment timing, audit rights, logging duties, recourse design, public communication, rollback triggers, or the decision to delay.

Evidence Discipline

Public worked examples should be built from public sources, public procurement or policy documents where available, tender documents, official notices, regulator materials, parliamentary or committee records, CAG audit reports, RTI responses, published audits, credible media, and secondary literature.

Unverified figures are too weak to carry the argument. If a case depends on a disputed claim, the worked example should either verify it before publication or mark the uncertainty directly. The method can work from public evidence without requiring confidential access, private government information, or institutional endorsement.

First Worked Example

The first public worked example is a reconstructive governance stress-test of Telangana's Samagra Vedika welfare-eligibility system.

Samagra Vedika is useful because it shows the method in a hard public setting. The system promised better welfare targeting, duplicate detection, and fraud reduction. Public reporting and Amnesty International's technical explainer later linked its entity-resolution logic to wrongful exclusions from food-security benefits. The stress-test reconstructs the questions that should have been answered before algorithmic matches could alter entitlements.

The example asks whether the institution could evaluate match accuracy, inspect the evidence behind a denial, preserve human override authority, provide meaningful recourse, and learn from error before the system became routine welfare administration.

The output is a Revised Governance Decision: what institutional conditions should exist before an algorithmic match can shape a public entitlement.