What affordable oversight looks like

By Nathan Donaldson

A calm, orderly field of many identical translucent glass cards receding into soft darkness, representing a complete record of thousands of routine case decisions. One single card is lifted into a shaft of soft light and edged in red, held apart from the field for close inspection.
Affordable oversight is not watching everything; it is being able to watch anything, and choosing where to look.

What affordable oversight looks like

Picture an agency that pays income support. It has granted an agent the authority to progress routine cases: thousands of decisions a week, each one appealable, each one someone’s money. The plan promises efficiency. The question that decides whether the plan survives is quieter: who watches the agent, how closely, and what the watching costs.

Oversight is a capability that runs, not a gate you pass once set out the evidence that point-in-time oversight stops working once agents do the work. It ended on the question that decides the economics: can continuous oversight actually be paid for, or does the cost of oversight quietly eat the efficiency gain? Boost thinks it can be paid for, but in fewer places than the enthusiasm suggests.

The frame is Boost’s 5-layer model: substrate at layer 1, internal coordination at layer 2, the citizen interface at layer 3, work-performance at layer 4, oversight at layer 5. Layers 2 to 4 are where agents do the work; layers 1 and 5 decide whether that work lands safely. The model is set out in full in the layers-and-planes piece. Everything that follows sits at layer 5. None of it changes what the agents do; each of the five approaches below is a way of setting up the watching so that an agency can afford to keep doing it.

Why the cost cannot be wished down

Two facts set the constraint. The first is that oversight of running agents is an operating cost, not a project cost. Government already knows what an operating cost looks like: the eleven most critical legacy systems in the US federal government are between 23 and 60 years old and cost roughly US$754 million a year to operate (GAO-25-107795). Oversight of agents belongs in the same category: a cost that runs for as long as the work does.

The second fact is that the size of the cost is tied to the autonomy granted. “The Stochastic Gap” (Pal and Bhattacharya, arXiv 2603.24582) formalised the tie: “the same quantities that delimit statistically credible autonomy also determine expected oversight burden”. In plain words, the dial that sets how much an agent may decide on its own is the same dial that sets what its watching costs. The paper is a v1 preprint, not yet peer-reviewed; its mathematical result stands on its own, but its evidence so far is one simulated agent on one procurement workflow, so it carries that caveat here as it did in the previous piece. The mapping from its result to the model’s layers is Boost’s, not the authors’.

The design consequence is uncomfortable and useful at the same time. An agency cannot make the watching cheap by declaring the agent trustworthy; the arithmetic does not care what the procurement document says. The only real lever is where the watching effort goes. An agency that spends its oversight evenly, the same intensity on every decision, pays the highest price on offer. The affordable approach spends unevenly, on purpose.

The record comes first

None of the approaches below works without one build requirement: the work-performance layer writes down what it did as it does it. What was called, by whom, on whose authority, with what result. Not a logging module bolted on after go-live; a property of the layer-4 system, priced in from the first line of the build.

Governments already run systems that treat the record this way. Estonia’s X-Road logs every message it carries, on both sides of the exchange, digitally signed and time-stamped, and those logs hold evidentiary value in court; that property covers roughly 2.2 billion transactions a year across more than 3,000 services. The United Kingdom’s Algorithmic Transparency Recording Standard makes a structured public record of each algorithmic tool a standing obligation for central departments, with over a hundred records published so far. Canada’s Directive on Automated Decision-Making requires departments to document client feedback, human overrides and system failures, and to feed what they document into corrective action. The record-by-default principle is not a proposal; it is in production.

The cost argument is just as plain.

A record kept by default costs close to nothing at the moment of the work and cannot be bought at any price six months later.

Every cheaper form of watching described below, sampling, rule checks, appeal investigation, works by consulting that record. Skip it and the only oversight left is the expensive kind: people reviewing work in real time, which is the cost the whole plan was supposed to retire.

Price autonomy per decision class, not per system

The most expensive mistake available at layer 5 is granting autonomy to a system instead of to a class of decisions. A single autonomy dial for “the eligibility agent” means the watching has to be set for the riskiest thing the agent might ever touch, and that price is then paid across the entire caseload.

Canada already runs the alternative. Its Algorithmic Impact Assessment is a published questionnaire, 65 questions on risk and 41 on mitigation, that scores each automated decision system into an impact level, and the Directive attaches graduated requirements to that level: more documentation, more human involvement and peer review as impact rises, with independent audit and continuous monitoring at the top. Departments must re-score when a system’s scope or behaviour changes. That is oversight intensity priced per risk tier, running today in a peer government. Singapore’s Model AI Governance Framework for Agentic AI (January 2026) extends the same tiered thinking to agents specifically.

A hand-drawn sketch-note diagram in Boost navy ink on white. On the left, a rising arrow tracked by an ascending row of bars, showing that oversight cost climbs in lock-step with the autonomy granted. On the right, a bracket sorts decision classes into three tiers: a lightly-watched top band, a middle band, and a heavily-hatched bottom band picked out in red for the low-autonomy tier where a person decides.
Autonomy and its oversight cost rise together; decision classes are sorted into tiers so the high bills attach only where they are earned.

The model treats autonomy the same way: granted per decision class, and priced there. A renewal that repeats last year’s decision with no change in circumstances, reversible within a day if wrong, sits in a high-autonomy tier and gets light-touch watching. A first-time decline, a debt established, anything that cuts someone’s money off, sits in a low-autonomy tier: the agent prepares, a person decides. The tiers are not a maturity ladder. The same agent can hold high autonomy in one class and none in another, and a class moves between tiers as the record accumulates evidence about how the agent behaves there. The cost of oversight still rises with autonomy, exactly as the arithmetic says it must. What changes is that the high costs attach only to the decision classes that earn them.

Pre-authorise the agent, not every action

Schmitz, Rystrøm and Batzner looked at how public-sector organisations govern their systems today and found “siloed compliance units and episodic approvals rather than continuous, integrated supervision” (arXiv 2506.04836, REALM workshop at ACL 2025). Episodic approval fails because it checks a system once and then trusts it while it changes. The opposite instinct, a committee in front of every action, fails the other way: once agents act at machine frequency the communication costs are, in their words, “already prohibitive”.

Between those two failures sits an approach older industries settled on long ago. Aviation certifies an aircraft type once, then watches the data from every flight for as long as the type flies; when the data shows a problem, regulators issue directives that ground or fix the type. Medicines regulators approve a drug once, then monitor its benefits and harms for as long as it is prescribed, hunting for signals in the reports that come in. Financial supervisors are moving the same way, from periodic returns toward continuous feeds. Authorise the actor once; watch the activity continuously.

The model carries that approach as two layer-5 components doing paired work. An agent registry records what each agent is, what it may touch, and on whose authority it acts. Rules-as-code turns eligibility criteria and procedural constraints into a form a machine can check the agent’s record against, at machine speed, because the check itself is now machine-readable. Together they replace the unaffordable question, was every action correct, with two affordable ones: is this agent still the thing that was authorised, and does its record still follow the rules that were encoded. Machines run that check on everything. People handle what falls outside it.

Spend the human attention where the contest is

Continuous does not mean exhaustive. No team can read thousands of decisions a week, and the approaches above exist so that nobody has to. The human share of the watching looks deeply at a deliberately chosen few.

Sampling first. Draw cases from the record at random, review them end to end, widen the draw the moment something looks off. Factories have checked quality this way for around a century: pull a few off the line, inspect them properly, trust what the batch tells you.

In government the open question about sampling is legal rather than statistical, and it is real. A decision made about a person carries that person’s right to contest it: GDPR Article 22 gives an explicit right to contest decisions made solely by automated processing, and procedural fairness in the common-law world requires that each affected person can understand and challenge the decision that affects them. The legal literature treats audit and appeal as complementary, not substitutes: sampling gives systemic assurance, and it does not discharge the duty owed to the individual. What a court would make of an agency that closely examines one decision in two hundred, while every individual keeps full appeal rights, has not been tested anywhere Boost can point to. Boost thinks sampling is necessary, and that the legal design for it does not exist yet. Both are true at once.

Appeals second, and this is the one the enthusiasm most often misses. A citizen who contests a decision is doing oversight work: selecting, out of the whole caseload, exactly the case they believe the system got wrong, and attaching the evidence the record alone cannot hold. An appeal is the most information-rich review trigger the agency has.

Australia paid for the proof. Robodebt raised debts automatically and left review hard to reach: around 470,000 debts unlawfully raised, AU$751 million repaid, a further AU$112 million in compensation, and a Royal Commission whose 57 recommendations include making review easier to get and better resourced. Suppressing appeals did not save oversight money. It hid the failure until the failure was national.

So the design runs against instinct: make appealing cheap for the citizen, route every appeal into the oversight record as a signal about the decision class it came from, and treat a rising appeal rate in a class as the trigger that moves the class down a tier. Canada’s Directive already requires overrides and client feedback to be documented and acted on. The Algorithm Charter for Aotearoa New Zealand commits signatory agencies to “provide a channel for challenging or appealing of decisions informed by algorithms”, though the Charter has no enforcement power behind that commitment. An agency that makes appealing hard is not saving oversight money. It is switching off its best sensor.

The shape the five approaches share

Set them side by side. A record kept by default. Autonomy priced per decision class. Agents pre-authorised against encoded rules. Random samples reviewed deeply. Appeals routed as signals. None of them watches everything; watching everything would be the old cost in new clothes. All of them protect the same ability: any single decision in the caseload, picked for any reason, can be reconstructed, checked and acted on while it still matters.

That is the design answer to the question the previous piece left standing. Affordable oversight is not watching everything. It is being able to watch anything, and choosing where to look.

It is also why affordable oversight is narrow. The approaches fail in predictable places: where decisions are high-stakes and contested at a rate no appeal routing can absorb, where the rules resist being written as code, where the record cannot capture what actually mattered to the decision. In those places the arithmetic does not offer a cheaper way of watching. It offers less autonomy: the agent prepares, a person decides. Boost thinks that is the constraint working as intended, not the design failing. Narrow oversight an agency can pay for beats wide oversight it cannot.

The open questions

Where has continuous supervision of high-volume automated decisions actually been priced? The pattern runs in production elsewhere. Financial supervisors take in reporting continuously, aviation safety runs on monitoring and incident learning, and medicines are watched for as long as they are prescribed. Those systems use the same approaches described here. None of them offers a costed template for a government agency running autonomous casework, and until a government publishes real numbers, “affordable” remains a design argument rather than an observed fact.

Who does the watching? Schmitz, Rystrøm and Batzner argue oversight should be “centrally coordinated, but diffused”: placed with the operational teams whose work the agents are doing. The five approaches say what those teams would look at. Nobody has yet described the job itself, who staffs it, what they can see, what they can stop, and the authors’ own survey of current practice says that machinery does not exist yet.

What would prove this wrong

Two kinds of evidence would do it, measured from July 2026. If a government runs high-autonomy casework affordably under uniform, episodic oversight, the five approaches are unnecessary, and Boost will revise the design. If a government builds all five and still finds the cost of oversight eating the efficiency gain, there is no affordable version of continuous oversight, the model’s answer is less autonomy across the board, and Boost will say so. Boost expects the second to arrive somewhere, in some decision class, before the first arrives anywhere. That expectation is revisable on the same terms as the rest.

Sources and further reading

Make a bigger impact tomorrow