By Nathan Donaldson
Boost's 5-layer model is a map of Agentic Government. This post is about layer 5, and it sets out the position the model holds most strongly.
Agentic AI is software that can chase goals on its own, not just answer one prompt at a time. The five layers, in brief:
The numbering is about parts, not steps. Today, layer 5.
Layer 5 is oversight and governance. Audit logs. A register of which agents are allowed to do what. Rules written as code. A human in the loop. A real right to appeal. An explanation a person can actually read.
The building analogy lands hardest here. Layer 5 is the building code, the fire inspection, and the certificate that says people are allowed in. A tower's height was never set by how high concrete can be poured. It was set by what the inspector will sign off. That is the point of this post. The thing that holds agentic government back is not how clever the models get. It is whether the oversight is good enough to let them run.
On this model, the binding constraint is oversight maturity, not model capability. Most of the debate, hopeful or worried, turns on whether the models are good enough. That is the wrong axis.
Start with the bar that already exists. In Estonia, every data exchange between agencies has been authenticated, signed, and logged for around twenty-five years. Who touched a record, and why, is traceable. That is a high bar, and it is the foundation's bar. It is built for data moving between systems.
Now raise the stakes. When an agent starts making the call on a person's benefit, or visa, or tax, the record of the agent's decision needs to be at least as trustworthy. Not “the data moved”, but “here is what the agent decided, on what basis, and here is how a person challenges it”. As far as Boost has been able to find, nothing at that standard runs in production anywhere, for agent decisions.
The ambition is on paper. The Agentic State paper describes a governance layer with an agent registry, kill switches, rollback history, and rules published as code next to the law they encode. The Tony Blair Institute offers a set of principles: predictability, explainability, accountability, reversibility, and sensitivity. These are good. Neither is a running production system.
What governments actually have divides into two kinds of thing. There are duties to disclose what a system is, or to assess it before it runs. Several countries have those, and in Canada and the United Kingdom they are mandatory. Then there are binding audit and appeal rights that attach to a particular decision. That is the part that is thin, and it is thinnest of all for a decision an agent made.
Canada has had a mandatory federal Directive on Automated Decision-Making since 1 April 2019, with compliance required by 1 April 2020. It requires departments to publish an Algorithmic Impact Assessment before an automated decision system goes into production. That is a real duty with a real deadline, and it is not lighter than the ambition. It is an assessment-and-disclosure duty, though. It does not audit what a particular system decided, and it does not give the person a right of appeal against that decision.
The United Kingdom has a transparency register for algorithms. The government made it a requirement for departments on 6 February 2024, and a policy published on 17 December 2024 set out the scope: central government departments, plus the arm’s-length bodies that provide public or frontline services or routinely deal with the public. A few dozen records were published in the first year. It is the most developed thing of its kind. But it is a disclosure list, not a binding audit or a right of appeal. The United Kingdom also has a voluntary AI playbook.
The permission underneath moved the other way in 2026. On 5 February 2026, section 80 of the Data (Use and Access) Act 2025 came into force. It took Article 22 out of the UK GDPR and put Articles 22A to 22D in its place. A decision made entirely by a computer, where the consequences for the person are serious, is now allowed as long as safeguards are in place, where the data is not special-category data. For special-category data the tighter restriction stays. A bill that would have put binding duties on public bodies did not survive: the Public Authority Algorithmic and Automated Decision-Making Systems Bill passed every stage in the House of Lords by 7 February 2025, never got its first reading in the Commons, and did not become law.
Aotearoa New Zealand has both kinds of thing. The Algorithm Charter, launched in 2020 as a world first, is signed by 29 agencies according to the page that carries the list. That page's content was last reviewed on 15 June 2023, so the figure may not be current. The Charter is voluntary, and it does include a commitment to a channel for challenging decisions, which is the most specific promise of its kind. New Zealand also has a Better Rules programme, the most developed rules-as-code effort by intent, still at the discovery stage.
New Zealand also has binding law, in one area. Since 1 July 2023, Subpart 5A of the Social Security Act 2018 has been in force. It covers social security decisions, not government as a whole. Three things in it matter for layer 5. Appeal and review rights are preserved (section 363D(1)). A human has to be available before an automated system can be approved at all (section 363A(4)(d)). And the automated decision takes legal effect as the decision of a named official (section 363A(7)). Nothing in the subpart limits the technology. The decisions MSD has said it will automate are rules-based, and a rules-based decision has an explicit rule behind every outcome. Boost thinks that is the easy case to get layer 5 right on, and a good place to start.
Singapore has a voluntary governance framework for agentic AI. Estonia, as an EU member, will be bound by the EU AI Act's rules for high-risk systems once it deploys them, including human oversight and a right to an explanation. Bound, but not yet triggered, because the material agentic deployment has not happened.
Add it up. The duties to disclose a system and to assess it before it runs are real, and in Canada and the United Kingdom they are mandatory. Binding law that attaches to the decision itself is real too, in New Zealand's social security system since 2023. What is missing is any of that reaching an agent's decision: an audit of what the agent decided, and a right of appeal against it, at the bar a citizen-affecting decision deserves. As far as Boost has been able to find, no jurisdiction has that in production. That gap is layer 5. On this model, that gap binds.
The United Kingdom is worth a careful word, because it is easy to over-read. It has put AI into benefits work at scale, in tools that help humans rather than decide alone. It has also picked up real accountability trouble while doing so: findings of biased outcomes, complaints about opacity, and reports from rights groups about harm to disabled and marginalised people.
That is not proof of anything about the agentic stage, because the United Kingdom has not reached the agentic stage. These are human-assist systems. The model treats the United Kingdom as an early warning, not as evidence of the agent layer. If the failures already show up at the gentler stage, the case for getting oversight right before the agent stage gets stronger.
Finland, of the places reviewed for the model, is the one that answered the oversight question with law first, before deploying at agentic scale. Its 2023 rules let public bodies use automated decision-making, but only where the decision rules are set in advance and officially approved.
That requirement quietly rules learning AI out of public-sector decisions. A learning system changes its own rules as it goes, so it cannot meet a set-in-advance-and-approved test. The Finnish customs agency says it plainly on its own page: learning AI techniques are not used. The chain is worth reading carefully. The law regulates automated decisions and the rules behind them. The exclusion of learning AI is how the agencies and legal scholars read and apply that law. Finland is not a “no AI” country. It uses AI in plenty of other ways. It has drawn a clear line around the one place that matters most: decisions about people.
New Zealand did not cross this line in 2026. It crossed it in 2023. Subpart 5A of the Social Security Act 2018 was inserted by the Child Support (Pass On) Acts Amendment Act 2023 and came into force on 1 July 2023. That is where the power to make automated decisions in social security comes from, and the safeguards described earlier come with it. The Amendment Act of 28 May 2026 widens that existing power rather than granting a new one: it extends the scope to an “administrative programme”, which means social assistance that does not sit in or under legislation.
The Ministry was clear it means rules-based decisions, not generative AI. The model treats this as the easy case, and a good place to get layer 5 right. Rules-based decisions have an explicit rule behind every outcome. That is more inspectable than a human queue, not less. The opportunity is to turn that hidden inspectability into something a claimant, an auditor, or a tribunal can actually see and contest. The strongest points the opposition made in the debate were about transparency, and that is exactly the layer-5 work to do well.
The strength of a position depends on whether it can be checked. Here is the test.
Measured from 29 July 2026, Boost is wrong if an autonomous agent starts making decisions that go against people, refusing or suspending a benefit, refusing a visa, raising a tax assessment, at material scale, and keeps doing that for at least eighteen months with no serious accountability failure, and the same pattern then holds in two or more comparable countries. Only decisions that go against the person count. A system that only ever says yes does not count.
What counts as a serious accountability failure was pinned in advance, on 29 July 2026, so the bar cannot be set after the evidence arrives. It means one of four things and only these four: a senior court finds the system unlawful or its decisions invalid; a statutory oversight office, an ombudsman, an auditor-general, a privacy or information commissioner, makes an adverse finding about the system as a whole; a minister or a department suspends or withdraws it; or there is a class remediation or compensation programme. Media criticism, academic criticism, one-off successful appeals, and Boost’s own assessment do not count.
Three things short of that would shift the position, and Boost states each one publicly when it sees it.
The first is a single country. If an autonomous agent is deciding against people at material scale, and has been for at least eighteen months, that counts whether or not anything has gone wrong yet. If one jurisdiction crosses that line, the model stops treating this as a position it holds and starts treating it as contested.
The second tests the other half of the claim, which is that where oversight is mature it narrows what a machine is allowed to decide. A country that has the layer-5 machinery in place, the audit trails, the register of what each agent may do, the appeal route, and that still permits autonomous decisions against people at any scale, disconfirms that half in its own terms. The narrowing is the claim, so permitting it at all is enough.
The third is about timing. A country that changes its law to allow decisions made entirely by machine, and then puts such a system into production within twenty-four months of that law starting, triggers an immediate re-test of whether this constraint is binding or only temporary.
One case sits outside the test. An autonomous agent making favourable-only decisions at material scale, sustained for eighteen months in one country, does not make the position wrong. It does weaken it, and Boost will say so publicly.
A few cautions on the argument itself.
The test is strict. Eighteen months, material scale, real decisions. A critic could say the bar is set so high it cannot fail. The reply is that it is the same bar the model would be judged against. If the bar is wrong, the whole position is wrong, and the model would rather be checkable than safe.
“Material scale” is a judgement call, and reasonable people will draw it in different places, especially for chat-style front doors that serve many people without making real decisions.
Proving a negative is hard. Saying a country does not have an agent registry takes a real search that comes up empty, not just a search not yet done. The model marks the difference where it can.
And the view is narrow. The countries the model leans on, Estonia, the United Kingdom, Canada, New Zealand, Finland, plus Singapore, are mostly European or Anglo. A wider scan might show a different pattern. This is where the evidence actually sits.
On this model, layer 5 is the part that decides whether any of the rest works safely, and it is the layer most commentary skips. This is also where Boost has standing to talk. Boost builds government digital services for a living, under real accountability rules. That is what production-grade oversight looks like up close, and it is what makes it possible to tell when a proposed system does not have it.
The series ends here. The model stands as five parts, not five rungs. Whether agentic government works depends most on the two layers that get the least attention: the foundation underneath, and the inspector at the top.
2026-07-29. Corrected: this post described New Zealand’s protections around automated decisions as voluntary only, when Subpart 5A of the Social Security Act 2018 has been binding law in social security since 1 July 2023; and its conclusion that what governments have is lighter than the ambition counted Canada’s and the United Kingdom’s mandatory duties as part of that gap.
2026-07-30. Corrected: this post dated the United Kingdom’s requirement to publish algorithmic transparency records to March 2024, when the government made it a requirement for departments on 6 February 2024 and set out the scope in a policy published on 17 December 2024; March 2024 was the start of the phased rollout, not the date the requirement was made.
2026-07-30. Also updated: the test this post publishes for what would prove this position wrong. The earlier test, measured from 22 May 2026, counted any citizen-affecting decision at material scale. The test now measured from 29 July 2026 counts only decisions that go against the person, and it names in advance what Boost will count as a serious accountability failure. The change makes the position harder to disprove, so it is recorded here rather than made quietly.