
The federal government is adopting AI faster than almost anyone predicted a year ago. Chief AI Officers are being appointed across agencies, generative tools are reaching tens of thousands of public servants, and the APS AI Plan 2025 has set the policy direction for all of it. The Prime Minister has now gone a step further, announcing an Office of AI inside his own portfolio, the Department of Prime Minister and Cabinet, to draw the whole effort into a single national framework. Everyone is asking how quickly agencies can adopt. Far fewer can say whether any of it is working.
That second question is the one that gets asked at Senate Estimates and tested in an ANAO performance audit, and it is the one most agencies are least equipped to answer. They are set up to count adoption, when the thing worth knowing is impact.
Adoption is the easy number. It is also the wrong one.
Ask an agency how its AI rollout is going and you will usually get a usage figure: seats activated, prompts run, logins this month. Those numbers are easy to pull from an admin console and close to meaningless as a measure of value. Heavy use can still be shallow use, polishing emails while the casework goes untouched, and a usage chart cannot tell you which one you have.

This is not a new problem. It is an old one in new clothes.
A decade ago, research in this country was judged largely by counting outputs: publications produced, citations accrued and grant dollars won. The numbers were real, but they were silent on the question that mattered most, which was whether any of the research had changed anything in the world.
So the sector built a different tool. The Australian Research Council’s Engagement and Impact Assessment set out to capture the two things the output counts missed. Engagement meant how deeply researchers worked with the people outside academia who would actually use their work. Impact meant the difference that work made beyond the university.
I spent the better part of a decade inside one of Australia’s leading universities while that shift was underway, responsible for the academics whose work was judged against it. That lesson translates almost word for word to AI in government, where the gap between activity and impact is the same as the gap between usage and engagement. The agencies that learn to close that gap get the most from what they have built.
Measuring AI well is less about a bigger dashboard than a truer one, and three moves get you there.
The first is to separate activity from engagement. Activity is the seat count. Engagement is whether AI has been taken up inside the casework and the decisions themselves, not in the meeting notes and tidier emails around them. The Commonwealth has already run this experiment: the 2024 whole-of-government Copilot trial put licences in front of more than 7,000 public servants, yet only about a third use the tool daily. Most of what they did with it was summarising and first drafts. Shallow uptake and deep integration look identical in a usage report and could not be more different in their worth.
The second is to track impact against known workflows instead of AI in the abstract. The questions are concrete ones: whether a veteran’s claim or a grant application is moving faster, or whether a class of assessments has gained accuracy. Behind those sit harder questions: how much capacity has been freed for the work only people can do, and what has actually changed for the citizen at the other end. These are benefits-realisation questions, and government already owns the discipline for them in the DTA’s Benefits Management Policy. It has simply never been pointed at AI.
The third is to stay honest. When a new tool is introduced or a process is improved, part of the gain belongs to the people not just the new tool: teams reshape the process around it, or lift their game because the work is being measured differently. AI is rarely the only thing changing a benefit, so a credible reading treats it as a contribution and never the sole cause, because the fastest way to lose a benefits story in an audit is to overclaim it. Half-a-dozen measures tied to decisions an executive already has to make, such as whether to renew the licences or extend the tool into the next workflow, will beat fifty dashboard metrics that inform no decision at all.
Measurement is the clearest example of a wider pattern: the parts of the AI shift that neither the technology nor the vendors will do for you. They rely on the human behaviours and governance system the AI runs inside.
Start with the workforce question: as AI disrupts the work, an agency has to decide which roles grow and evolve, and whether the roles are redesigned deliberately or settled by attrition and accident.
Next, accountability: once AI drafts a brief or scores an application, the question is who is accountable for the result, and Robodebt is the standing reminder of what it costs when the answer is ‘nobody’. The other question is whether an agency’s ‘human in the loop’ survives contact with the automated-decision transparency rules that take effect in December 2026. The answer to both is a failure-tolerant design: assume the system will sometimes get things wrong, catch the error close to the work, redress the outcome of the error for those it affected, and run a short feedback loop from the person who spotted the mistake to the fix that stops it recurring. A scheme built to repair its mistakes quickly is a safer pair of hands than one built to deny it makes any.
Then the operating model: whether AI-augmented work survives the APS Code of Conduct and an ANAO audit.
And finally capability, which is far bigger than a how-to guide on prompting.
None of this needs a new model or a new platform. All of it needs people who understand both government and change. That is the unglamorous middle of the AI story, and it is where most of the value, and most of the risk, actually sit.

The technology is the easy part; the hard part is the human and governance system around it.
At Parbery, this is where we work: how AI is measured, how the workforce is designed around it, how accountability and operating models are settled and how real capability is built. This is the work we already do: benefits realisation, workforce and operating-model design, change leadership, applied to a new and fast-moving variable.
Take measurement, the subject most agencies ask us about first. It is the kind of work we set up, and the approach we favour starts nowhere near the AI. It begins with a process that has run long enough to have a real baseline, an enduring, scaled workflow where the numbers already exist: the mean time to process a grant application is a good one. You introduce the tool at a known point and hold everything else steady, then watch the trend over time rather than a single before-and-after baseline. Where the change happened matters more than whether some metric came down: did the augmented step get faster on its own, or did the improvement carry through the whole process instead of shifting the bottleneck further along? Instrumenting the exact spot where the tool enters, and reading it as workstream velocity, is what separates a real efficiency gain from a number that looks better on a slide.
Across a stable reporting period, the trend tells you what a usage chart never can: whether the people doing the work picked the tool up at all, and whether it improved anything worth improving: the time taken to assess grants; the error rate in the assessments; and how much of the tool’s output still gets re-checked before any decisions are made. That is efficiency, accuracy and confidence, measured against the workflow.
We will be honest that our own position is still forming. Like the rest of the sector, we are still working out what that position should be, and we would rather do that thinking in the open than pretend it is settled. But one part of it is already settled: the value of government AI will be decided less by how much of it gets switched on than by how well the organisation around it is built to use it, and to prove that it did.
If you were asked tomorrow what impacts your AI had actually changed, could you defend your answer with confidence? That question now has an email address. With an Office of AI sitting at the centre of government, someone will eventually put that question to every agency, and a seat count is not going to cut it.
Move from ideas to outcomes with confidence
Reach out to discuss how we can partner in designing, delivering and embedding practical solutions that create real impact from day one.