The Hon Andrew Leigh MP
Assistant Minister for Productivity, Competition, Charities and Treasury
How government can learn what works
Published in The Mandarin
8 October 2026
For years, JobSeekers using Workforce Australia Online have generally faced a monthly mutual-obligation target of 100 points.
JobSeekers can meet their obligations through job applications, training, study, and other activities. The idea behind the target is that requiring more effort should, up to a point, increase the chances of finding a job.
But is 100 points the right number? Intuition and ideology don’t tell us much. If the target is too high, it seems reasonable to think that the obligation might stop encouraging useful job search and begin consuming time that could be spent finding work in other ways.
To shed light on the question, the Australian government ran an experiment. Almost 40,000 job-ready JobSeekers were randomly divided into three groups. One group remained under the existing arrangements, with a target of 100 points each month, including a minimum of four job applications. For the other two groups, the requirements were reduced to 70 points or 50 points, with three and two job applications, respectively.
Six months later, the employment results were similar. The three groups had worked an average of 10.4 weeks, 10.3 weeks, and 10.5 weeks. Their earnings and income support payments were also very similar.
Researchers also interviewed participants. People facing lighter requirements tended to report less stress. Some said the change gave them more time to pursue activities they thought would actually help them find work.
Before the trial, there were plausible arguments in both directions. Tougher requirements might strengthen the incentive to get a job. Lighter requirements might allow people to search smarter, focusing on the particular job where they were most likely to get hired. The experiment allowed government to find out.
Public policymakers should be judged not only by the quality of the decisions they make, but also by their ability to learn and improve over time. A better feedback loop is one that’s grounded in evidence, not gut feel. It helps ensure policies become a little better every year.
Governments already draw on data, expert advice, international experience, and public consultation. Randomised trials add another tool. By comparing groups that differ only in the policy treatment they receive, they help distinguish what caused an outcome from what merely happened alongside it.
Randomised trials are like pilot programs, but with a control group. That’s why medicine has used randomised trials for decades. The same techniques can help governments test and refine public policy.
Consider a Centrelink phone call. Services Australia often sends a text message before an officer calls a customer. Earlier analysis suggested that the text increased the chance of successful contact. But other factors could have produced some of that difference.
So Services Australia worked with the Australian Centre for Evaluation to run a randomised trial. Across nearly 5,000 outbound contacts, some customers received a text before the call, and others did not. A pre-call text increased successful contact from 52% to 64%.
A text message before a phone call is not the frontier of technological innovation. It is unlikely to feature in the next season of Black Mirror (which is probably a good thing). Yet a 12-percentage point improvement, applied across a large service system, can save thousands of hours for customers and staff.
Another trial concerned the annual reporting obligations of charities. Each year, registered charities need to submit an Annual Information Statement to the Australian Charities and Not-for-profits Commission. The Australian Centre for Evaluation tested an additional reminder sent directly to a charity’s responsible person.
Around 15,000 charities that had yet to lodge were randomly allocated between the intervention and control groups. The extra reminder raised on-time submission from 56% to 62%, with charities also lodging around three days earlier.
Much of government consists of millions of interactions involving forms, letters, websites, and deadlines. A small improvement repeated across a huge administrative system can deliver a substantial return.
Experiments are useful even when the results don’t show that the treatment makes a big difference. Recent banking trials tested ways to encourage customers to seek a better deal on their home loan. A fairly generic prompt, tucked away in a banking app, had almost no effect. A second trial made the message prominent and personalised, telling customers they might be eligible for a lower rate on their particular loan. Contact with the bank rose from 5% to 7%.
The world is more complicated than ‘prompts work’ or ‘prompts fail’. Experiments can show where a promising idea loses power in the real world.
Economists sometimes talk about ‘option value’: the benefit of keeping options open so a better decision can be made later when more information is available. When a decision is hard to reverse, uncertainty is expensive. A trial preserves the ability to change course.
If a program can be tested on 20,000 people before being extended to two million, the experiment itself has an economic return. Its value rises with the scale of the eventual program and the uncertainty around its effects.
There is also a balance between exploiting the information we already know and exploring to find something better. To see this, imagine choosing repeatedly between restaurants. One strategy is to stick with your favourite. Another is to try the place your foodie friend keeps recommending. The first exploits current knowledge. The second explores.
Public policy faces the same problem. Once an approach looks promising, there are good reasons to extend it. Yet a system that always makes the same choice risks failing to improve.
Australia is doing more evaluation than we once did. The 2026 State of Evaluation report, produced by the Australian Centre for Evaluation, identified 847 evaluations underway, planned or recently completed across 41 Commonwealth agencies. Among the 219 impact evaluations were 14 randomised trials conducted across eight agencies.
Other countries are using rigorous evaluation in impressive ways. Britain has brought more evaluation into the policy development and spending process. Canada asks departments bringing forward major proposals to identify gaps in the evidence and consider whether a pilot, randomised trial or other rigorous evaluation could fill them.
We can also do more to get the benefits of Australia’s federation. States and territories already make different choices in schooling, hospitals, transport, planning, and regulation. That variation can generate useful evidence when evaluation is built into policy design.
Eight different policies become eight experiments only when the variation is structured for learning. If we are going to have eight systems, we may as well learn eight times as much.
Randomised trials are not suitable for every policy question. Some questions can be answered from existing evidence. Others require observational studies or qualitative research. Trials are most useful when genuine uncertainty remains and different approaches can fairly be compared.
There is also a democratic benefit. Randomised trials offer a form of public humility. Government can say that we have a view, while remaining willing to test, learn, and adapt. Clinical trials might help explain why public trust in medical research remains higher than trust in many other institutions.
Governments will always have to make decisions with incomplete information. But we will make better choices over time if we are confident enough to test our ideas and curious enough to learn from the results.
This article draws on the assistant minister’s Shann Memorial Lecture, hosted by the UWA Economics Department and the Economic Society of Australia (WA Branch).
Andrew Leigh is the Assistant Minister for Productivity, Competition, Charities and Treasury.
Ends