AI in Government: Earning Trust Before Shipping Features
Most AI products are sold on what they can do. In government, that pitch falls flat almost immediately. A village administrator does not want to hear about model capabilities or benchmark scores. They want to know one thing: when this system makes a decision that affects a citizen’s pension, land record, or welfare payment, can it be trusted, and can it be explained? Building Apli for last-mile governance taught us that in the public sector, trust is not a feature you add after shipping. It is the thing you have to earn before you are allowed to ship at all.
Why the public sector is the hardest test
Enterprise AI has a comfortable escape hatch: if a system is wrong, there is usually a human with the authority and the context to catch it, and a commercial relationship that can absorb the mistake. Government has neither luxury. The people most affected by a public-sector system are frequently those with the least power to contest a bad outcome, and the decisions carry the weight of the state behind them. A hallucinated answer in a consumer chatbot is a bad joke. A wrong eligibility decision in a welfare system is a family that does not eat.
In government, the cost of being wrong is not a refund. It is a citizen who was failed by their own state.
What earning trust actually requires
We stopped asking "what can we automate" and started asking "what can we automate in a way that a sceptical citizen, an auditor, and an administrator would all accept." That reframing led us to a few hard commitments:
- Explainability over cleverness. Every consequential output has to be traceable to the rule or record that produced it. A decision no one can explain is a decision no one should trust, however accurate it is on average.
- The human stays in charge. The system is built to assist the administrator, not replace their judgement. It surfaces, flags, and drafts, the person decides, especially where the outcome is irreversible.
- Accessibility is not optional. Last-mile governance means serving people on low-end phones, patchy networks, and in their own language. A system that only works for the digitally fluent deepens the divide it was meant to close.
- Auditability by design. If a regulator or citizen asks why a decision was made, the answer must already exist in the logs. You cannot reconstruct trust after the fact.
The counterintuitive lesson
The feature we were proudest of was not the most capable one. It was the one that admitted uncertainty, the system saying, in effect, "this case is ambiguous, a human should look at it" rather than producing a confident answer it could not justify. Administrators trusted the whole system more because of the cases it refused to decide. Knowing where it would not go was what made them comfortable with where it would.
There is a broader lesson here for anyone building AI in high-stakes domains. The instinct is to lead with capability, because capability demos well. But in the places where AI could do the most good, public services, healthcare, the systems people cannot opt out of, capability is table stakes and trust is the product. Earn the trust first. The features will have somewhere to land.
Want to go deeper?
Talk to the team building this. We'd love to hear about the problems you're trying to solve.
Get in touch →