Don’t Just Buy the Model:
When AI in State Government Is a Team-Design Problem
AI adoption in government extends far beyond procurement. Anyone who has seen old tech stacks piling up, or been forced to use an overengineered solution, knows that adopting AI is not the same as integrating it effectively. We can issue the RFP, score vendors on benchmark performance, buy the most capable model, and bolt it onto the existing processes. That instinct may move us in the right direction, but it is an incomplete solution.
For expanded state capacity, we ought also to consider the roles staffed by people around a model, where their judgment is exercised, and how decisions are recorded and contested. A body of evidence suggests that the return a state captures from AI depends far less on which model it buys than on the workflow and organization into which that model is placed. States that get the org chart wrong will underperform states that bought a weaker model and designed the work around it well.
A great deal of government work is routine, high-volume, and low-stakes, involving intake forms, data extraction, first-pass document classification, and status notifications. Automating workflows entirely composed of these types of tasks increases state employees’ capacity to instead focus on tasks requiring human judgement and expertise. However, for a particular class of tasks, that is, discretion-heavy, non-routine, consequential decisions made under genuine uncertainty, how human and machine roles are organized matters at least as much as raw model capability, and frequently more. Fraud triage, benefits eligibility, child welfare screening, audit selection, and licensing exceptions all sit in that class. The trick is knowing which bucket a task belongs in and designing the human–AI team accordingly. And with the rightly understood upside of AI, designing the best method of integration could mean the difference between an ever-expanding bureaucratic state and a revolution in efficiency and public trust in governance.
The Economics Points at Tasks, Not Jobs
Labor economics shows us that AI exposure typically automates tasks rather than entire occupations, and while most jobs contain some automatable tasks, at this point, very few jobs are entirely composed of them.[1],[2] According to the standard task framework, while productivity gains will come in part from substitution, much of the gain will come from complementarity, workflow reorganization, and the creation of new tasks.[3],[4] Instead of human–AI interaction as a vestige of incomplete automation, we can anticipate it being at the locus of where much of the value is actually produced.
The productivity evidence is real but uneven, highlighting the importance of team design. Brynjolfsson, Li, and Raymond find a 15 percent average improvement among customer support agents supported by AI assistance, with that improvement concentrated among less-experienced workers.[5] Noy and Zhang find that generative AI assistance raises output on mid-level writing tasks by roughly 40 percent while compressing the skill distribution.[6] And Dell’Acqua and colleagues, in a field experiment with consultants, find substantial gains on tasks inside the model’s frontier and substantial losses on tasks outside of it, with workers unable to reliably tell when the model was performing better or worse.[7] Realizing the anticipated gains will be contingent on the match among task, worker, and system, and that match is largely an integration design choice, rather than a property of the model.
“Keep a Human in the Loop” Is Weaker Than It Sounds
The reflexive safeguard for AI implementation is to add a human approver at the end. For discretion-heavy tasks, that design tends to fail in predictable ways. Automation bias, or the tendency to defer to automated outputs even against contradictory evidence, is one of the best-established findings in human-factors research.[8] AI-provided explanations do not reliably fix the bias, as these AI explanations frequently increase overreliance because they sound persuasive, rather than meaningfully helping the user separate good outputs from bad ones.[9] And when the human is positioned only as a terminal gate, the self-checking automated pipeline hands them a fluent, internally consistent artifact under time pressure, with the residual error concentrated in framing, which is the hardest error to catch from a finished product. Purported oversight degenerates into post-hoc approval theater, in that it changes who is accountable without improving the decision.
For government specifically, two failure modes are especially costly. In Sweden’s Public Employment Service, opacity combined with efficiency pressure led frontline staff to disengage from an AI decision-support system, rather than utilizing it to do the slower, substantive work required.[10] The AI system’s intended human teammates became passive. Furthermore, institutional trust itself is at stake with Wuttke, Rauchfleisch, and Jungherr finding that AI-driven performance improvements can simultaneously raise short-term trust and erode perceived control, with trust then also declining once people recognize they can no longer contest or reverse a result.[11] In public administration, contestability and auditability are not usability niceties; they are part of what makes a decision legitimate.
Design the Team, Not Just the Tool
The alternative is to design the team, not just the tool. One such design is to treat structured disagreement as a feature of the workflow rather than a defect to be smoothed away. The Agonism Hypothesis holds that, for high-uncertainty, non-routine, consequential decisions, human–AI teams that embed bounded, task-focused friction at points of epistemic uncertainty outperform both full automation and approval-only oversight. The friction must earn its cost, and it is warranted only where premature convergence, automation bias, or weak accountability would otherwise be expensive. Friction should surface evidence and reasons, preserve live alternatives, calibrate reliance on the machine, and leave an auditable record of how the decision was reached. It generalizes a scattered literature on cognitive forcing functions and productive friction to the level of team architecture.[12]
What This Means for State Policy
A state’s AI strategy is an organizational design strategy, and its procurement documents should reflect this with the following three components. First, specify the workflow and the human roles, including where judgment is exercised, what information gets logged, and how a decision can be contested after the fact. A procured model’s parameters are important, and we want to attend to them, just not at the expense of intentional integration. Second, budget for the complements, including role redesign, training, and decision procedures. These are where partial automation becomes reliable performance. Third, evaluate on more than accuracy. Throughput and cost are easy to measure; error correction, calibrated reliance, auditability, and the maintenance of staff expertise over time will determine whether the system still supports a state’s needs years from now.
[1] Tyna Eloundou, Sam Manning, Pamela Mishkin, and Daniel Rock, “GPTs Are GPTs: An Early Look at the Labor Market Impact Potential of Large Language Models,” working paper, 2023, https://arxiv.org/abs/2303.10130.
[2] Daron Acemoglu, “The Simple Macroeconomics of AI,” NBER Working Paper 32487 (Cambridge, MA: National Bureau of Economic Research, May 2024), https://www.nber.org/papers/w32487.
[3] Daron Acemoglu and Pascual Restrepo, “Artificial Intelligence, Automation and Work,” NBER Working Paper 24196 (Cambridge, MA: National Bureau of Economic Research, January 2018), https://doi.org/10.3386/w24196.
[4] Daron Acemoglu and Pascual Restrepo, “Automation and New Tasks: How Technology Displaces and Reinstates Labor,” Journal of Economic Perspectives 33, no. 2 (2019): 3-30, https://doi.org/10.1257/jep.33.2.3.
[5] Erik Brynjolfsson, Danielle Li, and Lindsey R. Raymond, “Generative AI at Work,” Quarterly Journal of Economics 140, no. 2 (2025): 889-942, https://doi.org/10.1093/qje/qjae044.
[6] Shakked Noy and Whitney Zhang, “Experimental Evidence on the Productivity Effects of Generative Artificial Intelligence,” Science 381, no. 6654 (2023): 187-192, https://doi.org/10.1126/science.adh2586.
[7] Fabrizio Dell’Acqua et al., “Navigating the Jagged Technological Frontier: Field Experimental Evidence of the Effects of Artificial Intelligence on Knowledge Worker Productivity and Quality,” Organization Science 37, no. 2 (2026): 403-423, https://doi.org/10.1287/orsc.2025.21838.
[8] Raja Parasuraman and Dietrich H. Manzey, “Complacency and Bias in Human Use of Automation: An Attentional Integration,” Human Factors 52, no. 3 (2010): 381-410, https://doi.org/10.1177/0018720810376055.
[9] Gagan Bansal et al., “Does the Whole Exceed Its Parts? The Effect of AI Explanations on Complementary Team Performance,” in Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems (New York: Association for Computing Machinery, 2021), 1-16, https://doi.org/10.1145/3411764.3445717.
[10] Alexander Berman, Karl de Fine Licht, and Vanja Carlsson, “Trustworthy AI in the Public Sector: An Empirical Analysis of a Swedish Labor Market Decision-Support System,” Technology in Society 76 (2024): 102471, https://doi.org/10.1016/j.techsoc.2024.102471.
[11] Alexander Wuttke, Adrian Rauchfleisch, and Andreas Jungherr, “Artificial Intelligence in Government: Why People Feel They Lose Control,” preprint, arXiv:2505.01085, 2025, https://doi.org/10.48550/arXiv.2505.01085.
[12] Zana Buçinca, Maja Barbara Malaya, and Krzysztof Z. Gajos, “To Trust or to Think: Cognitive Forcing Functions Can Reduce Overreliance on AI in AI-Assisted Decision-Making,” Proceedings of the ACM on Human-Computer Interaction 5, no. CSCW1, article 188 (2021): 1-21, https://doi.org/10.1145/3449287.