Federal statistics enters the age of AI — cautiously

Machine Learning


Federal statistics occupy an important place in public life, but they are often invisible. Created and maintained with taxpayer dollars, this data informs policy, guides business decisions, supports research, helps people make everyday choices in their lives, and understands events across the country.

Many federal statistical agencies are beginning to implement artificial intelligence to improve the speed and efficiency of their systems. However, to maintain the accuracy and reliability of public data, government agencies must assess the effectiveness of AI, determine necessary guardrails, and use the technology in ways that are ethical, transparent, replicable, and useful.

Federal Statistics AI Day 2026 explored the opportunities and potential risks of using AI in government data. The event, convened by the National Academies’ Committee on National Statistics, the Federal Committee on Statistical Methods, and the National Institute of Statistical Sciences, reviewed active experimentation with AI across several federal agencies, including the Census Bureau, Centers for Disease Control and Prevention, Bureau of Economic Analysis, NASA, and the National Center for Health Statistics.

The statistician in the driver’s seat

David Matteson, director of the National Institute of Statistical Sciences, said that when AI is deployed in federal agencies, statisticians need to be involved early in the implementation and decision-making process.

“Statisticians must be at the heart of AI leadership,” he argued. “Statisticians add evaluation, discipline, inference, validity, and accountability to decisions made under uncertainty.”

“Statisticians should not be downstream reviewers asking you to bless the system afterwards.” [AI tools have] It is deployed. They should be partners early and often, and they should be the leader from the beginning, or a partner at arm’s length from the leader. ” A major pitfall is modernization without rigorous evaluation, which can give the appearance of progress while introducing new risks.

“Without a statistical framework, you run the risk of optimizing the wrong thing very well,” Matheson says.

Successful conduct and evaluation of research

Surveys are central to federal statistics, but they are expensive, complex, and increasingly difficult to conduct. AI reduces the burden of manual coding associated with processing survey data, while improving the handling of complex text in response fields. Linda Laughlin, director of the U.S. Census Bureau’s Division of Industry and Occupational Statistics, discussed how moving from traditional automated coding systems (autocoders) to large-scale language model-based systems can support better interpretation of industry and employment data.

“While the legacy system served us well, it was highly maintenance-dependent and performance degraded as the language evolved and the quality of the responses degraded,” Laughlin explained. “Advances in large-scale language models have greatly expanded what is possible with coding textual data.”

Traditional coders are still valuable in data evaluation, but large-scale language models (LLMs) can speed up your work. In 2012, approximately 30 percent of cases were automatically coded and 70 percent were sent to the Census Bureau’s National Processing Center for administrative coding. Using LLM, the Census Bureau has been able to code more efficiently, resulting in approximately 500,000 fewer cases being sent for manual coding each year, Laughlin said.

However, this process relies on careful verification. Agencies need to know not just whether a model produces an answer, but also where errors are concentrated, how those errors affect published statistics, and when human review is needed, Laughlin said.

Christina Grigorich, assistant professor of computer science at Johns Hopkins University, focused on another promising but sensitive use case in surveying: response simulation. From pre-testing questions to simulating respondents to improving sampling, LLMs can potentially enhance and speed up parts of the survey lifecycle.

“Maybe you can decide who to ask what questions and in what order,” Grigorich surmised. She went on to pose a central systematic question. “How can a government agency know whether an AI-generated response is providing valid information or just adding noise?”

AI systems may also reproduce what researchers expect to hear, a risk sometimes described as social dissonance.

“A respondent who says what we want to hear isn’t actually a good respondent, right?” Grigorich said.

The success of applying LLM in survey emulation depends not on whether the simulated responses sound plausible, but on whether the inferences are improved, bias is reduced, and they can be evaluated against known benchmarks.

Gizem Korkmaz, vice president of data science and AI at Westat, emphasized how principles such as validity, reliability, fairness, transparency, and objectivity need to be translated into steps that project teams can mobilize. He highlighted elements of effective governance and implementation, including the need for practical strategies such as human review, documentation, reproducibility standards, model validation, monitoring, and lifecycle governance.

institutional approach

To facilitate the successful adoption of these standards, Zach Whitman, chief AI and data officer at the General Services Administration, described USAi, a platform aimed at providing ready-to-use AI tools, capabilities, and services to federal users. He said that the research community’s access to these tools not only ensures a streamlined way to access federal statistical data, but also facilitates adoption by academia, private sector innovators, and other federal agencies.

Benjamin Rogers, CDC’s acting deputy chief officer for AI, said the agency’s strategy is focused on supporting, enhancing, advancing, and empowering staff in the use of AI. The impact was felt across the agency, he said, with a 527 percent return on investment, a $3.7 million reduction in personnel costs, and approximately 41,460 hours of CDC staff time that could be redirected to higher-value work.

The public is also using AI tools for self-diagnosis, changing the way Americans interact with their health care providers.

“Gallup recently published a study showing that as many as 14 million Americans may have canceled a provider visit in the past month because of information from the GenAI tool,” Rogers said. “This is important to me at the CDC, but you can see the impact across the board.”

Privacy is paramount

Speakers emphasized that privacy and confidentiality should be the main focus when using AI in statistical institutions. “We need to take into account the fact that AI and machine learning models can retain this information. [AI] We use it to improve performance,” said Cordell Golden, a data scientist at the CDC’s National Center for Health Statistics.

Lisa Mirel, program director at the National Science Foundation’s National Center for Science and Technology Statistics, addressed privacy-enhancing technologies in the National Security Data Service and emphasized the need for a path to implementation that focuses on risk management and privacy and confidentiality. This concern for data permeates important early considerations for production workflow processes. Who (or what) chooses what data is eligible for the model, how is uncertainty measured, how is output reviewed, when results are rejected, and how are limitations communicated to users?

Mr. Golden discussed incorporating AI and machine learning into data collaboration workflows. Improving efficiency and data quality must be balanced with privacy protection and analytical validity.

continuous

These topics are not new to the National Academies. Previous reports, etc. Aiming for a national data infrastructure for the 21st century report series and A roadmap for avoiding disclosure in income and program participation research.we have considered how statistical systems can expand access to and usefulness of data while protecting individuals and maintaining trust. But AI exacerbates these concerns, as models make it easier to discover, combine, summarize, and reuse data, participants said, in some cases allowing it to be reused in ways not anticipated by existing governance frameworks.

Planning is worthless, but planning is everything.

AI can help increase the efficiency and capacity of federal statistical systems. But to meet that potential, agencies must maintain the qualities that make federal statistics valuable in the first place: rigor, fairness, transparency, confidentiality, and public trust. The future of AI in federal statistics will depend not just on what models can do, but how carefully agencies decide what to do and proceed intentionally, iteratively, and visibly.

“Accountability, transparency, reproducibility and public trust are at the forefront of the development of these AI tools,” Matheson emphasized.



Source link