StatGPT: AI for Official Statistics
Departmental Papers, March 10, 2026
Source details
- Canonical URL
- StatGPT: AI for Official Statistics
Other formats
Bibliographic details
- Authors: James Tebrake, El Bachir Boukherouaa, Jeff Danforth, Miss Nivashini Harikrishnan
- Published: March 10, 2026
- Series: Departmental Papers
- DOI: https://doi.org/10.5089/9798229036863.087
Overview and motivation
- National statistical systems produce statistics that underpin policy, economic analysis, and public trust.
- Two persistent challenges limit impact: data accessibility and interpretability.
- Large language models (LLMs) and GenAI applications (examples cited: ChatGPT and Gemini) enable natural-language retrieval but, according to testing described in the paper, “frequently provide dangerously ‘reasonable’ but incorrect figures.”
- StatGPT is an IMF Statistics Department initiative that leverages LLMs not to generate statistics, but to generate structured queries against APIs of official statistical agencies, ensuring users receive the exact published figures while benefiting from natural language interaction.
Key findings from the paper
- Off-the-shelf GenAI applications:
- Excel at synthesizing text.
- Perform poorly at delivering official statistics, producing plausible but incorrect numeric answers.
- StatGPT approach:
- Uses LLMs to translate natural-language requests into structured API queries.
- Ensures retrieval of exact published figures “every time.”
- The paper examines limitations of generic GenAI for statistical use and demonstrates how StatGPT overcomes them.
Proposed roadmap and requirements for AI-readiness of official statistics
- Open data access:
- Emphasizes the need for APIs of official statistical agencies to be available for structured querying.
- Enriched metadata standards:
- Calls for improved metadata to enable correct interpretation and unambiguous querying of statistical series.
- Strengthened data governance:
- Recommends governance frameworks that align technological innovation with statistical rigor to preserve authoritativeness and trust.
Policy recommendations and implications
- Align technological innovation with statistical rigor to keep official statistics authoritative, trusted, and universally accessible in an AI-driven world.
- Adopt StatGPT-style architectures that separate natural-language understanding from statistical production—i.e., translate queries into official API calls rather than allow LLMs to synthesize numeric outputs directly.
- Improve data accessibility, metadata quality, and governance to make official statistics “AI-ready.”
Content in this bundle
- StatGPT: AI for Official Statistics; Departmental Paper No. 26/04; March 2026