## Big Data: Potential, Challenges and Statistical Implications

_Staff Discussion Notes, September 13, 2017_

## Source details

**Canonical URL:** [Big Data: Potential, Challenges and Statistical Implications](https://www.imf.org/en/publications/staff-discussion-notes/issues/2017/09/13/big-data-potential-challenges-and-statistical-implications-45106)

## Other formats

- [Markdown version](/en/publications/staff-discussion-notes/issues/2017/09/13/big-data-potential-challenges-and-statistical-implications-45106/index.md)
- [Structured JSON version](/en/publications/staff-discussion-notes/issues/2017/09/13/big-data-potential-challenges-and-statistical-implications-45106/index.json)
- [Bundle manifest](/en/publications/staff-discussion-notes/issues/2017/09/13/big-data-potential-challenges-and-statistical-implications-45106/bundle-manifest.json)

## Bibliographic details
- Authors: Cornelia Hammer, Diane C Kostroch, Gabriel Quiros-Romero
- Published: September 13, 2017
- Series: Staff Discussion Notes
- DOI: https://doi.org/10.5089/9781484310908.006

---

### Summary and framing
- Big data are part of a paradigm shift that is significantly transforming statistical agencies, processes, and data analysis.
- Administrative and satellite data are already well established; the statistical community is now experimenting with:
  - structured and unstructured human-sourced data,
  - process-mediated data,
  - machine-generated big data.
- The Staff Discussion Note (SDN) sets out a typology of big data for statistics and highlights that opportunities to exploit big data for official statistics will vary across countries and statistical domains.
- The SDN provides examples from a diverse set of countries to illustrate opportunities.
- The SDN discusses key challenges associated with proprietary data from the private sector regarding accessibility, representativeness, and sustainability.
- The SDN concludes by discussing implications for the statistical community going forward.

### Typology and data types (as presented)
- Structured human-sourced data
- Unstructured human-sourced data
- Process-mediated data
- Machine-generated data
- Administrative data (well established)
- Satellite data (well established)

### Opportunities and illustrative examples
- Opportunities to exploit big data for official statistics differ across:
  - countries,
  - statistical domains.
- The SDN presents examples from a diverse set of countries to illustrate how opportunities vary (examples summarized qualitatively in the SDN).

### Key challenges identified
- Accessibility
  - Proprietary data from the private sector may be difficult to access for statistical agencies.
- Representativeness
  - Big data sources may not be representative of the population or economic activity of interest.
- Sustainability
  - Reliance on private-sector data sources raises concerns about long-term availability and continuity.

### Implications for the statistical community
- The SDN discusses implications going forward for statistical agencies, processes, and data analysis in light of the paradigm shift toward big data.
- The discussion emphasizes adapting statistical practices and institutional arrangements to:
  - evaluate and integrate new data sources,
  - address issues of access, representativeness, and sustainability,
  - exploit machine-generated and human-sourced data where appropriate.

---

## Content in this bundle

- **Staff Discussion Note**
  - [Staff Discussion Note (Markdown version)](/-/media/files/publications/sdn/2017/sdn1706-bigdata.pdf.md){rel="alternate" type="text/markdown"}
  - [Staff Discussion Note (PDF)](/-/media/files/publications/sdn/2017/sdn1706-bigdata.pdf){rel="external" type="application/pdf"}

---

_Source: https://www.imf.org/en/publications/staff-discussion-notes/issues/2017/09/13/big-data-potential-challenges-and-statistical-implications-45106_
