## BIG DATA: POTENTIAL, CHALLENGES, AND STATISTICAL IMPLICATIONS (SDN1706)

## Source details

**Canonical URL:** [BIG DATA: POTENTIAL, CHALLENGES, AND STATISTICAL IMPLICATIONS (SDN1706)](https://www.imf.org/-/media/files/publications/sdn/2017/sdn1706-bigdata.pdf)

## Other formats

- [Markdown version](/-/media/files/publications/sdn/2017/sdn1706-bigdata.pdf.md)
- [Structured JSON version](/-/media/files/publications/sdn/2017/sdn1706-bigdata.pdf.json)

---

### What big data means
- Characterizations:
  - Often characterized by the 3Vs—high-volume, high-velocity, and high-variety.
  - Expanded list: the 5Vs: Volatility, Variety, Velocity, Veracity, and Volume.
- Nature and scope:
  - Differs from statistical (“made”) data: it is a byproduct “found” in business and administrative systems, social networks, and the internet of things.
  - Important categories for macroeconomic and financial statistics follow an adapted UNECE big data classification (social networks; traditional business systems; internet of things), with additional subcategories such as administrative data, business websites, and online news noted as relevant.
  - Excluded for current macroeconomic and financial statistics use are categories such as videos, medical records, and pictures, although they may become relevant later.

### Potential of big data for macroeconomic and financial statistics
- Three principal features through which big data benefits statistics and policymaking:
  - By answering new questions and producing new indicators.
  - By bridging time lags in the availability of official statistics and supporting the timelier forecasting of existing indicators.
  - By providing an innovative data source in the production of official statistics.
- Applications and illustrative initiatives:
  - Predict stock market liquidity and construct sentiment indices using Google web searches and Facebook posts.
  - Predict socioeconomic patterns, population movements, credit-risk profiles, climate-smart agriculture, and crisis-related stress using mobile phones, call detail records, satellite imagery, and apps.
  - IMF in-house top six proof-of-concept ideas:
    - (1) Using SWIFT data to monitor global financial flows
    - (2) a sentiment-based early-warning system
    - (3) nowcasting GDP using Google trends data
    - (4) automating and expanding the Week @ the Beach Index
    - (5) pooling government cash flow data to enhance surveillance and policy analysis
    - (6) applying analytics for better tax and customs administration
  - Example: M-Pesa in Kenya (established in 2007) — more than 85 percent of households use M-Pesa; some 78,000 agents are distributed everywhere in the country, reaching almost every village.
    - Potential uses of M-Pesa data: measure current transfers between users, estimate consumption patterns, produce regional demand by industry, and provide early signals of changes in demand patterns.
- Asymmetry of opportunities:
  - Opportunities will be asymmetric across countries and depend on country characteristics and the availability of systems and networks generating big data.
  - Big data is especially promising for low-income countries (LICs) where official data are scarce or outdated.

### Nowcasting, forecasting, and illustrative tools
- Nowcasting uses and evidence:
  - Widely used in private and public sectors for indicators including tourism, unemployment, retail trade, and trade flows.
  - Examples: The Billion Prices Project (web-scraped daily price data); Twitter used to nowcast food prices in Indonesia; Predata condensing web data into signals for geopolitical risk.
- SWIFT as a case study:
  - SWIFT used by about 11,000 financial institutions in more than 200 countries; more than 25 million SWIFT messages are sent daily; SWIFT has more than 100 different message types.
  - Available per message type: total number of messages sent and total value of payments.
  - Uses: build indicators of global financial flows by region and currency; construct a SWIFT Index as a mirror for economic activity; nowcast international trade statistics.
  - Limitations: subscription required; limits on single download sizes; publishing controls requiring SWIFT approval; need for large-volume processing capacities; staff training beyond Excel.
- Example index:
  - Week @ the Beach Index: constructs a price index comparing cost of beach holidays using hotel rates, taxi fares, and prices of meals/beverages from Expedia, TripAdvisor, Worldcabfares, and Numbeo. Recorded quarterly; applications include competitiveness for tourism-dependent economies and an alternative measure of equilibrium real exchange rates.

### Main challenges with big data
- Data quality (veracity and volatility):
  - Veracity: noise and bias in the data are among the biggest validity challenges.
  - Volatility: changing technology or business environments can lead to invalid analyses and fragility as a data source.
  - Representativeness concerns: many big data types are nonrandom samples; social media underrepresents nonusers.
  - Time series continuity: short time spans, outliers, and lack of guaranteed continuity require transformation, cleaning, and outlier handling.
  - Need to prioritize research and compilation of statistical techniques that address veracity and volatility.
  - Practical stance: “good enough” information available “now” can be used in some policy contexts even if full quality standards are not met.
- Access and legal issues:
  - Big data often originates outside control of public institutions; access typically requires agreements with private data owners while preserving independence and legal privacy/confidentiality frameworks.
  - Traditional confidentiality techniques may reach limits; new technologies and methods are required to safeguard independence, reputation, and public trust.
  - Licensing and proprietary costs may be significant; big data is increasingly a commercial asset.
- Privacy, confidentiality, and cybersecurity risks:
  - Large amounts of sensitive personal information create risks of cyberattacks, profiling, resale to third parties, and legal restrictions on use without prior knowledge.
  - Potential loss of data can lead to reputational damage and loss of consumer trust.
  - Need for thorough privacy protection procedures, cryptography, anonymization, and user access control to prevent re-identification.
  - Consumer rights may require strengthening via regulatory intervention.

### Statistical implications and institutional considerations
- Institutional responses:
  - International statistical cooperation is key to overcoming challenges and building lasting partnerships among national and international statistical agencies, users, and data owners.
  - Statistical agencies should increase involvement in big data projects and collaborate with international organizations leading methodological development.
  - Participate in networks such as the United Nations Global Working Group on Big Data and the UNECE Sandbox.
- Methodological and operational implications:
  - Incorporation of big data as new data sources—either supplementing or substituting traditional sources—imposes methodological, organizational, and budgetary challenges.
  - Big data projects succeed by establishing environments of people and processes, not by implementing specific technologies alone.
  - Big data provides an opportunity to break internal organizational silos between users and producers of data and statistics.
- Governance and project lifecycle:
  - Typical early-stage institutional actions:
    - (1) identify standard technology platforms for multiple projects;
    - (2) establish governance around use, monitoring, and costs;
    - (3) establish a network of experts with specific skills;
    - (4) make business and budgetary provisions to fund future costs.
  - Cloud technology recommended for flexibility and scale; governance required to manage “pay for what you use” models.

### Technology, costs, and skills
- Cloud budgeting and costs:
  - Typical first-year cloud budgetary allocation: about $250,000 for the first year.
  - Cloud costs thereafter follow “pay for what you use” and need strong governance.
  - The cost of data itself is unpredictable; many sources require recurring licensing.
- Team composition and scale:
  - Multidisciplinary teams needed: data scientists, IT architecture specialists, data visualization specialists, and subject matter professionals.
  - Core team size in initial stages: three to four staff members.
- Infrastructure examples:
  - UNECE Sandbox technical/current terms:
    - set of five servers;
    - reliable storage of up to 56 terabytes;
    - high-performance concurrent processing across 80 CPU cores;
    - unified login portal connected via a high-bandwidth network connection;
    - current annual fee is €10,000.

### Task forces, pilot workflow, and capacity development
- Task force workflow and pilot process:
  - Specialized task forces start with brainstorming to identify desired outputs.
  - Phases:
    - Phase 1–2: conceptualization and specification by task forces within time targets set by cross-departmental groups.
    - Phase 3: technical implementation led by IT in collaboration with task force experts.
    - User assessment for “fit for purpose” and adequacy for surveillance.
    - If fit, pilot, incorporate into surveillance, expand scale, and potentially include in capacity development activities.
- Dynamic guidance and capacity building:
  - Produce dynamic and interactive guidance notes assessing methodological soundness, data reliability, and policy relevance.
  - Capacity development should provide expertise on data quality judgment, institutional change management, strategic partnerships, and user communication.
  - Leverage networks in academia, private sector, and international statistical community (for example, UNSD Global Working Group).

### Fixed sensors and machine-generated data (Internet of Things)
- Classification and sensor types:
  - Fixed sensors are a subset of Internet of Things (machine-generated data): 311. Fixed sensors include:
    - 3111. Home automation
    - 3112. Weather/pollution sensors
    - 3113. Traffic sensors/webcam
    - 3114. Scientific sensors
- Potential indicators derivable from fixed sensors:
  - Home automation / Smart meters:
    - nonoccupancy rates
    - household consumption
    - electricity supply and consumption
    - price differentials
    - household structure and size
    - Relevant statistical domains: Environmental and energy statistics; National accounts; Price statistics; Demographic and social statistics; Transportation statistics; Geospatial statistics; Agricultural statistics; Rural and population statistics
    - Potential classification tags: 1,2,3
  - Weather/pollution sensors:
    - environmental and climate-related indicators useful for agricultural and land-use measurement
  - Traffic sensors / webcam:
    - traffic intensity
    - commuting time
    - incoming/outgoing traffic
    - proxy of economic growth/health
    - Relevant statistical domains: National accounts; External sector statistics; Transport statistics; Tourism statistics; Mobility statistics
    - Potential classification tags: 1,2,3
  - Scientific sensors:
    - domain-specific measurements feeding specialized statistics and research applications
- Linkages with other geospatial sources:
  - Satellite imagery, GPS positioning, and traffic sensors can be integrated for improved localization, sampling frames, land use, crop area measurement, trip duration, remoteness indices, and tourism/travel statistics.
  - Potential classification tags for linked sources: 1,3 (satellite imagery); 1,2,3 (GPS positioning/tracking).

### Conclusions and work ahead (selected findings)
- Strategic needs and priorities:
  - Many national and international statistical organizations view big data as a potentially strategic asset requiring a vision and plan.
  - Organizations need strategic plans to select promising applications that complement official statistics and add value such as improved timeliness, support for forecasting, and production of new indicators.
  - Begin with proofs of concept/pilot projects; operationalize only after findings prove valuable and feasible organizationally—especially in low-income countries.
  - Sound partnerships, legal clarity, and the right skills and technologies are as important as statistical expertise and data representativeness.
  - Best practices for partnerships and legal frameworks are being developed and field-tested.
  - Big data offers greater promise for statistics on flows and transactions, insights, correlations, trends, and sentiments; currently less promise for statistics on stocks or breakdowns of flows into revaluations and other changes.
- Institutional role:
  - International organizations responsible for official statistics should cooperate closely with user departments and integrate big data considerations into capacity development activities.
  - The IMF’s Statistical Department will continue its mission to provide leadership in application of sound statistical practices and to foster the use of big data for macroeconomic and financial statistics.

*Source: sdn1706-bigdata — BIG DATA: POTENTIAL, CHALLENGES, AND STATISTICAL IMPLICATIONS (excerpt).*

### EXECUTIVE SUMMARY _______________________________________________________________________________ 4

### EXECUTIVE SUMMARY

### What big data means
- The term is often characterized by the 3Vs—high-volume, high-velocity, and high-variety.
- The list of Vs has expanded to include veracity and volatility, yielding the 5Vs: Volatility, Variety, Velocity, Veracity, and Volume.
- Big data differs from statistical (“made”) data: it is a byproduct “found” in business and administrative systems, social networks, and the internet of things.
- Important categories for macroeconomic and financial statistics follow a classification based on the UNECE big data classification (see Appendix I), with additional subcategories such as administrative data, business websites, and online news noted as relevant.
- Excluded for current macroeconomic and financial statistics use are categories such as videos, medical records, and pictures, although they may become relevant later.

### Potential of big data for macroeconomic and financial statistics
- Big data can benefit policymaking through at least three features:
  - By answering new questions and producing new indicators.
  - By bridging time lags in the availability of official statistics and supporting timelier forecasting of existing indicators.
  - As an innovative data source in the production of official statistics.
- Big data can provide innovative, real-time, and more granular insight for economic and financial analysis.
- Opportunities from big data will be asymmetric across countries and depend on country characteristics and the availability of systems and networks generating big data.
- Numerous individual applications and pilot studies are already being carried out by both users and compilers of data and statistics.

### Main challenges with big data
- Data quality concerns:
  - Veracity: noise and bias in the data are among the biggest challenges to validity.
  - Volatility: changing technology or business environments can lead to invalid analyses and fragility as a data source.
- Access to big data:
  - Partnerships with private firms and data owners are necessary; voluntary data sharing examples exist but are not universal.
- New skills and technologies:
  - Handling big data requires a network of human skills, advanced technologies, and data access infrastructure.
  - Statistical agencies and policymaking organizations face challenges in building these capabilities.

### Statistical implications and institutional considerations
- International statistical cooperation is key to overcoming big data challenges and building lasting partnerships among national and international statistical agencies, users, and data owners.
- Implications for organizations include the need to build required skills and technologies internally.
- For national statistical agencies, incorporating big data as new data sources—either supplementing or substituting traditional sources—imposes methodological, organizational, and budgetary challenges.
- Big data projects succeed not by implementing specific technologies but by establishing an environment of people and processes to operationalize innovations.
- Big data provides an opportunity to break internal organizational silos, including between users and producers of data and statistics.

### Recommended approach and next steps
- Start with a proof of concept; operationalize projects only after findings prove valuable and organizationally feasible.
- Statistical agencies should decide on a case-by-case basis and select the most promising big data projects to complement existing statistics.
- Proactively search for big data sources to address urgent research needs.
- Incorporate selected big data projects into capacity development activities to help members build capabilities to benefit from available big data sources.
- Prioritize research and compilation of best practices—especially statistical techniques and methodologies that address veracity and volatility.

### Overarching assessment
- Big data is evolutionary and dynamic; systems and networks generating big data will continue to evolve, changing the possibilities, challenges, and limitations for statistics.
- The assessment provided needs periodic revisiting as the worlds of big data and official statistics evolve.

*Source: EXECUTIVE SUMMARY, BIG DATA: POTENTIAL, CHALLENGES, AND STATISTICAL IMPLICATIONS (SDN1706).*

### Box 1. Adapted UNECE Big Data Classification

### Box 1. Adapted UNECE Big Data Classification

### 1. Big data classification (adapted UNECE)
- 1. Social Networks (human-sourced information)
  - 1100. Social Networks: Facebook, Twitter, LinkedIn
  - 1200. Blogs and comments
  - 1600. Internet searches on search engines (Google)
  - 1700. Mobile data content: text messages, Call Detail Record, Data Detail Record, Location update, Radio coverage updates; Online news
- 2. Traditional Business systems (process-mediated data)
  - 21. Data produced by public agencies; Administrative data
  - 22. Data produced by businesses
    - 2210. Commercial transactions
    - 2220. Banking/stock records
    - 2230. E-commerce
    - 2240. Credit cards; Business websites; Scanner data
- 3. Internet of Things (machine-generated data)
  - 31. Data from sensors
    - 311. Fixed sensors
      - 3111. Home automation
      - 3112. Weather/pollution sensors
      - 3113. Traffic sensors/webcam
      - 3114. Scientific sensors
    - 312. Mobile sensors (tracking)
      - 3121. Mobile phone location (GPS)
      - 3122. Cars
      - 3123. Satellite images

### 2. Country-specific considerations and prerequisites
- Availability of social networks, traditional business/administrative systems, and the internet of things will vary by country; big data opportunities depend on country characteristics.
- To assess big data opportunities for policymaking one must consider:
  - availability of data and related tools,
  - users’ capabilities,
  - privacy and security issues,
  - legal and technological systems.

### 3. Three features through which big data benefits statistics and policymaking
- 1. By answering new questions and producing new indicators
- 2. By bridging time lags in the availability of official statistics and supporting the timelier forecasting of existing indicators
- 3. By providing an innovative data source in the production of official statistics

### 4. Feature A — Answering new questions and producing new indicators
- Big data emphasizes patterns and correlations; may alert that something is happening even if causality is not established (Mayer-Schönberger and Cukier 2014).
- Potential applications and examples:
  - Predict stock market liquidity and construct sentiment indices using Google web searches and Facebook posts (Arouri and others 2014; Karabulut 2013).
  - Predict socioeconomic patterns, population movements, credit-risk profiles, climate-smart agriculture, and crisis-related stress using mobile phones, call detail records, satellite imagery, and apps.
  - Support measurement for Sustainable Development Goals (SDG) indicators, including gender equality; LinkedIn publishing gender diversity statistics and providing training on gender statistics (Karani 2017).
- IMF in-house initiatives and proof-of-concept projects (top six ideas approved for development):
  - (1) Using SWIFT data to monitor global financial flows
  - (2) a sentiment-based early-warning system
  - (3) nowcasting GDP using Google trends data
  - (4) automating and expanding the Week @ the Beach Index
  - (5) pooling government cash flow data to enhance surveillance and policy analysis
  - (6) applying analytics for better tax and customs administration
- Additional IMF uses mentioned:
  - Use of big administrative data for the IMF Fiscal Affairs Department Revenue Administration Gap Analysis Program to determine the VAT compliance gap.
  - IMF Research Department’s use of big geophysical data to measure effects on low-income countries’ economic activity.
- Big data can improve measurement of financial inclusion and its effects on poverty and growth via electronic money systems (example: M-Pesa).
  - M-Pesa: established in 2007 in Kenya; in Kenya, more than 85 percent of households use M-Pesa; some 78,000 agents are distributed everywhere in the country, reaching almost every village.
  - Potential uses of M-Pesa data:
    - Measuring current transfers between M-Pesa users to produce more accurate estimates of remittances, disposable income, and financial inclusion.
    - Estimating consumption patterns to provide regional demand by industry and early signals of changes in demand patterns.
- Big data is especially promising for low-income countries (LICs) where official data are scarce or outdated; human-sourced information and counterparty business data can help fill household-level data gaps.
- Big data can cross-check third-party indicators (TPIs) used in IMF staff reports and serve as benchmarks to assess TPIs’ quality and methods.
- Big data can prompt rethinking of economic modeling and increase adoption of machine learning and natural language processing techniques at the IMF (examples cited: IMF working paper on nowcasting in Lebanon; “Text Mining, Inform and Connect” seminar).

### 5. Feature B — Bridging time lags and supporting forecasting (nowcasting)
- Faster insights: financial and price data observable almost instantaneously; timely data are essential for policy-relevant statistics.
- SWIFT highlighted as a promising data source:
  - SWIFT used by about 11,000 financial institutions in more than 200 countries; more than 25 million SWIFT messages are sent daily; SWIFT has more than 100 different message types.
  - Available for each message type: total number of messages sent and total value of payments.
  - Example applications of SWIFT data:
    - Build indicators capturing global financial flows by region and currency.
    - The SWIFT Index: uses SWIFT traffic (volume) as a mirror for economic activity to provide early GDP growth estimates.
    - Use in IMF work: data contributed to the 2015 SDR review as an indicator of renminbi internationalization; IMF proof-of-concept work in 2016 used SWIFT to nowcast international trade statistics.
  - Benefits from SWIFT data:
    - Monitor currency compositions.
    - Nowcast export/import indicators by distinguishing trade from financial transactions.
    - Assess risks to correspondent banking relationships and support banks’ due diligence measures.
  - Limitations and considerations with SWIFT data:
    - Access restrictions: subscription required; limits on size of single downloads.
    - Publishing controls: SWIFT approval needed for external publication to maintain confidentiality.
    - Data processing: need to build processing capacities for large volumes.
    - Staff cost and training: familiarity required; Excel insufficient for large-data work.
- Big data enables almost-real-time extraction of economic signals to nowcast series before official figures, strengthening surveillance, stability monitoring, and early-warning for fast-moving markets and international monetary system risks.

### 6. Illustrative indices and tools
- Week @ the Beach Index:
  - Constructs a price index comparing cost of beach holidays using hotel rates, taxi fares, and prices of meals/beverages; data drawn from Expedia, TripAdvisor, Worldcabfares, and Numbeo.
  - Recorded quarterly with applications including: competitiveness indicator for tourism-dependent economies, alternative measure of equilibrium real exchange rates, and explanatory variable in empirical work.
  - Sharing-economy data (e.g., Airbnb) could be incorporated to improve market and price data.

*Source: sdn1706-bigdata - Box 1. Adapted UNECE Big Data Classification*

### 26.      Nowcasting (or real-time forecasting) is already widely used in the private and

### Nowcasting (or real-time forecasting) is already widely used in the private and public sector

### Nowcasting: examples and applications
- Nowcasting is used across private and public sectors for many indicators (FRBNY 2017; Banbura and others 2013).
- Specific examples:
  - The Billion Prices Project at MIT: web scraping price data from online retailers to monitor price changes and inflation at a daily frequency and turning points in inflation trends.
  - Twitter used to nowcast food prices in Indonesia by filtering and modeling tweets with price information (UNGP 2014).
  - Predata (2016): condenses web data into signals for geopolitical risk by monitoring digital conversations across open-source social and collaborative media; machine learning applied to anticipate asset price changes, civil protests, labor strikes, electoral results, and national security outcomes.
- Other prominent nowcasting uses: economic series on tourism, unemployment, retail trade, and trade flows.
- Statistical compilers are experimenting with social networks and satellite imagery:
  - The Netherlands: Facebook and Twitter data to estimate consumer sentiment.
  - China and Italy: web scraping to approximate job vacancy rates.
  - Google Trends: nowcasting unemployment, consumer sentiment, car sales, and property sales (Pavlicek and Kristoufem 2015).

### Benefits of big data for forecasting and nowcasting
- Big data can, in some cases, provide better predictors than traditional leading indicators (OECD 1987) and improve forecasting accuracy because of the larger volume of available data.
- More detailed and granular big data can improve the quality of estimated economic series (Galbraith and Tkacz 2013).
- Forecasting exercises use available information to assess recent past developments relative to expectations.
- Big data complements existing statistical series by providing:
  - More granular and higher-frequency observations.
  - Socioeconomic detail and a larger number of observations.
- Potential outcome: improved forecasting accuracy through harnessing and increasing observations from big data.

### Complementarity with official statistics and continuity requirements
- Even when used for forecasting, consistent and harmonized historical time series remain necessary.
- Big data indicators should be compared with existing official time series to ensure robustness; big data measures “insights” and correlations (movements, trends, sentiments) vs. “actual information” (positions of outstanding debt at the end of the year).
- The two sources—official statistics and big data—are complementary, not substitutes in many cases.

### Big data integration into official statistics: strategies and pilot uses
- International recognition and commitments:
  - UN Statistical Commission (45th session in 2014): “Big Data constitute a source of information that cannot be ignored” (UNGWG 2017).
  - 2013 Scheveningen Memorandum: European national statistical offices committed to integrating big data strategies and road maps.
  - Eurostat: typical use of big data will be in combination with existing sources (EC 2017).
- Integration strategies range from partially to entirely replacing existing sources and from providing improved to complementary or new statistical outputs (Florescu and others 2014).
- Drivers: limited budgets and declining survey response rates (for example, Meyer, Mok, and Sullivan 2015).
- Pilots and documented uses by official compilers:
  - Mobile phone data for tourism, transportation, and urban statistics (Eurostat, Belgium, Brazil, Indonesia, Israel, Italy, World Bank in Nigeria, Poland).
  - Web scraping for price indices, labor market indicators, enterprise profiling (Eurostat, China, Ecuador, Finland, Germany, Hungary, Japan).
  - Smart meters for energy and environmental statistics (Eurostat, Belgium, Canada).
  - Credit card, cash register, and scanner data for price and other economic statistics.
- Big data sources mapped to official statistical domains are presented in Table 1 in Section V (based on the UN classification and tailored to IMF needs).

### Box 5 — Mobile Positioning Data for International Travel Service Statistics (summary)
- Background:
  - In 2010 Statistics Estonia stopped using border surveys; Eesti Pank explored alternatives and found mobile positioning data simplest, least costly, and timeliest. No legal obstacles to access in Estonia.
- Uses and methodology:
  - Mobile positioning estimates trip duration, inbound and outbound international travelers, and monthly travel services exports/imports.
  - Cell phone location data provide SIM card ID, date and time, location, and country code. Residency assumed from SIM card residency.
  - Methodology based on anonymized roaming pattern analysis:
    - Nonresidents in resident networks estimate inbound travelers and country.
    - Roaming activity reports from resident operators to partner nonresident operators estimate outbound travelers and country.
    - Algorithms adjust for short/long visits, transit travel (harbors, airports, transit roads), permanent workers from other countries, and border noise (ship traffic and random switching).
    - Length of stay derived from days between first and last roaming activity; mapping of multi-country trips included.
- Benefits compared with conventional surveys and administrative sources:
  - Enhanced quality: improved accuracy, quality, and comprehensiveness; captures travelers staying with relatives/friends and nonregistered accommodations; more detailed breakdowns by region/country of residence.
  - Timeliness: data available almost in real time; monthly and quarterly estimates with practically no time lag.
  - Cost: data readily available; costs much lower than surveys and administrative data collection and processing.
- Source cited in box: http://www.oecd.org/trade/its/46287481.ppt.

### Administrative data and its relation to big data (Box 6 summary)
- Administrative data: a distinctive form of big data, generated as byproducts of large-scale administrative systems for nonstatistical purposes; extensively used by statistics agencies.
- Examples: tax data, public finance transaction data, social security records.
- Experience with administrative data has paved the way for statistical agencies to use big data, provided data quality, methodological soundness, and privacy/confidentiality are ensured.
- Uses of administrative records:
  - Sole source for some indicators (customs data for merchandise trade; International Transactions Reporting Systems for balance of payments).
  - Combined with other sources (tax records to measure production).
  - Verification and calibration (compare survey estimates with administrative estimates).
  - Indirect estimation (benchmarking).
  - Survey design (leveraging detail richness).
- Nordic countries benchmark: extensive use of registers for statistics (Population and Housing censuses based on registers in Denmark 1980, Finland 1990, Norway and Sweden 2011).
- Registers support:
  - Statistics based on one register.
  - Combining multiple registers.
  - Combining registers with survey data.
  - Sampling frames.
  - Quality control.
- Preconditions for effective administrative data use:
  - National statistical office access rights to administrative sources for official statistics.
  - Unique identification for all individuals and businesses to combine data.
  - Adherence to confidentiality (use only aggregated form).
  - Influence of statistical agencies in design and generation of administrative data.
- Source cited in box: UNWDF 2017.

### Key challenges with big data
- Data quality (Section IV.A):
  - Quality assessment of big data-derived indicators is crucial to minimize governance, political, and reputational risks.
  - Newly produced indicators need benchmarking to existing series during experimental phases to meet minimum data quality standards for real, external, fiscal, monetary, and financial statistics.
  - Suitability assessment should consider accuracy, sustainability, and methodological soundness; metadata are key for interpretation and assessment.
  - No assurance that private-sector-produced big data frameworks will persist; availability, comparability, and coherence of time series are at risk (Kitchin 2015).
  - Representativeness issues: many big data types are nonrandom samples; social media underrepresents nonusers; users must assess representativeness and suitability.
  - Big data indicators often have short time spans, contain outliers, and lack guaranteed continuity; transformation, cleaning, outlier replacement, and handling missing observations are needed (Eurostat 2016).
  - The 5V characteristics: big data is complex, incomplete, noisy, and can contain outliers and extreme events. Lack of methodological frameworks and best practices could undermine statistical quality standards.
  - Research and compilation of statistical techniques addressing veracity and volatility should be prioritized.
  - Practical stance: “good enough” information available “now” can be used for specific actions even if it does not meet comprehensive quality standards (Meyer and others 2013); sentiment analysis can uncover meaningful early signals.
- Access to big data (Section IV.B):
  - Big data origin/creation mostly outside control of national/international institutions, except administrative data.
  - Access typically requires agreements with private data owners while preserving independence and ensuring legal frameworks for privacy and confidentiality.
  - Users should be transparent about data source legitimacy and purpose when access is granted.
  - Traditional confidentiality techniques may reach limits; new technologies and methods required to safeguard independence, reputation, and public trust.
  - Legislative bodies may need to negotiate access to private-sector big data to ensure compiling agencies are not excluded from the big data revolution.

*BIG DATA: POTENTIAL, CHALLENGES, AND STATISTICAL IMPLICATIONS, INTERNATIONAL MONETARY FUND*

### 42.      Privacy, confidentiality, and cybersecurity risks are a major concern when using big

### 42. Privacy, confidentiality, and cybersecurity risks are a major concern when using big data

### Key risks and implications
- Big data contains large amounts of sensitive and personal information that may be exposed to privacy, confidentiality, and cybersecurity risks.
- If not sufficiently protected, information can be vulnerable to cyberattacks, used to profile individuals, and sold to third parties.
- Information, depending on the national legal framework, may not be used legally without a person’s prior knowledge.
- The potential loss of data can lead to reputational damages as well as loss of consumers’ trust.
- International institutions and public agencies need to ensure data sources and indicators were obtained without any violation of privacy or confidentiality regimes.

### Technical and procedural mitigations
- Companies, public agencies, and third-party data users must establish thorough privacy protection procedures.
- Invest in security layers and adapt traditional IT techniques such as cryptography, anonymization, and user access control to big data characteristics to safeguard privacy and prevent reconstruction and re-identification.
- Consumer rights may need strengthening through regulatory intervention, detailing what companies can do with collected information.

### Institutional arrangements and access
- Public-private partnerships are being established, but access rights are sometimes unaddressed and unclear.
- The UN Global Working Group is preparing “access recommendations” and best practices for partnerships between official statistics producers and data owners.
- What is needed are standardized, ethical, and stable data sharing tools and protocols with private companies.
- Risk of volatility: no assurance that a company and its data will exist in the future, creating continuity risks for analytical and surveillance applications.

### Costs and market dynamics
- The price of big data will not necessarily be low or negligible; licensing or proprietary costs for accessing information may be significant.
- These costs add to acquiring processing and storage technology for in-house use or access through external vendors.
- Big data is becoming a major asset for companies; private firms increasingly sell their data for profit and may generate data as an objective rather than a byproduct.
- The statistical community and official users will face complex negotiation processes with private data owners.

### Technology selection and implementation guidance
- Typical early-stage institutional actions for establishing a big data practice:
  - (1) identify standard technology platforms that can be provisioned by teams working on various projects;
  - (2) establish governance around the use, monitoring, and costs associated with the environment;
  - (3) establish a network to tap into experts with specific skills;
  - (4) make business and budgetary provisions to fund future costs.
- Cloud technology offers flexibility to scale infrastructure and is the preferred industry approach.
- Typical first-year cloud budgetary allocation: about $250,000 for the first year.
- Cloud costs thereafter follow “pay for what you use” and require strong governance to ensure efficient use of funds.
- Expectations:
  - Cloud-based systems will see wider adoption and storage/computing costs will continue to drop.
  - The IMF has started provisioning cloud service and expects associated costs to grow incrementally with the incorporation of big data into analytic practices.
- The cost of data itself is hard to predict; some sources may not exist as discrete products or may require licensing from commercial sources. Both cloud services and commercial data are typically licensed on a recurring basis with administrative budget implications.

### Skills, teams, and governance
- Multidisciplinary teams are needed: data scientists, IT architecture specialists, data visualization specialists, and subject matter professionals.
- On average, a big data practice in initial stages starts with a core team of three to four staff members with a mix of technical and business skills.
- Governance implications highlighted in Box 7 (Rethinking IT & IT Governance):
  - Need to rethink long-standing business practices and legacy technology.
  - Information security and privacy: understand data provenance and ensure national and international standards for data privacy.
  - Governance of research projects: address data retention, data destruction, cost of running large models, and sharing of software code and algorithms.
  - Managing research projects through their life cycle: big data research is iterative and may produce new data sets, models, and visualizations rather than production systems.
  - Open-source software: prevalent in big data research; adopting similar support at the IMF would be initially countercultural.
  - Cloud computing: reluctance to adopt the cloud must be addressed given low cost, deployment speed, and maturing service and security models.

### Statistical implications and recommendations
- Statistical agencies should increase involvement in big data projects and collaborate with international organizations leading methodological development.
- Participate in existing big data networks (for example, the United Nations Global Working Group on Big Data) and contribute to robust governance frameworks that promote international and national cooperation.
- Statistical agencies need to inventory projects that could enrich official macroeconomic and financial statistics and cultivate links with big data expert networks in academia, private sector, and international community to streamline training, skills, and capacity development.
- The UNECE High-Level Group for the Modernization of Official Statistics (HLG-MOS) created the Sandbox as a shared platform and data repository (subject to confidentiality constraints) for testing tools, techniques, workflows, and methodologies; participation in the Sandbox is recommended, especially for countries with less-developed statistical systems.
  - Sandbox technical/current terms noted: set of five servers; reliable storage of up to 56 terabytes; high-performance concurrent processing across 80 CPU cores; unified login portal connected via a high-bandwidth network connection.
  - The current annual fee is €10,000.
  - Statistics Netherlands (CBS) used the Big Data Sandbox as a prototype for its internal platform.
- Identify the most promising big data sources for macroeconomic and financial statistics over time, recognizing asymmetric opportunities across countries.
- Develop new data quality concepts and expand existing frameworks to incorporate big data opportunities and challenges; preserve comparability over time and across countries where possible.
- Coordinate efforts through standing committees on statistics and expert groups; update international statistical manuals and guides (for example, IMF BPM6 Compilation Guide and UNECE Guide to Measuring Global Production) as needed.
- Inside organizations, statistics and user departments should collaborate to identify a core set of early-warning and snapshot indicators for policymaking. Transparency of methodology and data origin is essential.
- Establish specialized task forces (topic-related and IT experts) to pilot projects, report to cross-departmental groups, and help break silos by concentrating multidisciplinary expertise on selected big data initiatives.

### Selected numeric and technical details preserved from source
- Typical first-year cloud budgetary allocation: about $250,000 for the first year.
- Core team size in initial stages: three to four staff members.
- Historical estimates and milestones (Box 7):
  - “the world produces between 1 and 2 exabytes of unique information per year”
  - “roughly 250 megabytes for every man, woman, and child on earth.”
- UNECE Sandbox configuration and fee:
  - set of five servers;
  - reliable storage of up to 56 terabytes;
  - high-performance concurrent processing across 80 CPU cores;
  - current annual fee is €10,000.

*International Monetary Fund — BIG DATA: POTENTIAL, CHALLENGES, AND STATISTICAL IMPLICATIONS (excerpt).*

### 61.      These specialized task forces could start with a brainstorming phase (Figure 4) to

### 61.      These specialized task forces could start with a brainstorming phase (Figure 4) to

### Task force workflow and pilot process
- Specialized task forces could start with a brainstorming phase (Figure 4) to identify the desired outputs.
- Phases 1 and 2 could lie with the established task forces, working within the time targets set by the cross-departmental group.
- Phase 3 would be the technical implementation led by IT teams in collaborations with the task force experts.
- Simultaneously, the new products would have to be assessed by users for their “fit for purpose” (adequacy) for surveillance work.
- Statistics departments should provide the methodological guidance.
- If regarded as “fit for purpose” the new products could be:
  - piloted,
  - incorporated into surveillance work,
  - expanded to a larger scale.
- In the last phase, institutions could decide whether to include the new products in capacity development activities (training, technical assistance).
- Footnote/example: 24 Such as the International Agency Group on Economic and Financial Statistics, the IMF Balance of Payments Statistics Committee (BOPCOM), Government Finance Statistics Advisory Committee (GFSC), and the Inter Secretariat Working Group on National Accounts (ISWGNA).

### Dynamic and interactive guidance notes
- An outcome of the specialized task forces could be dynamic and interactive guidance notes.
- Guidance notes would entail assessments of pilot projects for:
  - methodological soundness,
  - reliability of data source,
  - potential to serve as an input to policy analysis for both international and national organizations.
- Guidance notes should be:
  - dynamic (timely and high-frequency updates),
  - interactive (online based documents, with interactive links and search functions).
- For some projects, new methodologies may have to be developed.

### Capacity development and technical assistance
- Capacity Development and Technical Assistance considerations for institutions such as the IMF:
  - Results for statistics and for indicators directly usable in bilateral and multilateral surveillance could be incorporated into capacity development programs after further exploration of potential and challenges.
  - Capacity development could encompass provision of expertise in:
    - judgment of data quality issues and methodologies,
    - conveyance of best practices on institutional change management,
    - building of strategic partnerships,
    - communication with users.
  - For the inventory of best practices, international institutions need to cultivate strong links to networks of big data experts in:
    - academia,
    - the private sector,
    - the international statistical community.
  - Use the network to streamline international efforts on training, skills, and capacity development in member countries.
- Existing expert group example: The UNSD Global Working Group, composed of nine other international and regional organizations and 22 countries, is an already established expert group on big data in the field of official statistics.

### Conclusions and work ahead (selected findings)
- 64. Many national and international statistical organizations have recognized that big data is not just a buzzword, but a potentially strategic asset that requires a vision and a plan.
- Big data necessitates a strategy to select the most promising applications to complement official statistics and to bring added value, such as:
  - improved timeliness,
  - supporting forecasting of existing data sets,
  - producing new indicators.
- 58. Organizations need to go beyond the existing individual and scattered applications of big data; require a strategic organizational plan to deliver measurable and high-scale results.
- Before large investments, especially in low-income countries, organizations should begin with a proof of concept or pilot project; operationalization should follow only after findings prove valuable and feasible organizationally.
- Sound partnerships, legal issues, and the right skills and technologies are as important as:
  - statistical expertise,
  - data representativeness and methodological accuracy,
  - effective collaboration between data scientists and subject matter economists.
- 59. Best practices are being developed and tested for partnerships between official statistical agencies and data owners; legal questions are being clarified; best-use cases are field-tested.
- Big data success depends on creating an environment of people and processes, not implementing a single technology; projects are an opportunity to break institutional silos.
- 65. Opportunities offered by big data to macroeconomic and financial statistics vary considerably across statistical domains:
  - promising opportunities for statistics on flows and transactions, insights, correlations, trends, and sentiments,
  - currently less promise for statistics on stocks or breakdowns of flows into transactions, revaluations, and other volume changes.
- Detailed country-by-country time series in accordance with internationally agreed standards remain crucial for measuring and monitoring countries’ economic performance and policies over time.
- 66. International organizations responsible for official statistics should work in close cooperation with user departments; consider a new dimension for international statistical coordination and cooperation, including incorporation into capacity development activities.
- The IMF’s Statistical Department will continue “its mission to provide strong leadership for the development and application of sound statistical practices” and reach out to the rest of the IMF and the community at large to foster the use of big data for macroeconomic and financial statistics as input for policymaking.
- 67. Big data is dynamic; systems and networks generating it will continue to evolve, as will the opportunities, challenges, and statistical implications.

### Appendices and classifications (selected excerpts)
- Appendix I. Classification Developed by the UNECE Task Team on Big Data, (UNECE Wiki June 2013) 25
  - 1. Social Networks (human-sourced information): Data are loosely structured and often ungoverned.
    - 1100. Social Networks: Facebook, Twitter, Tumblr etc.
    - 1200. Blogs and comments
    - 1300. Personal documents
    - 1400. Pictures: Instagram, Flickr, Picasa, etc.
    - 1500. Videos: YouTube etc.
    - 1600. Internet searches
    - 1700. Mobile data content: text messages
    - 1800. User-generated maps
    - 1900. E-mail
  - 2. Traditional Business Systems (process-mediated data): Highly structured; includes transactions, reference tables, relationships, and metadata.
    - 21. Data Produced by Public Agencies
      - 2110. Medical records
    - 22. Data Produced by Businesses
      - 2210. Commercial transactions
      - 2220. Banking/stock records
      - 2230. E-commerce
      - 2240. Credit cards
  - 3. Internet of Things (machine-generated data): Well-structured but large and fast; includes sensors and system logs.
    - 31. Data from Sensors
      - 311. Fixed sensors
        - 3111. Home automation
        - 3112. Weather/pollution sensors
        - 3113. Traffic sensors/webcam
        - 3114. Scientific sensors
        - 3115. Security/surveillance videos/images
      - 312. Mobile sensors (tracking)
        - 3121. Mobile phone location
        - 3122. Cars
        - 3123. Satellite images
    - 32. Data from Computer Systems
      - 3210. Logs
      - 3220. Web logs

*Source: sdn1706-bigdata - 61.      These specialized task forces could start with a brainstorming phase (Figure 4) to*

### 311. Fixed sensors

### 311. Fixed sensors

### Overview
- Fixed sensors are a subset of Internet of Things (machine-generated data) used in statistical compilation and analysis.
- They are listed under the broader "Internet of Things" category alongside mobile sensors (tracking) such as Mobile phone location, Cars, and Satellite images.
- The classification follows an adapted UN big data classification and is linked to the framework where:
  - 1 = Big data to answer "new questions" and produce new indicators
  - 2 = Big data to bridge time lags in the availability of official statistics and supporting the more timely forecasting of existing indicators
  - 3 = Big data as an innovative data source in the production of official statistics

### Sensor types (as enumerated)
- 3111. Home automation
- 3112. Weather/pollution sensors
- 3113. Traffic sensors/webcam
- 3114. Scientific sensors

### Potential indicators derivable from fixed sensors
- Home automation / Smart meters (energy consumption measures)
  - nonoccupancy rates
  - household consumption
  - electricity supply and consumption
  - price differentials
  - household structure and size
  - Relevant statistical domains: Environmental and energy statistics; National accounts; Price statistics; Demographic and social statistics; Transportation statistics; Geospatial statistics; Agricultural statistics; Rural and population statistics
  - Potential classification tags: 1,2,3
- Weather/pollution sensors
  - Research and mapping of weather and climate data (noted in adjacent satellite imagery entry)
  - Environmental and climate-related indicators useful for agricultural and land-use measurement (see Satellite images linkage)
- Traffic sensors / webcam
  - traffic intensity
  - commuting time
  - incoming/outgoing traffic
  - proxy of economic growth/health
  - Relevant statistical domains: National accounts; External sector statistics; Transport statistics; Tourism statistics; Mobility statistics
  - Potential classification tags: 1,2,3
- Scientific sensors
  - Provide domain-specific measurements that can feed into specialized statistics and research applications

### Contextual linkages to other sensor and geospatial sources
- Fixed sensors integrate with other machine-generated and geospatial data to produce or augment indicators:
  - Traffic/road sensors serve as a proxy of economic growth/health and link to travel/tourism statistics.
  - Satellite imagery supports:
    - improved geographical localization of statistical units and assets
    - spatial sampling frame for output measurement
    - land use and geostatistical cartography
    - crop planting area, land use, and agricultural output
    - population and asset location as proxy for SDG "Gender Equality"
    - Relevant statistical domains: National accounts; Price statistics; External sector statistics; Demographic and social statistics; Transport statistics; Agricultural statistics; Demographic and urban statistics
    - Potential classification tags: 1,3
  - GPS positioning/tracking data support:
    - travel services exports/imports
    - trip duration
    - inbound/outbound international travelers
    - remoteness index
    - traffic intensity
    - Relevant statistical domains: National accounts; External sector statistics; Demographic statistics; Transport statistics; Urban statistics; Tourism statistics; Population statistics
    - Potential classification tags: 1,2,3

### Statistical potential and applications
- Fixed sensors can:
  - Produce high-frequency indicators (bridging time lags — tag 2).
  - Offer new types of measures and proxies for traditional statistics, such as:
    - nonoccupancy and household consumption (from smart meters)
    - traffic intensity and commuting measures (from traffic sensors/webcams)
  - Serve as innovative data sources in official statistics production (tag 3) and to answer new questions (tag 1).
- These sensor-derived indicators feed multiple statistical domains including: National accounts; Price statistics; External sector statistics; Demographic and social statistics; Transport statistics; Environmental and energy statistics; Agricultural statistics; Geospatial and urban statistics.

*Source: sdn1706-bigdata - 311. Fixed sensors*

---


_Source: https://www.imf.org/-/media/files/publications/sdn/2017/sdn1706-bigdata.pdf_
