## THE HIDDEN PRICE OF DATA

## Source details

**Canonical URL:** [THE HIDDEN PRICE OF DATA](https://www.imf.org/-/media/files/publications/fandd/article/2025/12/veldkamp.pdf)

## Other formats

- [Markdown version](/-/media/files/publications/fandd/article/2025/12/veldkamp.pdf.md)
- [Structured JSON version](/-/media/files/publications/fandd/article/2025/12/veldkamp.pdf.json)

---

### Overview
- Data is the input that powers artificial intelligence algorithms and is produced as a by-product of everyday activities: searches, clicks, app use, and transactions.
- Data lacks an explicit market price and does not have an observable production cost like government services; firms incur costs to process and analyze data but the underlying raw data is ambient.
- Consumers effectively receive a net price that bundles monetary payment and the implicit sale of their data; firms have incentives to discount observable prices to generate more transactions and thus more data.

### Bundling and consumer impact
- Transactions in the digital economy are typically bundled: consumers simultaneously buy goods/services and sell their data.
- Because the price of the data component is hidden, consumers cannot observe the discount they receive for selling their data.
- Bundling prevents consumers from learning fair data prices over time, keeping them in the position of "first-day tourists" who consistently undersell their data.
- Regulatory unbundling—requiring firms to post both the price that includes rights to use transaction data and the price for a private transaction—would reveal the data discount and allow consumers to choose whether to supply data.

### Five approaches to valuing data
- Market prices approach
  - Uses prices from open data marketplaces (examples: Snowflake, Datarade) where data sets are traded.
  - Limitation: marketplace data is not representative because firms typically do not sell their most valuable, competitive data.
- Revenue approach
  - Treats data as a productive asset worth the extra revenue it generates.
  - Requires counterfactual modeling: estimating what profits would have been without the data.
  - Feasible in settings like finance where data-driven investment decisions are measurable; harder when data has multiple, less observable uses.
- Complementary inputs approach
  - Infers the value of a firm's data stock from the resources devoted to managing and exploiting data (labor, computing power).
  - Premise: firms spend real money on inputs only when the underlying data is valuable.
- Correlated behavior approach
  - Measures the alignment between actions and rewards to infer the informational content of data.
  - Examples: accuracy of recommendations matching purchases; firm's ability to stockpile goods that will sell well.
  - High covariance between actions and payoffs implies valuable data.
- Cost-accounting approach
  - Counts the bills accountants pay for data; the United Nations System of National Accounts counts purchased data sets as assets.
  - Shortcoming: most data is bartered (consumers pay with information), and implicit discounts are not itemized on books.
  - Requires imputing the value of dollars or cents discounted to elicit more transactions and data revelation.
  - Would be facilitated by unbundling transactions and requiring separate pricing for transactions with and without data-use rights.

### Toward quantification
- The five approaches capture different facets of data value: labor devoted, revenue earned, precision of actions, market prices, and implicit cost.
- No single approach is infallible or universally feasible; measurement of data value will remain imperfect.
- To craft sound policy and inform decisions, data must be moved from intuition into quantification.
- Unbundling data and goods transactions and requiring separate pricing would make cost-accounting approaches more feasible and illuminate the implicit discounts consumers give up.

### Policy implications and recommendations
- Require firms to unbundle transactions and post separate prices: one price for the right to use transaction data and one for a private transaction without data rights.
- Reveal the implicit data discount to allow consumers to decide whether to supply data or withhold it unless adequately compensated.
- Develop and adopt measurement tool kits that combine multiple approaches to estimate data value in different contexts.
- Recognize that measurement is difficult but necessary to prevent opaque extraction of value by firms and to enable consumers to become active suppliers demanding fair value.

*laura veldkamp is the Leon G. Cooperman Professor of Finance & Economics at Columbia University’s Graduate School of Business and author of The Data Economy: Tools and Applications.*

---


_Source: https://www.imf.org/-/media/files/publications/fandd/article/2025/12/veldkamp.pdf_
