## RPPI Practical Guide

## Source details

**Canonical URL:** [RPPI Practical Guide](https://www.imf.org/-/media/files/data/guides/rppi/rppi-guide.pdf)

## Other formats

- [Markdown version](/-/media/files/data/guides/rppi/rppi-guide.pdf.md)
- [Structured JSON version](/-/media/files/data/guides/rppi/rppi-guide.pdf.json)

---

### Overview and purpose
- The Guide sets out practical advice on the compilation of a Residential Property Price Index (RPPI) based on the conceptual approach described in the Handbook on Residential Property Price Indices (Handbook), published in 2013.
- Prepared by the IMF Statistics Department using experience from technical assistance and training to RPPI compilers from more than 80 countries over the past six years.
- Target audience: data compilers in statistical offices, central banks, housing agencies, or other statistical agencies; assumes background in statistics and some knowledge of econometrics.
- Includes practical exercises using a Synthetic House Price Dataset generated by CSO Ireland and R scripts prepared for RPPI compilation.

### Key uses and context
- Principal uses of RPPIs:
  - macro-economic indicator of economic growth;
  - input to monetary policy and inflation targeting;
  - estimating housing as component of wealth;
  - financial stability / soundness indicator to measure risk exposure;
  - deflator in the national accounts;
  - input to consumer decision making on buy/sell;
  - input into the consumer price index for wage bargaining and indexation;
  - inter-area and international comparisons.
- Policy context:
  - The Second Phase of the G-20 Data Gaps Initiative and guidance on Financial Soundness Indicators identify residential property price changes as critical for financial stability and macroprudential analysis.
  - The IMF provides training and capacity building; compilers requested practical guidance complementing the Handbook.

### Data requirements, coverage, and source selection
- Recommended target coverage:
  - national monthly or quarterly index covering monetary transactions (mortgage and cash purchases) of all residential property types (single and multi-family dwellings, new and existing dwellings) in line with Handbook concepts.
  - If constraints exist, reduce coverage/frequency (e.g., urban markets only, registered transactions).
- Essential data elements:
  - transaction price corresponding to the date the transaction is finalized (transaction price preferred over asking, valuation, or declared prices).
  - detailed dwelling characteristics to allow constant-quality indices because dwellings are infrequently sold and heterogeneous.
- Common dwelling characteristics that affect price:
  - dwelling type, floor area, number of rooms, district, region, property type, age, vintage (new or existing), completion state, elevator, usable floor area, renovations, central heating, etc.
- Location and aggregation guidance:
  - Acquire granular location descriptors (province/state, city, sub-city areas) to reduce aggregation bias and unobserved sub-area trends.
  - Two limitations when granular location data are unavailable:
    - high- and low-value areas within a city are lumped into one category and sub-area trends may be missed;
    - different sub-areas may have different price trends and levels averaged into one trend.
- Typical data sources and typical price concepts:
  - Websites — asking prices; Availability Good; Timeliness Best (Within the day); Coverage Partial.
  - Agents / Developers — asking prices via surveys; Availability Good (costly); Coverage Partial.
  - Land registries / Tax authorities — declared prices; Availability Good; Coverage Complete; Timeliness Can be rather late.
  - Lending institutions — valuation prices or lending-record prices; Availability Must be negotiated; Coverage Excludes cash transactions.
  - Lawyers / Notaries — registered transaction prices if electronic; Availability Good if mandatory; Timeliness Good.
- Source selection recommendations:
  - Obtain metadata from administrative registers; adjust/filter when concepts differ.
  - Verify mortgage application prices reflect final selling price before using mortgage data.
  - For web-scraped data, verify listings represent sales rather than market intentions.

### Data management, validation, and confidentiality
- Data-exchange agreements should specify:
  - secure transmission procedures;
  - provision of metadata;
  - agreed schedule for transmission;
  - validation reports.
- Data management best practices:
  - Maintain a data management system to store, process, and back up files.
  - Keep raw, cleaned, and processed data archives for each compilation cycle to guarantee replicability.
- Validation checks on receipt:
  - Visual inspection for corrupted files, incorrect time scope, unexpected number of observations, unrecognizable characters, binary variables outside {0,1}, structural stability.
  - Produce standard validation reports automatically upon receipt.
- Currency handling:
  - Record prices in the official local currency; if other currencies present use daily exchange rates from the national central bank website to convert to local currency.

### File preparation and formatting (Box 4)
- Text variables: no spaces, accents, or misleading characters; use underscores (e.g., “floor_area”, “semi_detached”).
- Numbers formatted as “number”; qualitative variables as “text”.
- Do not use commas or dots to split thousands.
- Variable names must be identical across files (case-sensitive).
- Replace all missing values with “NA”.
- Include a date column plus quarter/month columns for sub-annual indices.

### Source data quality assessment (minimum dimensions)
- Assess: accuracy, fit for purpose, structure, coverage, availability.
  - Accuracy: errors, missing observations.
  - Fit for purpose: whether price reflects target concept; coverage adequacy; timeliness (monthly, quarterly, annually).
  - Structure: number of variables per observation.
  - Coverage: share of observations relative to total required.
  - Availability: reliable acquisition schedule and data format accessibility.
- Integration challenges to anticipate:
  - differing formats and variable definitions;
  - variables with same name but different meanings;
  - lack of common keys.

### Cleaning, outlier detection, and imputation
- Visual inspections and summary statistics to detect outliers and errors.
  - Example Price summary for Period 1Q2008:
    - Minimum: 31,400
    - q1: 188,600.00
    - Median: 249,200
    - Mean: 282,320
    - q3: 331,600
    - Maximum: 1,799,900
    - Na: 0
  - Example Floor Area summary for Period 1Q2008:
    - Minimum: 43
    - q1: 90.00
    - Median: 112
    - Mean: 120
    - q3: 139
    - maximum: 313
    - Na: 0
  - Example missing-values in synthetic dataset: 332 missing values for period 1Q2008 representing 12.39 percent of total observations.
- Outlier detection:
  - Define plausible ranges per variable/stratum using statistical techniques or expert judgment and test limits for at least one year.
  - Example Price per Area (P_area) for 1Q2008:
    - Minimum: 131
    - q1: 1,703.00
    - Median: 2,307
    - Mean: 2,488
    - q3: 3,066
    - Maximum: 16,564
  - Using amplification technique, 120 records identified as outliers for extreme P_area.
- Imputation guidance:
  - Numerical variables: impute with the average.
  - Qualitative variables: impute with the mode.
  - Use expert judgement where appropriate and recognize imputations introduce bias.

### Weights, stratification, and recommended practices
- Weight concepts:
  - Stock weights: positions in time (e.g., number of properties on a date).
  - Flow weights: reflect transactions over a period (one year recommended).
- Typical weight sources:
  - Stock weights: census data (may be outdated; often every 10 years).
  - Flow weights: administrative records (transaction or loan data); if sample-based (e.g., web scraping), complementary source needed unless representative.
- Practical weighting recommendations:
  - Derive weights from one year or average of years (three-year average general rule).
  - Update weights annually and keep constant during the year for sub-annual indices.
  - If values unavailable, use volumes (number of transactions, square feet) as substitutes.
- Example stratum weights for Year 2008:
  - Total_W = 1
  - W_New = 626111900 / 2631870500 = 0.24
  - W_Existing = 2005758600 / 2631870500 = 0.76
  - Conclusion: 24 percent of dwellings transacted during 2008 are new; 76 percent are existing.
- Example aggregated sums 2008:
  - SUM of Prices for New Dwellings during 2008 = 626,111,900
  - SUM of Prices for Existing Dwellings during 2008 = 2,005,758,600

### Methods for RPPI compilation — overview and guidance
- Methods illustrated:
  - Method 1: Stratification with simple averages (median with stratification).
  - Method 2: Hedonic methods — Time dummy, Imputations, Characteristics.
- Method selection guidance:
  - Median with stratification — use if prices and at least one characteristic (e.g., location) are available.
  - Time dummy hedonic — use if few observations per period or monthly index required.
  - Imputations hedonic — use with detailed data and geospatial data.
  - Characteristics hedonic — preferred with large number of transactions per period; equivalent to Imputations when log-linear and no geospatial data.
  - General rule: the lower the number of observations, the more appropriate the time dummy method with rolling window; larger datasets permit imputations or characteristics methods.
- Methods excluded:
  - Repeated sales method — excluded due to data intensity, potential bias from depreciation/renovations, sample bias, exclusion of first-time sales, and revision issues.
  - Sales Price Appraisal Ratio (SPAR) method — applicable only where properties are frequently reassessed.

### Stratification with median — mechanics and example
- Median preferred over mean due to positive skewness of price distributions.
- Typical strata: location, vintage (new vs existing), dwelling type.
- Avoid over-stratification; ensure enough observations per stratum.
- Worked example (vintage stratification 2008–2010):
  - RPPI quarterly indices (reference 1Q2008):
    - RPPI: 1Q2008 = 105.0; 2Q2008 = 103.3; 3Q2008 = 98.3; 4Q2008 = 93.4
    - New: 1Q2008 = 106.9; 2Q2008 = 101.7; 3Q2008 = 96.2; 4Q2008 = 95.2
    - Existing: 1Q2008 = 104.4; 2Q2008 = 103.8; 3Q2008 = 98.9; 4Q2008 = 92.9
  - Example median prices (reference 1Q2008):
    - Median New = 213900
    - Median Existing = 242300
  - Example sub-index calculation (2Q2008 vs 1Q2008):
    - New sub-index = 220600 / 231900 = 95.1 → New decreased 4.9 percent compared to 1Q2008.
    - Existing sub-index = 241000 / 242300 = 99.5 → Existing decreased 0.5 percent compared to 1Q2008.
  - Aggregation (Laspeyres) with weights New = 0.24, Existing = 0.76 produced quarterly RPPI indices:
    - 1Q2008 = 100.0; 2Q2008 = 98.4; 3Q2008 = 93.6; 4Q2008 = 89.0
  - Re-referencing to annual base (2008=100) used mean of quarterly indices (mean = 0.95) and produced RPPI re-referenced:
    - 1Q2008 = 105.0; 2Q2008 = 103.3; 3Q2008 = 98.3; 4Q2008 = 93.4
  - Production-year steps for transition to 2009 include reference price from last quarter previous year, calculation of sub-indices, aggregation, and chaining.

### Hedonic methods — rationale, forms, and diagnostics
- Hedonic regressions estimate marginal contributions (“shadow” prices) of characteristics and are used for quality adjustment.
- Log-linear hedonic form recommended:
  - reduces skewness and heteroscedasticity;
  - allows multiplicative relationships and parabolic variable effects.
- General log-linear specification:
  - ln p_nt = sum_{k=1}^K beta_k_t Z_nk_t + epsilon_nt
- Data depth guidance:
  - Ensure sufficient observations relative to variables; practice suggests around 20 observations per variable as a rule of thumb.
- Model selection:
  - Use ANOVA with AIC for model selection; keep selected model for at least one year; document any methodological change.
- Example diagnostic outcomes (1Q2008 final model):
  - Regression p-value very close to zero.
  - R-squared example: 54.88 percent of variability of ln(price) explained by model.

### Time Dummy Hedonic — mechanics, rolling window, and example
- Pool observations across several periods with time dummy variables; index derived from estimated time-dummy coefficients:
  - ln p_nt = beta_0 + sum_{t=1}^T delta_t D_nt + sum_{k=1}^K beta_k Z_nk_t + epsilon_nt
  - I_t = exp(delta_hat_t) * 100
- Rolling-window approach to limit revisions:
  - For quarterly indices use current quarter + previous three quarters (12 months total) and keep shadow prices fixed for at least one year.
  - Chain indices across overlapping windows using last overlap period.
- Example OLS selected coefficients for New and Existing dwellings (selected rows shown; Std.Error and t.value reported in Guide):
  - New: Period2Q2008 Estimate -0.06; Period3Q2008 Estimate -0.10; Period4Q2008 Estimate -0.13
  - Existing: Period2Q2008 Estimate -0.04; Period3Q2008 Estimate -0.05; Period4Q2008 Estimate -0.10
- Example sub-indices from time-dummy coefficients:
  - New: 1Q2008 100.0; 2Q2008 exp(-0.06) * 100 = 94.2; 3Q2008 exp(-0.10) * 100 = 90.8; 4Q2008 exp(-0.13) * 100 = 88.2
  - Existing: 1Q2008 100.0; 2Q2008 exp(-0.04) * 100 = 96.5; 3Q2008 exp(-0.05) * 100 = 95.2; 4Q2008 exp(-0.10) * 100 = 90.3
- Aggregation (Laspeyres) and re-referencing procedures mirror other methods; example re-referenced RPPI series given in Guide.

### Characteristics Hedonic — mechanics and example
- Constructs a “typical” dwelling by averaging numerical characteristics and frequency for categorical variables, then computes index comparing price of typical property across periods using estimated shadow prices.
- Example typical-property averages (selected):
  - Dwelling_TypeApartment: 0.17
  - Dwelling_TypeDetached: 0.17
  - Dwelling_TypeSemi_Detached: 0.67
  - Floor_Area: 111.17
  - Neighborhood_Affluence: 27.77
  - Neighborhood_TypeUrban: 0.50
- Index formula (log-linear):
  - I_0t = exp( Σ_k (β̂_k_t − β̂_k_0) Z̅_k_0 )
- Worked example results:
  - New Dwellings 2Q2008 sub-index = 94.5 → 5.5 percent decrease since 1Q2008.
  - Aggregated RPPI 2Q2008 = 96.1 (Laspeyres aggregation with weights).

### Imputations Hedonic — mechanics and example
- Imputation method imputes prices for each property for each period and aggregates imputed prices (geometric mean) to form indices; suitable with detailed characteristics and geospatial data.
- Imputed price formulas (log-linear):
  - p̂_l0(0) = exp(β̂_0_0) exp( Σ_n β̂_n_0 z_l n 0 )
  - p̂_l t (0) = exp(β̂_0_t) exp( Σ_n β̂_n_t z_l n 0 )
- Index from imputed prices using geometric means (Jevons):
  - I_t = [ ∏_{n∈S(0)} (p̂_n t (0))^{1/N(0)} ] / [ ∏_{n∈S(0)} (p̂_n 0)^{1/N(0)} ]
- Worked example highlights:
  - Jevons of imputed prices for 2008 (selected):
    - New: 1Q2008 233573; 2Q2008 220722; 3Q2008 213389; 4Q2008 206634
    - Existing: 1Q2008 239609; 2Q2008 231316; 3Q2008 227946; 4Q2008 216885
  - New sub-indices (Table 49):
    - 1Q2008 100; 2Q2008 94.5; 3Q2008 91.4; 4Q2008 92.9
  - Aggregated RPPI 4Q2008 = 91.1 (Laspeyres aggregation).
  - Re-referencing so 2008 = 100 example:
    - RPPI re-referenced 1Q2008 = 100 / 0.95 = 104.9; 2Q2008 = 96.1 / 0.95 = 100.7; 3Q2008 = 94.2 / 0.95 = 98.8; 4Q2008 = 91.1 / 0.95 = 95.5

### Comparison of methods — numeric series and interpretation
- Key findings from comparative series (2008–2015) across four methods:
  - Characteristics and Imputations methods provide similar results.
  - Time dummy method is less volatile because of the rolling four-quarter window smoothing.
  - Stratification (median) is the most volatile due to limited within-stratum quality adjustment.
- Representative RPPI index values (selected ranges from Guide):
  - Stratification (1Q2008–4Q2011): 105.0, 103.3, 98.3, 93.4, 88.2, 86.3, 81.3, 76.9, 70.7, 71.5, 71.7, 70.0, 60.7, 64.6, 67.4, 67.9
  - Time dummy (1Q2008–4Q2011): 105.2, 101.0, 99.1, 94.7, 89.2, 85.7, 80.8, 76.1, 71.7, 70.8, 71.0, 70.1, 66.5, 68.0, 70.0, 70.5
  - Characteristics (1Q2008–4Q2011): 105.2, 101.0, 99.1, 94.7, 89.2, 85.7, 81.2, 76.5, 72.5, 71.2, 71.4, 70.6, 67.0, 68.8, 70.7, 71.6
  - Imputations (1Q2008–4Q2011): 105.2, 101.0, 99.1, 94.7, 89.2, 85.7, 81.2, 76.5, 72.5, 71.2, 71.4, 70.6, 67.0, 68.8, 70.7, 71.6
  - Stratification (1Q2012–4Q2015): 64.1, 67.6, 72.7, 70.1, 70.3, 75.7, 80.5, 80.4, 80.8, 81.6, 87.7, 86.9, 91.2, 89.6, 98.2, 97.5
  - Time dummy (1Q2012–4Q2015): 70.0, 74.3, 77.5, 78.0, 78.6, 83.0, 85.7, 86.1, 87.5, 89.7, 93.1, 94.3, 97.0, 97.4, 104.4, 105.6
  - Characteristics (1Q2012–4Q2015): 71.5, 76.2, 79.0, 79.4, 80.0, 84.4, 87.3, 87.4, 88.9, 91.1, 94.5, 95.9, 98.7, 99.0, 106.1, 107.7
  - Imputations (1Q2012–4Q2015): 71.5, 76.2, 79.0, 79.4, 80.0, 84.4, 87.3, 87.4, 88.9, 91.1, 94.5, 95.9, 98.7, 99.0, 106.1, 107.7

### Recommendations for compilation, testing, and dissemination
- Method selection:
  - Hedonic regression–based methods are generally considered superior when data permit.
  - Use median with stratification if hedonic data unavailable.
- Experimental testing:
  - Develop experimental indices for at least one year to vet accuracy, relevance, timeliness, and business processes before official release.
- Dissemination timing and practice:
  - Release quarterly data within 90 days after the reference period.
  - Make data available in machine-readable form (e.g., .csv) and publish according to an advance release calendar.
  - Publish indices with two decimals and rates with one decimal.
  - Accompany releases with a short, clear technical note describing aim, coverage, stratification, weighting system, reference periods, and methods used.
  - Organize stakeholder briefings and consider social media/videos to explain interpretation.
- Publication and visualization guidance:
  - Provide time-series graphs of total RPPI and sub-indices with reference period clearly indicated.
  - Graph contributions of strata to annual change (YoY), comparisons with number of dwellings sold, and benchmarking with CPI.
  - Encourage maps, heat maps, and interactive graphics for public interest (e.g., price-per-square-meter heat maps).
- Complementary indicators to publish:
  - Sales counts, regional price developments, price-per-square-meter, house price-to-income ratios, city-level affordability measures.
- Operational planning:
  - Plan dissemination prior to index development considering user needs and visualization advances.
  - Embed RPPI in a broader residential property information framework including prices, sales, stocks, quality, and affordability.

### Practical tools and software
- R (open-source, R-Studio) recommended:
  - Hardware/software requirements similar to STATA, SAS, or E-views.
  - R scripts accompanying the Guide allow compilers to apply methods to their own data for monthly or quarterly indices with adaptations; guidance provided within scripts.

*Source: RPPI Practical Guide, International Monetary Fund.*

### INTRODUCTION _________________________________________________________________________________ 6

### rppi-guide - INTRODUCTION _________________________________________________________________________________ 6

### Overview and Purpose
- The Guide sets out practical advice on the compilation of a Residential Property Price Index (RPPI) based on the conceptual approach described in the Handbook on Residential Property Price Indices (Handbook), published in 2013.
- The Guide was prepared by the IMF Statistics Department, building on experience gathered during technical assistance and training missions delivered to RPPI compilers from more than 80 countries over the past six years.
- The Guide is written for data compilers in statistical offices, central banks, housing agencies, or other statistical agencies and demonstrates how to compile and disseminate RPPIs.
- Using the Guide requires a background in statistics and some knowledge of econometrics.
- The Guide includes several practical exercises delivered through step-by-step instructions that make use of a synthetic data set built by the Central Statistics Office Ireland (CSO) adapted by the IMF and R scripts developed specifically for the RPPI compilation.

### Key Uses of RPPIs (as noted in the Handbook)
- as a macro-economic indicator of economic growth;
- for use in monetary policy and inflation targeting;
- as an input into estimating the value of housing as a component of wealth;
- as a financial stability or soundness indicator to measure risk exposure;
- as a deflator in the national accounts;
- as an input into an individual citizen’s decision making on whether to buy (or sell) a residential property;
- as an input into the consumer price index, which in turn is used for wage bargaining and indexation purposes;
- for use in making inter-area and international comparisons.

### Context and Demand
- The Second Phase of the G-20 Data Gaps Initiative and guidance on Financial Soundness Indicators identify real estate statistics, in particular price changes on residential property, as a critical input into financial stability policy analysis and macroprudential measures.
- Data users in IMF member countries have expressed a need for RPPIs and the IMF has provided training and capacity building to many member countries.
- Compilers have expressed the need for practical guidance to complement the theoretical content of the Handbook, particularly when transitioning from less to more sophisticated methods.

### Target Audience and Capacity Focus
- The Guide targets organizations with capacity constraints (lack of econometricians, programmers, or IT support) compiling new RPPIs and aims to help developing countries close data gaps and improve data quality.
- Countries already compiling an RPPI can also use the Guide to improve methodologies or processes.

### Structure of the Guide
- The Guide is divided into three sections:
  - First section: illustrates the type of data commonly used to construct RPPIs, how to evaluate the quality of source data, how to prepare or pre-process data prior to compiling an RPPI, and the importance of selecting appropriate weights and the different type of data that can be used to construct the weights.
  - Second section: outlines the different methods that can be used to compile RPPIs.
  - Third section: provides guidance on dissemination practices and suggests additional indicators compilers may want to construct to complement the RPPI.

### Synthetic House Price Dataset and Access Conditions
- The Synthetic House Price Dataset was generated by CSO Ireland by applying several statistical masking techniques to create a dataset which is anonymous and materially different from the Irish RPPI raw data, but which preserves the aggregate correlations between the variables which affect house prices.
- The techniques used in constructing the synthetic data prevent the identification of individual transactions or dwellings. As the Synthetic House Price Dataset is materially different to the Irish RPPI dataset, it cannot be used to replicate the Irish RPPI results.
- The IMF will provide the Synthetic House Price Dataset, upon request to compilers, researchers and students who agree to the following conditions of use:
  - The Synthetic House Price Dataset will only be used for training and educational purposes.
  - The CSO Ireland is to be acknowledged as the source of the Synthetic House Price Dataset.
  - In all coursework material, training workshops and related reports the dataset is referred to as the Synthetic House Price Dataset, or simply as a synthetic dataset, not as Ireland’s Residential Property Price Index dataset.
  - Acknowledgement that the CSO Ireland is not liable in any way for the usage of the Synthetic House Price Dataset. All analyses, inferences and interpretations of the data and all consequences thereof are entirely the responsibility of the data user.
  - Neither the Synthetic House Price Dataset, full or partial copies of the dataset, are to be made available to any third party.
- Requests for access to the synthetic dataset should be emailed to the Division Chief of the Real Sector Division of the Statistics Department (rppi@imf.org) of the IMF and should include information on the reason for access (compilation, research, study), organization details (if any) and other relevant information. A register of users who receive the Synthetic House Price Dataset will be maintained by the IMF.

### Practical Tools: R Software (Box 1)
- R is a programming language and free open-source (R-Studio) software environment for statistical uses, with a large online community that supports its availability and reliability and provides technical support.
- The hardware and software requirements are the same as for other econometric software such as STATA, SAS or E-views.
- The R software is open which means that any person can use it free of charge.
- The R scripts accompanying the Guide have been prepared such that compilers should be able to use them with their own data to compile a monthly or quarterly index, with a few and easy-to-perform adaptations.
- Specific guidance on how to make these adaptations is provided in each R script.

### Guidance on Target Coverage and Data Requirements
- When initially designing an RPPI program, compilers should target a national monthly or quarterly index covering monetary transactions (mortgage and cash purchases) of all types of residential properties (single and multi-family dwellings, new and existing dwellings) in line with the concepts outlined in the Handbook.
- In cases where available data, operation of the market or other constraints prevent complete alignment with the Handbook, compilers may need to reduce the coverage and/or frequency of the index (for example, an index that only covers urban markets or registered transactions).
- To compile an RPPI conforming to the recommended best practices data related to the transaction price as well as key characteristics of the dwellings are required to assure a constant-quality index.
- Given infrequent sale (the same residential property may only be sold once every 20 years) and heterogeneity of residential properties, quality adjustment techniques are required to derive measures of pure price change; this implies extensive data requirements and detailed information about each property to follow the price trend by comparing “like-with-like”.

*Source: RPPI Practical Guide, International Monetary Fund.*

### 16.      The most commonly available characteristics are the dwelling type, floor area, number of

### Key characteristics, price concepts, data sources, and data-quality guidance for compiling a Residential Property Price Index (RPPI)

### Key dwelling characteristics that affect price
- Commonly available characteristics: dwelling type, floor area, number of rooms, district, region, property type, age, vintage (new or existing), completion state of dwelling, existence of an elevator, etc.
- Other structure/location attributes (key characteristics): usable floor area, number of rooms, district, region, property type, age, new or existing, completion state, renovations, etc.
- Market variation in relevance:
  - Heating is a key variable in northern European countries but irrelevant in Africa.
  - Size is important in all markets, but measurement is multi-dimensional (number of rooms, floor area, site/lot size for houses).
- Recommendation: Acquire granular location descriptors (province/state, city, sub-city areas) where possible; greater granularity reduces risks of aggregation bias and unobserved sub-area trends.

### Location and aggregation risks
- Location is the most important characteristic and should be described with categorical variables for province (state), city, and sub-city areas.
- Two limitations when granular location data are not available:
  - High- and low-value areas within a city are lumped into one category, so systematic variation by sub-area over time may be missed (e.g., if low-value areas trade more over time, this may be measured as a decline in prices).
  - Different sub-areas may have different price trends and levels that are averaged into one trend when aggregated.
- Advanced analysis can include location-specific amenities such as access to transport facilities.

### Transaction timing and price concepts (Box 2)
- The same dwelling can have multiple prices at distinct points in time (asking price, transaction price, declared price). The RPPI compiler should ideally use the agreed-upon price corresponding to the date the transaction is finalized (the transaction price).
- Price concepts and fitness for purpose:
  1. Asking price – Price asked when the dwelling is put on the market; may change; often higher than target price but can be a suitable proxy to capture changes.
  2. Transaction price – Price agreed after negotiations; the target price (price actually paid). Can be difficult to obtain because it may not be in a formal register at time of agreement.
  3. Valuation price – Appraisal by the lending institution; can be higher, lower or equal to the transaction price. Suitable only if it equals the transaction price.
  4. Declared price – “Legal” price used for tax purposes; registration normally happens later than transaction date and may differ from transaction price. Suitability depends on the size of these differences.

### Data sources: features, typical price concepts, advantages and limitations
- Common data sources and their typical associated price concepts:
  - Websites — typically provide asking prices; wide coverage; detailed descriptions; timeliness advantage (price as soon as property enters market).
  - Agents and developers — often provide asking prices; can be obtained via surveys.
  - Lawyers, notaries, land registers, tax authorities — often provide declared prices (legal/tax records).
  - Lending institutions — provide valuation prices and may provide declared or transaction prices via lending records.
- Advantages and drawbacks:
  - Websites:
    - Advantages: wide coverage, detailed descriptions, asking price, timeliness (within the day).
    - Drawbacks: listings can be confined to higher-end properties; asking price often differs from final transaction price; delisting does not guarantee a sale.
  - Surveys of agents/developers:
    - Advantages: compiler control over questionnaire; can obtain precise information.
    - Drawbacks: costly, response burden, difficulty obtaining representative samples.
  - Land registries / property assessment files:
    - Advantages: often complete coverage and good quality; used for legal and tax purposes; low cost and low respondent burden.
    - Drawbacks: registration may occur three months or more after the transaction (introducing time bias); definitions/classifications may differ from statistical needs; may lack detailed property characteristics for hedonic methods.
  - Mortgage data:
    - Advantages: timely, broad geographic coverage, rich dwelling characteristics.
    - Drawbacks: cash transactions excluded; weak coverage in less developed countries; mortgage application values may reflect an upper borrowing limit rather than final selling price; banking sector fragmentation may require standardization.
  - Lawyers/notaries:
    - Advantages: contracts registered after sale can give detailed location and actual transaction price; notaries are normally regulated and required to record sales in a common database.
    - Drawbacks: where lawyers are not required to register electronically, compiling complete electronic files is difficult and costly; client confidentiality restrictions can limit access; databases may lack detailed dwelling characteristics.
- Table 1 comparative features summarized (qualitative):
  - Web sites: Availability Good. Immediate and free. Coverage Partial. Timeliness Best (Within the day). Detail for Quality Adjustment Must be checked on each website. Limitations: Asking prices often differ from transaction prices.
  - Agents: Availability Good (requires survey therefore costly). Coverage Partial. Timeliness Good. Detail for Quality Adjustment Good. Limitations: Asking prices often differ from transaction prices.
  - Developers: Similar profile as Agents.
  - Lending institutions: Availability Must be negotiated. Coverage Excludes cash transactions. Timeliness Good. Detail Good. Limitations: Value can differ from transaction price.
  - Lawyers or notaries with electronic data set: Availability Good. Coverage Good if mandatory or very common. Timeliness Good. Detail Must be checked. Limitations: May suffer from underreporting of values.
  - Lawyers or notaries without electronic data set: Availability Costly. Coverage Sample. Timeliness Good. Detail Must be checked. Limitations: May suffer from underreporting of values.
  - Land register and Tax authorities: Availability Good. Coverage Complete. Timeliness Can be rather late. Detail Must be checked. Limitations: May suffer from underreporting of values.

### Designing the RPPI: data selection and process guidance
- First step: Acquire an in-depth understanding of the real estate transaction process within the country to identify which information is generated and which data are useful for RPPI compilation.
- Aim to acquire the transaction price corresponding to the date the transaction is finalized.
- Recommended actions when selecting sources:
  - Obtain metadata describing source concepts and definitions when using administrative registers.
  - When concepts differ from RPPI needs, perform necessary adjustments or filtering.
  - For mortgage data, verify the mortgage application price reflects the final selling price before use.
  - For web-scraped or API-provided data, verify whether listings represent actual sales or merely market intentions.

### Assessing and ensuring source data quality
- Minimum assessment dimensions (if no country-specific quality framework exists): accuracy, fit for purpose, structure, coverage, availability.
  - Accuracy: determine occurrence of errors and missing observations.
  - Fit for purpose: whether price reflects target price concept; coverage adequacy (national vs local); timeliness and frequency requirements (monthly, quarterly, annually).
  - Structure: number of variables/attributes per observation (price + city vs price + exact location + size + type + age).
  - Coverage: share of observations included in source as a share of total required; regulated markets tend to have higher coverage; unregulated markets may have low coverage.
  - Availability: reliability of acquisition schedule; data format accessibility (e.g., paper vs digital; varied regional systems requiring standardization).
- Examples of specific integration challenges:
  - dealing with different formatting;
  - varied variables;
  - variables with the same name but diverse meanings;
  - lack of a common key to integrate the information.

### Data-exchange agreements, data management, confidentiality, and validation
- Enter into a formal data exchange agreement with data providers including:
  - procedures for secure transmission of sensitive data;
  - provision of metadata;
  - an agreed-upon schedule for transmission;
  - validation reports to ensure correct transmission.
- Data management best practices:
  - Ensure a data management system to store, process, and back up files.
  - Keep track of files used in each RPPI calculation to guarantee replicability.
  - Maintain and archive raw data, cleaned data, and processed data for each compilation cycle.
- Confidentiality considerations:
  - Source data may be sensitive and access to micro-data may be limited.
  - Establish a memorandum of understanding to govern information exchange and responsibilities for protecting confidentiality.
- Validation practices on receipt of data:
  - Develop standard validation reports generated upon receipt.
  - Perform a brief visual inspection to detect severe errors such as:
    - an empty or corrupted file;
    - incorrect time scope of the data set;
    - an unexpected number of observations;
    - a file with unrecognizable characters;
    - binary variables with numbers other than zero or one;
    - confirm that the structure of the data remains stable.
- Currency handling:
  - Prices should usually be recorded in the official local currency.
  - If original data show other currencies, use the daily exchange rates from the national central bank website to convert the transaction value into the local currency.

*Source: RPPI Practical Guide, International Monetary Fund.*

### Box 4. Preparing Data for Processing

### Box 4. Preparing Data for Processing

### File preparation and formatting guidelines
- Text variables and qualitative occurrences should not have spaces, accents, or misleading characters like minuses (-). Use underscores (e.g., “floor_area”, “semi_detached”).
- Numbers must be formatted as “number” and qualitative variables as “text.”
- Commas or dots to split the thousands should not be used.
- Variable names must be identical across files (e.g., “detached_house” ≠ “Detached_house”; “Floor_Area” ≠ “FloorArea”).
- Replace all missing values with “NA” since software often does not read empty cells correctly.
- Include a date column and also add a quarter (for quarterly indices) or month (for monthly indices) column, because subsequent files may include transactions that refer to previous periods.

*The RPPI Practical Guide recommends software-specific preparations while offering the above general advice.*

### Exercise 1: Assessing source data quality — workflow and checks
- Visual inspection of the raw data; identify problem cases requiring further investigation.
- Check size of the file: compare size, variables, and number of observations with previous files; variables should always be the same.
- Analyze summary statistics for economic reasonability and to detect outliers.
  - Example summary statistics for Price for Period 1Q2008 (Table 3):
    - Minimum: 31,400
    - q1: 188,600.00
    - Median: 249,200
    - Mean: 282,320
    - q3: 331,600
    - Maximum: 1,799,900
    - Na: 0
  - Example summary statistics for Floor Area for Period 1Q2008 (Table 4):
    - Minimum: 43
    - q1: 90.00
    - Median: 112
    - Mean: 120
    - q3: 139
    - maximum: 313
    - Na: 0
- Verify the date field is properly formatted and in the expected range (e.g., if expecting first quarter 2008 but finding fourth quarter 2008 observations, investigate).
- Produce frequency counts for categorical variables to assess representativeness and guide grouping/stratification.
  - Example frequency for BER in 1Q2008 (Table 5):
    - A: 15
    - B: 547
    - C: 1,010
    - D: 367
    - E: 257
    - F: 314
  - Example aggregation for BER in 1Q2008 (Table 6):
    - A+B+C: 1322
    - D+E+F: 747
- Analyze coverage (geographic, dwelling types, financing). Example coverage for Neighborhood_Type in 1Q2008 (Table 7):
  - Rural: 466
  - Urban: 2,055
- Check for duplicated records (e.g., same selling date and address). Synthetic dataset accompanying the Guide does not contain duplicated records.
- Check for missing values and compute share of missing observations.
  - Example: synthetic data set contains 332 missing values for period 1Q2008, representing 12.39 percent of total observations.
  - Example missing-value extract (Table 8) shows specific records with NA in Dwelling_Type or Building_Levels.
- Identify errors vs outliers; establish automated error and outlier detection routines to run on each new file.

### Outlier detection and handling
- Define plausible ranges per variable/stratum using statistical techniques or expert judgment; test different limits and use chosen limits for at least one year.
- Statistical rule example using interquartile amplification (from Table 9, Price per Area for 1Q2008):
  - Minimum: 131
  - q1: 1,703.00
  - Median: 2,307
  - Mean: 2,488
  - q3: 3,066
  - Maximum: 16,564
- Using the amplification technique, 120 records were identified as outliers due to extreme price per floor area (P_area).
- Example outlier records (excerpt, Table 10) include P_area values:
  - 4,735.82; 4,621.50; 5,172.73; 5,336.95; 5,720; 6,179.55
- Visualization: histograms of P_area and by vintage (New, Existing) help identify long right tails before cleaning and show reduced tails after deleting duplicates, missing, and outliers.
- Additional threshold examples (adapt to country market):
  - Table 11: Allowable number of bedrooms by property type (e.g., Detached House between 2 and 6; Apartment/Flat between 1 and 4).
  - Table 12: Allowable floor area (meters squared) by property type and number of bedrooms (explicit ranges per cell).
  - Table 13: Allowable number of rooms by property type and bedrooms (explicit ranges per cell).

### Imputation guidance
- If imputations are required (especially for matched-model indices), suggested options:
  - Impute numerical values with the average.
  - Impute qualitative variables with the mode.
  - Use expert knowledge to determine the most likely value.
- Note: imputations introduce bias; compilers should seek as much observed data as possible.

### Weights — concepts and examples
- Weights aggregate prices from primary levels (strata) into meaningful aggregates. Strata group like items together.
- Weights choice and quality significantly affect index results.
- Two types of weights:
  - Stock weights: positions in time (e.g., number of properties in a geographic region on a specific date).
  - Flow weights: reflect transactions over a period (one year recommended).
- Typical sources:
  - Stock weights often derived from census data; challenges include lack of detail and outdated frequency (often every 10 years).
  - Flow weights generally from administrative records (real estate transactions, loan data). When only a sample is available (e.g., web scraping), a complementary data source is needed to estimate weights unless sample is representative.
- Practical recommendations:
  - Weights can be derived from one year of data or an average of years (three is the general rule) for robustness.
  - Update weights annually and keep them constant during the year when constructing sub-annual indices.
  - If values are unavailable for weights, volumes (e.g., number of transactions, square feet) are an adequate substitute.
- Example weight aggregation from synthetic data (after cleaning):
  - SUM of Prices for New Dwellings during 2008 = 626,111,900
  - SUM of Prices for Existing Dwellings during 2008 = 2,005,758,600

*RPPI Practical Guide — Box 4, Preparing Data for Processing*

### 68.      Second, the stratum weights are derived by dividing the total value of all properties in

### RPPI Practical Guide — Methods for Weights, Stratification, and Hedonic Approaches

### Stratum weights and composition (Year 2008)
- The RPPI Practical Guide derives stratum weights by dividing the total value of all properties in the stratum by the total value of all properties.
- Weights for 2008:
  - Total_W = 1
  - W_New = 626111900 / 2631870500 = 0.24
  - W_Existing = 2005758600 / 2631870500 = 0.76
- The Guide concludes that 24 percent of the dwellings transacted during 2008 are new dwellings and 76 percent are existing dwellings.

### Overview of RPPI compilation methods
- The Guide illustrates two commonly used methods:
  - (1) Stratification with simple averages (median with stratification)
  - (2) Hedonic methods (three variants: Time dummy, Imputations, Characteristics)
- Method selection guidance (Table 17 summary):
  - Median with stratification — used if prices and at least one characteristic (e.g., location) are available.
  - Time dummy hedonic method — used if few price observations per period are available or a monthly index is required.
  - Imputations hedonic method — used if more detailed data are available; can be used with geospatial data.
  - Characteristics hedonic method — used if more detailed data are available; equivalent to Imputations when the model is log-linear and geospatial data are not used.
- General rule: the lower the number of observations, the more appropriate the time dummy method with a rolling window; the larger the dataset, the more feasible the imputations or characteristics methods.

### Methods excluded from coverage and reasons
- The Guide does not address:
  - Repeated sales method (e.g., Case-Shiller Index)
    - Reasons: very data intensive (particularly historical data), difficult to implement in many countries; potential bias if depreciation and renovations are not considered; sample bias risk; excludes dwellings sold for the first time; published index values are subject to constant revisions as new linked-sale data become available.
  - Sales Price Appraisal Ratio (SPAR) method
    - Applicable only where properties are reassessed frequently (normally for tax purposes).

### Method 1 — Median with Stratification: rationale and guidance
- Rationale:
  - Comparing average prices across periods is straightforward but discouraged due to extreme heterogeneity of properties.
  - Given positive skewness of house price distributions, the median is preferred over the mean for estimating the price of the “typical” property.
- Stratification guidance:
  - Use a small set of observable characteristics (e.g., location, vintage new/existing, type) to form strata.
  - Avoid over-stratification that produces more strata than data points; ensure enough observations per stratum for robust estimates.
  - Location is generally the first level of stratification; vintage (new vs. existing) is a common secondary stratum.
  - Strata produce sub-indices that have economic meaning and aid validation.
- Distinction: strata (used in stratification) differ from regression dimensions (used in hedonic methods); regression can exploit many characteristics even when strata would be too narrow.

### Worked example — Stratification by vintage (New vs. Existing) for 2008–2010
- Index structure (Table 18):
  - RPPI: 1Q2008 = 105.0; 2Q2008 = 103.3; 3Q2008 = 98.3; 4Q2008 = 93.4
  - New: 1Q2008 = 106.9; 2Q2008 = 101.7; 3Q2008 = 96.2; 4Q2008 = 95.2
  - Existing: 1Q2008 = 104.4; 2Q2008 = 103.8; 3Q2008 = 98.9; 4Q2008 = 92.9
- Example median prices (reference period 1Q2008, Table 19):
  - Median New = 213900
  - Median Existing = 242300
- Example median prices (current period 2Q2008, Table 20):
  - Median New = 220600
  - Median Existing = 241000
- Sub-index calculation (2Q2008 vs 1Q2008, Table 21):
  - New sub-index = 220600 / 231900 = 95.1
  - Existing sub-index = 241000 / 242300 = 99.5
  - Interpretation in example: New dwellings decreased 4.9 percent compared to 1Q2008; Existing dwellings decreased 0.5 percent compared to 1Q2008.
- Aggregation (Laspeyres) using base year weights (Table 22):
  - Weights: New = 0.24; Existing = 0.76; RPPI total weight = 1.0
  - Aggregated RPPI quarterly indices: 1Q2008 = 100.0; 2Q2008 = 98.4; 3Q2008 = 93.6; 4Q2008 = 89.0
- Re-referencing sub-indices to annual base (2008=100) (Table 23):
  - RPPI re-referenced:
    - 1Q2008: 100 / 0.95 = 105.0
    - 2Q2008: 98.4 / 0.95 = 103.3
    - 3Q2008: 93.6 / 0.95 = 98.3
    - 4Q2008: 89.0 / 0.95 = 93.4
    - (Mean of quarterly RPPI indices = 0.95 when expressed relative to annual base)
  - New re-referenced:
    - 1Q2008: 100 / 0.94 = 106.9
    - 2Q2008: 95.1 / 0.94 = 101.7
    - 3Q2008: 90.0 / 0.94 = 96.2
    - 4Q2008: 89.0 / 0.94 = 95.2
    - (Mean of quarterly New indices = 0.94)
  - Existing re-referenced:
    - 1Q2008: 100 / 0.96 = 104.4
    - 2Q2008: 99.5 / 0.96 = 103.8
    - 3Q2008: 94.7 / 0.96 = 98.9
    - 4Q2008: 88.9 / 0.96 = 92.9
    - (Mean of quarterly Existing indices = 0.96)
- Production-year compilation steps (example transitioning from 2008 to 2009):
  - Step 1 — Reference price: median of last quarter of previous year (4Q2008):
    - New = 206500
    - Existing = 215500
  - Step 2 — Median for current period (1Q2009):
    - New = 195000
    - Existing = 203350
  - Step 3 — Sub-indices (1Q2009 compared to 4Q2008) (Table 26):
    - New = 195000 / 206500 * 100 = 94.4
    - Existing = 203350 / 215500 * 100 = 94.4
    - Interpretation: both New and Existing decreased 5.6 percent since 4Q2008.
  - Step 4 — Aggregation (Laspeyres) using base year weights (Table 27):
    - RPPI 1Q2009 = 94.4 (Weights: New 0.24; Existing 0.76; total 1.0)
  - Step 5 — Chaining the indices:
    - Chain current-period indices by the chained index of the last period of the previous year to maintain the published annual base.
    - Example chaining results (Table 28):
      - Unchained indices: 1Q2008 = 105.0; 2Q2008 = 103.3; 3Q2008 = 98.3; 4Q2008 = 93.4
      - Chained indices example:
        - 1Q2009 unchained = 94.8; chained = 94.8 * 93.4 / 100 = 88.2
        - 2Q2009 unchained = 92.4; chained = 92.4 * 93.4 / 100 = 86.3
        - 3Q2009 unchained = 87.1; chained = 87.1 * 93.4 / 100 = 81.3
        - 4Q2009 unchained = 82.3; chained = 82.3 * 93.4 / 100 = 76.9
        - 1Q2010 unchained = 91.9; chained = 91.9 * 76.9 / 100 = 70.7
        - 2Q2010 unchained = 93.0; chained = 93.0 * 76.9 / 100 = 71.5
        - 3Q2010 unchained = 93.3; chained = 93.3 * 76.9 / 100 = 71.7
        - 4Q2010 unchained = 91.0; chained = 91.0 * 76.9 / 100 = 70.0
- Notes and practical considerations:
  - Stratification controls for quality differences between strata but not within strata.
  - Hedonic regressions within a stratum can control for within-stratum quality-mix changes and generate sub-indices that stratified medians cannot.

### Hedonic methods (overview)
- The Guide presents hedonic methods as preferred by most compilers when data permit.
- Hedonic variants:
  - Time dummy
  - Imputations
  - Characteristics
- Use cases:
  - Time dummy hedonic method is suitable for small samples per period or when monthly indices are required (often used with a rolling window).
  - Imputations and Characteristics methods require more detailed data and can exploit geospatial information (Imputations can be used with geospatial data; Characteristics is equivalent to Imputations when log-linear and without geospatial data).

*Source: RPPI Practical Guide, INTERNATIONAL MONETARY FUND*

### 96.      To measure pure price changes, it is essential to measure the same product over time,

### rppi-guide - 96.      To measure pure price changes, it is essential to measure the same product over time,

### Hedonic regression: concept and purpose
- Hedonic regressions treat a dwelling as a bundle of characteristics whose marginal contributions to price (so-called “shadow” prices) can be estimated even though the characteristics are not sold separately.
- Quality adjustments for RPPIs require detailed information about property characteristics and can be made using the hedonic method.
- Quoted definition (Handbook): “The hedonic regression method recognizes that heterogeneous goods can be described by their attributes or characteristics. ... Regression techniques can be used to estimate those marginal contributions or shadow prices.”

### Functional form and rationale
- The Guide uses the log-linear hedonic form because:
  - The distribution of real estate prices is usually positively skewed.
  - The log-linear form reduces the impact of skewness and heteroscedasticity.
  - The log-linear form allows for parabolic relationships of variables and a multiplicative association between characteristics.

### Log-linear hedonic specification
- General log-linear specification presented:
  - ln p_nt = sum_{k=1}^K beta_k_t Z_nk_t + epsilon_nt
  - Notation:
    - t - period
    - n – number of dwellings in period t
    - k – each of K characteristics
    - ln p_nt – price logarithm
    - beta_k_t – “shadow” price of characteristic k in period t
    - Z_nk_t – the value/quantity of characteristic k in dwelling n in period t
    - epsilon_nt – error term

### Simple illustrative example: decomposing prices
- Two neighboring houses sold in consecutive periods:
  - House A sold for 465,000 (period t).
  - House B sold for 400,000 (period t+1).
- Characteristics and estimated “shadow” prices (Table 29 example):
  - Size: 350,000 (House A), 350,000 (House B)
  - Bathrooms: 45,000 (House A), 30,000 (House B)
  - Garage: 20,000 (House A), 20,000 (House B)
  - Balcony: 10,000 (House A), (none for House B)
  - Total implied shadow-price decomposition: 425,000 (House A), 400,000 (House B)
- Key point: naïve comparison of 465,000 vs 400,000 would misrepresent price change (apparent fall of 16 percent) because characteristics differ; hedonic regression decomposes prices into component shadow prices to allow correct pure-price measurement.

### Data depth and degrees of freedom
- Reliable coefficient estimation requires sufficient observations relative to variables:
  - Example caution: 20 observations and 25 variables is logically and statistically impossible to estimate each effect.
  - Established practice and statistical theory suggest a minimum of around 20 observations per variable to enable the regression to estimate the effect of a variable (this level can vary by period).

### Model selection and implementation
- The R-script “fit model” is used to find the best regression model; the selected model should be kept for at least one year.
- Any change in the model is a fundamental methodological change that must be documented and communicated to users.
- Model selection uses ANOVA with the AIC (Akaike criterion); the lower the AIC value the better the fit.
- Example Final Model for 1Q2008 (selected variables):
  - log(Price) ~ Dwelling_TypeApartment + Dwelling_TypeDetached + Dwelling_TypeSemi_Detached + Floor_Area + Neighborhood_Affluence + BERHigh + Neighborhood_TypeRural + Central_HeatingMains_Gas + Year_Built_Agg1950_1979 + County_Agg11_14 + County_Agg22_31 + County_Agg42_52 + County_Agg53_63 + County_Agg71_81

### Model diagnostic summaries (example diagnostics)
- p-value: closeness to zero indicates low likelihood that results are due to random variation; the example regression p-value is very close to zero.
- R-squared: example value 54.88 percent of the variability of the price logarithm is explained by the model. The Guide notes that while this might be low elsewhere, for house price indices this is rather good due to many unobserved housing characteristics.
- F-statistic: used to evaluate fitness; normally the higher the F-statistic the better the fit, but it should be analyzed with the table of F-Statistics.

### Method 2 — Time Dummy Hedonic Method: when and how
- Use case: large number of characteristics but few transactions per period.
- Pool observations across several periods and include time dummy variables for each period; index follows from estimated time-dummy parameters.
- Log-linear time-dummy specification:
  - ln p_nt = beta_0 + sum_{t=1}^T delta_t D_nt + sum_{k=1}^K beta_k Z_nk_t + epsilon_nt
  - Notation:
    - beta_0 – intercept
    - delta_t – coefficient of the time dummy variable that will generate the index
    - D_nt – time dummy variables
- Index derivation for period t:
  - I_t = exp(delta_hat_t) * 100  (equation (3))
- Reference period (0) index = 100; no dummy required for reference period.

### Rolling-window approach to prevent revisions
- Adding new-period data can revise previous delta_hat estimates; to avoid index revisions a rolling window is used.
- Practice for quarterly indices:
  - “Shadow” prices of characteristics (beta_hat) are kept fixed for at least a year.
  - Use data from current quarter and previous three quarters (12 months total).
  - Every quarter estimate regression with those 4 quarters; chain indices from overlapping 12-month windows using the last overlap period.

### Worked quarterly example (rolling-window time-dummy)
- Chaining example table (selected rows):
  - Quarter windows shown for 1Q2010 to subsequent windows; method demonstrates chaining by overlap.
- Ireland practice note:
  - O’Hanlon (2011) recommends using data from a 12-month window for a monthly index to mitigate seasonality; a larger window of 15 months may also be used.

### Exercise 3: step-by-step for Time Dummy Hedonic Quarterly Index
- Step 1: Obtain data for four quarters.
- Step 2: Generate dummy variables for each period and categorical variables (example dataset variables listed; example first observation: urban semi-detached dwelling sold in 1Q2008 with 109 square meters, built between 2000 and 2009, building with 2 or more levels, county 62, region 6, heated with gas).
- Step 3: Build strata.
- Step 4: Calculate “shadow” prices by running regression with ln(price) as dependent variable and all other variables as explanatory variables.
  - There will be no coefficient for the first quarter (reference), and index for first quarter = 100 by definition.
- Example OLS results (selected coefficients shown):

  - Time-dummy OLS results for New Dwellings (Table 32, selected rows):
    - (Intercept) Estimate 11.82 Std.Error 0.06 t.value 194.34
    - Period2Q2008 Estimate -0.06 Std.Error 0.02 t.value -3.74
    - Period3Q2008 Estimate -0.10 Std.Error 0.02 t.value -6.09
    - Period4Q2008 Estimate -0.13 Std.Error 0.02 t.value -7.67
    - Dwelling_TypeDetached Estimate 0.10 Std.Error 0.02 t.value 4.42
    - Floor_Area Estimate 0.00 Std.Error 0.00 t.value 22.49
    - Neighborhood_Affluence Estimate 0.01 Std.Error 0.00 t.value 14.07
    - BERHigh Estimate -0.27 Std.Error 0.04 t.value -6.11
    - County_Agg11_14 Estimate -0.19 Std.Error 0.03 t.value -6.86
    - County_Agg71_81 Estimate 0.06 Std.Error 0.03 t.value 2.08

  - Time-dummy OLS results for Existing Dwellings (Table 33, selected rows):
    - (Intercept) Estimate 11.10 Std.Error 0.03 t.value 428.57
    - Period2Q2008 Estimate -0.04 Std.Error 0.01 t.value -3.41
    - Period3Q2008 Estimate -0.05 Std.Error 0.01 t.value -4.93
    - Period4Q2008 Estimate -0.10 Std.Error 0.01 t.value -9.87
    - Dwelling_TypeApartment Estimate -0.13 Std.Error 0.02 t.value -7.72
    - Dwelling_TypeDetached Estimate 0.15 Std.Error 0.01 t.value 10.60
    - Floor_Area Estimate 0.00 Std.Error 0.00 t.value 39.70
    - Neighborhood_Affluence Estimate 0.02 Std.Error 0.00 t.value 44.08
    - Neighborhood_TypeRural Estimate -0.15 Std.Error 0.01 t.value -13.39
    - County_Agg71_81 Estimate 0.39 Std.Error 0.02 t.value 25.65

- Step 5: Calculate sub-indices using delta coefficients (equation (3)):
  - Example sub-indices (Table 34):
    - New: 1Q2008 100.0; 2Q2008 exp(-0.06) * 100 = 94.2; 3Q2008 exp(-0.10) * 100 = 90.8; 4Q2008 exp(-0.13) * 100 = 88.2
    - Existing: 1Q2008 100.0; 2Q2008 exp(-0.04) * 100 = 96.5; 3Q2008 exp(-0.05) * 100 = 95.2; 4Q2008 exp(-0.10) * 100 = 90.3
  - Example interpretation: prices of new dwellings decreased 9.1 percent from 1Q2008 to 3Q2008 (as shown by the sub-index change).

- Step 6: Aggregation (Laspeyres) using property-value weights:
  - Example (Table 35):
    - RPPI: 1Q2008 100.0; 2Q2008 96.0; 3Q2008 94.1; 4Q2008 89.8; Weights 1
    - New: 1Q2008 100.0; ...; Weights 0.24
    - Existing: 1Q2008 100.0; ...; Weights 0.76

- Step 7: Re-reference sub-indices so that 2008 = 100 by dividing quarterly indices by their mean:
  - Example re-referencing (Table 36, selected results):
    - RPPI re-referenced: 1Q2008 105.3 (100 / 0.950 = 105.3); 2Q2008 101.0 (96.0 / 0.950 = 101.0); 3Q2008 99.1 (94.1 / 0.950 = 99.1); 4Q2008 94.6 (89.8 / 0.950 = 94.6)
    - New re-referenced: 1Q2008 107.2 (100 / 0.933 = 107.2); 2Q2008 100.9 (94.2 / 0.933 = 100.9); 3Q2008 97.3 (90.8 / 0.933 = 97.3); 4Q2008 94.6 (88.2 / 0.933 = 94.6)
    - Existing re-referenced: 1Q2008 104.7 (100 / 0.955 = 104.7); 2Q2008 101.1 (96.5 / 0.955 = 101.1); 3Q2008 99.7 (95.2 / 0.955 = 99.7); 4Q2008 94.5 (90.3 / 0.955 = 94.5)
  - Example chaining across windows for overall index (Table 37, selected rows):
    - 1Q2008 105.3
    - 2Q2008 101.0 (chained)
    - 3Q2008 99.1 (chained)
    - 4Q2008 94.6 (chained)
    - 1Q2009 example chaining: 88.9 = 94.6 / 93.7 x 88.1
    - 2Q2009 example chaining: 85.7 = 88.9 / 89.6 x 86.4
  - Note: “User may not obtain the same result due to rounding” appears in example tables.

### Method 3 — Characteristics Hedonic Method (brief)
- Measures price evolution of a “typical” property by averaging characteristics of properties in the stratum.
- Preferred when the compiler has access to a large number of transactions for each period.
- Procedure:
  - Compute average (typical) characteristics for the stratum in period t and in reference period 0.
  - Estimate shadow prices for each characteristic in current period (t) and reference period (0).
  - Index = price of typical property in period t compared with price of typical property in period 0.

*RPPI Practical Guide, International Monetary Fund*

### 127.      The mean is calculated for each numerical variable and the frequency for the categorical

### rppi-guide - 127.      The mean is calculated for each numerical variable and the frequency for the categorical

### Example: Calculation of the “Typical” Dwelling
- Six dwellings (A–F) used to compute the “typical” property by taking:
  - Mean for each numerical variable.
  - Frequency for each categorical/dummy variable.
- Selected mean/frequency results (Typical Property):
  - Dwelling_TypeApartment: 0.17
  - Dwelling_TypeDetached: 0.17
  - Dwelling_TypeSemi_Detached: 0.67
  - Dwelling_TypeTerraced: 0.00
  - Central_HeatingElectricity: 0.17
  - Floor_Area: 111.17
  - Neighborhood_Affluence: 27.77
  - Year_Builtor_less_1910: 0.00
  - BERHigh: 0.00
  - BERLow: 0.17
  - Neighborhood_TypeUrban: 0.50
  - Building_Levels1: 0.67
  - Building_Levels2_or_more: 0.33

### Method 1 (Characteristics Hedonic Method) — Model and Index Formula
- Log-linear specification for each stratum:
  - lnp_nt = β0 + Σ_k β_k z_nk_t + ε_nt   (equation (4))
- “Shadow” prices: estimated regression coefficients (β̂) from separate regressions for reference period (0) and current period (t).
- Index calculation:
  - I_0t = exp( Σ_k (β̂_k_t − β̂_k_0) Z̅_k_0 )   (equation (5))
  - Z̅_k_0 = average characteristic in reference period (0).

### Characteristics Hedonic — Step-by-step (Exercise 4)
- Step 1: Obtain data.
- Step 2: Build strata.
- Step 3: Calculate “shadow” prices — run regression with log(price) dependent; transform categorical variables into dummies.
  - Example coefficients for New Dwellings (1Q2008):
    - (Intercept): 11.88 Std.Error 0.12 t.value 99.53
    - Dwelling_TypeApartment: 0.00 Std.Error 0.03 t.value 0.12
    - Dwelling_TypeDetached: 0.15 Std.Error 0.04 t.value 3.61
    - Dwelling_TypeSemi_Detached: 0.14 Std.Error 0.03 t.value 4.73
    - Floor_Area: 0.00 Std.Error 0.00 t.value 12.53
    - Neighborhood_Affluence: 0.01 Std.Error 0.00 t.value 7.47
    - BERHigh: -0.28 Std.Error 0.09 t.value -3.30
    - Central_HeatingElectricity: -0.14 Std.Error 0.07 t.value -1.98
    - Central_HeatingMains_Gas: 0.07 Std.Error 0.06 t.value 1.12
    - Central_HeatingOil: -0.15 Std.Error 0.06 t.value -2.49
    - County_Agg11_14: -0.32 Std.Error 0.05 t.value -6.18
    - County_Agg22_31: -0.18 Std.Error 0.06 t.value -3.34
    - County_Agg31_41: -0.12 Std.Error 0.05 t.value -2.31
    - County_Agg42_52: -0.09 Std.Error 0.05 t.value -1.76
    - County_Agg53_63: -0.03 Std.Error 0.05 t.value -0.62
    - County_Agg71_81: -0.02 Std.Error 0.05 t.value -0.33
- Step 4: Calculate average characteristics for the “typical” property in reference period (1Q2008).
  - Selected CHR, Coef_P2s1, Coef_P1s1, Average (from Table 40):
    - BERHigh: Coef_P2s1 -0.24, Coef_P1s1 -0.28, Average 0.98
    - Building_Levels1: 0.00, 0.08, 0.05
    - Central_HeatingElectricity: 0.04, -0.14, 0.08
    - Central_HeatingMains_Gas: 0.08, 0.07, 0.54
    - Central_HeatingOil: -0.07, -0.15, 0.35
    - Dwelling_TypeDetached: 0.10, 0.15, 0.24
    - Dwelling_TypeSemi_Detached: 0.06, 0.14, 0.33
    - Floor_Area: 0.00, 0.00, 122.83
    - Neighborhood_Affluence: 0.01, 0.01, 23.90
    - Neighborhood_TypeRural: -0.10, -0.01, 0.27
- Step 5: Compile sub-indices using equation (5).
  - Result for New Dwellings, 2Q2008: Sub-index = Exp(sum(...)) * 100 = 94.5
  - Interpretation: decrease in prices of dwellings of 5.5 percent since 1Q2008.
- Step 6: Aggregation (Laspeyres) — weights by property values.
  - Table 42 (selected):
    - RPPI: 1Q2008 100, 2Q2008 96.1, 3Q2008 94.2, 4Q2008 90.0, Weights 1
    - New: 100, 94.5, 91.4, 88.5, Weights 0.24
    - Existing: 100, 96.5, 95.1, 90.5, Weights 0.76
- Step 7: Re-reference indices so that 2008 = 100 (mean of quarterly indices as denominator).
  - Example re-referencing results (from Table 43):
    - RPPI re-referenced:
      - 1Q2008: 100 / 0.95 = 105.2
      - 2Q2008: 96.1 / 0.95 = 101.0
      - 3Q2008: 94.2 / 0.95 = 99.1
      - 4Q2008: 90.0 / 0.95 = 94.7
    - New re-referenced:
      - 1Q2008: 100 / 0.94 = 106.9
      - 2Q2008: 94.5 / 0.94 = 101.0
      - 3Q2008: 91.4 / 0.94 = 97.6
      - 4Q2008: 88.5 / 0.94 = 94.5
    - Existing re-referenced:
      - 1Q2008: 100 / 0.96 = 104.7
      - 2Q2008: 96.5 / 0.96 = 101.0
      - 3Q2008: 95.1 / 0.96 = 99.6
      - 4Q2008: 90.5 / 0.96 = 94.7
- Ongoing production notes:
  - Update base average characteristics and coefficients yearly using last quarter of previous year (R procedure “base_characteristics.r”).
  - Steps 1–3 repeat each year; Step 4 uses “base_characteristics.r”.
  - Step 7 for multi-year series: chain indices by multiplying current period index by chained index of last period of previous year (see Tables 44–45).

### Method 4 (Imputations Hedonic Method) — Overview and Mechanics
- Description:
  - Imputation hedonic method imputes a price for each property for each period (including periods with no transaction) and calculates an index from imputed prices.
  - Requires detailed property characteristics and a large number of transactions; generally appropriate if geospatial data (coordinates) are used.
  - If log-linear form is used, characteristics and imputation hedonic methods are equivalent except when geospatial data are used.
- Log-linear specification:
  - ln p_lt = β0 + Σ_n β_n z_ln_t + ε_lt   (equation (6))
- Shadow prices: β̂ estimated separately for reference period (0) and current period (t).
- Imputed prices:
  - For reference period (0): p̂_l0(0) = exp(β̂_0_0) exp( Σ_n β̂_n_0 z_l n 0 )   (equation (7))
  - For current period (t) applied to dwellings sold in reference period (0): p̂_l t (0) = exp(β̂_0_t) exp( Σ_n β̂_n_t z_l n 0 )   (equation (8))
- Index from imputed prices:
  - I_t = [ ∏_{n∈S(0)} (p̂_n t (0))^{1/N(0)} ] / [ ∏_{n∈S(0)} (p̂_n 0)^{1/N(0)} ]   (equation (9))
  - Observed prices from reference period replaced by imputed values.

### Imputations Hedonic — Step-by-step (Exercise 5)
- Step 1: Obtain data.
- Step 2: Build strata.
- Step 3: Calculate “shadow” prices — regressions for each period; coefficients are shadow prices.
  - Table 46: Coefficients by quarter (selected entries)
    - Intcpt: 1Q2008 11.90, 2Q2008 11.80, 3Q2008 11.60, 4Q2008 11.72
    - BERHigh: 1Q2008 -0.28, 2Q2008 -0.24, 3Q2008 -0.25, 4Q2008 -0.30
    - Building_Levels1: 1Q2008 0.08, 2Q2008 0.00, 3Q2008 0.09, 4Q2008 0.02
    - Central_HeatingElectricity: 1Q2008 -0.14, 2Q2008 0.04, 3Q2008 0.01, 4Q2008 -0.12
    - Central_HeatingMains_Gas: 1Q2008 0.07, 2Q2008 0.08, 3Q2008 0.10, 4Q2008 0.08
    - Central_HeatingOil: 1Q2008 -0.15, 2Q2008 -0.07, 3Q2008 -0.05, 4Q2008 -0.07
    - Dwelling_TypeDetached: 1Q2008 0.15, 2Q2008 0.10, 3Q2008 0.12, 4Q2008 0.05
    - Floor_Area: 1Q2008 0.00, 2Q2008 0.00, 3Q2008 0.00, 4Q2008 0.00
- Step 4: Impute prices for each property and each period using regression coefficients.
  - Example (Table 47): Three dwellings (ID 6, ID 18, ID 23) imputed log prices and imputed prices for 1Q2008:
    - ID 6: log imputed price 12.46 → imputed price 257699.00
    - ID 18: log imputed price 12.52 → imputed price 274896.24
    - ID 23: log imputed price 12.41 → imputed price 246011.28
- Step 5: Aggregate estimated prices by geometric mean for properties in each stratum and period.
  - Jevons of imputed prices for 2008 (Table 48):
    - New: 1Q2008 233573, 2Q2008 220722, 3Q2008 213389, 4Q2008 206634
    - Existing: 1Q2008 239609, 2Q2008 231316, 3Q2008 227946, 4Q2008 216885
- Step 6: Calculate sub-indices as ratio of geometric means (equation (10)).
  - Table 49 (sub-indices for 2008):
    - New: 1Q2008 100; 2Q2008 = 220722 / 233573 * 100 = 94.5; 3Q2008 = 213389 / 233573 * 100 = 91.4; 4Q2008 = 206634 / 233573 * 100 = 92.9
    - Existing: 1Q2008 100; 2Q2008 = 231316 / 239609 * 100 = 96.5; 3Q2008 = 227946 / 239609 * 100 = 95.1; 4Q2008 = 216885 / 239609 * 100 = 90.5
- Step 7: Aggregation (Laspeyres) — weighted average using property values as weights.
  - Table 50:
    - RPPI: 1Q2008 100, 2Q2008 96.1, 3Q2008 94.2, 4Q2008 91.1, Weights 1
    - New: 100, 94.5, 91.4, 92.9, Weights 0.24
    - Existing: 100, 96.5, 95.1, 90.5, Weights 0.76
- Step 8: Re-reference indices so 2008 = 100 by dividing quarterly indices by mean of quarterly indices.
  - Table 51 referencing factors (mean of indices) and results:
    - RPPI: (100 + 96.1 + 94.2 + 91.1) / 4 = 0.95
      - 1Q2008: 100 / 0.95 = 104.9
      - 2Q2008: 96.1 / 0.95 = 100.7
      - 3Q2008: 94.2 / 0.95 = 98.8
      - 4Q2008: 91.1 / 0.95 = 95.5
    - New: (100 + 94.5 + 91.4 + 92.9) / 4 = 0.95
      - 1Q2008: 100 / 0.95 = 105.6
      - 2Q2008: 94.5 / 0.95 = 99.8
      - 3Q2008: 91.4 / 0.95 = 96.5
      - 4Q2008: 92.9 / 0.95 = 98.1
    - Existing: (100 + 96.5 + 95.1 + 90.5) / 4 = 0.96
      - 1Q2008: 100 / 0.96 = 104.7
      - 2Q2008: 96.5 / 0.96 = 101.0
      - 3Q2008: 95.1 / 0.96 = 99.6
      - 4Q2008: 90.5 / 0.96 = 94.7
- Multi-year production notes:
  - Use R procedure “imputations_production.r” for following years.
  - Base price for subsequent years can be obtained with R procedure “base price.r”.
  - For 1Q2009 sub-index compilation (Table 52):
    - New: Jevons 4Q2008 206634; Jevons 1Q2009 199271.3 → 199271.3 / 206634 * 100 = 95.0
    - Existing: 216885 → 195873.1 → 195873.1 / 216885 * 100 = 94.0
  - Chain indices by multiplying current period index by chained index of last period of previous year (see Tables 53–54 for chained outcomes).

### Key numeric outcomes and illustrative interpretations
- Characteristics hedonic example:
  - Sub-index New Dwellings 2Q2008 = 94.5 → 5.5 percent decrease since 1Q2008.
  - Aggregated RPPI for 2Q2008 = 96.1 (from Table 42).
- Imputations hedonic example:
  - Jevons geometric means for New dwellings: 1Q2008 233573; 4Q2008 206634.
  - New sub-index 4Q2008 = 92.9 (Table 49).
  - Aggregated RPPI 4Q2008 = 91.1 (Table 50).
- Re-referencing adjustments ensure 2008 = 100 using the mean of quarterly indices:
  - Representative re-referenced RPPI 1Q2008 values range around 104.9–105.2 depending on method.

*RPPI PRACTICAL GUIDE, INTERNATIONAL MONETARY FUND*

### 168.      Table 55 and Figure 8 shows the RPPI compiled using the four methods described in this

### rppi-guide - 168.

### Comparison of RPPI compilation methods (Table 55 and Figure 8)
- Overview findings:
  - The characteristics and the imputations methods provide similar results since they are variants of the same method.
  - The time dummy method provides less volatile results because it uses a rolling window of four quarters which helps smooth the index.
  - The RPPI calculated with a simple stratification method using a median price is highly volatile due to the lack of proper quality adjustment.

- RPPI Index values compiled with four methods (quarters listed in order shown in Table 55):
  - Stratification (1Q2008–4Q2011):
    - 105.0, 103.3, 98.3, 93.4, 88.2, 86.3, 81.3, 76.9, 70.7, 71.5, 71.7, 70.0, 60.7, 64.6, 67.4, 67.9
  - Time dummy (1Q2008–4Q2011):
    - 105.2, 101.0, 99.1, 94.7, 89.2, 85.7, 80.8, 76.1, 71.7, 70.8, 71.0, 70.1, 66.5, 68.0, 70.0, 70.5
  - Characteristics (1Q2008–4Q2011):
    - 105.2, 101.0, 99.1, 94.7, 89.2, 85.7, 81.2, 76.5, 72.5, 71.2, 71.4, 70.6, 67.0, 68.8, 70.7, 71.6
  - Imputations (1Q2008–4Q2011):
    - 105.2, 101.0, 99.1, 94.7, 89.2, 85.7, 81.2, 76.5, 72.5, 71.2, 71.4, 70.6, 67.0, 68.8, 70.7, 71.6

  - Stratification (1Q2012–4Q2015):
    - 64.1, 67.6, 72.7, 70.1, 70.3, 75.7, 80.5, 80.4, 80.8, 81.6, 87.7, 86.9, 91.2, 89.6, 98.2, 97.5
  - Time dummy (1Q2012–4Q2015):
    - 70.0, 74.3, 77.5, 78.0, 78.6, 83.0, 85.7, 86.1, 87.5, 89.7, 93.1, 94.3, 97.0, 97.4, 104.4, 105.6
  - Characteristics (1Q2012–4Q2015):
    - 71.5, 76.2, 79.0, 79.4, 80.0, 84.4, 87.3, 87.4, 88.9, 91.1, 94.5, 95.9, 98.7, 99.0, 106.1, 107.7
  - Imputations (1Q2012–4Q2015):
    - 71.5, 76.2, 79.0, 79.4, 80.0, 84.4, 87.3, 87.4, 88.9, 91.1, 94.5, 95.9, 98.7, 99.0, 106.1, 107.7

- Visual summary:
  - Figure 8 shows the four method series plotted with 2008=100; the time dummy series is the smoothest and the stratification series is the most volatile.

### Recommendations for compilation, testing, and dissemination
- Method choice:
  - Hedonic regression–based methods are generally considered superior and provide higher quality estimates.
  - If necessary data for hedonic methods are unavailable, other methods such as median with stratification may be appropriate.

- Experimental testing before official release:
  - Compilers may consider developing experimental indices for at least one year to vet accuracy, relevance, and timeliness and to refine methods and business processes.
  - Experimental estimates allow testing of overall business processes prior to official release.
  - Once methodology and processes are finalized, the index should be compiled for several periods before disseminating official results.

- Dissemination strategy and timing:
  - RPPIs should be made public with a media release.
  - Compilers may organize a meeting or seminar with main stakeholders at release to explain appropriate interpretation.
  - Compilers may consider use of social media and videos to broaden awareness.
  - Compilers should target releasing quarterly data within 90 days after the reference period.
  - Compilers should make data available in machine readable form (e.g., .csv) and publish according to an advance release calendar.
  - Spreadsheets with the RPPI series should be available for download.

- Technical documentation:
  - Release should be accompanied by a technical note on sources and methods that is short, clear, and easy to read.
  - Suggested technical note content includes:
    - the aim of the index;
    - the coverage;
    - the stratification;
    - the weighting system;
    - the reference periods of the weights and the indices; and
    - the method used.
  - The technical note presentation should follow a standard format such as the IMF factsheet for the Data Quality Assessment Framework, the United Nations Generic National Quality Assurance Framework or the G20 Data Gaps Initiative template (Annex II).
  - Compilers should inform users of any changes over time to data sources and methods.

- Publication format guidance:
  - A general recommendation is to publish indices with two decimals and rates with one decimal.

### Visualisation and complementary indicators
- Graphical presentations recommended:
  - Time-series graphs of the total RPPI and sub-indices, clearly indicating the reference period (base). Example discussed uses 2008 as the reference period.
  - Graphs showing trend of the contribution of each stratum to the annual change (YoY) in the total RPPI (Figure 10); existing dwellings often contribute most due to larger weight.
  - Comparison of RPPI trends with number of dwellings sold (Figure 11); sales often follow RPPI trend from 3Q2010 but not before, with seasonal year-end peaks.
  - Benchmark RPPI with CPI to compare house price volatility with consumer prices (Figure 12); house prices can be more volatile.
  - Cross-country comparisons with caution: account for methodological differences between countries.
  - Visual vocabulary and good practice guidance suggested (e.g., Financial Times visual vocabulary, Figure 13).
  - Experimentation with maps, heat maps, and other interactive graphics (examples include city concentration of million-dollar dwellings and price-per-square-meter heat maps).

- Housing affordability and related indicators:
  - RPPI input data can be combined with income data to construct affordability measures such as house price-to-income ratios and city-level price-to-income ratios (Figures 15 and 16).
  - Release of indicators on the number of dwellings sold is encouraged (Second Phase of the G-20 Data Gaps Initiative recommends releasing sales indicators), as sales appeal to a broader audience than price indices.
  - Regional price development and prices per square meter/foot are useful for public interest when region-specific trends differ.

### Operational and dissemination strategy guidance
- Dissemination should be planned prior to index development and consider:
  - user needs;
  - latest technical advances in data visualization;
  - diverse uses of the information.
- Ideal dissemination is part of a broader residential property information framework including prices, sales, stocks, quality, and affordability.
- Encourage experimentation with novel visualizations to increase impact and accessibility.

---


_Source: https://www.imf.org/-/media/files/data/guides/rppi/rppi-guide.pdf_
