## 1uaeea2022004

## Source details

**Canonical URL:** [1uaeea2022004](https://www.imf.org/-/media/files/publications/cr/2022/english/1uaeea2022004.pdf)

## Other formats

- [Markdown version](/-/media/files/publications/cr/2022/english/1uaeea2022004.pdf.md)
- [Structured JSON version](/-/media/files/publications/cr/2022/english/1uaeea2022004.pdf.json)

---

### Summary of mission outcomes and priority recommendations
- Mission: Technical assistance (TA) to Dubai Statistics Center (DSC) on introducing hedonic methods for quality adjustments in the Consumer Price Index (CPI) and the Real Estate Price Index (REPI); mission took place remotely during January 16−20, 2022.
- Objective: Assist on introducing the use of hedonic methods for CPI and REPI; provide training and R code adapted to DSC sample data.
- Data source for REPI: Dubai Land Department (DLD) transaction data available since 2016 covering residential buildings, commercial buildings, and land (residential and commercial).
- Key methodological recommendation for RPPI:
  - Use the hedonics time dummy method with a rolling window of 12 months for the compilation of the Residential Property Price Index (RPPI).
  - Rationale: pools one year of data (less volatile indices), suitable when few observations are available, widely used for RPPI and for CPI compilation with web scraped data.
- Recommended change in index publication structure:
  - Publish a stand-alone RPPI following residential building price trends (a key indicator for financial stability).
  - Publish stand-alone indices for commercial properties and separate indices for residential land and commercial land (by type of activity).
- Data cleaning and processing:
  - R codes were created to perform data cleaning, analysis, preparation for modeling.
  - Cleaning steps: removing duplicates; removing observations with missing values in model variables; identification and removal of outliers; creating categories for number of bedrooms.
  - Outlier detection performed by strata (sectors) for each month using price per square meter; experiment with different interquartile multiplier values recommended.
- Weights and aggregation:
  - DSC currently calculates flow weights using transaction data to aggregate strata indices.
  - Recommendation: update weights annually with transaction data of the previous year, or of the previous three years, and keep those stable for the year (Laspeyres-type index).
- Communication and release:
  - Release of the new RPPI should include publication of a technical note and/or methodology paper; a draft technical note was provided by the mission.
  - Publicize the release in the media to inform potential users: real estate developers, financial institutions, households, tax office, Central Bank, National accounts staff, etc.
- Web scraping for CPI:
  - Web scraped data recommended to improve CPI sub-indices (better quality change capture, longer time coverage, reduced burden on retailers, higher coverage, automation efficiencies).
  - Typical web-scraped products: flights, electronics, clothes.
  - Sub-indices from web scraped data should be compiled using the time dummy hedonic method with 12-month rolling window.
  - Aggregation: compile sub-index per web site and product; aggregate web-site sub-indices (Laspeyres-type) using turnover weights; compile in-person CPI sub-index (Jevons) and aggregate web-scraped and in-person sub-indices using turnover weights of companies/shops included in each.
- Priority recommendations (as presented)
  - Use hedonic methods for the REPI compilation — Responsible: DSC.
  - Update the weights annually — Responsible: DSC.
  - Web scrape all available data for each product, three times per month — Responsible: DSC.

### Detailed technical assessment and recommended actions (selected milestones and dates)
- Priority Action/Milestone — Target Completion Date:
  - Compile experimental RPPI with time dummy hedonic method — February 2022.
  - Compile other experimental Real Estate Indices with time dummy hedonic method — March 2022.
  - Draft a new methodology paper/technical note to inform users and managers of the changes — April 2022.
  - Publish new REPI — Subject to DSC Management evaluation and approval.
  - Meet with real estate data providers to improve data quality and understand variables in the current data set — February 2022.
  - Begin web scraping data on mobile phones, TV, and other electronic products — April 2022.
  - Compile experimental indices for the web scraped products with time dummy hedonic method — Subject to DSC Management evaluation and approval.
  - Investigate data sources on companies’ turnover (weights) for the web scraped products — December 2022.
  - Draft a new methodology paper/technical note to inform users and managers of the changes (CPI release) — Subject to DSC Management evaluation and approval.
  - Publish new CPI — Subject to DSC Management evaluation and approval.

### A. The REPI — assessment and recommendations
- Current practice:
  - REPI compiled using stratification with simple averages.
  - Data from DLD include: value; flat or house indicator; location (area and sector); size in square meters; existence of a balcony; existence of car parking; information whether property is existing or off plan; procedure type.
  - Built year has a high number of missing values.
  - Some variable meanings (e.g., procedure type) remain unclear and require clarification with DLD.
- Recommended index structure changes:
  - Maintain an overall REPI but publish a stand-alone RPPI for residential building trends.
  - Publish stand-alone indices for commercial properties (each type of activity).
  - Publish separate indices for residential land and commercial land (by type of activity).
- Data processing and quality control:
  - Complete a data-structure table (per RPPI Practical Compilation Guide) after clarifying variables with DLD.
  - Improve future data collection by requesting built year and other meaningful variables from DLD.
  - Data cleaning steps reiterated: remove duplicates; remove observations with missing model variables; detect/remove outliers; create categories for number of bedrooms.
  - Outlier detection: performed by strata (sectors) for each month using price per square meter; experiment by changing interquartile multiplier.
- Hedonic method specifics:
  - Preferred method: hedonics time dummy with rolling window of 12 months.
  - Advantages: more stable (less volatile) indices; pools one year of data; suitable for limited observations; widely used for RPPI and CPI with web scraped data.
- Revision policy:
  - RPPI can be revised up to two quarters prior to the reference date.
  - Revision policy should be publicized on the website and in the methodology note/paper.
- Recommended actions (concise):
  - Meet with main users and stakeholders (Central Bank, DLD, others) to share methodology and future plans.
  - Meet with DLD to clarify current data and improve future collection.
  - Adapt REPI structure to include RPPI, commercial property indices, residential land and commercial land indices.
  - Experiment with outlier options by varying the interquartile multiplier.
  - Update weights annually and keep them stable for the year.
  - Use the hedonics time dummy method with a rolling window of 12 months.
  - Release new RPPI with a technical note and/or methodology paper.

### B. Web Scraping — assessment and recommendations
- Rationale and benefits:
  - Web scraped data provides price information over longer periods (not just one day per month).
  - Better source for inclusion of new items.
  - Potentially reduces administrative burden on retailers and cost of price collection.
  - Expected to increase retailer and item coverage and enable greater automation and production efficiency.
- Operational recommendations:
  - Perform web scraping once per week during the first three weeks of the month.
  - After collection, join all data pertaining to one month and one product into a single dataset.
  - Retrieve all available data: all varieties and all characteristics of each variety.
  - Scrape more than one web site for each product.
- Data cleaning and modeling:
  - Clean web scraped data for outliers and missing values prior to index compilation.
  - Methodology for cleaning/analysis follows same approach as for real estate indices for one month of data, for each product.
  - Convert categorical characteristics into dummy variables; create categories when many instances exist (more than five).
- Index compilation and aggregation:
  - Compile sub-indices with web scraped data using the time dummy hedonic method with 12-month rolling window.
  - For each product: compile sub-index per web site, aggregate web-site sub-indices (Laspeyres-type) using turnover weights.
  - Continue in-person price collection for sampled varieties where applicable; compile in-person sub-index (Jevons).
  - Aggregate web-scraped and in-person sub-indices using turnover weights of companies/shops included in each.
  - Use turnover concept/coverage as close as possible to product concept; ensure consistency across companies to avoid bias.
  - Obtain turnover data from business register and/or tax offices (often available for national accounts).
- Frequency and coverage recommendation:
  - Data collection should be done in both online and physical shops if prices or varieties differ.

### Section 2 — Web scraping and price collection procedures; Aggregation, indices, and weighting
- Communication and metadata:
  - The introduction of this new form of data collection does not imply a series break and should be communicated to users by updating all metadata and methodology documents.
  - Estimate the volume of turnover for online and offline.
  - Otherwise, the in-person price collection can be replaced by web scraping.
- Recommended collection frequency and scope:
  - Perform web scraping once per week during the first three weeks of the month.
  - Retrieve all available data, i.e., all varieties and all characteristics of each variety.
  - Web scrape more than one web site for each product.
  - Use the time dummy hedonic method with 12-months rolling window.
  - Use turnover weights to aggregate product sub-indices by web site and further to aggregate the web scraped sub-indices with the in-person price collection subindices, by product.
  - Investigate potential data sources for turnover with national accounts, tax authorities, business statistics, etc.
  - Keep data collection in both online and physical shops if the prices or the varieties are different.
- Aggregation and index options listed:
  - Jevons Index
  - Weighted average with Turnover of the shops included in the web scraping
  - Weighted average with Turnover of the shops
  - Indices per shop/website using time dummy hedonic method
- Example product sub-index workflow (Mobile phones Sub-index):
  - In-person price collection: Shop 1; Shop 2
  - Webcraped: Shop 3 (website); Shop 4 (website)
  - Aggregation options shown:
    - Jevons Index
    - Weighted average with Turnover of the shops included in the web scraping
    - Weighted average with Turnover of the shops
    - Indices per shop/website using time dummy hedonic method
- Recommended condensed actions:
  - Perform web scraping once per week during the first three weeks of the month.
  - Retrieve all available data, i.e., all varieties and all characteristics of each variety.
  - Web scrape more than one web site for each product.
  - Use the time dummy hedonic method with 12-months rolling window.
  - Use turnover weights to aggregate product sub-indices by web site and to aggregate web-scraped sub-indices with in-person price collection subindices, by product.
  - Investigate potential data sources for turnover with national accounts, tax authorities, business statistics, etc.
  - Keep data collection in both online and physical shops if the prices or the varieties are different.

### Officials met during the mission
- Maryam Al-Mulla DSC
- Nusaiba Al Marzooqi DSC
- Adnan Alafari DSC
- Rola Rashad DSC
- AbdelHamid AbdulHadi DSC
- Sherif Bayoumi DSC
- Maryam AlMarri DSC
- Amro Mohammed DSC
- Mohamid Elfateh DSC

*Source: IMF Technical Assistance Report—Hedonic Methods for Price Indices Mission (United Arab Emirates), mission completed January 2022.*

### Section 1

### 1uaeea2022004 - Section 1

### Summary of mission outcomes and priority recommendations
- Mission: Technical assistance (TA) to Dubai Statistics Center (DSC) on introducing hedonic methods for quality adjustments in the Consumer Price Index (CPI) and the Real Estate Price Index (REPI). Mission took place remotely during January 16−20, 2022.
- Objective: Assist on introducing the use of hedonic methods for CPI and REPI; provide training and R code adapted to DSC sample data.
- Data source for REPI: Dubai Land Department (DLD) transaction data available since 2016 covering residential buildings, commercial buildings, and land (residential and commercial).
- Key methodological recommendation for RPPI:
  - Use the hedonics time dummy method with a rolling window of 12 months for the compilation of the Residential Property Price Index (RPPI).
  - Rationale: pools one year of data (less volatile indices), suitable when few observations are available, widely used for RPPI and for CPI compilation with web scraped data.
- Recommended change in index publication structure:
  - Publish a stand-alone RPPI following residential building price trends (a key indicator for financial stability).
  - Publish stand-alone indices for commercial properties and separate indices for residential land and commercial land (by type of activity).
- Data cleaning and processing:
  - R codes were created to perform data cleaning, analysis, preparation for modeling.
  - Cleaning steps: removing duplicates; removing observations with missing values in model variables; identification and removal of outliers; creating categories for number of bedrooms.
  - Outlier detection performed by strata (sectors) for each month using price per square meter; experiment with different interquartile multiplier values recommended.
- Weights and aggregation:
  - DSC currently calculates flow weights using transaction data to aggregate strata indices.
  - Recommendation: update weights annually with transaction data of the previous year, or of the previous three years, and keep those stable for the year (Laspeyres-type index).
- Communication and release:
  - Release of the new RPPI should include publication of a technical note and/or methodology paper; a draft technical note was provided by the mission.
  - Publicize the release in the media to inform potential users: real estate developers, financial institutions, households, tax office, Central Bank, National accounts staff, etc.
- Web scraping for CPI:
  - Web scraped data recommended to improve CPI sub-indices (better quality change capture, longer time coverage, reduced burden on retailers, higher coverage, automation efficiencies).
  - Typical web-scraped products: flights, electronics, clothes.
  - Sub-indices from web scraped data should be compiled using the time dummy hedonic method with 12-month rolling window.
  - Aggregation: compile sub-index per web site and product; aggregate web-site sub-indices (Laspeyres-type) using turnover weights; compile in-person CPI sub-index (Jevons) and aggregate web-scraped and in-person sub-indices using turnover weights of companies/shops included in each.

Priority recommendations (as presented)
- Use hedonic methods for the REPI compilation — Responsible: DSC.
- Update the weights annually — Responsible: DSC.
- Web scrape all available data for each product, three times per month — Responsible: DSC.

### Detailed technical assessment and recommended actions (selected milestones and dates)
- Priority Action/Milestone — Target Completion Date:
  - Compile experimental RPPI with time dummy hedonic method — February 2022.
  - Compile other experimental Real Estate Indices with time dummy hedonic method — March 2022.
  - Draft a new methodology paper/technical note to inform users and managers of the changes — April 2022.
  - Publish new REPI — Subject to DSC Management evaluation and approval.
  - Meet with real estate data providers to improve data quality and understand variables in the current data set — February 2022.
  - Begin web scraping data on mobile phones, TV, and other electronic products — April 2022.
  - Compile experimental indices for the web scraped products with time dummy hedonic method — Subject to DSC Management evaluation and approval.
  - Investigate data sources on companies’ turnover (weights) for the web scraped products — December 2022.
  - Draft a new methodology paper/technical note to inform users and managers of the changes (CPI release) — Subject to DSC Management evaluation and approval.
  - Publish new CPI — Subject to DSC Management evaluation and approval.

### A. The REPI — assessment and recommendations
- Current practice:
  - REPI compiled using stratification with simple averages.
  - Data from DLD include: value; flat or house indicator; location (area and sector); size in square meters; existence of a balcony; existence of car parking; information whether property is existing or off plan; procedure type.
  - Built year has a high number of missing values.
  - Some variable meanings (e.g., procedure type) remain unclear and require clarification with DLD.
- Recommended index structure changes:
  - Maintain an overall REPI but publish a stand-alone RPPI for residential building trends.
  - Publish stand-alone indices for commercial properties (each type of activity).
  - Publish separate indices for residential land and commercial land (by type of activity).
- Data processing and quality control:
  - Complete a data-structure table (per RPPI Practical Compilation Guide) after clarifying variables with DLD.
  - Improve future data collection by requesting built year and other meaningful variables from DLD.
  - Data cleaning steps reiterated: remove duplicates; remove observations with missing model variables; detect/remove outliers; create categories for number of bedrooms.
  - Outlier detection: performed by strata (sectors) for each month using price per square meter; experiment by changing interquartile multiplier.
- Hedonic method specifics:
  - Preferred method: hedonics time dummy with rolling window of 12 months.
  - Advantages: more stable (less volatile) indices; pools one year of data; suitable for limited observations; widely used for RPPI and CPI with web scraped data.
- Revision policy:
  - RPPI can be revised up to two quarters prior to the reference date.
  - Revision policy should be publicized on the website and in the methodology note/paper.
- Recommended actions (concise):
  - Meet with main users and stakeholders (Central Bank, DLD, others) to share methodology and future plans.
  - Meet with DLD to clarify current data and improve future collection.
  - Adapt REPI structure to include RPPI, commercial property indices, residential land and commercial land indices.
  - Experiment with outlier options by varying the interquartile multiplier.
  - Update weights annually and keep them stable for the year.
  - Use the hedonics time dummy method with a rolling window of 12 months.
  - Release new RPPI with a technical note and/or methodology paper.

### B. Web Scraping — assessment and recommendations
- Rationale and benefits:
  - Web scraped data provides price information over longer periods (not just one day per month).
  - Better source for inclusion of new items.
  - Potentially reduces administrative burden on retailers and cost of price collection.
  - Expected to increase retailer and item coverage and enable greater automation and production efficiency.
- Operational recommendations:
  - Perform web scraping once per week during the first three weeks of the month.
  - After collection, join all data pertaining to one month and one product into a single dataset.
  - Retrieve all available data: all varieties and all characteristics of each variety.
  - Scrape more than one web site for each product.
- Data cleaning and modeling:
  - Clean web scraped data for outliers and missing values prior to index compilation.
  - Methodology for cleaning/analysis follows same approach as for real estate indices for one month of data, for each product.
  - Convert categorical characteristics into dummy variables; create categories when many instances exist (more than five).
- Index compilation and aggregation:
  - Compile sub-indices with web scraped data using the time dummy hedonic method with 12-month rolling window.
  - For each product: compile sub-index per web site, aggregate web-site sub-indices (Laspeyres-type) using turnover weights.
  - Continue in-person price collection for sampled varieties where applicable; compile in-person sub-index (Jevons).
  - Aggregate web-scraped and in-person sub-indices using turnover weights of companies/shops included in each.
  - Use turnover concept/coverage as close as possible to product concept; ensure consistency across companies to avoid bias.
  - Obtain turnover data from business register and/or tax offices (often available for national accounts).
- Frequency and coverage recommendation:
  - Data collection should be done in both online and physical shops if prices or varieties differ.

*Source: IMF Technical Assistance Report—Hedonic Methods for Price Indices Mission (United Arab Emirates), mission completed January 2022.*

### Section 2

### Section 2

### Web scraping and price collection procedures
- The introduction of this new form of data collection does not imply a series break and should be communicated to users by updating all metadata and methodology documents.
- In this case you will need to estimate the volume of turnover for online and offline.
- Otherwise, the in-person price collection can be replaced by web scraping.
- Perform web scraping once per week during the first three weeks of the month.
- Retrieve all available data should, i.e., all varieties and all characteristics of each variety.
- Web scrape more than one web site for each product.
- Use the time dummy hedonic method with 12-months rolling window.
- Use turnover weights to aggregate product sub-indices by web site and further to aggregate the web scraped sub-indices with the in-person price collection subindices, by product.
- Investigate potential data sources for turnover with national accounts, tax authorities, business statistics, etc.
- Keep data collection in both online and physical shops if the prices or the varieties are different.

### Aggregation, indices, and weighting
- Jevons Index
- Weighted average with Turnover of the shops included in the web scraping
- Weighted average with Turnover of the shops
- Indices per shop/website using time dummy hedonic method

Example product sub-index workflow (as presented)
- Mobile phones Sub-index
  - In-person price collection: Shop 1; Shop 2
  - Webcraped: Shop 3 (website); Shop 4 (website)
  - Aggregation options shown:
    - Jevons Index
    - Weighted average with Turnover of the shops included in the web scraping
    - Weighted average with Turnover of the shops
    - Indices per shop/website using time dummy hedonic method

### Recommended actions (condensed)
- Perform web scraping once per week during the first three weeks of the month.
- Retrieve all available data, i.e., all varieties and all characteristics of each variety.
- Web scrape more than one web site for each product.
- Use the time dummy hedonic method with 12-months rolling window.
- Use turnover weights to aggregate product sub-indices by web site and to aggregate web-scraped sub-indices with in-person price collection subindices, by product.
- Investigate potential data sources for turnover with national accounts, tax authorities, business statistics, etc.
- Keep data collection in both online and physical shops if the prices or the varieties are different.

### Officials met during the mission
- Maryam Al-Mulla DSC
- Nusaiba Al Marzooqi DSC
- Adnan Alafari DSC
- Rola Rashad DSC
- AbdelHamid AbdulHadi DSC
- Sherif Bayoumi DSC
- Maryam AlMarri DSC
- Amro Mohammed DSC
- Mohamid Elfateh DSC

*Source: 1uaeea2022004 - Section 2*

---


_Source: https://www.imf.org/-/media/files/publications/cr/2022/english/1uaeea2022004.pdf_
