## _wp14236

## Source details

**Canonical URL:** [_wp14236](https://www.imf.org/-/media/websites/imf/imported-full-text-pdf/external/pubs/ft/wp/2014/_wp14236.pdf)

## Other formats

- [Markdown version](/-/media/websites/imf/imported-full-text-pdf/external/pubs/ft/wp/2014/_wp14236.pdf.md)
- [Structured JSON version](/-/media/websites/imf/imported-full-text-pdf/external/pubs/ft/wp/2014/_wp14236.pdf.json)

---

### Introduction and hypothesis
- Fact: developing countries have a thick left tail in the size distribution of manufacturing establishments relative to developed countries.
- India (2005-06): about 60 percent of workers are employed by establishments of size less than five.
- US (2006): corresponding number is less than 2 percent.
- Across Indian states (2005-06, 15 largest states covering 96.5 percent of manufacturing workforce):
  - Poorest states: almost 90 percent of manufacturing workforce employed in establishments of size five or less.
  - Richest states: about 40 percent of workforce in establishments of size five or less.
  - Average plant size in the richest states is about two times the average in the poorest states.
- Cross-state regression (share of employment in plants of size five or less on log per-capita state NDP):
  - Coefficient (15 states): −0.320 (0.0553).
  - Coefficient (all states): −0.319 (0.0568).
  - Controlling for 2-digit industry composition: −0.274 (0.0417).
- Hypothesis: demand-side differences driven by income and non-homothetic preferences over quality explain a large part of the size-income relation:
  - Poor households demand low quality products produced efficiently in small establishments with small fixed investments.
  - Rich households demand higher quality goods requiring larger fixed investments and larger establishments.
  - As income rises, demand shifts to higher quality → production shifts to larger plants → share employment in small plants falls.

### Consumer-side empirical evidence
- Data: Consumer Expenditure Survey 2004-05 (NSS), about 125,000 households; prices computed for 209 goods (188 goods used in regression).
- Regression ln(P_h,g) = α_g,state,rural + β ln(c_h) + ε_h,g (c_h = per-capita expenditure excluding durables).
- Key estimates (Table 1, dependent variable log(price)):
  - β = 0.112*** (Column 1), 0.111*** (Column 2), 0.106*** (Column 3).
  - Price Ratio (75th to 25th %tile): 1.091 (Column 1), 1.090 (Column 2), 1.086 (Column 3).
  - Price Ratio (95th to 5th %tile): 1.249 (Column 1), 1.246 (Column 2), 1.234 (Column 3).
  - Observations: 5,348,463.
  - Number of products: 188.
  - Clusters: 124,635.
- Interpretation: richer households tend to pay higher unit price for the same good, consistent with buying higher-quality products.
- Robustness:
  - Winsorize 1% (Column 2) and excluding own-good expenditure (Column 3) produce similar β.
  - Kernel-smoothed residual regressions fit a constant elasticity well (Figure 3).
- Opportunity cost of time checks (2003 survey):
  - Controlling for “non-worker present” indicator leaves coefficient on log(per-capita expenditure) essentially unchanged.
  - Selected Table 2 coefficients:
    - log(per-capita expenditure): 0.102***, 0.094***, 0.105***, 0.104***, 0.099*** across columns.
    - non-worker present: 0.020***, −0.059***, 0.017***, −0.065** in various columns.
    - (non-worker present)*pce: 0.012***, 0.011** in some columns.
    - Observations: 1,822,762 (Columns 1-3); 219,390 (Columns 4-6).
    - Number of products: 169; Clusters: 41,013 and 6,161.

### Producer-side empirical evidence (output prices, inputs, and skills)
- Data sources:
  - ASI covers plants employing ten or more workers (twenty or more if not using power).
  - SUM covers plants employing less than ten workers.
  - US County Business Patterns Database (2006) used for tradability concordance.
- Output price regressions (ln(P_f,g) on ln(L_f), product and state×rural fixed effects):
  - ASI only: log(labor) = 0.096*** (0.0087)
    - Price Ratio (Size 50 to 5): 1.247
    - Price Ratio (Size 500 to 5): 1.556
    - Observations: 46,704; Number of products: 1,217; Clusters: 1,078
  - SUM only: log(labor) = 0.053*** (0.0192)
    - Price Ratio (Size 50 to 5): 1.130
    - Price Ratio (Size 500 to 5): 1.276
    - Observations: 28,457; Number of products: 2,739; Clusters: 2,731
  - Combined: log(labor) = 0.106*** (0.0133)
    - Price Ratio (Size 50 to 5): 1.276
    - Price Ratio (Size 500 to 5): 1.629
    - Observations: 75,161; Number of products: 3,181; Clusters: 3,042
- Input price regressions (ln(P_f,i) on ln(L_f)):
  - ASI only: log(labor) = 0.077*** (0.0072)
    - Price Ratio (Size 50 to 5): 1.194
    - Price Ratio (Size 500 to 5): 1.426
    - Observations: 107,325; Number of products: 2,189; Clusters: 1,502
  - SUM only: log(labor) = 0.033* (0.0193)
    - Price Ratio (Size 50 to 5): 1.079
    - Price Ratio (Size 500 to 5): 1.164
    - Observations: 105,422; Number of products: 4,316; Clusters: 4,241
  - Combined: log(labor) = 0.050*** (0.0104)
    - Price Ratio (Size 50 to 5): 1.122
    - Price Ratio (Size 500 to 5): 1.259
    - Observations: 212,747; Number of products: 5,257; Clusters: 4,569
  - Notes: units misreporting in ASI inputs addressed; winsorize 1% tails of prices and plant size.
- Worker education by establishment size (Employment-Unemployment Survey 2004-05, ~600,000 individuals) — shares by education within size categories:
  - L ≤ 5: No School 0.43; Grade 1 to 9 0.41; Grade 10 to 12 0.13; > Grade 12 0.03
  - 5 < L ≤ 10: No School 0.34; Grade 1 to 9 0.41; Grade 10 to 12 0.17; > Grade 12 0.08
  - 10 < L ≤ 20: No School 0.33; Grade 1 to 9 0.41; Grade 10 to 12 0.16; > Grade 12 0.10
  - L > 20: No School 0.23; Grade 1 to 9 0.32; Grade 10 to 12 0.22; > Grade 12 0.22
- Major takeaway: larger plants charge higher output prices, pay higher prices for inputs, and employ more educated workers—consistent with producing higher-quality goods.

### Model structure (summary)
- Households:
  - Mass L of households; share h skilled earning w_S, share 1−h unskilled earning w_U; w_U normalized to 1.
  - Discrete quality set Q = {q1,...,qN}, qn > qm for n>m.
  - Utility: u_{j,qn}(c_{j,qn}, ε_{j,qn}) = a_{qn} + qn log(c_{j,qn}) + ε_{j,qn}.
  - Random utility ε i.i.d. Gumbel Type 1 → logit choice:
    - ρ(qn | w) = e^{a_{qn}} (w / P_{qn})^{qn} / ∑_{i=1}^N e^{a_{qi}} (w / P_{qi})^{qi}.
  - Non-homotheticity via complementarity between quality and quantity: richer households more likely to choose higher quality.
- Final goods producers:
  - N competitive final goods producers, each with M_{qn} horizontally differentiated varieties.
  - CES aggregator with scaling 1 / M_{qn}^{1/(σ−1)} so baseline final prices independent of M_{qn}.
  - Demand for variety i of quality qn: x_{i,qn} (Equation 5).
- Intermediate producers:
  - Technology combining unskilled and skilled labor with parameter θ_{qn} and elasticity σ_su (Equation 6).
  - Marginal cost κ(A_{i,qn}) decreasing in idiosyncratic productivity A_{i,qn}.
  - Pricing p(A_{i,qn}) = (σ / (σ−1)) κ(A_{i,qn}).
  - Entry requires f_{qn} units of labor with α_{qn} share skilled; entrants draw log(A) ∼ N(μ_{qn}, ν^2).
  - Free entry condition equates expected profits to entry labor cost (Equation 8); M_{qn} adjusts so free entry holds.
- Equilibrium defined by consumers’ choices, firms’ demands and pricing, free entry, and market clearing for goods and labor.

### Calibration choices and targets
- Production-side calibration:
  - Definition: individual with less than ten years of education = unskilled.
  - h = 0.24 (share skilled).
  - σ_su = 1.75.
  - σ = 5 ⇒ markup of 25 percent for intermediate producers.
  - N = 12 quality levels.
  - Fixed costs f_{qn} set so average employment across qualities: size_qn = {1.25, 2.5, 5, ..., 2560} (each higher quality doubles average size).
  - θ_{qn} chosen to match wage premium target w_S = 1.6 and eleven relative unskilled-to-skilled ratios from Employment-Unemployment Survey.
  - μ_{qn} chosen to match price-size slope of 0.1 (each higher quality charges log price 0.1 * log(2) higher than previous).
  - α_{qn} chosen to match entry skill shares to production skill shares.
  - ν^2 chosen to match standard deviation of log employment = 0.64.
- Utility-side calibration:
  - q1 = 1 and qn = q_{n−1} + Δ; Δ chosen so price-income elasticity is 0.1: average log price for skilled households is 0.1 * log(w_S / w_U) more than unskilled.
  - a_{qn} chosen so model size distribution matches India 2005-06 manufacturing size distribution (combined ASI and SUM).

### Counterfactual experiments and quantitative results
- Counterfactuals vary productivity μ_{qn}, share skilled h, and θ_{qn} (adjusted to keep wage premia unchanged) to simulate per-capita income variation across Indian states while holding wage premia and relative prices of qualities unchanged.
- Cross-section of Indian states (baseline model):
  - Baseline model share of employment in small plants (size ≤ 5): 63.9 percent.
  - Poorer counterfactual (per-capita income 0.39 times baseline): share in small plants = 75.6 percent.
  - Richer counterfactual (per-capita income 1.57 times baseline): share in small plants = 56.3 percent.
  - Data: poorest states ≈ 91.9 percent employment in small plants; richest ≈ 47.2 percent.
    - Data difference across states = 44.7 percentage points.
    - Model predicts a 19.3 percentage points difference across the same income variation.
    - Model explains about 43 percent of the cross-state difference (19.3 percentage points / 44.7 percentage points).
  - Pooled three poorest vs three richest states:
    - Data difference ≈ 36 percentage points.
    - Model predicts ~15 percentage points difference → about 42 percent explained.
- India over time (1989–2009):
  - Data: share in small plants fell from 77 percent (1989) to 58 percent (2009) — decline ≈ 19 percentage points.
  - Per-capita income and skills:
    - 1989 per-capita income = 0.54 times 2005 level; share with ≥10 years schooling = 14 percent.
    - 2009 per-capita income = 1.30 times 2005 level; share with ≥10 years schooling = 31 percent.
  - Model predicts share in small plants = 72 percent for 1989 (data 77 percent).
  - Overall, model explains about 65 percent of the historical change between 1989 and 2009.
- Love of variety sensitivity:
  - Final goods aggregator generalized with love of variety parameter η.
  - Table 8 percent of cross-state difference explained:
    - q1 = 1, η = 1/(σ−1): 43.1%
    - q1 = 0.1, η = 1/(σ−1): 43.1%
    - q1 = 1, η = 0: 71.2%
    - q1 = 0.1, η = 0: 53.1%
  - Interpretation: allowing full love of variety (η = 0) amplifies counterfactual changes; baseline (η = 1/(σ−1)) explains 43.1 percent.
- India vs US counterfactual:
  - Simulate per-capita income 17 times India (US 2005 level).
  - Result: share of employment in plants with 5 or less falls from 64 percent (calibrated baseline) to 13 percent in the US counterfactual.
  - Caveat: calibration local to India; extrapolation to US may be biased.

### Tradability and inter-state trade checks
- Two tradability measures at 3-digit industry level:
  1. Herfindahl index of geographical concentration in the US (H-index) mapped from 6-digit NAICS to 3-digit NIC via concordance.
  2. Export-Import index: (exports + imports) as share of gross production for that industry in India (HS 2002 → ISIC Rev 3 concordance).
- Measures weakly positively correlated: rank correlation coefficient = 0.25.
- Regression sd_{i,s,t} = α_{i,t} + α_{s,t} + γ ln(SNDP_{s,t}) * tradability_i + ε_{i,s,t} (sd = share employment in plants ≤ 5).
  - Positive γ implies relation is stronger for non-tradables.
- Table 9 interaction coefficients (selected):
  - Column 1 (H-index, median cutoff): γ = 0.068 (s.e. 0.0351) — positive, marginally significant (*p<0.1).
  - Column 2 (H-index, quartile cutoff): γ = 0.052 (s.e. 0.0394).
  - Column 3 (Exp-Imp index, median cutoff): γ = −0.010 (s.e. 0.0469).
  - Column 4 (Exp-Imp index, quartile cutoff): γ = 0.000 (s.e. 0.0498).
  - Observations: Column 1: 3,885; Column 2: 1,826; Column 3: 3,899; Column 4: 1,959.
  - Notes: data from ASI and SUM (1989,1994,2000,2005,2010); industry×time and state×time fixed effects; observations weighted; standard errors clustered at state level.
- Conclusion: size-income relation across states is not stronger for tradables; inter-state trade unlikely to be dominant driver of cross-state size-income relation.

### Units misreporting in ASI and corrections
- Problem: some ASI product quantities reported in inconsistent units (e.g., liters vs kiloliters), producing large computed price differences (example: milk log-price gap ≈ 7 log points → exp(7) = 1,096).
- Manual correction:
  - Reviewed ~1,000 product categories; split affected products into two categories using price cutoffs.
  - Effect on price-size elasticity (Table A.8):
    - ASI corrected: log(labor) = 0.096*** (Observations 46,704; Number of products 1,217).
    - ASI not corrected: log(labor) = 0.155*** (same observations; Number of products 1,077).
    - ASI+SUM corrected: log(labor) = 0.106*** (Observations 75,161; Number of products 3,181).
    - ASI+SUM not corrected: log(labor) = 0.125*** (same observations; Number of products 3,041).
- Algorithmic detection alternative:
  - Stepwise rules comparing max/min price ratios, jumps of factor ≥ 20, R-square comparisons.
  - Algorithmic elasticity in ASI = 0.1037; varying R-square threshold to 0.8 and 0.7 yields 0.1091 and 0.0986 respectively.
- Finding: correcting units reduces estimated price elasticity with respect to employment, indicating misreporting correlated with size.

### Data sources and key survey details
- Surveys used:
  - ASI waves: 1989-90, 1994-95, 2000-01, 2005-06, 2009-10.
  - SUM waves: 1989-90, 1994-95, 2000-01, 2005-06, 2010-11.
  - Consumer Expenditure Surveys: 2003 (59th round) and 2004-05 (61st round).
  - Employment-Unemployment Survey: 2004-05 (61st round).
- ASI coverage: factories registered under Sections 2m(i) and 2m(ii) of Factories Act, 1948; primary ASI used = 2005-06.
- SUM coverage: unregistered manufacturing enterprises; primary SUM used = 2005-06 (62nd Round).
- CES: 2004-05 interviewed about 125,000 households; prices computed for 209 goods (156 food items, 10 fuel and light, 24 clothing and footwear, remainder durables).
- Employment-Unemployment Survey: ~125,000 households (about 600,000 individuals); education and establishment size used to define skilled (≥ Grade ten) and calibration of wage premium.
- Wage premium from Mincerian regression (Table A.4):
  - Skilled dummy coefficient = 0.450 (Column 1) and 0.445 (Column 2).
  - Implied wage premium = 1.568 and 1.560 (rounded to 1.6 in calibration).
  - Observations: 11,003.

### Key quantitative summary statistics (preserve exact figures)
- Consumer β estimates: 0.112***, 0.111***, 0.106***; Observations: 5,348,463; Products: 188; Clusters: 124,635.
- Price Ratios (consumer): 75th/25th = 1.091, 1.090, 1.086; 95th/5th = 1.249, 1.246, 1.234.
- Output price-size elasticities: ASI 0.096***; SUM 0.053***; Combined 0.106***.
- Input price-size elasticities: ASI 0.077***; SUM 0.033*; Combined 0.050***.
- Employment education shares (L ≤ 5 vs L > 20): > Grade 12 shares 0.03 vs 0.22 respectively.
- Calibration targets:
  - h = 0.24; σ_su = 1.75; σ = 5; N = 12.
  - size_qn = {1.25,2.5,5,...,2560}.
  - Price-size slope target = 0.1.
  - Std dev of log employment target = 0.64.
  - Wage premium target w_S = 1.6.
- Cross-state model result: model accounts for 19.3 percentage point fall in share of employment in plants of size five or less across income variation, explaining about 43 percent of observed cross-state difference (data difference = 44.7 percentage points).
- Historical result: model explains about 65 percent of the 1989–2009 change in share of employment in small plants.
- Love-of-variety sensitivity: baseline explains 43.1 percent; allowing η = 0 and q1 = 1 explains 71.2 percent.
- India→US counterfactual: share in plants ≤ 5 falls from 64 percent to 13 percent.

### Policy-relevant implications (drawn from model and evidence)
- A significant fraction of the missing-middle (left tail of establishment size distribution) can be explained by low income and non-homothetic demand for quality rather than solely by size-dependent policies or regulatory distortions.
- Model suggests substantial part of cross-state and historical differences in plant-size distributions are a natural consequence of income and skill accumulation:
  - Counterfactuals indicate raising incomes and skill supply shifts production toward higher-quality, larger plants, reducing employment share in very small establishments.
- Sensitivity to love of variety implies that policies affecting variety and entry at higher-quality segments can amplify structural changes in size distribution.
- Inter-state trade is unlikely to be the dominant explanation for cross-state size-income patterns based on two tradability measures and regression evidence.

*Source: _wp14236 - References (PDF).*

### References .............................................................................................................

### _wp14236 - References

### Introduction
- Cross-country and cross-state fact: developing countries have a thick left tail in the size distribution of manufacturing establishments relative to developed countries.
  - India (2005-06): about 60 percent of workers are employed by establishments of size less than five.
  - US (2006): corresponding number is less than 2 percent.
- Across Indian states (2005-06, 15 largest states covering 96.5 percent of manufacturing workforce):
  - Poorest states: almost 90 percent of manufacturing workforce employed in establishments of size five or less.
  - Richest states: about 40 percent of workforce in establishments of size five or less.
  - Average plant size in the richest states is about two times the average in the poorest states.
- Cross-state regression evidence:
  - Regressing share of employment in plants of size five or less on log of per-capita state NDP:
    - Coefficient (standard error) restricting to 15 states: -0.320 (0.0553).
    - Coefficient (standard error) including all states: -0.319 (0.0568).
    - Controlling for industry composition at the 2-digit level (weighting by all India industry composition) changes slope to -0.274 (0.0417).
- Hypothesis: demand-side differences driven by income and non-homothetic preferences over quality explain a large part of the size-income relation:
  - Poor households have high demand for low quality products, efficiently produced in small establishments that require small fixed investments.
  - Rich households demand higher quality goods, requiring larger fixed investments and larger establishments.
  - As income rises, demand shifts to higher quality products → production shifts to higher quality and larger plants → share employment in small plants falls.

### Empirical Findings (Consumer and Producer Evidence)
- Consumer-side evidence (from Consumer Expenditure Surveys, 2004-05):
  - Richer households tend to pay a higher unit price for the same good, consistent with richer households buying higher quality products.
  - Table 1 regression (Dependent Variable: log(price)):
    - Coefficient on log(per-capita expenditure): 0.112*** (Column 1), 0.111*** (Column 2), 0.106*** (Column 3).
    - Price Ratio (75th to 25th %tile): 1.091 (Column 1), 1.090 (Column 2), 1.086 (Column 3).
    - Price Ratio (95th to 5th %tile): 1.249 (Column 1), 1.246 (Column 2), 1.234 (Column 3).
    - Winsorize 1%: Y in Column 2.
    - Exclude product from RHS: Y in Column 3.
    - Observations: 5,348,463.
    - Number of products: 188.
    - Clusters: 124,635.
    - Notes: regressions include fixed effects for the interaction of each good, state, rural-urban cell. Standard errors clustered at the household level. ***p<0.01.
- Producer-side evidence (ASI and SUM):
  - Larger plants charge a higher unit price for the same good compared to smaller plants; relation holds within formal sector (ASI) and when pooling formal and informal (SUM) plants.
  - Larger plants use higher price material inputs, consistent with higher quality inputs.
  - Household surveys show larger plants hire more skilled workers.
  - The data sources:
    - ASI covers plants employing ten or more workers (twenty or more if not using power).
    - SUM covers plants employing less than ten workers.
    - US data from County Business Patterns Database (2006).
- Tradability and inter-state trade:
  - Constructed two measures of tradability at the 3-digit industrial level.
  - The size-income relation across states is not stronger for tradables compared to non-tradables (for one measure, non-tradables have a stronger negative relation), suggesting inter-state trade is unlikely to be the dominant force behind the cross-state size-income relation.

### Model Structure and Calibration
- General equilibrium model setup:
  - Households choose from a finite number of discrete quality levels; choice modeled as a discrete-choice problem.
  - Preferences are non-homothetic with respect to quality due to complementarity between quality and quantity: marginal utility from additional quantity is larger for higher quality goods; richer households can consume more quantity and thus are more likely to choose higher quality.
  - Production: higher quality goods use skilled labor more intensively and require higher fixed costs; free entry implies higher quality producers are larger on average.
- Calibration strategy:
  - Model parameters chosen to match micro-facts documented in consumer and producer surveys.
  - Quality-size relation on producer side matched to relation between prices and plant size from producer surveys.
  - Degree of non-homotheticity chosen to match the price-income relation seen in consumer surveys.

### Counterfactual Exercises and Quantitative Results
- Simulated income variation to match cross-state income variation:
  - As income increases in the model (via productivity and skill level changes), demand shifts to higher quality goods, reducing low quality producers and increasing high quality producers; average producer size rises.
  - Quantitative results:
    - Share of employment in plants of size five or less falls by 19.3 percentage points in the model when income varies by the same extent as across Indian states.
    - This 19.3 percentage point fall accounts for about 43 percent of the cross-state difference observed in the data.
- Historical comparison for India (1989 to 2009):
  - Empirical decline in share of employment in plants of size five or less: about 20 percentage points.
  - The model can explain about 65 percent of this historical change.
- Extensions:
  - Section 5.1 explores implications of the model for the entire establishment size distribution.
  - Section 6 discusses inter-state trade considerations in greater detail.

### Relation to Literature
- Contrasts with literature focusing on size-dependent policies and regulatory distortions (e.g., Little, Mazumdar, and Page Jr (1987); De Soto (1989); Loayza (1996); Djankov and others (2002); Loayza, Oviedo, and Serven (2005); Loayza, Serven, and Sugawara (2009); Garicano, LeLarge, and Van Reenen (2013)).
  - Notes that distortion-based explanations are unlikely to account for all cross-country/state differences; Tybout (2000) and Hsieh and Klenow (2012) provide evidence that policies alone do not explain missing middle.
- Aligns with views of dual-sector informal/formal segmentation (La Porta and Shleifer (2008); Banerjee and Duflo (2011)) but emphasizes heterogeneity in quality and non-homothetic demand.
- Empirical parallels:
  - Deaton and Dupriez (2011) and Dikhanov (2010) document richer Indian households buy higher price goods.
  - Kugler and Verhoogen (2012) find larger Colombian plants produce higher price goods and use higher price inputs; this paper documents analogous facts for India including very small plants.
  - Fajgelbaum, Grossman, and Helpman (2011) and others develop non-homothetic preference models with respect to quality; this paper’s non-homotheticity arises from complementarity between quantity and quality.

*Source: _wp14236 - References*

### 3.  Larger plants use higher price material inputs and hire more skilled labor

### 3.  Larger plants use higher price material inputs and hire more skilled labor

### Richer households buy higher price goods
- Data: Consumer Expenditure Survey of 2004-05 (NSS), about 125,000 households, prices computed for 209 goods (188 goods used in regression).
- Regression specification:
  - ln(P_h,g) = α_g,state,rural + β ln(c_h) + ε_h,g
  - c_h is per-capita expenditure excluding durables.
- Key estimates and implications:
  - β = 0.112 (Column 1, Table 1).
  - This implies the average price paid by the 95th percentile household is 24.9 percent more than the price paid by the 5th percentile household (the 95th percentile household’s per-capita expenditure is about seven times that of the 5th percentile household).
  - Winsorizing 1 percent tails for per-capita expenditure and prices does not change results substantially (Column 2).
  - Using c_h with the good’s own expenditure subtracted (log(c_h − P_h,g Q_h,g / household size)) produces very similar results (Column 3).
- Non-parametric evidence:
  - Kernel-smoothed local linear regression of residualized log prices on residualized log per-capita expenditures fits a constant elasticity well (Figure 3). Epanechnikov kernel, bandwidth 0.13; top and bottom 1 percent of residualized log per-capita expenditure excluded.
- Opportunity cost of time (2003 survey):
  - Regressions controlling for a “non-worker present” indicator (proxy for low opportunity cost of time) show the coefficient on log(per-capita expenditure) remains essentially unchanged.
  - Table 2 highlights estimated coefficients (selected):
    - log(per-capita expenditure): 0.102*** (Columns 1 and 2), 0.094*** (Column 3), 0.105*** (Column 4), 0.104*** (Column 5), 0.099*** (Column 6).
    - non-worker present: 0.020*** (Column 2), -0.059*** (Column 3), 0.017*** (Column 5), -0.065** (Column 6).
    - (non-worker present)*pce: 0.012*** (Column 3), 0.011** (Column 6).
    - Observations: 1,822,762 (Columns 1-3); 219,390 (Columns 4-6).
    - Number of products: 169; Clusters: 41,013 (Columns 1-3), 6,161 (Columns 4-6).

Major takeaway: richer households tend to pay higher unit prices for the same goods, consistent with higher-quality consumption rather than being driven away by opportunity cost of time.

### Larger plants produce higher price goods
- Data: Annual Survey of Industries (ASI) 2005-06 and Survey of Unorganized Manufacturing (SUM) 2005-06; samples combined to represent manufacturing sector.
- Regression specification:
  - ln(P_f,g) = α_g + α_state,rural + γ ln(L_f) + ε_f,g
  - L_f is number of workers at plant f; α_g product fixed effect; α_state,rural state × urban-rural fixed effect.
- Table 3 estimates (log(output price) on log(labor)):
  - ASI only (Column 1): log(labor) = 0.096*** (0.0087)
    - Price Ratio (Size 50 to 5): 1.247
    - Price Ratio (Size 500 to 5): 1.556
    - Observations: 46,704; Number of products: 1,217; Number of clusters: 1,078
  - SUM only (Column 2): log(labor) = 0.053*** (0.0192)
    - Price Ratio (Size 50 to 5): 1.130
    - Price Ratio (Size 500 to 5): 1.276
    - Observations: 28,457; Number of products: 2,739; Number of clusters: 2,731
  - Combined (Column 3): log(labor) = 0.106*** (0.0133)
    - Price Ratio (Size 50 to 5): 1.276
    - Price Ratio (Size 500 to 5): 1.629
    - Observations: 75,161; Number of products: 3,181; Number of clusters: 3,042
- Non-parametric evidence:
  - Kernel-smoothed local linear regression of residualized log prices on residualized log employment (Epanechnikov kernel, bandwidth 0.502; top and bottom 1 percent excluded) shows near log-linear relation (Figure 4).

Major takeaway: larger plants charge higher prices for their outputs—consistent with larger plants producing higher-quality goods.

### Larger plants use higher price inputs and hire more skilled labor
- Relation between plant size and input prices:
  - Regression specification:
    - ln(P_f,i) = α_i + α_state,rural + γ ln(L_f) + ε_f,i
    - P_f,i is price paid by plant f for input i; α_i input fixed effect.
  - Table 4 estimates (log(input price) on log(labor)):
    - ASI only (Column 1): log(labor) = 0.077*** (0.0072)
      - Price Ratio (Size 50 to 5): 1.194
      - Price Ratio (Size 500 to 5): 1.426
      - Observations: 107,325; Number of products: 2,189; Number of clusters: 1,502
    - SUM only (Column 2): log(labor) = 0.033* (0.0193)
      - Price Ratio (Size 50 to 5): 1.079
      - Price Ratio (Size 500 to 5): 1.164
      - Observations: 105,422; Number of products: 4,316; Number of clusters: 4,241
    - Combined (Column 3): log(labor) = 0.050*** (0.0104)
      - Price Ratio (Size 50 to 5): 1.122
      - Price Ratio (Size 500 to 5): 1.259
      - Observations: 212,747; Number of products: 5,257; Number of clusters: 4,569
  - Notes: units misreporting in ASI inputs addressed as in outputs; winsorize 1% tails of prices and plant size.
- Plant size and worker education (Employment-Unemployment Survey 2004-05, ~600,000 individuals):
  - Table 5: share of individuals by establishment size category and education level (columns sum to 1 across each row):
    - L <= 5: No School 0.43; Grade 1 to 9 0.41; Grade 10 to 12 0.13; > Grade 12 0.03
    - 5 < L <= 10: No School 0.34; Grade 1 to 9 0.41; Grade 10 to 12 0.17; > Grade 12 0.08
    - 10 < L <= 20: No School 0.33; Grade 1 to 9 0.41; Grade 10 to 12 0.16; > Grade 12 0.10
    - L > 20: No School 0.23; Grade 1 to 9 0.32; Grade 10 to 12 0.22; > Grade 12 0.22
  - Interpretation: a larger share of workers in bigger establishments have higher education; e.g., 22 percent of workers in establishments of size more than 20 have graduated high school, versus 3 percent in establishments of size less than 6.

Major takeaway: larger plants both pay higher prices for material inputs and employ more educated workers—consistent with larger plants producing higher-quality goods that require higher-quality inputs and more skilled labor.

*Source: _wp14236 - 3.  Larger plants use higher price material inputs and hire more skilled labor*

### 3.    Model

### 3.    Model

### 3.1. Households
- Population and labor:
  - Mass L of households indexed by j.
  - Share h are skilled and earn wage w_S (determined endogenously); share 1−h are unskilled and earn wage w_U.
  - Unskilled wage w_U is the numeraire and is normalized to 1.
- Quality space:
  - There are N quality levels. Q = {q1, q2, ..., qN} with qn > qm ∀ n > m; q1 is lowest quality, qN highest quality.
- Utility and choice:
  - Utility from consuming quality qn by household j:
    - u_{j,qn}(c_{j,qn}, ε_{j,qn}) = a_{qn} + qn log(c_{j,qn}) + ε_{j,qn} ∀ qn ∈ Q.
  - Random utility ε_{j,qn} i.i.d. Gumbel Type 1 with density f(ε) = e^{-ε} e^{e^{-ε}}.
  - Households choose exactly one quality level and spend entire income on it, yielding indirect utility:
    - v_{j,qn}(w_j, P_{qn}, ε_{j,qn}) = a_{qn} + qn log(w_j / P_{qn}) + ε_{j,qn} ∀ qn ∈ Q.
  - Choice probabilities (logit form):
    - ρ(qn | w) = exp(a_{qn} + qn log(w / P_{qn})) / ∑_{i=1}^N exp(a_{qi} + qi log(w / P_{qi}))
    - Equivalent form: ρ(qn | w) = e^{a_{qn}} (w / P_{qn})^{qn} / ∑_{i=1}^N e^{a_{qi}} (w / P_{qi})^{qi}.
  - Elasticity of ρ(qn | w) w.r.t. wage w:
    - γ_{ρ(qn),w} = ∂ log[ρ(qn | w)] / ∂ log(w) = qn − ∑_{i=1}^N qi ρ(qi | w).
  - Implications:
    - Non-homotheticity operates on the extensive margin: as wages increase, richer households are more likely to choose higher quality goods (a positively sloped “quality Engel curve”).
    - Lowest quality share always has negative elasticity; highest quality share always has positive elasticity.
- Parameterization example:
  - Quality indexes can be parameterized with q1 = 1 and qn = q_{n−1} + Δ; Δ determines steepness of the quality Engel curve.
  - Illustrative N = 2 example used in Figure 5:
    - P_{q1} = 1, P_{q2} = 1.5, q1 = 1, q2 = 1 + Δ.
    - For each Δ, a_{q2} chosen so 30 percent of households with w = 1 choose high quality q2.
    - If Δ = 0, no change in share with wage; for Δ > 0, share buying high quality increases with wage, more so for larger Δ.
- Aggregate demand for quality qn given skilled and unskilled wages:
  - C_{qn} = N h ρ(qn | w_S) (w_S / P_{qn}) + N (1−h) ρ(qn | w_U) (w_U / P_{qn}) ∀ qn ∈ Q.
  - First term: demand from skilled households; second term: demand from unskilled households.
- Summary point:
  - Complementarity between quantity and quality in preferences generates richer households buying higher price (higher quality) goods, matching empirical patterns.

### 3.2. Final Goods Producers
- Market structure:
  - N competitive final goods producers, one per quality level qn.
  - Within each quality qn there is horizontal differentiation: M_{qn} varieties (plants), indexed by i.
- Production technology (CES over varieties with scaled love-of-variety factor):
  - Y^s_{qn} = (1 / M_{qn}^{1/(σ−1)}) (∑_{i=1}^{M_{qn}} x_{i,qn}^{(σ−1)/σ})^{σ/(σ−1)} ∀ qn ∈ Q.
  - σ is elasticity of substitution between varieties of same quality.
  - The multiplicative factor 1 / M_{qn}^{1/(σ−1)} scales out love of variety so final prices P_{qn} are independent of M_{qn} under baseline.
- Cost minimization yields demand for intermediate variety i of quality qn:
  - x_{i,qn} = p_{i,qn}^{−σ} M_{qn}^{1/(σ−1)} Y^s_{qn} (∑_{i=1}^{M_{qn}} p_{i,qn}^{1−σ})^{(σ−1)/(1−σ)} ∀ qn ∈ Q. (Equation 5)
- Final producers earn zero profits and price charged to consumers:
  - P_{qn} = ∑_{i=1}^{M_{qn}} p_{i,qn} x_{i,qn} / Y^s_{qn} ∀ qn ∈ Q.
- Modeling choices:
  - Baseline assumes no love of variety (conservative choice; Section 5.3 explores allowing love of variety and sensitivity to q1).

### 3.3. Intermediate Goods Producers
- Production technology combining unskilled and skilled labor:
  - x(A_{i,qn}) = A_{i,qn} [ θ_{qn} (l_{U,i,qn})^{(σ_su−1)/σ_su} + (1−θ_{qn}) (l_{S,i,qn})^{(σ_su−1)/σ_su} ]^{σ_su/(σ_su−1)}. (Equation 6)
  - l_{U,i,qn} = unskilled labor hired by variety i of quality qn.
  - l_{S,i,qn} = skilled labor hired by variety i of quality qn.
  - σ_su = elasticity of substitution between skilled and unskilled labor.
  - A_{i,qn} = idiosyncratic productivity of variety i, quality qn. θ_{qn} is unskilled share parameter by quality.
- Marginal cost:
  - κ(A_{i,qn}) = (1 / A_{i,qn}) [ θ_{qn}^{σ_su} (1 / w_U)^{σ_su−1} + (1−θ_{qn})^{σ_su} (1 / w_S)^{σ_su−1} ]^{1/(σ_su−1)}.
  - Marginal cost is decreasing in A_{i,qn} and depends on skilled and unskilled wages.
- Pricing under monopolistic competition:
  - p(A_{i,qn}) = (σ / (σ−1)) κ(A_{i,qn}). (Equation 7)
- Entry and fixed costs:
  - To start an intermediate plant of quality qn requires f_{qn} units of labor; share α_{qn} of entry labor must be skilled.
  - Entrants receive productivity draw log(A_{i,qn}) ∼ N(μ_{qn}, ν^2).
- Free entry condition:
  - α_{qn} f_{qn} w_S + (1−α_{qn}) f_{qn} w_U = ∫ π(A_{i,qn}) g_{qn}(A_{i}) dA_{i} ∀ qn ∈ Q. (Equation 8)
  - π(A_{i,qn}) = [ p(A_{i,qn}) − κ(A_{i,qn}) ] x(A_{i,qn}) is flow profit.
  - M_{qn} adjusts so free entry holds. Higher f_{qn} implies larger average scale x(A_{i,qn}) for that quality in equilibrium.
  - If θ_{qn} > θ_{qm} ∀ n > m, higher quality producers use skilled labor more intensively and have higher costs.

### 3.4. Equilibrium
- An equilibrium is a set of prices (w_S, {p_{i,qn}}_{i∈M_{qn}}, P_{qn}), allocations ({c_{j,qn}}_{j∈L}, C_{qn}, {x_{i,qn}}_{i∈M_{qn}}, Y_{qn}), and masses M_{qn} such that:
  - Consumers choose optimal quality levels (equations 3 and 4 hold).
  - Final goods producers demand intermediate goods optimally (equation 5).
  - Intermediate producers maximize profits and charge markup price (equation 7).
  - Free entry conditions hold for all quality levels (equation 8).
  - Markets clear:
    - Y_{qn} = C_{qn} ∀ qn ∈ Q.
    - L(1−h) = ∑_{qn} M_{qn} ∫ l_{U}(A_{i,qn}) g_{qn}(A_{i}) dA_{i} + ∑_{qn} M_{qn} (1−α_{qn}) f_{qn}.
    - L h = ∑_{qn} M_{qn} ∫ l_{S}(A_{i,qn}) g_{qn}(A_{i}) dA_{i} + ∑_{qn} M_{qn} α_{qn} f_{qn}.
  - Last two equations are labor market clearing for unskilled and skilled labor respectively.

### 4. Calibration (beginning)
- Objective:
  - Calibrate model to cross-sectional facts from Section 2 and Indian data; then simulate counterfactuals varying per-capita income to study effects on size distribution.
  - Key determinants of size-distribution changes: consumer non-homotheticity Δ and producer price-size relation.
- Production-side calibration choices described:
  - Definition: individual with less than ten years of education = unskilled.
  - h (share skilled) = 0.24 (share of manufacturing workers with ≥ ten years of education in India 2004–05).
  - σ_su (elasticity of substitution between skilled and unskilled) = 1.75.
  - σ (elasticity of substitution between varieties) = 5 ⇒ implies a markup over cost of 25 percent for intermediate producers.
- Parameters to calibrate on production side (five sets):
  - f_{qn} fixed cost by quality.
  - θ_{qn} unskilled share in production by quality.
  - μ_{qn} mean of log productivity draw by quality.
  - α_{qn} share of skilled labor required for entry by quality.
  - ν^2 variance of productivity draw (common across qualities).
- Additional calibration notes:
  - N (number of quality levels) is set to 12.
  - Fixed costs f_{qn} determine average scale of intermediate producers for each quality level.

*Source: _wp14236 - 3.    Model*

### Section 2.2, larger plants tend to produce higher price products, which is indicative of higher quality goods

### Section 2.2, larger plants tend to produce higher price products, which is indicative of higher quality goods

### Calibration of producer-side parameters
- Fixed costs are chosen so that average employment (skilled plus unskilled) in intermediate producers of the lowest quality levels is 1.25 workers and each higher quality level has double the average size of the previous quality level; the average employment of intermediate producers of the different quality levels are:
  - size_qn = {1.25,2.5,5,...,2560}
- The θ′_qn levels determine demand for unskilled labor relative to skilled labor and inform the wage premium w_S.
- The ratio of skilled to unskilled workers in any quality level relative to the lowest quality is given by:
  - ratio_U,S_qn = (L_U_qn / L_S_qn) / (L_U_q1 / L_S_q1) = (θ_qn / (1−θ_qn))^{σ_us} / (θ_q1 / (1−θ_q1))^{σ_us}  ∀ qn ∈ Q. (equation (9))
- Twelve θ′_qn are chosen to match:
  - a target wage premium w_S = 1.6
  - eleven targets for unskilled-to-skilled ratios across quality levels relative to the lowest quality level.
- Targets for these moments are obtained from the Employment-Unemployment Survey (NSS) 2004-05.
- Table 6 (data) shows smaller plants have a much higher ratio of unskilled to skilled workers, indicating low quality producers have higher θ′_qn. Coarse size categories required extrapolation (with a minimum of 0.5) to compute eleven ratios for equation (9).
- μ_qn, the mean of the log of the productivity draw for each quality level, is chosen to match the price-size relation seen in Table 3.
  - Higher μ_qn → lower average price for that quality level (since p(A_i,qn) ∝ 1/A_i).
  - μ′_qn are chosen to match a price-size slope of 0.1. Specifically, each higher quality level charges a log price which is 0.1 * log(2) higher than the previous quality level's log price (given sizes double).
- α_qn, the share of skilled labor needed for entry for each quality level, is chosen to match the share of skilled labor used in production of that quality level; high quality producers use a more skill intensive production process (lower θ_qn) and have more skill intensive entry requirements.
- ν^2, the variance of the log of the productivity draw (common across qualities), is chosen to match the standard deviation of the log of employment in the combined ASI and SUM dataset:
  - Std dev of employment = 0.64

- Table 7: Calibration summary (parameter → target)
  - f_qn (Fixed costs) → size_qn = {1.25,2.5,5,...,2560}
  - μ_qn (Mean of productivity draws) → Price-size slope of 0.1
  - θ_qn (Share of U in production) → w_S = 1.6; and L_U_qn / L_S_qn across qualities
  - α_qn (Share of skilled in entry) → L_S_qn / (L_S_qn + L_U_qn) in production
  - ν^2 (Variance of productivity draw) → Std dev of employment = 0.64
  - q_n (∆) (Utility from quality) → Price-income slope of 0.1
  - a_qn (Constant in utility function) → Size distribution

### Utility-side parameters and demand calibration
- Utility function for quality qn:
  - u_{j,qn}(c_{j,qn}, ε_{j,qn}) = a_qn + qn * log(c_{j,qn}) + ε_{j,qn}  ∀ qn ∈ Q. (equation (10))
- Parameters to calibrate:
  1. qn, the quality indexes: q1 = 1 and qn = q_{n−1} + ∆.
     - ∆ determines steepness of the quality Engel curve (how quickly demand moves to higher quality as income increases).
     - ∆ chosen so the price-income elasticity in the model is 0.1: average log price paid by skilled households is 0.1 * log(w_S / w_U) more than for unskilled households.
     - w_S is calibrated to be 1.6; w_U is normalized to 1.
  2. a_qn, the quality-specific constant in the utility function:
     - a_qn pins down absolute levels of demand ρ(qn | w).
     - a_qn chosen so the model's size distribution matches India's manufacturing size distribution in 2005-06 (combined ASI and SUM).
- The model is calibrated so that the size distribution in the model matches the data closely (Figure 6), although it was not calibrated to match how the size distribution changes with income.

### Counterfactual exercises and main results
- Counterfactuals vary three sets of parameters while holding others fixed to simulate per-capita income variation across Indian states:
  1. Share of households who are skilled, h:
     - About 13 percent of manufacturing workers in poorest states are skilled vs 43 percent in richest states.
  2. θ_qn (share parameter of unskilled labor for intermediates) is changed across counterfactuals to keep wage premia unchanged (skill-biased technical change interpretation).
     - θ_q1 for each counterfactual is chosen to maintain wage premia w_S = 1.6; other θ′_qn picked via equation (9).
     - Entry skill shares α_qn are also adjusted to match production skill shares.
  3. μ_qn (mean productivity draws) changed to match differences in per-capita income across states while maintaining the price-size slope of 0.1.
     - Per-capita income of poorest state (Bihar) = 0.39 times India’s per-capita income (0.94 log points lower).
     - Per-capita income of richest state (Maharashtra) = 1.57 times India’s per-capita income (0.43 log points higher).
- Procedure ensures wage premia and relative prices of qualities remain unchanged across counterfactuals, isolating demand effects from income (non-homotheticity) on size distribution.

Key quantitative findings (cross-section of Indian states):
- Baseline model share of employment in small plants (size ≤ 5): 63.9 percent.
- When productivity and skill supply are lowered to make per-capita income 0.39 times baseline (poorer counterfactual): share in small plants increases to 75.6 percent.
- When productivity and skill supply are raised to make per-capita income 1.57 times baseline (richer counterfactual): share in small plants falls to 56.3 percent.
- Data: poorest Indian states ≈ 91.9 percent employment in small plants; richest ≈ 47.2 percent.
  - Data difference across states = 44.7 percentage points.
  - Model predicts a 19.3 percentage points difference across the same income variation.
  - Therefore, model explains about 43 percent of the difference in share of employment in small plants seen across Indian states.
- Across pooled three poorest vs three richest states (size distribution):
  - Data: poorest have about 36 percentage points more employment in plants of size five or less compared to richest.
  - Model predicts ~15 percentage points higher in poorer states for size ≤ 5, accounting for about 42 percent of the data difference.
- Mechanism described: higher real income (via higher productivity and skill supply) shifts demand toward higher-quality (higher-priced) goods due to non-homothetic preferences; production shifts to higher-quality producers and size distribution moves away from small plants (share in small plants falls).

India over time (1989–2009)
- SUM + ASI waves yield five data points for share of employment in plants of size ≤ 5 for years 1989, 1994, 2000, 2005, 2009.
- Data: share of employment in small plants fell from 77 percent in 1989 to 58 percent in 2009.
- Per-capita income and skills:
  - 1989 per-capita income = 0.54 times 2005 level; share of manufacturing workers with ≥10 years schooling = 14 percent.
  - 2009 per-capita income = 1.30 times 2005 level; share with ≥10 years schooling = 31 percent.
- Model predictions (when matching productivity and skill changes):
  - Model predicts 72 percent share in small plants for 1989 (compared to data 77 percent).
  - Model under-predicts the change from 2005 to 2009 by a small amount.
  - Overall, model predicts 65 percent of the change in share of employment in small plants seen in the data between 1989 and 2009.

### Parameter sensitivity: love of variety
- Final goods production function generalized with love of variety parameter η:
  - Y_s^qn = 1 / M_{qn}^{η_qn} ( Σ_{i=1}^{M_qn} x_{i,qn}^{(σ−1)/σ} )^{σ/(σ−1)}
- Baseline specification: η = 1/(σ−1) (no love of variety).
- Alternative: η = 0 (full love of variety) — this tends to increase changes in size distribution in response to income changes; results become more sensitive to q1 (quality index of lowest quality).
- Table 8: Percent of cross-state difference explained (rows: η values; columns: q1 values)
  - q1 = 1, η = 1/(σ−1): 43.1%
  - q1 = 0.1, η = 1/(σ−1): 43.1%
  - q1 = 1, η = 0: 71.2%
  - q1 = 0.1, η = 0: 53.1%
- Interpretation:
  - Baseline (no love of variety, q1 = 1) explains 43.1 percent of cross-state variation in share of employment in plants of size five or less.
  - Allowing full love of variety (η = 0) increases the percent explained (to as high as 71.2% for q1 = 1), showing higher sensitivity and larger predicted effects when love of variety is present.

*Italic: Source — IMF working paper section titled "Section 2.2, larger plants tend to produce higher price products, which is indicative of higher quality goods" (calibration, counterfactuals, and results).*

### 71.2 percent of the differences in size distribution between the rich and poor states.

### _wp14236 - 71.2 percent of the differences in size distribution between the rich and poor states.

### Why love of variety amplifies changes in size distribution
- In the model with love of variety (η=0) relative prices of different quality levels change in the counterfactual because of changes in the relative varieties of the different qualities.
- The CES price index (the price charged by the final producer to the consumer) for quality q_n is given by:
  P_q_n = M_q_n^{η−1/(σ−1)} ( (ˆ(p(A_i,q_n))^{1−σ} g_q_n(A_i) dA_q_i)^{1/(1−σ)} ) ∀ q ∈ Q.
- In the baseline specification, because η = 1/(σ−1), the price index for q_n was independent of the number of varieties M_q_n.
- When η = 0, P_q_n is inversely related to the number of varieties M_q_n available. As income increases in the counterfactual, demand shifts to higher quality, inducing more entrants of higher quality levels, which lowers relative prices for high quality and generates further demand and entry — a feedback loop absent when η = 1/(σ−1).
- Therefore, the change in size distribution in the counterfactual is larger with love of variety than in the baseline specification that abstracts from variety-induced relative price changes.

### Sensitivity to the lowest quality index q_1
- Allowing for love of variety makes the counterfactual size-distribution response more sensitive to the choice of q_1, the quality index for the lowest quality level.
- When q_1 = 1 and η = 0, the model counterfactual explains 71.2 percent of the difference in size distribution.
- When q_1 = 0.1 and η = 0, the model counterfactual explains only 53.1 percent of the difference in size distribution.
- Demand choice formula shown in Section 3.1:
  ρ(q_n | w) = e^{a_{q_n} (w / P_{q_n})^{q_n}} / ∑_{i=1}^N e^{a_{q_i} (w / P_{q_i})^{q_i}} ∀ q_n ∈ Q.
- Because P_{q_n} is raised to the power q_n in the numerator, absolute levels of q_n approximately determine the own-price elasticity of demand for a quality level. Lower absolute q_n imply demand is less sensitive to relative price changes, so lower q_1 makes the model less responsive to variety-induced relative price changes.

### India vs US counterfactual
- Counterfactual: simulate an economy with per-capita income level equivalent to that of the US in 2005 (seventeen times that of India).
- Vary productivity, supply of skill, and θ'_q to match per-capita GDP and supply of skill in the US while keeping the wage premium and relative prices unchanged.
- Result: share of employment in plants which employ 5 or less people falls from 64 percent in the calibrated baseline to 13 percent in the counterfactual.
- Important caveat: calibration was local to India’s level of development; extrapolation to the US may be biased.
  - Producer-side price-size and consumer-side price-income relations that hold across Indian states may differ for the US.
  - Example possibilities: producer-side relation between price and size could be flatter or negative in the US; elasticity of substitution between skilled and unskilled labor may differ.
- Conclusion: US counterfactual is interesting but should be interpreted with more caution than cross-state counterfactuals within India.

### Inter-state trade: potential confound and empirical checks
- Concern: richer states might attract large plants which then ship to poorer states, so observed smaller shares of employment in small plants in rich states could reflect plant location choices rather than local demand differences.
- Direct data on inter-state trade flows in India is not collected; the paper provides indirect evidence that inter-state trade is not fully driving the cross-state relation.
- Arguments and evidence:
  - Transportation costs in developing countries are often very high; intranational transport costs in two African countries are seven to fifteen times larger than similar estimates for the US (Atkin and Donaldson (2012)). Manufacturing production is extremely localized even in the US (Hillberry and Hummels (2008)).
  - If inter-state trade drove the cross-state relation, tradable industries should exhibit larger differences across states in share of employment in small plants than non-tradables. If states approximate closed economies, the relation should be stronger for non-tradables.
- Two tradability measures constructed at 3-digit NIC level:
  1. Herfindahl index of geographical concentration in the US (H-index): H_i = ∑_{c=1}^C (sh^L_{i,c})^2. Industries highly concentrated across US counties (high H-index) are considered tradable; low H-index considered non-tradable. Applied to India via concordance from 6-digit NAICS to 3-digit NIC.
  2. Degree of international trade in India: (exports + imports) in the industry as a share of gross production of that industry carried out by domestic plants in 2005-06. HS product classification converted to 3-digit NIC via WITS concordance.
- The two measures are weakly positively correlated; rank correlation coefficient = 0.25.
- Regression specification:
  sd_{i,s,t} = α_{i,t} + α_{s,t} + γ ln(SNDP_{s,t}) * tradability_i + ε_{i,s,t},
  where sd_{i,s,t} is share of employment in plants of size five or less in industry i in state s at time t, SNDP_{s,t} is per-capita NDP, tradability_i is a tradable industry dummy.
- Interpretation: positive γ implies the log(per-capita SNDP) – share relation is stronger for non-tradables (i.e., tradable industries have a less negative slope).
- Table 9 results summary:
  - Column 1 (Herfindahl index, median cutoff): interaction coefficient 0.068 with standard error (0.0351); positive and marginally significant at the 10 percent level.
  - Column 2 (Herfindahl index, quartile cutoff): coefficient 0.052 with standard error (0.0394).
  - Column 3 (Exp-Imp index, median cutoff): coefficient −0.010 with standard error (0.0469).
  - Column 4 (Exp-Imp index, quartile cutoff): coefficient 0.000 with standard error (0.0498).
  - Observations: Column 1: 3,885; Column 2: 1,826; Column 3: 3,899; Column 4: 1,959.
  - Notes: Data from five rounds of the ASI and SUM (1989, 1994, 2000, 2005, 2010). Regressions include industry×time and state×time fixed effects. Observations weighted by share of observations in the state-industry cell out of total ASI+SUM observations for the year. Standard errors clustered at the state level. *p<0.1.
- Conclusion from Table 9: size-income relation across states is not stronger for tradable industries compared to non-tradable industries; evidence does not suggest inter-state trade is the main driver of the cross-state size distribution relation.

### Conclusion and key quantitative findings
- The size distribution in developing countries has a thick left tail compared to developed countries; within India, richer states have a much smaller share of manufacturing employment in small plants.
- Hypothesis: low income countries/states have high demand for low quality products which can be produced efficiently in small plants.
- Evidence presented is consistent with the hypothesis from both consumer and producer sides: richer households buy higher price goods; larger plants produce higher price products and use higher price inputs.
- A calibrated model with non-homothetic preferences over quality is developed and matched to cross-sectional consumer and producer facts.
- Calibrated model indicates that up to 41 percent of the cross-state variation seen in the left tail of manufacturing plants in India can be explained by the model.
- Policy insight: a large part of differences in size distribution across countries and states may be a natural consequence of low income levels in developing countries and not necessarily due to policies that discriminate against large productive plants in favor of small unproductive plants.

*Source: _wp14236 - 71.2 percent of the differences in size distribution between the rich and poor states.*

### References

### References

### Appendix A — Data Sources (surveys used)
- Surveys from India used:
  - Annual Survey of Industries of 2005-06, 1989-90, 1994-95, 2000-01, and 2009-10
  - Survey of Unorganized Manufacturing of 2005-06, 1989-90, 1994-95, 2000-01, and 2010-11
  - Consumer Expenditure Survey of India of 2003 and 2004-05
  - Employment-Unemployment Survey of India of 2004-05

- A.1. Annual Survey of Industries (ASI)
  - Conducted by the Central Statistics Office of the Government of India every year.
  - Coverage: all factories registered under Sections 2m(i) and 2m(ii) of the Factories Act, 1948 (factories employing ten or more workers using power, and those employing twenty or more workers without using power).
  - Primary data used: 2005-06 ASI (reports data for the financial year ending March 2006).
  - Geographical coverage of 2005-06 ASI: all of India except Arunachal Pradesh, Mizoram, Sikkim and Union Territory of Lakshadweep.
  - Industrial classification: NIC 2004 (closely based on ISIC Rev 3.1).
  - Sample restriction used in analysis: plants reporting a 2-digit NIC between 15 and 36 (manufacturing sector).
  - For some figures, restricted to 15 large Indian states: Andhra Pradesh, Bihar, Gujarat, Haryana, Himachal Pradesh, Karnataka, Kerala, Madhya Pradesh, Maharashtra, Orissa, Punjab, Rajasthan, Tamil Nadu, Uttar Pradesh, and West Bengal.
  - Plant-level variables: total employment (sum of seven categories: male workers employed directly, female workers employed directly, child workers employed directly, workers employed through contractors, supervisory and managerial staff, other employees, and unpaid family workers); product-level quantities and rupee values (ASI product classification ASICC, about 5,500 product categories, up to ten main products).
  - Notes on prices: per-unit prices inferred from quantity and value; standardized units intended, but misreporting in units exists (discussed in Section C).
  - Additional ASI years used for employment data: 1989-90, 1994-95, 2000-01, and 2009-10.
  - Exclusions across years: Arunachal Pradesh, Mizoram, Sikkim, and Lakshadweep excluded for all years for comparability.
  - Different waves used different NIC versions (NIC 1987, NIC 1998, NIC 2004, NIC 2008); concordance created to map to 2-digit NIC 2004 between 15 and 36.
  - Table A.1 (in source) reports number of observations, estimated number of establishments (using sampling weights), and estimated total number of workers employed for all five ASI years used.
  - More details available on the Ministry of Statistics and Programme Implementation website (http://mospi.nic.in/).

- A.2. Survey of Unorganized Manufacturing (SUM)
  - Conducted by the National Sample Survey Office (NSS) of India; coverage: all manufacturing enterprises not registered under Sections 2m(i) and 2m(ii) of the Factories Act, 1948.
  - Typical periodicity: usually conducted every five years; last five waves: 1989-90, 1994-95, 2000-01, 2005-06, and 2010-11.
  - Primary data used: 2005-06 SUM (62nd Round of the NSS); survey period July 2005 to June 2006.
  - Geographical coverage: comprehensive, all States and Union-Territories except Leh and Kargil districts of Jammu and Kashmir and a few remote villages in Nagaland and Andaman and Nicobar Islands; Arunachal Pradesh, Mizoram, Sikkim and Lakshadweep dropped to match ASI coverage.
  - Industrial classification: NIC 2004; sample restricted to 2-digit NIC between 15 and 36; some figures restricted to the same 15 large states as ASI.
  - Plant-level variables: total employment; product-level quantities and rupee values.
  - Employment reporting: average number of hired workers, working owners, and other workers (part-time and full-time) for the reference period (for most plants this was one month); broadest definition of employment used.
  - Product reporting: uses same ASICC product classification as ASI; SUM plants can choose units for quantities and prices (units may differ across SUM plants). Concordance across ASI and SUM units performed:
    - If units are scalar multiples (e.g., kilograms and tonnes), SUM units converted to common unit (e.g., divide tonnes by 1000 to get kilograms).
    - If SUM product reported in incompatible unit (e.g., numbers vs kilograms for matchsticks), treated as a separate product category.
  - Note: three month difference in coverage period between ASI and SUM.
  - Additional SUM years used for employment data: 1989-90, 1994-95, 2000-01, 2005-06, and 2010-11.
  - Table A.1 (in source) reports sample size, number of establishments (using sampling weights), and total number of workers employed for all five SUM years used.
  - More details available on the Ministry of Statistics and Programme Implementation website (http://mospi.nic.in/).

- A.3. Consumer Expenditure Surveys (CES)
  - Conducted by the National Sample Survey Office (NSS); annual Consumer Expenditure Surveys (Schedule 1.0) and quinquennial larger-sample surveys since 1972-73.
  - Primary data used: 2004-05 (61st Round of the NSS) Consumer Expenditure Survey (quinquennial series), interviewed about 125,000 households; survey period July 2004 to June 2005.
  - Geographical coverage: comprehensive, all States and Union-Territories except Leh and Kargil districts of Jammu and Kashmir and a few remote villages in Nagaland and Andaman and Nicobar Islands.
  - Data captured: value of consumption for 339 different goods; quantities and rupee values separately for 209 goods (allowing price computation for these goods).
    - Of these 209 goods: 156 food items, 10 “fuel and light”, 24 clothing and footwear, remainder durables.
  - For food items: households report consumption out of home production (quantities and imputed rupee values) and total consumption (home production plus market purchases); computed price = total value of consumption divided by total quantity consumed (averaging across home and market consumption).
  - Reference periods:
    - Food items: 30 days (quantity and rupee values for last 30 days).
    - Clothing and footwear: both 30 days and 365 days reference periods; 365 day period used where 30 day purchases frequently zero.
  - Additional data used: 2003 (59th Round) Consumer Expenditure Survey (not quinquennial), interviewed about 41,000 households; survey period January 2003 to December 2003. Consumption items across 2003 and 2004-05 surveys very similar.
  - Table A.2 (in source) reports summary statistics for the 2004-05 CES: number of items and share of expenditure for five broad expenditure headings and share of expenditure within heading for which prices could be computed.
  - More details available on the Ministry of Statistics and Programme Implementation website (http://mospi.nic.in/).

- A.4. Employment-Unemployment Survey
  - Conducted by the NSS as part of the quinquennial series; primary data used: 2004-05 (61st Round), interviewed about 125,000 households (about 600,000 individuals); survey period July 2004 to June 2005.
  - Geographical coverage: comprehensive, all States and Union-Territories except Leh and Kargil districts of Jammu and Kashmir and a few remote villages in Nagaland and Andaman and Nicobar Islands.
  - Survey content: demographic characteristics (age, education), main industry of work, size of establishment where individuals work, wage earned in the last week.
  - Sample restriction: only individuals reporting a 2-digit NIC 2004 between 15 and 36 are used to maintain comparability with production surveys.
  - Variables used: education level of individuals and establishment size category.
    - Education response categories: illiterate, literate but not through formal schooling, primary, middle, secondary, higher secondary, diploma/certificate course, graduate, post graduate or above.
    - For model purposes: a person defined as skilled if finished at least secondary education (Grade ten).
    - Establishment size categories respondents could report: less than 6, between 6 and 9, between 10 and 19, 20 or greater, and unknown size.
  - Calibration detail: wage premium calibration in Section 4.1 uses wage data from this survey; Mincerian regression results and more details in Section D.
  - More details available on the Ministry of Statistics and Programme Implementation website (http://mospi.nic.in/).

- A.5. County Business Patterns Database (US)
  - Maintained by the US Census Bureau; provides level of employment for each 6-digit NAICS for each US county.
  - Employment level reference date: week of March 12th of that year.
  - Paper uses the 2006 release of the data.
  - For many industry-county cells exact employment not reported; dataset reports employment size class instead. In such cases, employment assigned the midpoint of the size class reported (example given: size class ‘B’ representing 20-99 employees assigned employment level of 60).
  - Data available at http://www.census.gov/econ/cbp/.

### References (selected bibliographic citations included in source)
- Aguiar, Mark, and Erik Hurst, 2007, “Life-Cycle Prices and Production,” American Economic Review, Vol. 97, No. 5, pp. 1533–1559.
- Alfaro, Laura, Andrew Charlton, and Fabio Kanczuk, 2009, “Plant-Size Distribution and Cross-Country Income Differences,” in NBER International Seminar on Macroeconomics 2008, NBER Chapters, pp. 243–272.
- Atkin, David, and Dave Donaldson, 2012, “Who’s Getting Globalized? The Size and Nature of Intranational Trade Costs,” Techn. rep., Yale University.
- Attanasio, Orazio P., and Christine Frayne, 2006, “Do the Poor Pay More?” Techn. rep.
- Banerjee, A.V., and E. Duflo, 2011, Poor Economics: A Radical Rethinking of the Way to Fight Global Poverty (PublicAffairs).
- Banerji, Arup, and Sanjay Jain, 2007, “Quality dualism,” Journal of Development Economics, Vol. 84, No. 1, pp. 234–250.
- Behar, Alberto, 2009, “Directed technical change, the elasticity of substitution and wage inequality in developing countries,” Economics Series Working Papers 467, University of Oxford, Department of Economics.
- Bils, Mark, and Peter J. Klenow, 2001, “Quantifying Quality Growth,” American Economic Review, Vol. 91, No. 4, pp. 1006–1030.
- Bloom, Nicholas, Raffaella Sadun, and John Van Reenen, 2012, “The Organization of Firms Across Countries,” The Quarterly Journal of Economics, Vol. 127, No. 4, pp. 1663–1705.
- Broda, Christian, and David E. Weinstein, 2006, “Globalization and the Gains from Variety,” The Quarterly Journal of Economics, Vol. 121, No. 2, pp. 541–585.
- Chanda, Areendam, 2011, “Accounting for Bihar’s Productivity Relative to India’s: What can we learn from recent developments in Growth Theory,” Techn. Rep. 11/0759, International Growth Centre.
- Choi, Yo Chul, David Hummels, and Chong Xiang, 2009, “Explaining import quality: The role of the income distribution,” Journal of International Economics, Vol. 77, No. 2, pp. 265–275.
- Dalgin, Muhammed, Devashish Mitra, and Vitor Trindade, 2008, “Inequality, Nonhomothetic Preferences, and Trade: A Gravity Approach,” Southern Economic Journal, Vol. 74, No. 3, pp. 747–774.
- De Soto, Hernando, 1989, The other path (Harper & Row New York).
- Deaton, Angus, and Olivier Dupriez, 2011, “Spatial price differences within large countries,” Working Papers 1321, Princeton University, Woodrow Wilson School of Public and International Affairs, Research Program in Development Studies.
- DiCecio, Riccardo, and Levon Barseghyan, 2010, “Entry Costs, Industry Structure, and Cross-Country Income and TFP Differences,” 2010 Meeting Papers 964, Society for Economic Dynamics.
- Dikhanov, Yuri, 2010, “Income Effect and Urban-Rural Price Differentials from the Household Survey Perspective,” Techn. rep., ICP Global Office.
- Djankov, Simeon, Rafael La Porta, Florencio Lopez-De-Silanes, and Andrei Shleifer, 2002, “The Regulation Of Entry,” The Quarterly Journal of Economics, Vol. 117, No. 1, pp. 1–37.
- Faber, Benjamin, 2012, “Trade Liberalization, the Price of Quality, and Inequality: Evidence from Mexican Store Prices,” Techn. rep., London School of Economics.
- Fajgelbaum, Pablo, Gene M. Grossman, and Elhanan Helpman, 2011, “Income Distribution, Product Quality, and International Trade,” Journal of Political Economy, Vol. 119, No. 4, pp. 721–765.
- Flam, Harry, and Elhanan Helpman, 1987, “Vertical Product Differentiation and North-South Trade,” American Economic Review, Vol. 77, No. 5, pp. 810–22.
- García-Santana, Manuel, and Josep Pijoan-Mas, 2010, “Small Scale Reservation Laws And The Misallocation Of Talent,” Working papers, CEMFI.
- Garicano, Luis, Claire LeLarge, and John Van Reenen, 2013, “Firm Size Distortions and the Productivity Distribution: Evidence from France,” Working Paper 18841, National Bureau of Economic Research.
- Ghani, Ejaz, Arti Grover Goswami, and William R. Kerr, 2012, “Is India’s manufacturing sector moving away from cities ?” Policy Research Working Paper Series 6271, The World Bank.
- Gollin, Douglas, 1995, “Do Taxes on Large Firms Impede Growth? Evidence from Ghana,” Bulletins 7488, University of Minnesota, Economic Development Center.
- Guner, Nezih, Gustavo Ventura, and Xu Yi, 2008, “Macroeconomic Implications of Size-Dependent Policies,” Review of Economic Dynamics, Vol. 11, No. 4, pp. 721–744.
- Hallak, Juan Carlos, 2006, “Product quality and the direction of trade,” Journal of International Economics, Vol. 68, No. 1, pp. 238–265.
- Hallak, Juan Carlos, and Jagadeesh Sivadasan, 2011, “Firms’ Exporting Behavior under Quality Constraints,” Working Papers 628, Research Seminar in International Economics, University of Michigan.
- Hasan, Rana, and Karl Robert L Jandoc, 2010, “The Distribution of Firm Size in India: What Can Survey Data Tell Us?” Techn. rep., Citeseer.
- Hillberry, Russell, and David Hummels, 2008, “Trade responses to geographic frictions: A decomposition using micro-data,” European Economic Review, Vol. 52, No. 3, pp. 527–550.
- Hsieh, Chang-Tai, and Peter J. Klenow, 2012, “The Life Cycle of Plants in India and Mexico,” Working Paper 18133, National Bureau of Economic Research.
- Hsieh, Chang-Tai, and Benjamin A. Olken, 2014, “The Missing "Missing Middle",” Journal of Economic Perspectives, Vol. 28, No. 3, pp. 89–108.
- Hummels, David, and Peter J. Klenow, 2005, “The Variety and Quality of a Nation’s Exports,” American Economic Review, Vol. 95, No. 3, pp. 704–723.
- Iacovone, Leonardo, and Beata Javorcik, 2012, “Getting Ready: Preparation for Exporting,” CEPR Discussion Papers 8926, C.E.P.R. Discussion Papers.
- Kugler, Maurice, and Eric Verhoogen, 2012, “Prices, Plant Size, and Product Quality,” Review of Economic Studies, Vol. 79, No. 1, pp. 307–339.
- La Porta, Rafael, and Andrei Shleifer, 2008, “The Unofficial Economy and Economic Development,” NBER Working Papers 14520, National Bureau of Economic Research.
- Little, Ian, Dipak Mazumdar, and John M. Page Jr, 1987, Small Manufacturing Enterprises: A Comparative Analysis of India and Other Economies (NY: Oxford U. Press).
- Loayza, Norman V., 1996, “The Economics of the Informal Sector: A Simple Model and Some Empirical Evidence from Latin America,” Carnegie-Rochester Conference Series on Public Policy, Vol. 45, No. 0, pp. 129–162.
- Loayza, Norman V., Ana Maria Oviedo, and Luis Serven, 2005, “The impact of regulation on growth and informality - cross-country evidence,” Policy Research Working Paper Series 3623, The World Bank.
- Loayza, Norman V., Luis Serven, and Naotaka Sugawara, 2009, “Informality in Latin America and the Caribbean,” Policy Research Working Paper Series 4888, The World Bank.
- Mandel, Benjamin R., 2010, “Heterogeneous firms and import quality: evidence from transaction-level prices,” International Finance Discussion Papers 991, Board of Governors of the Federal Reserve System (U.S.).
- Manova, Kalina, and Zhiwei Zhang, 2012, “Export Prices Across Firms and Destinations,” The Quarterly Journal of Economics, Vol. 127, No. 1, pp. 379–436.
- McFadden, Daniel F., 1974, Conditional Logit Analysis of Qualitative Choice Behavior, pp. 105–142 (Academic Press: New York).
- Mitra, Devashish, and Vitor Trindade, 2005, “Inequality and trade,” Canadian Journal of Economics, Vol. 38, No. 4, pp. 1253–1271.
- Nataraj, Shanthi, 2011, “The impact of trade liberalization on productivity: Evidence from India’s formal and informal manufacturing sectors,” Journal of International Economics, Vol. 85, No. 2, pp. 292–301.
- Restuccia, Diego, and Richard Rogerson, 2013, “Misallocation and productivity,” Review of Economic Dynamics, Vol. 16, No. 1, pp. 1–10.
- Schott, Peter K., 2004, “Across-product Versus Within-product Specialization in International Trade,” The Quarterly Journal of Economics, Vol. 119, No. 2, pp. 646–677.
- Train, Kenneth, 2009, Discrete Choice Methods with Simulation (Cambridge University Press).
- Tybout, James R., 2000, “Manufacturing Firms in Developing Countries: How Well Do They Do, and Why?” Journal of Economic Literature, Vol. 38, No. 1, pp. 11–44.

*Source: _wp14236 - References (PDF).*

### Section 6 uses a Herfindahl Index of employment concentration across US counties as a measure of trad-

### _wp14236 - Section 6 uses a Herfindahl Index of employment concentration across US counties as a measure of trad-

### Concordances and construction of tradability indexes
- Herfindahl Index of employment concentration across US counties (from US County Business Patterns Database using NAICS) is used as a measure of tradability for industries in India, but must be mapped to the Indian classification (NIC 2004).
- Concordance steps:
  - Created a concordance between 6-digit NAICS 2002 and 3-digit ISIC Rev 3.1 (NIC 2004 is a one to one match to ISIC Rev 3.1 at the 3-digit level).
  - Used the Census Bureau’s concordance file (concordances reduced from many-to-many to one-to-one by taking the 3-digit ISIC which was the closest fit for each 6-digit NAICS).
- Coverage and exclusions:
  - Of the 59 3-digit ISIC industries in the manufacturing sector, three industries (182, 231, and 233) were not represented in this concordance (i.e., none of the 6-digit NAICS industries mapped into these 3-digit ISIC industries).
  - These three industries employed only 0.16 percent of the total manufacturing workforce in India in 2005. These industries are dropped for all analysis using the Herfindahl Index.
- Export-Import Index concordance:
  - Export and import data for India were at the HS product level (HS 2002). WITS provides a one-to-one concordance from HS 2002 to ISIC Rev 3.
  - Two NIC04 industries (223 and 273) were not represented in this concordance (none of the HS codes mapped into these industries).
  - These two industries employed only 0.37 percent of the total manufacturing workforce in India in 2005. These industries are dropped for all analysis using the Export-Import Index.
  - Industry 233 (nuclear fuel) had some imports but no local production in India in the trade data; this industry was dropped from the Export-Import Index analysis.
- Concordances across NIC revisions:
  - ASI and SUM use different NIC revisions: NIC87 (1989 and 1994), NIC98 (2000), NIC04 (2005), NIC08 (2010).
  - A concordance from different NIC revisions to NIC04 at the 3-digit level was created using official concordance tables from MOSPI.
  - NIC04 industries 341 (Manufacturing of motor vehicles) and 342 (Manufacture of bodies of motor vehicles, trailers, and semi-trailers) cannot be separately identified in NIC87 and are merged into one industry group for all tradability regressions.

### Units misreporting problem in the ASI: detection, correction, and effects
- Nature of the problem:
  - Example: ASICC code 11401 (“milk”) — plants should report quantity in kiloliters (1000 liters). Dividing rupee values by quantity should yield price per kiloliter.
  - Observed: most plants report log price about ten, but a group reports log price about seven log points lower (exp(7) = 1096), indicating some plants reported quantities in liters instead of kiloliters (price computed per liter rather than per kiloliter).
  - Such misreporting can bias regressions of price on size if misreporting correlates with plant size.
- Manual correction:
  - Manually reviewed about 1000 product categories to identify unit-misreporting problems.
  - Split affected products into two product categories using sensible price cutoffs (e.g., for milk, plants charging a log price greater than six placed in a different product category).
  - Different product fixed effects allowed for the new categories; clustering when computing standard errors does not treat the new product category as separate (number of product fixed effects exceed number of clusters in regressions).
  - Table A.8 comparison:
    - Column 1 (ASI, units problem accounted for): log(labor) coefficient 0.096***, Observations 46,704, Number of products 1,217, Number of clusters 1,078.
    - Column 2 (ASI, units problem not accounted for): log(labor) coefficient 0.155***, Observations 46,704, Number of products 1,077, Number of clusters 1,078.
    - Column 3 (ASI+SUM, units problem accounted for): log(labor) coefficient 0.106***, Observations 75,161, Number of products 3,181, Number of clusters 3,042.
    - Column 4 (ASI+SUM, units problem not accounted for): log(labor) coefficient 0.125***, Observations 75,161, Number of products 3,041, Number of clusters 3,042.
    - Notes: 1 percent tails of prices (within a product) and plant size are winsorized. All regressions include product fixed effects and state times urban-rural fixed effects. Standard errors clustered at the product level. ***p<0.01.
  - Finding: price elasticity with respect to employment is smaller when the units problem is corrected, implying misreporting of units is correlated with size.
- Algorithmic detection (automated alternative to manual correction):
  - Step 1: If maximum price reported for a product is less than 50 times the minimum price, classify product as no units misreporting.
  - Step 2: Order prices ascending within a product. If two consecutive prices differ by a factor of at least 20, and the average price above the jump is between 500 and 2000 times the average price below the jump, classify product as units-misreporting and split it into two product categories.
  - Step 3: For a product category with N plants, run N regressions of log(price) on log(employment) plus a dummy that equals 1 for the lowest k price observations (k = 1..N). Compare the highest R-square obtained with dummies to the R-square with no dummy. If R-square difference > 0.75 and the mean price above the dummy (for the highest R-square) is at least 300 times higher than the mean price below the dummy, classify as units-misreporting and split into two product categories.
  - Algorithmic results:
    - When using the algorithm instead of manual correction, the elasticity of price to size is 0.1037 in the ASI.
    - Sensitivity: changing the threshold for the R-square in step 3 to 0.8 and 0.7 changes the estimated elasticity to 0.1091 and 0.0986 respectively.
  - The algorithm was also implemented for input prices regressions with similar results to the manual correction.

- Specific numeric example of misreporting magnitude:
  - Observed log-price difference of about 7 log points between two groups of plants for milk; exp(7) = 1096 (i.e., approximately 1,096 times difference in computed price due to unit mismatch).

### Calibration of production parameter θ_q_n (share of unskilled workers)
- Purpose:
  - θ_q_n chosen to match the wage premium and the ratio of unskilled to skilled workers for different qualities relative to the lowest quality level.
- Wage premium target:
  - Derived from Mincerian regression using the Employment-Unemployment Survey of 2004-5.
  - Wage calculation: average wage per individual computed by dividing total wage earned for each activity over the last seven days by intensity-days worked (full intensity = 1 day; half intensity = 0.5 days).
  - Regression: log(wages) on a dummy for skilled (ten or more years of education), controlling for potential experience (age minus years of education minus four) and its square, and dummies for each 4-digit industry, 2-digit occupation, state, sector (urban or rural), and sex. Sample restricted to manufacturing (2-digit NIC between 15 and 36) and ages 15–65.
  - Table A.4 results:
    - Coefficient on skilled dummy: 0.450 (Column 1) and 0.445 (Column 2).
    - Implied wage premium: 1.568 and 1.560 (i.e., 56.8 percent and 56.0 percent). The paper rounds the 56.8 percent up to 60 percent when calibrating the model.
    - Observations: 11,003. Column 2 winsorizes 1 percent tails of wages. Robust standard errors reported. ***p<0.01.
- Unskilled to skilled ratios for calibration:
  - Need N−1 ratios for equation 9 (unskilled to skilled ratio for different qualities relative to lowest quality).
  - Employment-Unemployment survey size categories are coarse; cannot compute eleven ratios directly.
  - Extrapolation approach:
    - Table 6 reports plants of size five or less have unskilled to skilled ratio of 5.05; plants of size 5 to 20 have ratio 2.92.
    - These two points are extrapolated to estimate ratios for larger sized plants, with the ratio bounded below at 0.5 (i.e., hire twice as many skilled as unskilled).
    - Extrapolated values used to compute equation 9 for different quality levels given average size of each quality level.

### Empirical regression highlights and key statistics
- Table A.8 (units problem and price-size regressions):
  - ASI only (units corrected): log(labor) = 0.096*** (Observations 46,704; Number of products 1,217; Number of clusters 1,078).
  - ASI only (units not corrected): log(labor) = 0.155*** (Observations 46,704; Number of products 1,077; Number of clusters 1,078).
  - ASI+SUM (units corrected): log(labor) = 0.106*** (Observations 75,161; Number of products 3,181; Number of clusters 3,042).
  - ASI+SUM (units not corrected): log(labor) = 0.125*** (Observations 75,161; Number of products 3,041; Number of clusters 3,042).
- Algorithmic elasticity results:
  - Elasticity of price to size in ASI using algorithm = 0.1037.
  - Varying step 3 R-square threshold to 0.8 and 0.7 yields elasticity 0.1091 and 0.0986 respectively.
- Wage premium regression (Table A.4):
  - Skilled coefficient = 0.450; Wage Premium = 1.568 (56.8 percent), rounded to 60 percent in calibration. Observations = 11,003.
- Concordance exclusions and workforce shares:
  - Missing 3 ISIC industries (182, 231, 233) mapped to 0.16 percent of total manufacturing workforce in India in 2005 — dropped for Herfindahl-based analysis.
  - Missing 2 NIC04 industries (223, 273) in HS→ISIC concordance mapped to 0.37 percent of total manufacturing workforce in India in 2005 — dropped for Export-Import Index analysis.
  - Industry 233 (nuclear fuel) had imports but no local production — dropped from Export-Import Index.

### Industry-level tradability classification summary (from Table A.5)
- The paper lists 3-digit NIC04 industries falling above and below the median for two tradability indexes (Herfindahl Index and Export-Import Index). Example listings (preserve source ordering):
  - Industries Below Median (Non-tradable) — Herfindahl Index: 151,152,153,154,155,171,201,202,210,221,222,241,242,251,252,261,269,272,273,281,289,291,292,311,312,313,314,315,319,321,322,323,331,332,333,341,342,343,351,352,353,359,361,369 (as in table organization).
  - Industries Above Median (Tradable) — Herfindahl Index: 160,172,173,181,191,192,223,232,243,271,293,300,319,321,322,323,331,332,333,341,342,351,352,353,359,369 (as in table organization).
- Note: Table A.5 gives parallel lists for the Export-Import Index; consult table for exact mapping and grouping.

### Additional empirical notes and figures
- Table A.1: ASI and SUM summary statistics (all numbers in thousands ’000) for five years (1989-90, 1994-95, 2000-05, 2005-06, 2009-10) report observations, plants, and employment (exact values provided in table).
- Table A.2: Consumer Expenditure Survey summary by category lists Items, Share of Expenses, Items with Prices, Share with Prices (exact numbers provided in table).
- Figure A.1 notes:
  - Uses ASI and SUM waves (1989, 2000, 2009).
  - First figure (1989 to 2009): slope of linear fitted line is -0.22 with a P-value of 0.077.
  - Second figure (2000 to 2009): slope of linear fitted line is 0.086 with a P-value of 0.580.

*Source: _wp14236 - Section 6 uses a Herfindahl Index of employment concentration across US counties as a measure of trad- (PDF chapter/section).*

---


_Source: https://www.imf.org/-/media/websites/imf/imported-full-text-pdf/external/pubs/ft/wp/2014/_wp14236.pdf_
