## 1.1 Relationship to the Literature — wpiea2022131-print-pdf

## Source details

**Canonical URL:** [1.1 Relationship to the Literature — wpiea2022131-print-pdf](https://www.imf.org/-/media/files/publications/wp/2022/english/wpiea2022131-print-pdf.pdf)

## Other formats

- [Markdown version](/-/media/files/publications/wp/2022/english/wpiea2022131-print-pdf.pdf.md)
- [Structured JSON version](/-/media/files/publications/wp/2022/english/wpiea2022131-print-pdf.pdf.json)

---

### Overview and motivation
- R&D tax credits introduced in 1981 at federal level in the US; estimated foregone federal revenues in 2019 exceed US$9 billion, "about 20% more than the National Science Foundation’s entire budget request for that year."
- Most US states adopted R&D subsidies → wide spatial dispersion of R&D tax credits.
- Core question: How does spatial dispersion of R&D tax credits affect aggregate innovation, spatial allocation of inventors and firms, and long-run growth?

### Conceptual contributions and model structure
- Adds a spatial dimension to Schumpeterian endogenous growth frameworks by:
  - Introducing endogenous city populations and agglomeration externalities in innovation.
  - Modeling cities c ∈ {0,1,...,C} with C → ∞, each with amenities α_c, permanent productivity ̄χ_c, and stochastic productivity shocks z_c(t) with χ_c(t) = ̄χ_c e^{z_c(t)}.
  - Agents: inventors (inelastic supply I), production workers (population L), and firms; inventors and production workers freely mobile across cities; hiring is local.
  - Innovation: quality-ladder framework (step size λ > 0), creative destruction, continuum of varieties j ∈ J ≡ [0,1].
  - Non-tradable goods produced with land fixed → decreasing returns and congestion; final tradable good and intermediate goods freely tradable.
  - Local R&D subsidies s_c modeled as transfers (share of R&D cost refunded).
- Key mechanism: inventor productivity increases with local inventor density Ĩ_c^η; agglomeration elasticity η ≥ 0. Congestion captured by elasticity parameter β.

### Agglomeration and trade-offs emphasized
- Interaction of:
  - pre-existing spatial heterogeneity (̄χ_c, α_c),
  - agglomeration externalities (η),
  - spatial variation in s_c,
  - congestion costs via land-fixed non-tradable production (β).
- Trade-off: agglomeration benefits (higher inventor productivity) vs congestion costs (higher non-tradable prices and land rents).

### Identification and estimation approach (three steps)
- Step 1 — Calibration: match parameters to data/literature (g = 0.02, ρ = 0.02, r = 0.038, ψ = 0.5, λ = 0.132, ε = 0.15, θ = 0.6, I = 1, L = 175, L_0 = 0.85).
- Step 2 — Linear regressions: estimate ψη and congestion elasticity via log-linear regressions; use a shift-share instrument I_{c,t,l} built from lagged industry shares and external industry growth rates (non-US inventor growth used as shifters).
- Step 3 — Moment matching: recover remaining parameters (κ, α_c, ̄χ_c, σ) by matching model moments (share of inventors and patents by city) and set growth g = λD = 0.02.

---

### Key quantitative findings and fit

### Empirical estimates and calibration (preserving reported numerics)
- Estimated ψ = 0.5 (elasticity of innovation w.r.t. R&D).
- Implied η (baseline for model applications): η = 0.20 (sensitivity considered at η = 0.15 and η = 0.25).
- Baseline congestion elasticity: β = 0.6 (sensitivity considered at β = 0.5 and β = 0.8).
- Fixed cost κ from moment-matching: I/N ≈ 21.07 ⇒ κ = 10.53 (using κ = (1−ψ) I / N with ψ = 0.5).
- Calibrated subsidy rates s_c in data: s_c ∈ [0.13,0.30] (effective R&D credit rates from Wilson (2009), averaged 1998–2006).
- Cities: C = 860 CBSAs with patents (917 total CBSAs; 57 treated as city 0 producing no innovation).

### Model fit to untargeted moments (correlations)
- Model-data correlations for untargeted variables (Panel A, Table 4):
  - Number of Firms: 0.97
  - Patents per Firm: 0.58
  - Employed Population: 0.81
  - Patents per Capita: 0.49
- Validation: model-data correlations for untargeted variables range from 0.49 to 0.97; correlations for patents per capita by city ≈ 0.49 and for share of firms by city ≈ 0.97.

### Elasticity estimation highlights (preserving reported numeric outcomes)
- Estimated ψη (second-stage coefficients on log(Inventors in City)):
  - OLS: 0.070∗∗∗ (0.014) ⇒ implied η = 0.140 (with ψ = 0.5).
  - IV (lags): examples — IV (l=5) → 0.104∗∗∗ (0.019) ⇒ implied η = 0.208; IV (l=10) → 0.105∗∗∗ (0.023) ⇒ implied η = 0.210.
  - Robust inference uses AKM SE and cluster SEs; F-statistics first-stage range from 33.94 up to 315.10 across IVs in Table 2.
- Congestion estimation (rent regressions) — implied β examples:
  - Using ZRI rents: implied β ≈ 0.608, 0.595, 0.493, 0.639 across IV specifications (Table 3).
  - Using housing prices per sq.ft: implied β ≈ 0.8 range (treated as upper bound).

---

### Mechanisms and analytic equilibrium results

### Closed-form spatial distribution and firm-level results (Proposition 2; preserve expressions)
- Inventor population I_c closed form (Equation (9); Θ = (1−β) θ − ψ η (1−θ)):
  - I_c = I × [ ̄χ_c^{1−s_c} ]^{(1−θ)/Θ} α_c^{θ/Θ} × [ Σ_{c=1}^C [ ̄χ_c^{1−s_c} ]^{(1−θ)/Θ} α_c^{θ/Θ} ]^{-1} × Z_c^{(1−θ)/Θ ( (1−θ)/Θ − 1 ) σ^2 / (4 φ) }.
- Firm-level arrival rate when investing (Equation (10)):
  - x_{f,c} = [ κ/ψ (1−ψ) ]^{ψ} ̄χ_c ˜I_c^{ψ η} Z_c.
- Inventors per investing firm:
  - i_{f,c} = ψ/(1−ψ) κ.
- Number of R&D-investing firms in city c (Equation (11)):
  - N_c = [ (1−ψ)/κ ] I_c.
- Aggregate creative destruction (Equation (13)):
  - D ∝ (1/C) Σ_{c=1}^C ̄χ_c ˜̄I_c^{1+ψ η}.
- Growth identity (Proposition 3):
  - g = λ D.

### Equilibrium properties and uniqueness
- SBGP existence requires r > g and sufficiently many cities so local shocks average out.
- Uniqueness requires Θ > 0; if Θ ≤ 0 agglomeration dominates congestion → multiple equilibria (all firms/people concentrate).

---

### Welfare analysis and policy counterfactuals

### Spatially homogeneous subsidy counterfactual (preserving numerical outcomes)
- Counterfactual: set s_c ≡ ̄s such that total subsidy spending unchanged; calibrated ̄s close to 19% (average subsidy under current distribution ≈ 16%).
- Effects of moving to homogeneous subsidy:
  - HHI index of city population shares: 0.027 → 0.025 (population slightly less concentrated).
  - Aggregate welfare falls by 0.77%.
  - Decrease in growth rate ≈ 0.03 percentage points.
  - Static baseline wage ̄w_i increases by 0.91%.
- Interpretation: current spatial heterogeneity in R&D credits appears welfare-enhancing relative to a spatially uniform subsidy given same resources.

### Approximate optimal subsidies (city-level and state-level; τ caps and reported gains)
- Functional-form approximation for optimal s_c: s_c = min{ ζ α_c^ξ ̄χ_c^ω, τ } with τ ∈ {0.3,0.4,0.5}; ζ chosen to satisfy budget constraint.
- City-level optimal subsidies (Panel A, Table 6) — examples:
  - τ = 0.3: ∆Welfare = 2.95%; ∆Baseline Wage = −3.60%; ∆Creative Destruction = 0.97 p.p.; ∆Rate of Growth = 0.13 p.p.
  - τ = 0.4: ∆Welfare = 5.23%; ∆Baseline Wage = −6.38%; ∆Creative Destruction = 1.70 p.p.; ∆Rate of Growth = 0.22 p.p.
  - τ = 0.5: ∆Welfare = 6.15%; ∆Baseline Wage = −7.75%; ∆Creative Destruction = 2.00 p.p.; ∆Rate of Growth = 0.26 p.p.
- State-level optimal subsidies (Panel B, Table 6) — examples:
  - τ = 0.3: ∆Welfare = 2.50%; ∆Baseline Wage = −3.65%; ∆Creative Destruction = 0.88 p.p.; ∆Rate of Growth = 0.12 p.p.
  - τ = 0.4: ∆Welfare = 3.06%; ∆Baseline Wage = −4.36%; ∆Creative Destruction = 1.07 p.p.; ∆Rate of Growth = 0.14 p.p.
  - τ = 0.5: ∆Welfare = 3.23%; ∆Baseline Wage = −4.71%; ∆Creative Destruction = 1.13 p.p.; ∆Rate of Growth = 0.15 p.p.
- Key policy implications:
  - Optimal city-level reallocations can raise aggregate welfare substantially (up to 6.15% under τ = 0.5 in baseline calibration).
  - Restricting subsidy variation to state level reduces gains (e.g., 6.15% → 3.23% at τ = 0.5).
  - Optimal reallocations typically shift part of inventor population from small/medium cities to large, more productive/high-amenity cities (e.g., San Jose, New York highlighted).
  - Higher creative-destruction/growth comes with lower static baseline wages (trade-off between long-run growth and static wage level).

### Empirical relevance: R&D tax credits explain historical spatial changes
- Using year-specific s_c series and holding other parameters fixed, model reproduces large part of spatial changes in inventor and patent shares since 1970s:
  - Decadal model-data correlations for inventor levels reported: 1970s Corr. = 0.86; 1980s Corr. = 0.92; 1990s Corr. = 0.96 (Panel A, Table 5).
  - Regression of decade changes: R^2 ≈ 0.39–0.45 for ∆Share of Inventors and ≈ 0.64–0.68 for ∆Share of Patents.

### Robustness and sensitivity (preserving reported numeric grid)
- Sensitivity grid: η ∈ {0.15,0.20,0.25} and β ∈ {0.5,0.6,0.8}, τ = 0.5 examples (Table F.6):
  - City-level ∆Welfare ranges (selected): 2.58% up to 8.42% across combinations; baseline example η = 0.20, β = 0.6 yields 6.15% welfare gain.
  - State-level ∆Welfare ranges (selected): 1.91% up to 10.68% across combinations; baseline η = 0.20, β = 0.6 yields 3.23% welfare gain.
- Main qualitative result robust: optimal reallocations increase spatial concentration, raise growth/creative destruction, and lower baseline wages.

### Extensions and limitations
- Extension: allow innovation to scale with firm size (Section G). Core qualitative predictions preserved; added complexities in identifying firm/product scales and stationary firm-size distributions.
- Limitations explicitly noted:
  - Model abstracts from moving costs and short-run adjustment costs; results characterize long-run equilibrium effects.
  - Other policy objectives (distributional impacts, local job losses) not modeled.
  - Approximate optimization (functional-form restriction) provides a lower bound on potential gains from full optimal reallocation.

---

### Policy takeaways (concise bullets)
- Spatially heterogeneous R&D subsidies can be welfare-improving relative to a uniform subsidy because they leverage agglomeration externalities.
- Existing US spatial variation in R&D tax credits appears to allocate larger credits to states relatively better at producing innovation; removing spatial variation reduces welfare (−0.77%) and concentration (HHI 0.027 → 0.025).
- Redeploying existing R&D subsidy revenue across locations can generate substantial long-run welfare gains:
  - City-level optimal reallocation: up to 6.15% welfare gain (τ = 0.5, baseline calibration).
  - State-level optimal reallocation: up to 3.23% welfare gain (τ = 0.5, baseline calibration).
- Trade-offs: higher long-run growth and creative destruction accompany lower static baseline wages (example: ∆Baseline Wage = −7.75% at τ = 0.5 city-level).
- Policy design matters by geographic scope: city-level targeting yields larger potential gains than state-level targeting under identical budget constraints.

*Italic: Source: IMF Working Paper — "Agglomeration, Innovation, and Spatial Reallocation" (wpiea2022131-print-pdf).*

### 1.1    Relationship to the Literature   .  .  .  .  .  .  .  .  .  .  .  .  .  .  .  .  .  .  .  .  .  .  .  .  .  .  . 

### 1.1    Relationship to the Literature

### Overview and motivation
- R&D tax credits: introduced in 1981 at the federal level in the US; estimated to amount to over US$9 billion in foregone federal revenues in 2019, which is "about 20% more than the National Science Foundation’s entire budget request for that year."
- Most US states have adopted subsidies to research activity, producing wide spatial dispersion of R&D tax credits.
- Research question: How does the dispersion of R&D tax credits across space affect aggregate innovation, spatial allocation of inventors and firms, and long-run growth?

### Conceptual contributions and link to existing literature
- Adds a spatial dimension to Schumpeterian endogenous growth frameworks by introducing endogenous city populations and agglomeration externalities in innovation.
- Contrasts with most spatial policy evaluations that are static (e.g., Kline and Moretti, 2014; Ossa, 2015; Gaubert, 2018; Fajgelbaum and Gaubert, 2020) and with long-run R&D policy studies that often omit local externalities or spatial heterogeneity.
- Emphasizes interaction between:
  - pre-existing spatial heterogeneity,
  - agglomeration externalities (inventor productivity increases with inventor population density), and
  - spatial variation in tax credit rates.
- Recognizes trade-off between agglomeration benefits and congestion costs (Henderson, 1974; Kline and Moretti, 2014).

### Model structure (high-level)
- Economy: system of cities c ∈ {0,1,...,C}, with C → ∞.
- City-specific endowments: amenity level, stochastic productivity for innovation, and an R&D tax credit (modeled as a direct subsidy).
- Agents: inventors (produce innovation), production workers (produce goods), and firms that employ them. Inventors and production workers are freely mobile across cities; hiring is local.
- Innovation: quality-ladder framework; firms innovate intermediate goods (continuum of varieties); successful innovation yields a technological leader producing under monopolistic competition; all innovation generates creative destruction; step size of quality improvement is fixed.
- Goods: intermediate goods (freely tradable), final tradable good, final non-tradable good (produced/consumed in each city; decreasing returns due to land fixed factor → generates congestion and limits city size).
- Government: central planner fully taxes land- and firm-owners to finance R&D subsidies and a public good; consumers derive utility from final tradable good, non-tradable good, and public good.
- Free entry of firms and free mobility of workers imply endogenous spatial distribution of population and firms.

### Agglomeration and spatial mechanisms emphasized
- Inventor productivity increases with local inventor density (agglomeration spillovers).
- Location choices of inventors and firms respond to spatial variation in R&D subsidies, amenities, and expected productivity shocks.
- Spatial distribution of inventors determines local agglomeration spillovers and thus equilibrium outcomes (including aggregate growth rate).
- Closed-form expression: share of inventors in each city can be expressed in closed form as a function of city-specific features and three parameters: the elasticity of agglomeration, the elasticity of congestion, and the share of expenditures on non-tradable goods.

### Identification and estimation approach (three steps)
- Step 1 — Calibration: match parameters that can be directly linked to data quantities.
- Step 2 — Linear regressions: estimate elasticity of agglomeration and elasticity of congestion by regressing outcomes on city inventor or worker population (with city and year fixed effects); use a shift-share style instrument leveraging industry-specific growth in inventor employment shares to address endogeneity.
- Step 3 — Moment matching: recover remaining parameters by matching model predictions to data moments related to innovation (e.g., share of inventors and patents by city); validate external predictions (correlations between model and data distributions range from 0.5 for patents per capita by city to 0.97 for share of firms by city).

### Key quantitative findings (preserving reported numerics)
- Removing spatial variation of R&D tax credits in the US produces a slightly less concentrated population: HHI index moves from 0.027 to 0.025.
- Welfare effect of removing spatial variation: welfare falls by 0.77%.
- Potential welfare gains from optimal reallocation of R&D subsidies:
  - If subsidies can vary by city, aggregate welfare increases by at least 6%.
  - If subsidies can vary only by state, gains are 3.2%.
- Quantitative implication: optimal reallocations imply shifting part of the inventor population from small and medium-sized cities to large, more productive cities.
- Model replicates spatial responses over time: changes in R&D tax credits over time can explain a large part of variation in the share of inventors and patents filed in each city since the 1970s.
- Correlation evidence for validation: model-data correlations for untargeted variables range from 0.5 to 0.97.

### Policy implications and interpretation
- Spatially heterogeneous R&D subsidies can be welfare-improving relative to spatially uniform subsidies because they can leverage agglomeration externalities.
- Current US spatial variation in R&D tax credits appears to place larger credits in states relatively better at producing innovation (removing variation reduces welfare and concentration).
- Substantial welfare gains are possible through redistributing existing R&D subsidy revenue across locations (city-level variation yields larger gains than state-level variation).
- Importance of jointly considering spatial and dynamic dimensions when evaluating R&D policy, since subsidies affect both location decisions and innovation investment over time.

*Source: IMF Working Paper — "Agglomeration, Innovation, and Spatial Reallocation" (section 1.1).*

### 1.1   Relationship to the Literature

### 1.1   Relationship to the Literature

### Contribution to existing literature
- Situates the paper within literature on spatial misallocation and optimal spatial policies; notes large potential gains from reallocating resources across space in the US.
- Connects to specific findings in the literature:
  - Hsieh and Moretti (2019): housing supply restrictions in productive US cities lowered growth between 1964 and 2009.
  - Fajgelbaum et al. (2019): tax dispersion across states leads to aggregate losses by distorting spatial allocation of resources.
  - Gaubert (2018) and Fajgelbaum and Gaubert (2020): general quantitative frameworks for optimal local subsidies to attract workers and firms to cities.
  - Ossa (2015): welfare effects of subsidy competition among states.
  - Kline and Moretti (2014): long-run effects of the Tennessee Valley Authority development program.
- Highlights complementary literature on policies to mitigate economic distress and inequality (Austin et al., 2018; Farrokhi, 2021) and discussion of geographically targeted innovation policies (Glaeser and Hausman, 2019).

### Gap addressed: static vs. dynamic spatial evaluation of R&D subsidies
- Argues that existing literature primarily analyzes spatial policies through static frameworks.
- Key point: R&D tax credits have both spatial and dynamic effects that interact.
  - A purely spatial model would ignore how changes in the rate of creative destruction alter firm incentives and therefore mis-evaluate welfare effects of spatially reallocating R&D subsidies.
  - A purely dynamic model would predict no relationship between the location of subsidies and aggregate growth.
- Conclusion: to evaluate how aggregate economy reacts to changes in the distribution of R&D subsidies, a framework must capture both dynamic and spatial effects.

### Theoretical framework and its innovations
- Based on endogenous growth models where innovation drives growth (examples: Aghion and Howitt, 1992; Klette and Kortum, 2004; Akcigit and Kerr, 2018).
- Novelty: nests a model of innovation via creative destruction into a spatial setting.
  - Retains features of growth literature while allowing for spatial heterogeneity, agglomeration spillovers, and an endogenous population distribution.
  - Links firm-level innovation (affected by location through agglomeration spillovers) to aggregate growth and welfare.
- Contrasts with related spatial-dynamic models:
  - Duranton (2007): embeds quality ladder model into urban structure but uses reduced-form externalities on city size.
  - Current model microfounds agglomeration and congestion externalities endogenously in equilibrium.
  - Desmet and Rossi-Hansberg (2014), Desmet et al. (2018), Caliendo et al. (2019): include realistic geography, trade and moving costs, and ordered locations; current paper abstracts from most geographical frictions to focus on firm-level innovation dynamics and tractability for mapping to patent data.

### Overview of model structure (as developed in sections 2–4)
- Model objective: allow spatial heterogeneity and agglomeration externalities to affect productivity of R&D investments by firms; spatial distribution of population matters for growth and is endogenously determined.
- Assumptions and setup:
  - Cities:
    - Number of cities: C + 1, indexed by c ∈ {0,1,...,C}.
    - Large number of cities: C → ∞.
    - City heterogeneity: amenities α_c, land ̄m_c, and stochastic time-varying city-specific productivity χ_c(t).
    - Normalization: ̄m_0 = 1 and ̄m_c = 1/C for all c ≥ 1.
    - City-specific productivity for innovation: χ_c(t) = ̄χ_c e^{z_c(t)}.
    - Process for z_c(t): Ornstein-Uhlenbeck (O-U) process dz_c(t) = φ(μ − z_c(t))dt + σ dW_c(t).
    - Assumed μ = −σ^2 / 4φ so that, under stationary distribution, E[e^{z_c}] = 1.
    - Interpretation: ̄χ_c captures permanent differences across cities; z_c(t) captures time-varying shocks. O-U process admits a stationary distribution, enabling Law of Large Numbers when number of cities is large.
  - Agents and mobility:
    - Two worker types: inventors and production workers.
    - Total population of inventors: I; total population of production workers: L. Both I and L are constant over time.
    - Supply of inventors is very inelastic (Goolsbee, 1998); model assumes no occupational switching in response to R&D subsidies.
    - Both worker types are freely mobile between cities and are hired locally.
    - Firms can enter any city but cannot move once located; firms are owned by absentee firm owners who are fully taxed by the social planner/government.
    - Land is fixed in each city and owned by absentee land owners, who are fully taxed by the government.
  - Goods and trade:
    - Three types of goods: a final non-tradable good, a final tradable good (referred to as “final good”), and intermediate goods.
    - Both the tradable final good and intermediate goods can be traded at no cost.
  - Preferences:
    - All consumers (inventors and production workers) share the same utility function but may differ in wages.
    - Utility of a worker of type h ∈ {i, ℓ} born at t_0:
      - U_h(t_0) = ∫_{t_0}^∞ e^{−ρt} max_{c(t)} { u_{h c(t)}(t) G(t) } dt.
      - u_{h c(t)}(t) = max_{n(t), y(t)} [ α_c n(t) ]^θ y(t)^{1−θ} subject to y(t) + p_{n,c}(t) n(t) ≤ w_{h c}(t).
    - All workers inelastically supply one unit of labor per period; income equals wage w_{h c}(t).
    - Consumers are “hand-to-mouth”: cannot borrow or save.
- Equilibrium concept: existence of a Balanced Growth Path where aggregate growth rate is constant while individual-city growth rates and populations can fluctuate; model introduces local shocks to R&D productivity but, given many cities, no aggregate uncertainty.

*Source: wpiea2022131-print-pdf — 1.1   Relationship to the Literature*

### 2.3   Technology

### wpiea2022131-print-pdf - 2.3   Technology

### 2.3.1 Non-tradable Good
- Representative non-tradable goods producer in each city chooses production workers ℓ_{n,c}(t) and land m_c(t) to solve:
  - max_{ℓ_{n,c}(t),m_c(t)} p_{n,c}(t) n(t) − w^ℓ_c(t) ℓ_{n,c}(t) − p^m_c(t) m_c(t)
  - subject to n(t) = ℓ_{n,c}(t)^β m_c(t)^{1−β}
- Definitions and roles:
  - p_{n,c}(t): price of the non-tradable good.
  - p_{m,c}(t): price of land.
  - w^ℓ_c(t): wage received by production workers in city c.
- Key mechanisms:
  - Land is fixed in each city → production has decreasing returns to scale in equilibrium.
  - Congestion costs arise because as city population increases, demand for the non-tradable good increases, pushing p_{n,c} up and reducing city attractiveness.
- Modeling note:
  - The production function implies the elasticity of supply of the non-tradable good with respect to population is the same across cities, though supply responds to land mass and amenities through wages.

### 2.3.2 Final Goods
- Representative final good producer uses production worker labor and all intermediate goods; location is freely chosen.
- If production occurs in city c^* the firm solves:
  - max_{ℓ_{y,c^*}(t), {k_j(t)}_{j∈J}} Y(t) − ∫_J p_j(t) k_j(t) dj − w^ℓ_{c^*}(t) ℓ_{y,c^*}(t)
  - subject to Y(t) = ℓ_{y,c^*}(t)^{ε/(1−ε)} ∫_J k_j(t)^{1−ε} q_j(t)^ε dj
- Variables:
  - ℓ_{y,c^*}: number of production workers hired by the final good producer.
  - k_j and q_j: quantity and quality of intermediate good j.
  - p_j: price at which each intermediate good is sold to the final good producer.
- Additional points:
  - Price of the final good normalized to 1.
  - Intermediate and final goods are freely tradable across cities; location of production does affect prices.

### 2.3.3 Intermediate Goods
- Product space:
  - Continuum of varieties j ∈ J ≡ [0,1]. Firms can produce any number of varieties; a firm’s portfolio equals products it has innovated on.
- Pricing game (Assumption 1):
  - Two-stage pricing game: (1) firms decide whether to pay a small fee to announce a price; (2) given entrants, firms propose prices.
  - Direct implication: only the technical leader (highest quality) pays the fee and enters stage two; leader behaves as monopolist for variety j.
- Cost and market structure conditions:
  - Crucial condition: all firms have the same cost of production regardless of location (excludes production functions using labor for intermediate goods).
  - Final good is used as input for intermediates; marginal cost ν > 0 for all firms.
- Firm production problem for variety j:
  - max_{k_j} p_j(k_j; q_j) k_j − ν k_j
  - Price of final good normalized to 1.
- Production independence from location:
  - No city-specific production costs for intermediates; location incentives come from R&D investment effects.

### 2.3.4 Research and Development
- Roles of innovation:
  - Increases quality of intermediate goods — drives growth.
  - Adds products to a firm’s portfolio by making it the quality leader for a variety.
- Innovation step:
  - If firm f innovates on product j, it produces quality (1 + λ) q_j(t), with λ > 0 (step size).
  - Innovator becomes new technical leader — “steals” product from previous producer.
- R&D investment and stochastic arrival:
  - Number of innovations per period ~ Poisson; in continuous time each firm produces at most one innovation per period.
  - Firms cannot target specific product lines; each innovation is uniformly random over j ∈ J.
  - Consequence: no strategic targeting across firms; probability a firm innovates its own product is zero.
- Arrival rate of innovation:
  - x_{f,c}(t) = χ_c(t) [Ĩ_c(t)^η i_{f,c}(t)]^ψ
  - Ĩ_c: population of inventors per unit of land in city c.
  - Inventor productivity ∝ Ĩ_c^η, with η ≥ 0; parameter η controls agglomeration externality strength.
  - ψ appears as exponent on the inventor-adjusted term.
- Fixed cost and scaling:
  - Fixed cost κ > 0 inventors must be hired each period a firm invests in R&D.
  - Fixed cost reflects managerial/maintenance R&D costs; paid only in periods firm invests.
  - Fixx: arrival rate parameter used in Poisson example; P(N = k) = e^{−Δx} (Δx)^k / k!.
- Modeling assumptions and implications:
  - Population of inventors is the relevant agglomeration measure (empirically tested).
  - Innovation does not scale with firm size → all firms in same city innovate at same rate.
  - Agglomeration spillovers identical for all inhabitants of a city and non-existent outside city borders.
  - No sorting into cities (inventors and firms homogeneous); ex-post productivity differences across cities captured by ̄χ_c.
  - Firms cannot target innovations at their own goods → all innovation generates creative destruction.
  - City 0: a city with ̄χ_0 = 0 to represent cities that never produced a patent; firms there can still produce goods but no inventors live there.

### 2.4 Local Policies and the Government
- Policy of interest: local R&D subsidies s_c that transfer a share of each firm’s R&D cost back to the firm; s_c varies by city.
- Public good G:
  - Government provides nationally available public good G; amount fixed at ̄G for all t.
  - Government fully taxes all firm- and land-owners to finance expenditures.
  - Public good provides redistribution of profits and land rents back to workers and balances the government budget given s_c from data.
- Other local taxes:
  - Corporate income and labor income taxes not explicitly modeled; their effects:
    - Corporate tax effects captured in ̄χ_c.
    - Labor income tax effects captured in α_c.
  - Note: time-varying local taxes correlated with evolution of local R&D subsidies could bias estimation; addressed in section 4.4 of the source.

### 3 The Balanced Growth Path Equilibrium
- Equilibrium restrictions imposed:
  - Solve for a Balanced Growth Path where aggregate variables grow at a constant rate.
  - Require each local shock z_c(t) follows its stationary distribution.
- Definition 1 (Stationary Balanced Growth Path Equilibrium) given initial {q_j(0) > 0}_{j∈J} and {s_c, α_c, ̄χ_c}_{c=0}^C with χ_c(t) = ̄χ_c e^{z_c(t)}, ̄χ_0 = 0 and dz_c = ψ(μ − z_c) + σ dW_c:
  - Equilibrium consists, for all t ≥ 0, of:
    - (a) allocation Y(t), {n_c(t)}_{c=0}^C, {k_j(t)}_{j∈J}
    - (b) spatial distribution {I_c(t), L_c(t), N_c(t)}_{c=0}^C of inventors, production workers, and firms
    - (c) prices {w^i_c(t), w^ℓ_c(t), p_{n,c}(t)}_{c=0}^C
  - Conditions:
    - (i) z_c(t) ~ Normal(mean μ, variance σ^2) for all t and c.
    - (ii) Final good Y(t), average quality Q(t) = ∫_{j∈J} q_j(t) dj, “baseline” wages (congestion-adjusted), and consumer utility grow at a constant rate.
    - (iii) Workers freely mobile and maximize utility over final and non-tradable consumption and city choice.
    - (iv) Final and non-tradable producers maximize profits; intermediate producers operate under monopolistic competition per product line j.
    - (v) Incumbent firms take location as given and choose R&D to maximize discounted profits; free entry to all cities with many potential entrants.
    - (vi) All labor and goods markets clear; public good ̄G balances government budget.

### 3.1 The Firm’s Static Problem
- Decomposition:
  - Static problem: firm chooses production of each good given its product set and qualities.
  - Dynamic problem: firm chooses R&D investment after observing local productivity shock.
- Final goods production location:
  - Final good production occurs in city 0 (lowest wages / least congestion): city 0 has χ_0(t) = 0 for all t, implying no R&D investment and no inventors there.
  - Final good producer’s profit maximization (time dropped for notation):
    - max_{ℓ_{y,0}, {k_j}_{j∈J}} ℓ_{y,0}^{ε/(1−ε)} ∫_J k_j^{1−ε} q_j^ε dj − ∫_J p_j k_j dj − w^ℓ_0 ℓ_{y,0}
  - First-order conditions:
    - [ℓ_{y,0}]: ε/(1−ε) ℓ_{y,0}^{ε−1} ∫_J k_j^{1−ε} q_j^ε dj = w^ℓ_0
    - [k_j]: p_j = (ℓ_{y,0} q_j k_j)^{ε}, ∀ j ∈ J
- Intermediate goods production given demand p_j:
  - Firm chooses k_j to max p_j(k_j; q_j) k_j − ν k_j.
  - Solution: k_j = q_j [(1 − ε)/ν]^{1/ε} ℓ_{y,0}
  - Profit from product j: π_j = ℓ_{y,0} [(1 − ε)/ν]^{1−ε} ε/ε q_j

_Italic: Source: wpiea2022131-print-pdf - 2.3   Technology_

### 3.2   Local Wages and Congestion Costs

### 3.2   Local Wages and Congestion Costs

### Local wages, mobility, and congestion
- Free mobility implies u_h^c(t) = u_h(t) for all c, t and h ∈ {i, ℓ} in equilibrium: utility levels of workers must be the same in all cities and periods.
- Wages adjust to compensate workers for variation in amenities or the price of the non-tradable good across cities.
- Wages differ between inventors and production workers due to differences in supply and demand for each worker type.
- Congestion costs arise because production of the non-tradable good uses a fixed factor (land) and thus displays decreasing returns to scale (DRS). Demand for non-tradable goods increases with city population, so larger cities face more expensive non-tradable goods, generating congestion costs that operate through prices (pecuniary externality).
- Congestion limits city size: as population rises, cost of living/producing rises and firms may locate elsewhere.

### Lemma 1 — wages, land rents, and population proportionality
- Use tilde notation to denote variables per unit of land: if I_c is population of inventors in city c, then ˜I_c = I_c / ̄m_c.
- Wages of inventors in cities c ∈ {1,...,C}:
  - w_i^c = w_i (˜I_c^{1−β} α_c)^{θ/(1−θ)}
  - where w_i = 1/(1−θ) [ u_i [(1−θβ)]^{θ} ]^{1/(1−θ)} [ I/(L−L_0) ]^{θβ/(1−θ)}.  (Equation (5) as presented)
- Wages of production workers:
  - w_ℓ^c = w_ℓ (˜L_c^{1−β} α_c)^{θ/(1−θ)}
  - where w_ℓ = 1/(1−θ) [ u_ℓ (θβ)^{θ} ]^{1/(1−θ)}.  (Equation (6) as presented)
- Land rents in city c:
  - p_{m,c} ̄m_c = (1−β) θ/(1−θβ) w_i^c I_c.
- City 0 (special city) hires ℓ_{n,0} = θβ L_0 production workers for non-tradable goods and ℓ_{y,0} = (1−θβ) L_0 for final good production.
  - Wage of production workers in city 0:
    - w_ℓ^0 = w_ℓ (˜L_0^{1−β} α_0)^{θ/(1−θ)} (θβ)^{(1−β)θ/(1−θ)}.  (Equation (7))
  - Total land rent in city 0: p_{m,0} ̄m_0 = (1−β) θ w_ℓ^0 L_0.
- Proportionality between populations:
  - Numbers of inventors and production workers in cities with innovation are proportional: I_c / I = L_c / (L − L_0). Thus population of production workers characterizes inventors and vice-versa.
- Alternative form for w_ℓ^0 by plugging (4) into F.O.C. of final good producer:
  - w_ℓ^0 = ε/(1−ε) [ (1−ε)/ν ]^{(1−ε)/ε} Q, where Q = ∫_J q_j dj is average quality of all intermediate goods.  (Equation (8))

---

### 3.3   The Firm’s Dynamic Problem

### Timing and agents
- Period timing:
  (i) shock z_c(t) realized and observed in all cities;
  (ii) potential entrants decide whether to enter and where to locate;
  (iii) entrants and incumbents decide inventors to hire;
  (iv) innovations realized (based on arrival rates x_{f,c}) and production occurs.
- Firms: entrants and incumbents. Incumbents cannot move; entrants choose city.

### Incumbents — HJB characterization (Lemma 2)
- Define q_f as multiset of qualities of products the firm produces; D is aggregate rate of creative destruction (also probability any product line is "stolen"); r is exogenous interest rate; A = (Q, w_i, D, L_0) is aggregate state; ̄π = (1−θβ) [ (1−ε)/ν ]^{1−ε} ε^{ } (definition as in text) so per-period profit π_j = ̄π L_0 q_j; Z_c = e^{z_c} and χ_c = ̄χ_c Z_c.
- HJB for incumbent in city c ∈ {1,...,C}:
  - r V_c(q_f, ˜I_c, Z_c, A) − ∂V_c/∂A ∂A/∂t
    = max_{x_{f,c}} { Σ_{q_j ∈ q_f} ̄π L_0 q_j + x_{f,c} E_j[ V_c(q_f ∪_+ {(1+λ) q_j}, ˜I_c, Z_c, A) − V_c(q_f, ˜I_c, Z_c, A) ]
      − (1−s_c) w_i^c (i_{f,c} + κ) − D Σ_{q_j ∈ q_f} [ V_c(q_f, ˜I_c, Z_c, A) − V_c(q_f \_{−} {q_j}, ˜I_c, Z_c, A) ] + R_c(q_f, ˜I_c, Z_c, A) }
  - Arrival rate: x_{f,c} = ̄χ_c Z_c (˜I_c^{η} i_{f,c})^{ψ}.
- Interpretation of HJB terms:
  - First term: profit flow from production and sale.
  - Second: expected gain from one more innovation (Poisson arrivals).
  - Third: cost of R&D investment (variable and fixed), subsidized at rate s_c.
  - Fourth: expected cost from loss of product line due to creative destruction (at rate D).
  - R_c(...) captures risk from city productivity shock.
- Corporate income taxes do not appear in HJB: taxing firm owners (profits including R&D expenditures) does not affect firm decisions; a full tax shifts firm owner share but leaves allocation unchanged.

### Entrants — two-stage problem and HJB
- Entrant problem:
  Step 1: choose city after observing {Z_c}_{c=1}^C:
    - V_e(A) = max_c V_e^c(˜I_c, Z_c, A).
  Step 2: choose level of innovation given city c:
    - r V_e^c(˜I_c, Z_c, A) − ∂V_e^c/∂A ∂A/∂t
      = max_{x_{f,c}} { x_{f,c} E_j[ V_c(q_j, ˜I_c, Z_c, A) − V_e^c(˜I_c, Z_c, A) ] − (1−s_c) w_i^c (i_{f,c} + κ) + R_e^c(˜I_c, Z_c, A) }
    - subject to x_{f,c} = ̄χ_c Z_c (˜I_c^{η} i_{f,c})^{ψ}.
- Second-stage HJB analogous to incumbent's but omits flow profits and creative-destruction losses from existing product lines.

### Proposition 1 — value functions in Stationary Balanced Growth Path (SBGP)
- In SBGP with final goods growth rate g < r:
  - Incumbent value:
    - V_c(q_f, ˜I_c, Z_c, A) = F(D, L_0) Σ_{q_j ∈ q_f} q_j + max{0, E_c(˜I_c, Z_c, w_i/Q, D, L_0) } Q,
    - where F(D, L_0) = ̄π L_0 / (r + D) is the "franchise value" of adding a new product and E_c is the entry value for firms in city c.
  - Entrant second-stage value:
    - V_e^c(˜I_c, Z_c, A) = max{0, E_c(˜I_c, Z_c, w_i/Q, D, L_0) } Q.
- Interpretation:
  - F is quality-adjusted franchise value of adding a new product.
  - E_c Q is the value at entry for firms in city c and does not depend on firms' product portfolios.
  - max{0, E_c Q} reflects option value of R&D investment: if E_c ≥ 0, firms invest (x_{f,c} > 0); if shocks are too low, firms may choose i_{f,c} = 0, x_{f,c} = 0, and avoid fixed cost w_i^c κ.
- Technical requirement g < r: ensures PDV of firms finite; if g > r, optimal infinite R&D leads to divergence.

---

### Free Entry and Spatial Equilibrium

### Free entry condition
- With large mass of potential entrants, free entry implies in equilibrium:
  - V_e^c(˜I_c, Z_c, A) = 0 for all c and t.
- Intuition:
  - If entry value positive, more entry increases congestion until entry value is zero.
  - If entry value negative, no entrants and incumbents refrain from R&D; reduced demand for inventors reduces congestion and raises entry value back to zero.
- Implication from Proposition 1: E_c = 0 regardless of state variables. Population of inventors must adjust so entry value is zero in all cities.

### Proposition 2 — population of inventors and firm-level rates
- Under (1) free entry, (2) labor market clearing for inventors and production workers, and (3) large number of cities C → ∞:
  - Population of inventors in city c:
    - I_c = I × [ ̄χ_c^{1−s_c} ]^{(1−θ)/Θ} α_c^{θ/Θ} × [ Σ_{c=1}^C [ ̄χ_c^{1−s_c} ]^{(1−θ)/Θ} α_c^{θ/Θ} ]^{-1} × Z_c^{(1−θ)/Θ ( (1−θ)/Θ − 1 ) σ^2 / (4 φ) }.
    - (Equation (9) as presented; Θ = (1−β) θ − ψ η (1−θ))
  - Arrival rate of innovation for firm f in city c (if firm invests):
    - x_{f,c} = [ κ/ψ (1−ψ) ]^{ψ} ̄χ_c ˜I_c^{ψ η} Z_c.  (Equation (10) as presented)
  - Number of inventors hired per firm (if investing):
    - i_{f,c} = ψ/(1−ψ) κ.
  - Number of firms in city c that invest in R&D each period:
    - N_c = [ (1−ψ)/κ ] I_c.  (Equation (11))
  - Population of production workers in city 0 is proportional to L (L_0 does not vary over time).
  - Baseline wage w_i is not affected by city-specific productivity shocks and scales with Q as:
    - w_i / Q ∝ ̄π L_0 / (r + D) ( 1/C Σ_{c=1}^C [ ̄χ_c^{1−s_c} ]^{1−θ/Θ} α_c^{θ/Θ} )^{Θ/(1−θ)}.  (Equation (12) as presented)

### Key implications from Proposition 2
- Closed-form solution for I_c in equation (9): cities with higher amenities (α_c), higher innovation productivity (̄χ_c), and higher R&D subsidies (s_c) have more inventors; Θ > 0 given parameter values in section 4.
- I_c responds to city productivity shocks Z_c: population of inventors increases in periods with larger Z_c.
- Θ interpretable as "net elasticity" of congestion: (1−β) captures elasticity of congestion w.r.t population; ψ η captures elasticity of innovation production w.r.t inventor population; weighted by θ.
- Baseline wage of inventors (and thus production workers) does not react to city-specific productivity shocks — shocks average out on aggregate allowing SBGP without aggregate uncertainty.
- Optimal arrival rate of innovation x_{f,c} is uniform across firms that make positive R&D investments in the same city. However, expected value of investing in R&D is null because of free entry, so some firms may invest and some may not; the number of inventors hired by investing firms is fixed and does not respond to productivity shocks. Adjustments to shocks occur on the extensive margin (incumbent investment decisions and firm entry/exit).

---

### 3.4   Determining the Growth Rate

### Aggregate creative destruction (Corollary 1)
- Because all firms in the same city choose identical R&D, aggregate D = Σ_{c=1}^C N_c x_{f,c}.
- Define ̄I_c = E[I_c] with respect to local productivity shocks as
  - ̄I_c = I [ ̄χ_c^{1−s_c} ]^{(1−θ)/Θ} α_c^{θ/Θ} × [ Σ_{c=1}^C [ ̄χ_c^{1−s_c} ]^{(1−θ)/Θ} α_c^{θ/Θ} ]^{-1}.
  - ˜̄I_c is expected density of inventors in city c.
- Aggregate rate of creative destruction:
  - D ∝ (1/C) Σ_{c=1}^C ̄χ_c ˜̄I_c^{1+ψ η}.  (Equation (13))
- Key takeaway: aggregate innovation rate depends on both total inventors and their geographic distribution. Local R&D subsidies that change spatial distribution of inventors can affect aggregate innovation even if average subsidy or total expenditure is unchanged.
- On SBGP, ̄χ_c and ˜̄I_c are fixed over time, so D is constant.

### Proposition 3 — growth rate and dynamics
- In SBGP where final good production grows at rate g:
  (1) Average quality Q and baseline wages w_i and w_ℓ all grow at rate g. Utility levels u_i and u_ℓ grow at rate (1−θ) g.
  (2) Growth rate satisfies g = λ D.
  (3) Let J_c(t) be set of intermediate goods produced in city c at time t, Q_c(t) = ∫_{J_c(t)} q_j(t) dj aggregate quality, and g_c(t) = ̇Q_c(t)/Q_c(t). For large t:
    - E[ ̇Q_c(t) ] / E[ Q_c(t) ] = g, equivalently E[g_c(t)] = g − Cov(g_c(t), Q_c(t)) / E[ Q_c(t) ].
- Interpretations:
  - Part (1): aggregate variables Q, w_i, w_ℓ grow at g; utilities grow at (1−θ) g.
  - Part (2): national growth rate equals innovation step-size λ times aggregate creative destruction D.
  - Part (3): ratio of expected variation in city aggregate quality to expected quality converges to national growth rate over time. Creative destruction generates mean reversion in city-level quality: cities with above-average quality tend to innovate on lower-quality products and vice-versa, preventing concentration of all economic activity in a single city.
- Results do not depend on initial spatial distribution as long as no single city dominates aggregate evolution.

*International Monetary Fund — IMF Working Papers, "Agglomeration, Innovation, and Spatial Reallocation" (section 3.2–3.4 as provided)*

### 3.5   Existence and Uniqueness

### 3.5   Existence and Uniqueness

### Existence conditions
- Two conditions required for existence of a solution for the SBGP equilibrium:
  - r > g so that present discounted values of profits are finite.
  - The number of cities must be large enough so that the local shocks do not generate aggregate uncertainty in the economy.
- Note: "An equilibrium could still exist if this condition is violated, but it would not be a balanced growth path."

### Uniqueness condition
- Uniqueness comes from the unique spatial distribution of inventors defined by equation (9).
- Caveat: the distribution is only unique if the net elasticity of congestion Θ > 0.
  - If Θ ≤ 0, agglomeration forces dominate congestion and it is always profitable for all firms to locate in the same place, generating multiple equilibria.
  - Example implication: if the entire population is located in city ĉ in the initial period, entrants have no incentive to locate elsewhere.

### Transition to estimation (beginning of Section 4)
- Identification and estimation proceed in three steps:
  1. Calibrate parameters that can be directly matched to data or literature (section 4.1).
  2. Use linear regressions to estimate the elasticity of the agglomeration spillover with respect to the population of inventors, and the elasticity of congestion with respect to the population of production workers (section 4.2).
  3. Identify remaining parameters by matching model moments to data moments (section 4.3).
- Data sources and construction:
  - Patent filings: USPTO Patent Dataset (patents, inventors, assignees, locations).
  - Demography and economic activity: County Business Patterns Dataset (CBP).
  - Prices of non-tradable goods: Zillow Rent Index (ZRI) — median rental value per square foot across counties.
  - Empirical city unit: CBSAs (core-based statistical areas) based on 2010 Census standards.
  - Estimation sample: panel of firms (and locations) who have filed patents between 1998 and 2016.
  - Appendix C.1: dataset construction details (referenced).

### 4.1 Calibration — key choices and logic
- Macroeconomic rates and literature choices:
  - g = 0.02 (rate of growth of the economy; annualized historic rate of growth in the US).
  - ρ = 0.02 (discount rate of consumers).
  - r = 0.038 (real interest rate; average interest rate in the US between 1961 and 2017 per World Bank).
- Innovation parameters:
  - ψ = 0.5 (curvature of the innovation production function; elasticity of patents w.r.t. R&D expenditure).
  - λ = 0.132 (innovation step-size; value estimated by Acemoglu et al. (2018)).
- R&D subsidies:
  - s_c set equal to effective R&D tax credit rate applying to the highest tier of R&D investments in each state, as computed by Wilson (2009).
  - Use average credit rate between 1998 and 2006 to represent s_c due to time variation in statutory rates.
  - For CBSAs spanning multiple states, match CBSA to the state of its largest urban center and apply that state's R&D credit rate to the CBSA.
- Production and preferences:
  - ε = 0.15 (elasticity of quality in production of the final good; coincides with profit/sales ratio for intermediate good producers, BEA).
  - θ = 0.6 (preference parameter; share of expenditure on non-tradable goods by consumers).
- Population and city counts:
  - Inventor population I constructed from unique inventor IDs in USPTO data: I_pat is average number of inventors authoring patents; I_pat = ψ I ⇒ I = I_pat / ψ.
  - Total employed population L + I set to match CBP total employed population.
  - Normalize total number of inventors in the economy to I = 1.
  - There are 917 CBSAs in the US (excluding Puerto Rico), of which 860 have filed at least one patent between 1998 and 2016.
  - Assume remaining 57 CBSAs produced no innovation over sample; set C = 860 (cities with positive expected productivity in innovation) and city0 represents remaining 57 CBSAs.

### Calibrated parameter table (values preserved)
- ψ 0.5 — Elast. innovation wrt R&D — Literature
- λ 0.132 — Innovation step size — Acemoglu et al. (2018)
- s_c [0.13,0.30] — Effective R&D credit rate — Wilson (2009)
- ρ 0.02 — Discount rate — Literature
- g 0.02 — Growth rate — Annualized growth rate
- r 0.038 — Real interest rate — Avg. real interest rate
- ε 0.15 — Elast. quality in final goods — Profit/sales ratio (BEA)
- θ 0.6 — Preference parameter — Share of expenditure in non-tradables (BLS; Bems, 2008)
- I 1 — Population of inventors — Avg. number of inventors residing in CBSA’s (1998 - 2016)
- L 175 — Population of production workers — Avg. employed population residing in CBSA’s (1998 - 2016)
- L_0 0.85 — Population of production workers in CBSA’s that do not innovate — Avg. employed population residing in CBSA’s w/ no patents filed (1998 - 2016)

### 4.2 Linear regressions — identification of η and β
- Objective: identify elasticities of agglomeration (η) and congestion (β) via linear regressions implied by the model.
- Elasticity of agglomeration — log-linear regression derived from the innovation production function:
  - Transform continuous-time equation (1) to yearly frequency (appendix C.2), yielding:
    - log(x_{f,c,t}) = ψ log(i_{f,c,t}) + ψ η log(I_{c,t}) + δ_c + z_{f,c,t}, (14)
      - x_{f,c,t} = number of innovations produced by firm f in city c during year t.
      - i_{f,c,t} = number of inventors hired by firm f during year t.
      - I_{c,t} = population of inventors in city c during year t.
      - δ_c = city fixed-effect.
      - z_{f,c,t} = function of firm- and city-specific productivity shocks.
- Measurement and controls:
  - Use number of patents filed by a firm in that year as proxy for x_{f,c,t}.
  - Include two controls to reduce mismatch between patents and true innovation:
    - Total number of citations that the patents filed by each firm jointly receive, interacted with a dummy for the year of the patent application (to adjust for fewer citations for recent patents).
    - Firm’s industry indicator.
  - Add a year fixed effect to capture aggregate variations over time.
- Endogeneity concern:
  - Model prediction: both i_{f,c,t} and I_{c,t} are correlated with local shock z_{f,c,t}, so OLS on regression (14) would not recover ψ or η.
    - i_{f,c,t} is correlated with z_{f,c,t} because productivity shocks affect the number of firms investing in R&D (low shock → fewer firms invest; high shock → more firms invest and entrants).
    - I_{c,t} depends on city productivity each period (via equation (9)), implying correlation with z_{f,c,t}.

*Source: IMF Working Paper (excerpt sections 3.5, 4.1–4.2).*

### appendix C.2 for more details). In practice what this means is that estimating the coefficients on

### wpiea2022131-print-pdf - appendix C.2 for more details). In practice what this means is that estimating the coefficients on

### Estimating the elasticity of agglomeration (η)
- Regression setup:
  - Starting point: regression (14) rearranged to construct left-hand side using known ψ from prior literature:
    - log(patentsf,c,t) − ψ log(i f,c,t) = ψ η log(I c,t) + X′f,c,t Γ + δc + δt + z f,c,t
  - Aggregated to city level (equation 15):
    - (1/N c,t) Σf [log(patentsf,c,t) − ψ log(i f,c,t)] = ψ η log(I c,t) + X′c,t Γ + δc + δt + z c,t
  - Dependent variable: average log production of patents per inventor in each firm; number of inventors per firm transformed by raising to the elasticity of labor in innovation.
- Instrument for I c,t (equation 16):
  - I c,t = Σk I k,c,t = Σk I k,c,t−l (1 + γ k,c,t−l→t)
  - Instrument: I c,t,l = Σk I k,c,t−l (1 + γ k,t−l→t)
  - γ k,t−l→t computed using inventors residing outside of the US (about 50% of all registered patents during sample period), industries defined by NBER patent subcategories (38 industries), inventor assigned by modal sub-class.
  - Instrument structure resembles shift-share design; I c,k,t−l is population level exposure; γ k,t−l→t are shifters.
- Standard errors and inference:
  - Two sets of SEs when IV used: (i) clustered across regions; (ii) adjusted using Adão et al. (2019) methods (AKM SE).
  - Regressions weighted by number of firms in each city; standard errors clustered at city (CBSA) level.

### Results for η (Table 2 summary)
- First-stage coefficient on log(I c,t,l):
  - Column (1) OLS: not applicable as first-stage reported for IV columns only.
  - IV columns: 0.546∗∗∗, 0.481∗∗∗, 0.398∗∗∗, 0.244∗∗∗ (standard errors in parentheses: (0.031),(0.033),(0.035),(0.042))
  - F-statistics: 315.10, 221.47, 130.93, 33.94
- Second-stage coefficients on log(Inventors in City):
  - OLS (col 1): 0.070∗∗∗ (0.014)
  - IV (l=5) (col 2): 0.104∗∗∗ (0.019)
  - IV (l=7) (col 3): 0.098∗∗∗ (0.020)
  - IV (l=10) (col 4): 0.105∗∗∗ (0.023)
  - IV (l=t−t90−95) (col 5): 0.104∗ (0.060) with AKM SE reported as (0.014),(0.017),(0.016),(0.001) for respective IV columns
- Observations by column: 11279, 11231, 11220, 11201, 11210
- Implied η from columns:
  - OLS: 0.140
  - IV (l=5): 0.208
  - IV (l=7): 0.196
  - IV (l=10): 0.210
  - IV (l=t−t90−95): 0.208
- Statistical significance:
  - Coefficients generally highly significant; notation: ∗, ∗∗, ∗∗∗ indicate significance at 10%, 5%, 1% respectively.
- Comparative context:
  - Estimated coefficients range between 0.07 and 0.10 for ψ η (implying η in reported implied values).
  - Literature comparison: Duranton and Puga (2014) report most studies find agglomeration elasticity between 0.02 and 0.05; Carlino et al. (2007) find elasticity approximately 0.19 for patents per capita.

### Identification conditions, robustness checks, and sensitivity for η
- Two interpretations of orthogonality for shift-share IV:
  - Goldsmith-Pinkham et al. (2020) condition: exposures I k,c,t−l must be uncorrelated with local shock z c,t — less plausible for small lags, more plausible for larger lags (e.g., 10-year lag) or when base-period industry levels fixed (e.g., 1990–1995).
  - Borusyak et al. (2022) condition: industry growth rates γ k,t−l→t asymptotically uncorrelated with E[I k,c,t−l z c,t ] across c; using non-US inventors for γ addresses many threats.
- Concentration check:
  - Re-run excluding industries whose employment share in any city exceeds 15% (thresholds 10–25% yield comparable results); estimates remain in line (appendix C.3.2).
- Zero-patent counts and selection:
  - Addressed by Poisson count-data specification (appendix C.3.3); resulting η ≈ 0.13−0.15 (slightly higher than log-log results).
- Alternative agglomeration sources:
  - Controls for number of firms (R&D investing), total employment, total establishments: coefficients either negative or not statistically significant after accounting for population of inventors.
- Sensitivity of η to ψ:
  - Using ψ = 0.4: estimated ψ η ≈ 0.11 (0.09 via OLS) → implied η = 0.275.
  - Using ψ = 0.6: estimated ψ η ≈ 0.09 (0.05 via OLS) → implied η = 0.15.
  - IV specifications remain statistically significant; implied η falls within range used in sensitivity analyses.

### Baseline choice for η in model applications
- Baseline used in section 5: η = 0.20.
- Sensitivity analyses in appendix F use η = 0.15 and η = 0.25.

### Estimating the elasticity of congestion (β)
- Theoretical linkage:
  - With consumer share on non-tradable good, β (returns to scale on non-tradable production) determines congestion elasticity.
  - With constant returns (β = 1) no congestion; as β → 0 congestion costs rise.
  - Empirical implication (equation 17):
    - log(p h c,t) = [(1 − β)/(1 − θ)] log(L c,t) + δc + δt + z h c,t
    - p h c,t: median rental value per square foot in city c, year t
    - L c,t: population of non-inventors in city c, year t
    - θ = 0.6 used to map coefficient to β
- Endogeneity and instrument:
  - L c,t correlated with z h c,t; use I c,t,l as instrument for L c,t (same instrument as for inventors).

### Results for β (Table 3 summary)
- First-stage coefficient on log(I c,t,l):
  - IV columns: 0.020∗∗∗, 0.023∗∗∗, 0.018∗∗∗, 0.023∗∗∗ (standard errors: (0.005),(0.008),(0.005),(0.008))
  - F-statistics: 14.58, 7.47, 12.95, 8.89
- Second-stage coefficients on log(Prod. Workers in City):
  - IV (l=5): 0.981∗∗∗ (0.326)
  - IV (l=7): 1.013∗∗∗ (0.394)
  - IV (l=10): 1.267∗∗∗ (0.392)
  - IV (l=t−t90−95): 0.902∗∗∗ (0.337)
  - AKM SE reported as (0.026),(0.023),(0.033),(0.013)
- Observations by column: 2855, 2849, 2845, 2846
- Implied β from columns:
  - 0.608, 0.595, 0.493, 0.639
- Comparison:
  - Behrens et al. (2014) find elasticity of rental prices w.r.t. population between 0.08 and 0.09; differences here partly due to inclusion of city fixed effects.
- Alternative housing price measure:
  - Using median housing price (series back to 1996) instead of rental values yields implied β ≈ 0.8 (housing prices less elastic to population than rents).

### Identification conditions, robustness checks, and sensitivity for β
- Identification mirrors that for η: instrument orthogonality conditions and concerns about concentrated industries addressed by lagging instrument and excluding industries with >15% inventor share in any city (appendix C.4).
- First-stage predictive power lower than for inventor instrument but F-statistics generally above common thresholds.
- Data limitation:
  - Rental values available in ZRI database only after 2010, explaining smaller sample size for Table 3.
- Sensitivity:
  - Additional IV specifications with lags up to 12 years and exclusion of highly concentrated industries reported in appendix C.4.

### Baseline choice for β in model applications
- Baseline used in section 5: β = 0.6.
- Sensitivity analyses use β = 0.5 and β = 0.8.

*Source: https://www.imf.org/-/media/files/publications/wp/2022/english/wpiea2022131-print-pdf.pdf*

### 4.3   Moment Matching

### 4.3 Moment Matching

### Fixed Cost of Innovation
- Equation (11) relates the number of firms in each city to the number of inventors in the city. Summing over cities and rearranging gives:
  - κ = (1−ψ) I / N
- Calibrated values and implication:
  - I/N ≈ 21.07
  - ψ = 0.5
  - κ = 10.53

### City-Specific Parameters (αc and ̄χc)
- Identification strategy:
  - Average share of inventors in city c over time converges (ergodic theorem) to:
    - (avg. share of inventors)c = [̄χc 1−sc]^{1−θ/Θ} αc^{θ/Θ} / Σ_{c=1}^C [̄χc 1−sc]^{1−θ/Θ} αc^{θ/Θ} ≡ ̄Ic / I
  - Average share of patents filed in city c:
    - (avg. share of patents filed)c = ̄χc ̄I_c^{1+ψη} / Σ_{c=1}^C ̄χc ̄I_c^{1+ψη}
- Normalizations and scale identification:
  - Normalize E_c[αc] = 1 (αc is a preference parameter).
  - Scale of ̄χc identified by imposing model growth rate g = λD = 2% (historic annualized US rate).
  - Amenity in city 0 found by matching population share L0/(I+L) to data.
- Note: the value of σ^2/4φ must be known before implementing these procedures.

### Law of Motion of the Productivity Shock (σ and φ)
- Only the ratio σ^2/φ matters for equilibrium; set φ = 1.
- σ is identified by matching the model-generated cross-sectional variance of the population of inventors between cities to the same moment in the data.
- Appendix D.2 derives the expression for this variance and shows how to identify σ.

---

### 4.4 Comparison to Untargeted Moments

### Model Fit to Untargeted Spatial Distributions
- Untargeted variables assessed: share of firms per city, average patents per firm, share of employed population per city, patents per capita.
- Correlations between model and data (Panel A of Table 4):
  - Number of Firms: 0.97
  - Patents per Firm: 0.58
  - Employed Population: 0.81
  - Patents per Capita: 0.49
- Observations on fit:
  - High fit for share of firms per city (correlation ≈ 0.97).
  - Patents per firm harder to match due to many cities with on average one patent per firm (vertical alignment in panel (b) of figure A.3).
  - Model tends to underestimate total population in cities with few inventors and overestimate population in cities with many inventors (model predicts inventor and production-worker populations proportional, but actual cities specialize).
  - Correlations in panel A of table 4 summarize these matches.

### Spatial Distribution by City-Size Quintiles (Panel B of Table 4)
- Cities ranked by average population of inventors (1998–2016) and divided into five equal bins; model vs data comparison for share of firms and avg. patents per firm:
  - Bin 1:
    - Share of Inventors (model): 0.002
    - Share of Inventors (data): 0.001
    - Share of Patents (model): 0.002
    - Share of Patents (data): 0.003
    - Share of Firms (model): 1.052
    - Avg. Patents/Firm (data): 0.81
  - Bin 2:
    - Share of Inventors (model): 0.006
    - Share of Inventors (data): 0.002
    - Share of Patents (model): 0.006
    - Share of Patents (data): 0.007
    - Share of Firms (model): 1.006
    - Avg. Patents/Firm (data): 0.875
  - Bin 3:
    - Share of Inventors (model): 0.013
    - Share of Inventors (data): 0.006
    - Share of Patents (model): 0.013
    - Share of Patents (data): 0.017
    - Share of Firms (model): 1.234
    - Avg. Patents/Firm (data): 0.993
  - Bin 4:
    - Share of Inventors (model): 0.040
    - Share of Inventors (data): 0.018
    - Share of Patents (model): 0.040
    - Share of Patents (data): 0.045
    - Share of Firms (model): 1.271
    - Avg. Patents/Firm (data): 1.171
  - Bin 5:
    - Share of Inventors (model): 0.939
    - Share of Inventors (data): 0.973
    - Share of Patents (model): 0.939
    - Share of Patents (data): 0.927
    - Share of Firms (model): 2.181
    - Avg. Patents/Firm (data): 2.894

---

### 4.4.1 Can R&D Tax Credits Shift the Spatial Distribution of the Economy?
- Motivation: assess how changes in spatial distribution of R&D tax credits over time affect the locations of inventors and firms.
- Empirical strategy:
  - Use model with year-specific R&D tax credit rates (all other parameters fixed) to construct city-level shares of inventors and patents, then average by decade.
  - Decade averages analyzed: 1970–1979, 1980–1989, 1990–1999.
  - Rationale for decades:
    - 1970s: no subsidies
    - 1980s: spatially uniform federal subsidy plus a few state subsidies
    - 1990s: most states adopted subsidies
- Panel A of Table 5: correlations between model and data by decade and quintile aggregation (selected highlights):
  - Correlations (Panel A overall): 1970s Corr. = 0.86; 1980s Corr. = 0.92; 1990s Corr. = 0.96 for inventors in levels (and other reported correlations for patents/time periods)
- Panel B of Table 5: regressions of differences in data on differences in model (decade changes relative to 2000), for ∆Share of Inventors and ∆Share of Patents:
  - Coefficients (∆Share of Inventors): 1970s: 2.27; 1980s: 1.78; 1990s: 2.15
  - Std. Error (∆Share of Inventors): 1970s: 0.09; 1980s: 0.06; 1990s: 0.08
  - R-squared (∆Share of Inventors): 1970s: 0.39; 1980s: 0.45; 1990s: 0.44
  - Coefficients (∆Share of Patents): 1970s: 3.27; 1980s: 2.77; 1990s: 3.16
  - Std. Error (∆Share of Patents): 1970s: 0.08; 1980s: 0.06; 1990s: 0.08
  - R-squared (∆Share of Patents): 1970s: 0.64; 1980s: 0.68; 1990s: 0.64
  - Observations: 860 for each regression
- Interpretation:
  - R&D tax credits are relevant for location decisions and production of innovation.
  - Changes in R&D tax credit rates explain about 40% of variation in changes in population shares and about 60% of variation in changes in patent production (R-squared evidence).
  - Correlations indicate model predictions often align with observed directional changes.

### R&D Tax Credits vs Corporate and Labor Income Taxes
- Concern: whether R&D tax credit changes move in tandem with corporate/labor taxes, confounding attribution.
- Data window: 1977–2006.
- Year-over-year changes computed for each state in:
  - statutory R&D tax credit rate
  - state corporate income tax rate
  - state labor income tax rate for top income bracket
- Correlations across states (1977–2006):
  - Correlation between changes in R&D tax credit rate and corporate income tax changes: −0.01
  - Correlation between changes in R&D tax credit rate and labor income tax changes: −0.03
- Statistical significance:
  - Correlations are very small and not statistically significant at the 10% level.
  - Multiple regression specifications (including lags/leads, fixed effects, adjusted recapture rates) produce small, statistically insignificant coefficients.

---

### 5 Welfare Effects of Spatial Policies

### Overview of Questions and Approach
- Two-part comparison for welfare implications of spatial R&D subsidy distribution:
  1. Compare current spatially heterogeneous U.S. distribution of R&D subsidies (decentralized state choices) with a spatially homogeneous subsidy implemented with the same total resources.
  2. Compare current distribution with the theoretical maximum welfare obtained by a central planner choosing all local R&D subsidies to maximize welfare (approximate solution computed).
- Aggregate welfare measure:
  - Sum of utility of all workers (firm- and land-owners fully taxed).
  - Cost of producing ̄G units of public good: γ(̄G) = ̄π ̄G Q(t).
  - Π(t) denotes aggregate flow of profits.

### Government’s Problem (static reduction)
- Dynamic maximization problem (continuous time) reduced to a static problem (see appendix E). Core structure (variables depend on vector of subsidies s = (s1,...,sC)):
  - Government objective involves terms reflecting:
    - population distribution ̃̄Ic and L0
    - wages ̄wi
    - rate of creative destruction D and growth λD
    - direct expenditure effects
- Important tradeoff highlighted:
  - Term ̄wi(s)^{1−θ} / [ρ − (1−θ)λD(s)] increases with D (higher growth raises PV of welfare).
  - But normalized wage ̄wi decreases with D (higher creative destruction raises firms’ discount rate r + D, reducing R&D investment and demand for inventors), lowering welfare through lower wages.
  - Corollary: D ∝ (1/C) Σ_{c=1}^C ̄χc ̃̄Ic^{1+ψη}_c — a more spatially concentrated population increases the rate of innovation, especially when concentrated in high-̄χc cities.

---

### 5.2 A Spatially Homogeneous Subsidy
- Counterfactual: states cannot compete and set sc ≡ ̄s for all c; ̄s determined by government budget constraint.
- Calibrated value and comparison:
  - Spatially homogeneous subsidy rate ̄s is close to 19% (average subsidy rate under current distribution ≈ 16%).
- Effects of moving to homogeneous subsidy:
  - HHI index of city population shares moves from 0.027 to 0.025 (population more evenly spread).
  - Aggregate welfare falls by 0.77%.
  - Decrease in growth rate of approximately 0.03 percentage points.
  - Static baseline wage ̄wi increases by 0.91%.
- Interpretation:
  - Decentralized adoption of R&D subsidies by states has produced higher welfare than a spatially neutral subsidy implemented with the same resources.
  - Suggests states offering largest R&D tax credits are comparatively better at producing innovation, leading to higher growth.

*Source: IMF Working Paper — Chapter/Section 4.3–5.2, "Agglomeration, Innovation, and Spatial Reallocation"*

### 5.3   Approximating the Optimal R&D Subsidies

### 5.3   Approximating the Optimal R&D Subsidies

### Methodology: functional-form approximation and parameterization
- The exact optimization problem (18) is non-convex with 860 choice variables (cities) and is computationally challenging.
- The approximate subsidy functional form imposed:
  - s_c = ( ζ α_c^ξ ̄χ_c^ω, if ζ α_c^ξ ̄χ_c^ω ≤ τ; τ, if ζ α_c^ξ ̄χ_c^ω > τ ).
  - Rationale: cities differ only by α_c or ̄χ_c, so differences in optimal subsidies are driven by α_c and ̄χ_c.
- Subsidy cap values considered: τ ∈ {0.3,0.4,0.5}.
  - The highest credit rate currently offered in the data (combining state and federal tax credits) coincides with the lowest cap value, at about 30%.
- Parameters ξ and ω are chosen in the interval [−5,10] to maximize aggregate welfare. The scale parameter ζ enforces the government budget constraint.

### Spatial pattern of optimal subsidies (cities)
- Aggregate-welfare maximization (example shown fixing τ = 0.4) indicates:
  - Welfare is maximized when innovation is concentrated in cities with high amenities and high productivity.
  - Optimal subsidy rates should move the economy toward greater concentration of innovation in such cities.
- Mapped results (Figure 2) identify heavy subsidy concentration in:
  - Silicon Valley (San Jose)
  - New York City
  - These are already the two largest producers of patents in the country.
- Mechanism:
  - High-productivity cities produce more innovation per worker → moving population there increases growth rate.
  - High-amenity cities induce workers to accept relatively lower wages → firms face less congestion cost.

### Welfare and macro effects: city-level subsidies (Table 6, Panel A)
- Gains from adopting optimal city-level subsidies (for each τ):
  - τ = 0.3:
    - ∆Welfare = 2.95%
    - ∆Baseline Wage = −3.60%
    - ∆Creative Destruction = 0.97 p.p.
    - ∆Rate of Growth = 0.13 p.p.
  - τ = 0.4:
    - ∆Welfare = 5.23%
    - ∆Baseline Wage = −6.38%
    - ∆Creative Destruction = 1.70 p.p.
    - ∆Rate of Growth = 0.22 p.p.
  - τ = 0.5:
    - ∆Welfare = 6.15%
    - ∆Baseline Wage = −7.75%
    - ∆Creative Destruction = 2.00 p.p.
    - ∆Rate of Growth = 0.26 p.p.
- Interpretation:
  - With τ = 0.5, total welfare would grow by at least 6.15% under the optimal city-level subsidy distribution.
  - The 6.15% welfare gain is generated in part by a 0.26 p.p. increase in the rate of growth.
  - Baseline wages fall by over 7.75% at τ = 0.5, reflecting higher creative-destruction-induced lowering of labor demand by innovative firms.

### Subsidies constrained to state level (Panel B and spatial implications)
- When subsidies are constrained to be constant within states, the approximate state-level subsidy:
  - s_c(S) = ( ζ (1/C(S)) Σ_{c=1}^{C(S)} α_c^ξ ̄χ_c^ω, if ζ (1/C(S)) Σ_{c=1}^{C(S)} α_c^ξ ̄χ_c^ω ≤ τ; τ, if ... > τ ).
- Panel B: Gains from adopting optimal state-level subsidies (for each τ):
  - τ = 0.3:
    - ∆Welfare = 2.50%
    - ∆Baseline Wage = −3.65%
    - ∆Creative Destruction = 0.88 p.p.
    - ∆Rate of Growth = 0.12 p.p.
  - τ = 0.4:
    - ∆Welfare = 3.06%
    - ∆Baseline Wage = −4.36%
    - ∆Creative Destruction = 1.07 p.p.
    - ∆Rate of Growth = 0.14 p.p.
  - τ = 0.5:
    - ∆Welfare = 3.23%
    - ∆Baseline Wage = −4.71%
    - ∆Creative Destruction = 1.13 p.p.
    - ∆Rate of Growth = 0.15 p.p.
- Key spatial result:
  - Optimal state-level subsidies differ from city-level results: the states most subsidized can shift (example: California and Idaho are identified as highly subsidized under state-level optimization, rather than California and New York).
  - Explanation: New York State contains multiple smaller cities that jointly produce a sizable share of innovation; subsidizing the state disperses effects across many cities. Idaho’s innovation is concentrated in Boise, so state-level subsidies more directly concentrate inventors in one city.
- Conclusion: The spatial distribution of optimal R&D subsidies can drastically change depending on the geographical scope (city vs. state) of the policy.

### Robustness and sensitivity
- The pattern of welfare gains from higher spatial concentration of innovation is robust to different values of agglomeration and congestion elasticities (sensitivity analyses in appendix F).

### Discussion: interpretation and limitations
- The welfare gains reported are the product of pure redistribution of existing R&D subsidies across space; expenditures on subsidies are kept constant across counterfactuals except for endogenous revenue changes from population reallocation.
- The computed gains are a lower bound because the imposed functional form may not capture the true policy that maximizes the government’s problem.
- Model limitations and omitted factors:
  - Moving costs are not included; their inclusion could materially affect welfare and optimal subsidy distribution.
  - Short-run adjustment costs (e.g., in R&D investments) are ignored; results represent long-run responses.
  - Other policy-relevant outcomes (e.g., income inequality, joblessness) that may be affected by spatial redistribution are outside the scope of this analysis.

### Implications (policy-relevant findings)
- Concentrating inventors in cities with both high amenities and high productivity:
  - Increases aggregate welfare substantially in the long run (example: up to 6.15% welfare gain at τ = 0.5 under city-level subsidies).
  - Raises the economy’s rate of growth (example: up to 0.26 p.p. increase at τ = 0.5).
  - Lowers baseline wages due to higher creative destruction (example: −7.75% at τ = 0.5).
- The geographical scope of subsidy policies (city-level vs. state-level) materially affects which locations should be subsidized to maximize welfare.

*Source: wpiea2022131-print-pdf — 5.3   Approximating the Optimal R&D Subsidies.*

### References

### References

### Bibliography highlights
- Extensive literature cited on agglomeration, innovation, R&D policy, spatial economics, and econometric methods, including (selection):
  - Acemoglu et al. (2018), American Economic Review, 108(11):3450–91.
  - Aghion and Howitt (1992), Econometrica, 60(2):323–351.
  - Autor, Dorn, and Hanson (2013), American Economic Review, 103(6):2121–2168.
  - Duranton and Puga (2014), Handbook of Economic Growth, volume 2, chapter 5, pages 781–853.
  - Goldsmith-Pinkham, Sorkin, and Swift (2020), American Economic Review, 110(8):2586–2624.
  - Borusyak, Hull, and Jaravel (2022), The Review of Economic Studies, 89(1):181–213.
  - Correia, Guimarães, and Zylkin (2019), ppmlhdfe: Fast Poisson Estimation with High-Dimensional Fixed Effects.
  - Wooldridge (2010), Econometric Analysis of Cross Section and Panel Data. The MIT Press, Cambridge, MA.
- Methodological and empirical references include works on shift-share designs, Bartik instruments, SBGPs, R&D taxation and incentives, patent data use, and spatial misallocation.

### Appendix A — Figures (contents and notes)
- Figure A.1–A.2: Spatial Distribution of R&D Tax Credit Rates; notes indicate average effective R&D tax credit rates in the US and discontinuities due to federal base computation changes and absence of federal credits in 1995.
- Figure A.3: Distributions in the Model and the Data: (a) Spatial distribution of firms; (b) Distribution of patents per firm; (c) Spatial distribution of employed population; (d) Distribution of patents per capita.
- Figure A.4: Correlation Between Changes City Population: Model vs Data. Note defines change in city population as difference between average population share in a decade minus average between 2000 and 2006.
- Figure A.5: Results from Welfare Maximization: (a) Welfare as a function of ξ and ω; (b) Optimal R&D tax credits/subsidies.
- Figure A.6: Changes in the spatial distribution of inventors and innovation; x-axis: share of inventors/patents in each percentile of city distribution (log scale); y-axis: expected change under optimal subsidy scheme.

### Appendix B — Proofs and Derivations: key analytical results and expressions
- City 0 (no inventors): production and market-clearing conditions yield
  - ℓn,0 = θβ L0 and ℓy,0 = (1−θβ) L0.
  - Land rent: p m,0 m0 = (1−β) θ wℓ0 L0.
  - Production worker wage:
    wℓ0 = 1
    1−θ
    [ uℓ (θβ)
    θβ
    ]^{1/(1−θ)}
    ̃L0^{1−β} α0^{θ/(1−θ)} .
- Cities 1,...,C (with inventors): supply/demand and wages
  - (B.2) Production worker wage in city c:
    wℓc = wℓ ̃L_c^{1−β} α_c^{θ/(1−θ)} θ^{(1−θ)/(1−θ)}
    with wℓ = 1
    1−θ
    [ uℓ (θβ)
    θ
    ]^{1/(1−θ)} .
  - (B.3) Inventor wage in city c:
    wic = wi ̃I_c^{1−β} α_c^{θ/(1−θ)} θ^{(1−θ)/(1−θ)}
    with
    wi = { u_i^{(1−θ)/(1−θ)} [ (1−θβ) (1/wℓ θβ/(1−θβ) wic)^{β(1−θ)/(1−θβ)} α^{θβ/(1−θβ)} ̃I_c^{-(1−β)/(1−θβ)} ]^{θ/(1−θβ)} }^{(1−θβ)/(1−θ)}.
  - Relation between L_c and I_c:
    L_c = [ θβ/(1−θβ) (wi/wℓ) ]^{(1−θ)/(1−θβ)} I_c.
  - Aggregate relation (B.4):
    wℓ = [ I/(L−L0) ]^{(1−θβ)/(1−θ)} (θβ/(1−θβ))^{(1−θβ)/(1−θ)} w i.
- Firm HJB and investment conditions (summary)
  - Firm HJB: r V_c = max_{x_fc} { flow profits + x_fc E[ΔV] − D Σ[...] − (1−s_c) w i c (i_fc + κ) + E[dV]/dt } subject to x_fc = ̄χ_c Z_c ( ̃I_c^{η} i_fc )^{ψ}.
  - First-order condition (interior): x_fc = ̄χ_c^{1/(1−ψ)} [ ψ F(1+λ) Q / (w_i α^{θ/(1−θ)} (1−s_c) ) ̃I_c^{η(1−θ)−(1−β)θ/(1−θ)} ]^{ψ/(1−ψ)} Z_c^{1/(1−ψ)}.
  - Corner solution: x_fc = 0 implies E_c = 0; incumbents/entrants have same FOC and same arrival rates.
- Staffing and firm counts (Proposition 2 results)
  - (B.7) Inventors per land in city c:
    ̃I_c = { ψ^{ψ} [ (1−ψ)/κ ]^{1−ψ} ̄χ_c (1+λ) ̄π L0/(r+D) (Q/w_i α^{θ/(1−θ)}_c (1−s_c) Z_c ) }^{(1−θ)/Θ},
    where Θ = (1−β)θ − ψ η (1−θ).
  - (B.8) Ratio w_i/Q given by:
    w_i/Q = 1/I Θ^{1−θ ψ^{ψ} [(1−ψ)/κ]^{1−ψ} (1+λ) ̄π L0/(r+D) (1/C Σ_{c=1}^C ( ̄χ_c (1−s_c) )^{(1−θ)/Θ} α_c^{θ/Θ} e^{(1−θ)/Θ ( (1−θ)/Θ −1) σ^2/(4φ)} )^{Θ/(1−θ)} }.
  - (B.9) ̃I_c closed form (equivalent to equation (9)):
    ̃I_c = I × ( ̄χ_c (1−s_c) )^{(1−θ)/Θ} α_c^{θ/Θ} / [ (1/C) Σ_{c=1}^C ( ̄χ_c (1−s_c) )^{(1−θ)/Θ} α_c^{θ/Θ} ] × Z_c^{(1−θ)/Θ} e^{(1−θ)/Θ ((1−θ)/Θ −1) σ^2/(4φ)}.
  - Inventors per firm when investing:
    i_fc = ψ/(1−ψ) κ.
  - Firms per city:
    I_c = (κ/(1−ψ)) N_c  ⇒ N_c = (1−ψ)/κ I_c.
- Aggregate dynamics and growth (Proposition 3)
  - Rate of creative destruction D = Σ_{c=1}^C N_c x_fc.
  - Balanced growth: ̇Q/Q = g and w_i and w_ℓ grow at rate g.
  - Identity: g = λ D.
  - Cross-city quality dynamics: asymptotic relation (B.11)
    lim_{t→∞} E[ Q_c(t) ] = lim_{t→∞} E[ N_c x_fc ] / D × Q(t).
- Note on corporate income taxes: introducing corporate tax τ_{πc} alters firm value v_c but V_c = v_c/(1−τ_{πc}) implies identical R&D choice x_fc; location and R&D decisions unaffected as long as R&D is treated as profit expenditure.

### Appendix C — Linear Regressions: data, identification, and estimation details
- Data sources and construction:
  - USPTO PatentsView platform used for patent-level data; year of patent = filing year; drop patents never granted.
  - Assignees restricted to corporations; assign patent to CBSA of assignee; if multiple owners split equally.
  - "Completion" procedure for inventors and firms: add observations with zero patents for missing years between entry and exit in CBSA.
  - Industry classification: patents → 6 broad categories and 37 subcategories; firm/inventor industry = mode of subcategories; “industry 0” for non-unique mode.
  - Additional data: County Business Patterns (CBP); Zillow Rent Index (ZRI) and Zillow Home Value Index (ZHVI); aggregated to C BSA via NBER county-CBSA crosswalk.
  - Final combined dataset used for estimation: 2,217,577 patents; 1,191,418 inventors; 136,124 firms (assignees); 860 CBSAs; years 1998–2016.
- From model to regression (derivation summary)
  - Production function of innovation yields empirical regression:
    log(x_{f,c,t}) = ψ log(i_{f,c,t}) + ψ η log(I_{c,t}) + δ_c + z_{f,c,t},
    where x_{f,c,t} = innovations by firm f in city c in year t, i_{f,c,t} = inventors hired by firm f in year t, I_{c,t} = population of inventors in city c in year t, δ_c city fixed effect, z residual.
- Estimation of elasticity of agglomeration (η) — shift-share / Bartik-style instrument
  - Instrument construction: lagged national industry employment growth shares across K industries with sectoral shifters γ_{k,t−l→t}, excluding industries with any city employment share >15% at any time to mitigate concentration-driven endogeneity (excludes 6 of 38 industries).
  - First-stage coefficients (selected): table C.1 First Stage log(I_{c,t,l}) estimates: 0.532, 0.467, 0.389, 0.260 for IVs with lags l = 5,7,10 and industry fixed at 1990–95 (columns 1–4); corresponding F-statistics: 196.20, 137.93, 97.70, 40.71.
  - Second-stage implied η (from table C.1): 0.218 (l=5), 0.206 (l=7), 0.184 (l=10), 0.180 (fixed 1990–95).
- Robustness and Poisson specifications
  - To include firms with zero patents, a Poisson pseudo-ML approach (ppmlhdfe) is used with control for endogeneity via Wooldridge (2010) control function.
  - Table C.2 (Poisson): First-stage log(I_{c,t,l}): 0.536, 0.469, 0.388, 0.248 (lags l=5,7,10 and fixed 1990–95). Second-stage implied η: 0.394, 0.244, 0.190, 0.058, −0.002 (some imprecise).
- Multiple potential agglomeration sources (table C.3)
  - Regression including log(Inventors in City), log(Innov. Firms in City), log(Employment in City), log(Establishments in City).
  - Results (OLS / Poisson):
    - log(Inventors in City): 0.064 (OLS, SE 0.017, significant), 0.203 (Poisson, SE 0.034, significant).
    - Other agglomeration measures generally not significant in presence of inventors measure.
- Estimation of congestion elasticity (β)
  - Derived regression: log(p_hc,t) = [(1−β)/(1−θ)] log(L_{c,t}) + δ_c + δ_t + z^h_{c,t}, with p_hc,t proxied by median rent value per square foot (ZRI) or housing price per sq. ft (ZHVI).
  - Instrument I_{c,t,l} used as in agglomeration estimation; industries with >15% concentration excluded.
  - Selected IV estimates (table C.4, using rental prices):
    - First-stage log(I_{c,t,l}): 0.013, 0.017, 0.013, 0.022 (lags l=5,7,10 and fixed 1990–95). F-stats: 11.20, 5.67, 8.31, 9.56.
    - Second-stage log(Prod. Workers in City): 1.336 (SE 0.413), 1.063 (SE 0.459), 1.384 (SE 0.479), 1.033 (SE 0.368).
    - Implied β: 0.466, 0.575, 0.446, 0.587.
  - Alternative using housing prices (table C.5):
    - First-stage log(I_{c,t,l}): 0.039, 0.037, 0.035, 0.032 (F-stats: 42.94, 36.00, 40.04, 22.42).
    - Second-stage log(Prod. Workers in City): 0.335, 0.322, 0.418, 0.426 (implied β: 0.866, 0.871, 0.833, 0.830). These housing-price-based estimates used as upper bound in sensitivity analysis.
- Instrument validity discussion and threats to identification
  - Two interpretations for exogeneity condition E_c[I_{c,t,l} z_{c,t}] = 0:
    1. Lagged industry employment shares uncorrelated with current shocks (ω_{k,t} = 0), achievable with sufficiently long lags l.
    2. Asymptotic orthogonality across a large number of industries K; potential problem if an industry is both highly concentrated in a city and large internationally — mitigated by excluding industries with >15% city share.
  - Robustness: results similar when excluding concentrated industries; lag lengths up to 10 years and using 1990–95 average employment shares considered.

### Appendix D — Matching Moments and parameter identification
- Scale of productivity χ_s identification:
  - From corollary 1 and D expressions:
    χ_s = g/λ × ( ψ^{ψ} [(1−ψ)/κ]^{1−ψ} e^{(1−θ)(1+ψη)/Θ ((1−β)θ/Θ +1) σ^2/(4φ)} (1/C Σ_{c=1}^C ˆχ_c ̃ ̄I_c^{1+ψη}) )^{−1}.
- α_0 (amenity/scale in city 0) identified from population share equation (B.10) where L0/L−L0 ratio (observable) determines α_0 given other parameters.
- Law of motion parameter σ identified from cross-sectional variance of inventor population across cities:
  - Var(I_{c,T}) = (1/C) Σ_{c=1}^C ̄I_c^2 [ ∫_0^1 ∫_0^1 exp( (1−θ)/Θ σ^2/(2φ) exp(−φ|s−t|) ) dtds − 1 ] with φ set to 1; observed Var(I_{c,T}) identifies σ.

### Appendix E — Government’s problem (aggregated and reduced form)
- Government objective and budget constraints reduced using model relationships:
  - Objective reduces to maximization over {s_c} of:
    (L−L0)^{θβ} [ 1 + θβ L0/(L−L0) ]^{(1−θ)/(1−θβ)} ̄w_i^{1−θ} ∫_0^∞ e^{−ρt} Q(t)^{1−θ} dt
  - Budget constraint in reduced form equates spatially integrated subsidies plus public good cost ̄G ̄π Q(t) to land rents and aggregate firm profits (present value F(D,L0) Q(0) etc.); uses law of large numbers to replace cross-city averages with expectations.
  - With Q(t) = Q(0) e^{g t} and ρ > (1−θ) g, integral ∫_0^∞ e^{−ρt} Q(t)^{1−θ} dt = Q(0)^{1−θ}/(r − (1−θ) g), producing equation (18) in main text.

### Appendix F — Sensitivity analyses (setup)
- Notes that various specifications used to estimate agglomeration and congestion elasticities feed into sensitivity analysis; housing-price-based β estimates treated as upper bound.

*Document: References and appendices (figures, proofs, regressions, data construction, matching moments, government problem, sensitivity notes) from wpiea2022131-print-pdf.*

### section 4 produced different point estimates, sometimes significantly different from one another.

### wpiea2022131-print-pdf - section 4 produced different point estimates, sometimes significantly different from one another.

### Sensitivity analysis: alternative elasticity values and counterfactuals (Table F.6, τ = 0.5)
- Setup:
  - Consider all combinations of η ∈ {0.15, 0.20, 0.25} and β ∈ {0.5, 0.6, 0.8}.
  - Subsidy cap fixed at 50%.
  - Re-estimate step three and re-run counterfactual experiments using those elasticity values.
- Main qualitative finding:
  - Despite wide variation in η and β, the optimal subsidies follow the same pattern: they increase spatial concentration of the population, increase the rate of creative destruction/growth, and decrease the baseline wage of workers.
  - Intuition: higher η (greater agglomeration elasticity) and higher β (lower congestion elasticity) raise gains from spatial concentration.
- City-level subsidies — gains from adopting optimal subsidies, τ = 0.5 (rows show η, β, ∆Welfare, ∆Baseline Wage, ∆Creative Destruction, ∆Rate of Growth):
  - 0.15 0.5 3.97% −7.29% 1.33p.p. 0.18p.p.
  - 0.15 0.6 5.45% −7.13% 1.80p.p. 0.24p.p.
  - 0.15 0.8 5.23% −6.56% 3.53p.p. 0.47p.p.
  - 0.20 0.5 4.12% −7.20% 1.54p.p. 0.20p.p.
  - 0.20 0.6 6.15% −7.75% 2.00p.p. 0.26p.p.
  - 0.20 0.8 8.42% −7.05% 4.09p.p. 0.54p.p.
  - 0.25 0.5 4.61% −7.60% 1.68p.p. 0.22p.p.
  - 0.25 0.6 6.97% −8.05% 2.18p.p. 0.29p.p.
  - 0.25 0.8 2.58% −7.59% 4.76p.p. 0.63p.p.
- State-level subsidies — gains from adopting optimal subsidies, τ = 0.5 (rows show η, β, ∆Welfare, ∆Baseline Wage, ∆Creative Destruction, ∆Rate of Growth):
  - 0.15 0.5 1.91% −4.34% 0.82p.p. 0.11p.p.
  - 0.15 0.6 2.86% −4.48% 1.03p.p. 0.14p.p.
  - 0.15 0.8 8.67% −5.01% 2.24p.p. 0.30p.p.
  - 0.20 0.5 2.14% −4.50% 0.88p.p. 0.12p.p.
  - 0.20 0.6 3.23% −4.71% 1.13p.p. 0.15p.p.
  - 0.20 0.8 10.68% −5.83% 2.68p.p. 0.35p.p.
  - 0.25 0.5 2.39% −4.66% 0.95p.p. 0.13p.p.
  - 0.25 0.6 3.65% −4.83% 1.23p.p. 0.16p.p.
  - 0.25 0.8 3.79% −6.70% 3.30p.p. 0.44p.p.

### Extensions of the model: allowing innovation to scale with firm size (Section G overview)
- Motivation:
  - Original model predicts identical arrival rates of innovation for all firms in the same city regardless of size; empirical reality differs (big and small firms produce different expected numbers of patents).
  - Extension allows innovation production and fixed innovation costs to scale with firm size (in the spirit of Klette and Kortum (2004)).
- Key modeling choices:
  - Define p_f(t) = 1 + |q_f(t)| as number of product lines plus one (entrants: q_f = ∅ so p_f = 1).
  - Innovation production function for firm f in city c:
    - x_f,c(t) = χ̄_c(t) [Ĩ_c(t)^η i_f,c(t)]^ψ p_f(t)^{1−ψ}  (equation G.1 uses same notation as source).
  - Fixed cost of innovation scales as κ p_f inventors.
  - Assumptions from sections 2 and 3 kept unless explicitly changed.

### Extended model results and equilibrium characterization (G.1, propositions and corollaries)
- Incumbent HJB (Lemma G.1) and entrant analogue describe firm dynamic optimization with size-scaling innovation and fixed costs.
- Value function solution (Proposition G.1):
  - V_c(q_f, Ĩ_c, Z_c, A) = F(D, L_0) Σ_{q_j ∈ q_f} q_j + max{0, E_c(Ĩ_c, Z_c, w_i/Q, D, L_0)} p_f Q_o,
  - with F(D, L_0) = χ̄π L_0 /(r + D) as franchise value (notation retained).
  - Entrant second-stage value: V_e_c(Ĩ_c, Z_c, A) = max{0, E_c(Ĩ_c, Z_c, w_i/Q, D, L_0)} Q_o.
- Interior optimal arrival rate (x_f,c) expression (as given in text):
  - x_f,c = p_f χ̄_c^{1/(1−ψ)} { ψ[F(1 + λ) + E_c] Q / [w_i α θ^{1−θ}_c (1−s_c) Ĩ^{η−(1−β)θ/(1−θ)}_c] }^{ψ/(1−ψ)} Z_c^{1/(1−ψ)}  (symbolic expression preserved as in source).
- Franchise value identity:
  - F(D, L_0) = χ̄π L_0 /(r + D)  (equation G.2).
- Entry value equation (implicit) determining E_c (equation G.3), repeated verbatim structure in source.
- With free entry (E_c = 0), population of inventors and spatial distribution (Proposition G.2):
  - I_c = I × [χ̄_c / (1−s_c)]^{1−θ/Θ^{α θ/Θ}_c} × Z_c^{1−θ/Θ} e^{(1−θ/Θ)(1−θ/Θ −1) σ^2/(4φ)}  (structure preserved; Θ = (1−β)θ − ψ η (1−θ); equation G.4).
  - Arrival rate of innovation for firm f in city c:
    - x_f,c = p_f (κ/ψ^{1−ψ})^{ψ} χ̄_c Ĩ_c^{ψ η} Z_c  (equation G.5 preserved in structure).
  - Inventors hired per firm in city c:
    - i_f,c = ψ/(1−ψ) κ p_f.
  - Firm and product counts relationship:
    - N_c + J_c = ((1−ψ)/κ) I_c  (equation G.6).
  - Production-worker population and wage expression (equation G.7) retained in structure; w_i/Q ∝ χ̄π L_0 /(r + D) (1/C Σ_c (χ̄_c/(1−s_c))^{1−θ/Θ α θ/Θ}_c )^{Θ/(1−θ)} as stated.
- Aggregate creative destruction (Corollary G.1):
  - D ∝ (1/C) Σ_{c=1}^C χ̄_c Ĩ_c^{1+ψ η}  (equation G.8 preserved).
  - Proof: aggregate D = Σ_c ∫_{F_c} x_f,c df and reduces to same structure as simplified model.

### Estimation challenges for the extended model (G.2)
- Core identification issue:
  - Equation (G.6) links number of firms N_c, number of products J_c, and population of inventors I_c; J_c appears explicitly so cannot normalize product measure to 1 without affecting scales of N_c and I_c.
- Firm size distribution dynamics:
  - Let μ_c(q, t) be measure of firms with q products in city c at time t.
  - Dynamics: ∂μ_c(q,t)/∂t = x_f(q−1),c μ_c(q−1,t) + (q+1)D μ_c(q+1,t) − x_f(q),c μ_c(q,t) − qD μ_c(q,t).
  - Implication: the measure of firms with q products in city c is non-stationary and depends on local shocks via x_f,c; average products per firm varies over time.
- Practical consequence:
  - Hard to pin down relative scales of N_c and J_c without additional structure or data on firm-size/product distributions.
  - Simplified model (main text) avoids this identification issue because number of products does not affect inventors hired; both models make same predictions on how R&D subsidies affect spatial distribution and aggregate growth, so the simplified model is used for counterfactual policy exercises.

*International Monetary Fund — Working Paper: Agglomeration, Innovation, and Spatial Reallocation (Working Paper No. WP/2022/131).*

---


_Source: https://www.imf.org/-/media/files/publications/wp/2022/english/wpiea2022131-print-pdf.pdf_
