## INTRODUCTION

## Source details

**Canonical URL:** [INTRODUCTION](https://www.imf.org/-/media/files/publications/wp/2022/english/wpiea2022259-print-pdf.pdf)

## Other formats

- [Markdown version](/-/media/files/publications/wp/2022/english/wpiea2022259-print-pdf.pdf.md)
- [Structured JSON version](/-/media/files/publications/wp/2022/english/wpiea2022259-print-pdf.pdf.json)

---

### Overview
- Deep reinforcement learning (DRL) combines deep learning (DL) and reinforcement learning (RL) to solve complex problems through interactive learning.
- DRL models can learn from multi-dimensional and continuous state-action spaces and non-stationary complex environments; they can in principle improve learning via meta learning (“learning to learn”, “learning from experience”, or “learning by doing”).
- Applications of RL and DRL in economics have been limited but are gaining traction due to difficulties in modeling environments, designing reward functions, and solving high-dimensional economic models.
- This chapter introduces RL and DRL theory in relatively non-technical terms, surveys recent applications in economics with a focus on macroeconomics, and discusses prospects and open issues.

### Reinforcement learning (RL) and deep reinforcement learning (DRL): key concepts
- RL: agents learn by interacting with an environment through trial-and-error, receiving rewards that guide behavior.
- DRL: RL that uses deep neural networks to approximate value or policy functions.
- Distinction from other ML paradigms:
  - Supervised learning requires labelled datasets.
  - Unsupervised learning finds structure in unlabeled data.
  - RL does not require a pre-existing labeled dataset; it learns from interactions with an environment.
- Core RL terminologies:
  - State, s ∈ S: representation of an environment.
  - Action, a ∈ A: behavior of the RL agent.
  - Reward, R(s,a): stimulus due to action and state.
  - Policy function, π(s): mapping from state to action or distribution over actions (stochastic or deterministic).
  - Value function, Q(s,a): expected cumulative rewards for a state-action pair.
- Objective: find optimal policy π* = argmax_π Q(s,a).

### Brief history (selected milestones)
- Origins in behaviorism and links to psychology and neuroscience; early contributors include Alan Turing, Norbert Wiener, Richard Bellman.
- 1950s–1960s: trail-and-error ideas (Minsky 1954, Farley and Clark 1954); term “reinforcement” introduced by Minsky (1961); Bellman (1957) formalized dynamic programming and the Bellman equation.
- 1970s–1980s: temporal-difference learning and trial-and-error linked to psychology; Klopf (1972); Sutton and Barto (1981); RL combined with control theory and adaptive control.
- Late 1980s: Watkins (1989) developed Q-learning (first modern RL algorithm).
- 1990s–2000s: RL applied to inverted pendulum, backgammon (Tesauro 1995), connections to neuroscience (Schultz et al 1997); combination with function approximation and neural networks.
- Late 2000s onward: RL combined with deep learning → modern deep RL applied broadly in robotics, healthcare, finance, etc.

### Theory and solution structure
- RL solves Markov Decision Processes (MDPs); environment often modeled as a Markov Chain where next state depends only on current state.
- Typical DRL workflow:
  - Initiate memory.
  - Define state and action spaces and specify reward function.
  - Agent-environment interactions produce transitions (state, action, next state, reward).
  - Store transitions in memory and sample from memory to update value or policy functions.
  - Repeat until task solved or terminal state reached.
- Exploration vs Exploitation:
  - Trade-off between exploring new actions and exploiting known good actions.
  - Two common strategies: ε-greedy (with probability ε choose actions randomly; with probability 1−ε choose greedy action) and noise perturbation (a_t = π(s_t) + N_t).
- TD (temporal-difference) updating:
  - Learns from intermediate rewards and bootstraps from next-state value; TD-error central to updates.
  - Two main TD methods: SARSA and Q-learning.

### Main algorithm classes and examples
- Model-based vs Model-free:
  - Model-based (planning) requires environment model (e.g., dynamic programming).
  - Model-free learns from experience without a model (e.g., temporal-difference, policy gradient, actor-critic).
- Value-based vs Policy-based:
  - Value-based learns value function (e.g., Q-learning).
  - Policy-based learns policy parameters directly via gradient ascent on performance J(θ) (e.g., REINFORCE).
- On-policy vs Off-policy:
  - On-policy learns from the current policy.
  - Off-policy can learn from data generated by other policies.
- Prominent DRL algorithms listed:
  - DQN (Deep Q Networks) — Q-learning combined with ANNs; off-policy; applied to Atari games.
  - DDPG (Deep Deterministic Policy Gradient) — off-policy actor-critic for continuous control.
  - TD3 (Twin Delayed DDPG) — addresses DDPG Q-value overestimation.
  - PPO (Proximal Policy Optimization), TRPO (Trust Region Policy Optimization).
  - A2C (Advantage Actor-Critic) and A3C (Asynchronous Advantage Actor Critic).
- Actor-critic methods:
  - Jointly learn value function and policy; reduce variance of policy gradient estimates but can introduce bias and instability.
- Experience replay:
  - Stores past transitions and replays them to decorrelate samples and improve learning.

### Multiagent RL
- Multiagent problems generalize MDPs to stochastic games (Markov games) with n agents and joint action space A = A1 × A2 × ... × An.
- Challenges:
  - Non-stationary environment for each agent due to interdependent agents.
  - Private information and strategic interactions complicate equilibrium selection.
  - Centralized training vs decentralized training trade-offs; hybrid centralized/decentralized approaches emerging.
- Use cases and studies:
  - Curry et al (2022): multi-agent deep RL in real-business cycle models with 100 consumer-workers, 10 firms, and a government; discover ε-Nash equilibria via meta-game best-response analyses.
  - Johnson et al (2022): agents produce, trade, and consume in a spatial world; emergent production, consumption, and pricing behaviors including arbitrage.
  - Applications to ABMs, mechanism design, and game-theoretic problems cited.

### Recent applications in economics and macroeconomics
- DRL used across domains: AlphaGo, AlphaFold, WaveNet, energy optimization in data centers, personalized recommendations, trading strategies, automated diagnosis and therapy.
- Macro applications split into two main areas:
  1. Using DRL as a solution method to find optimal policy or policy response functions:
     - Hinterlang and Tänzer (2022): DRL to find optimal interest rate reaction functions using ANNs to estimate transition equations from a three-equation New Keynesian model.
     - Castro et al (2022): REINFORCE algorithm for banks’ initial liquidity choice in a real-time gross settlement payments system (daily episodes divided into T intraday periods); agents risk neutral and payment demand profile assumed common knowledge; show RL can learn policies reducing liquidity costs.
     - Curry et al (2022): multi-agent DRL identifies stable ε-Nash equilibria in RBC models; evaluate equilibria by holding other agent-type policies fixed and testing best-response improvements.
  2. Modeling bounded rationality, learning, and transition dynamics:
     - Chen et al (2021): DDPG in a monetary model with multiple equilibria; RL agent locally converges to steady states.
     - Shi (2021a, 2021b): DDPG and DDPG-like studies on consumption-saving behavior and money demand with transaction costs and policy regime changes; show exploration levels affect beliefs, convergence, and welfare; agents with experience of regime shifts adjust better.
     - Hill et al (2021): solve rational expectations models with discrete heterogeneous agents across topics like precautionary savings and ‘epi-macro’ models.
- Other cited studies include Jirnyi and Lepetyuk (2011); Zheng et al (2020); Covarrubias (2022); Calvano et al (2020).

### Prospects (selected avenues)
- Forecasting: potential to apply DRL to forecasting macro variables and strategic behaviors.
- Macro simulation and ABM integration: simulate many interacting agents to study cooperative/competitive behavior; RL agents could share information to improve welfare.
- Rational expectations and learnability: study local vs global convergence in models with multiple equilibria (Chen et al 2021).
- Transfer learning: reuse models trained on one country/region to improve learning in another; alleviate data hunger of deep models.
- Meta learning: train meta-models that can be fine-tuned quickly for subproblems; Zhang et al. (2022) DRL meta-learning example reduces training time for many submodels.
- Automated hyperparameter optimization and model selection/construction: RL frameworks for automatic model selection and repair.
- Quantum reinforcement learning (QRL): potential future speedups with quantum computing.
- Super integrated policy frameworks and interdisciplinary integration: DRL could enable highly granular, interconnected policy models, integrating natural language processing and social-science inputs for more holistic welfare-focused macro models.
- Combined NLP-DRL approaches: DRL4NLP could ingest textual data (e.g., central bank minutes, news) to enhance macro models.
- Causal deep reinforcement learning (CRL & CDRL): combine RL with causal inference to simulate counterfactual policy scenarios where experiments are infeasible.

### Issues and challenges
- Training costs and scalability:
  - DRL requires long simulation times and high computational power.
  - Scalability problematic for large environments and multiagent settings; ANNs are data- and computation-intensive.
  - Transfer learning may help.
- Reward function design:
  - Critical and often difficult to design; multiagent reward alignment can be problematic.
- Stability and sensitivity:
  - Instabilities during training and sensitivity to hyperparameters (batch size, exploration level, episode length).
- Interpretability and algorithm choice:
  - ANNs are opaque; many algorithms exist and algorithm selection depends on state/action design and reward structure.
- Ergodicity:
  - Non-ergodic processes where some states/actions are unreachable impede learning from all experiences.
- Credit assignment and sparse rewards:
  - Identifying which actions lead to long-term goals is difficult, especially with delayed or sparse rewards.
- Long-term planning:
  - Hard to account for long-term consequences and future uncertainty.
- Bias, fairness, transparency:
  - Potential for bias in learned policies; fairness issues in multiagent settings; opacity undermines transparency.
- Benchmarking and safety:
  - Lack of standardized benchmarks; safety guarantees are difficult in stochastic RL systems—critical in industrial, health, and policy contexts.
- Computational and identifiability concerns:
  - Identifiability of heterogeneous agent models and robustness of learning to environmental changes remain open problems.

### Conclusion (summary points)
- DRL applications in macroeconomics have grown rapidly but remain limited in scope.
- Two primary uses in macroeconomics:
  - As flexible, model-free solution methods for optimal policies and general equilibrium problems.
  - To model bounded rationality, learning dynamics, and convergence properties.
- DRL holds promise across forecasting, simulation, transfer learning, automated model construction, interdisciplinary modeling, NLP integration, causal simulation, and potentially quantum-enhanced methods.
- Significant technical and conceptual challenges persist—computational demands, reward design, stability, interpretability, ergodicity, credit assignment, fairness, transparency, benchmarking, and safety—requiring focused research.

*Source: wpiea2022259-print-pdf - Introduction*

### INTRODUCTION ...........................................................................................................

### INTRODUCTION

### Document structure and major sections
- INTRODUCTION ............................................................................................................................................ 4
- I. WHAT IS REINFORCEMENT LEARNING AND DEEP REINFORCEMENT LEARNING? ....................... 4
  - A. Brief History ............................................................................................................................... 5
  - B. Theory ........................................................................................................................................ 6
  - C. Recent Applications ................................................................................................................. 16
- II. ECONOMIC DEEP REINFORCEMENT LEARNING: APPLICATIONS AND EMERGING  
  TRENDS IN MACROECONOMICS ............................................................................................................. 16
  - A. Solution Methods ....................................................................................................................... 17
  - B. Bounded Rationality, Learning and Convergence ...................................................................... 19
- III. DEEP REINFORCEMENT IN MACROECONOMICS:  PROSPECTS AND ISSUES ............................ 19
  - A. Prospects ................................................................................................................................. 20
  - B. Issues ...................................................................................................................................... 22
- IV. CONCLUSION ........................................................................................................................................ 24
- ANNEX
  - Full Algorithms ..................................................................................................................................... 25
- REFERENCES ............................................................................................................................................. 26

### Figures and Table (listed)
- FIGURES
  - 1. A Markov Decision Process (A Reinforcement Learning Problem) ............................................................ 6
  - 2. Comparison of RL and other methods ....................................................................................................... 8
  - 3. RL algorithm overview ................................................................................................................................ 9
  - 4. A feedforward deep ANN ......................................................................................................................... 13
  - 5. Workflow of deep RL algorithms .............................................................................................................. 14
  - 6. Multiagent RL ........................................................................................................................................... 15
- TABLE
  - 1. Terminologies in reinforcement learning .................................................................................................... 7

### Glossary (abbreviations as presented)
- AC Actor-Critic
- A2C Advantage Actor-Critic
- A3C Asynchronous Advantage Actor Critic
- ANNs Artificial Neural Networks
- CDRL Causal Deep Reinforcement Learning
- DL Deep Learning
- DRL Deep Reinforcement Learning
- DDPG Deep Deterministic Policy Gradients
- DQN Deep Q Networks
- MDP Markov Decision Processes
- ML Machine Learning
- RL Reinforcement Learning
- PPO Proximal Policy Optimization
- QRL Quantum reinforcement learning
- SARSA State-Action-Reward-State-Action
- TD Temporal Difference
- TRPO Trust Region Policy Optimization

*IMF Working Paper: Deep Reinforcement Learning: Emerging Trends in Macroeconomics and Future Prospects — INTRODUCTION (pages and listings as in source).*

### Introduction

### Introduction

### Overview
- Deep reinforcement learning (DRL) is a subset of artificial intelligence (AI) combining deep learning (DL) and reinforcement learning (RL) to solve complex problems through interactive learning.
- DRL models can learn from multi-dimensional and continuous state-action spaces and non-stationary complex environments; they can in principle improve learning via meta learning (“learning to learn”, “learning from experience”, or “learning by doing”).
- Applications of RL and DRL in economics have been limited but are gaining traction due to difficulties in modeling environments, designing reward functions, and solving high-dimensional economic models.
- This chapter introduces RL and DRL theory in relatively non-technical terms, surveys recent applications in economics with a focus on macroeconomics, and discusses prospects and open issues.

### Reinforcement learning (RL) and deep reinforcement learning (DRL): key concepts
- RL: agents learn by interacting with an environment through trial-and-error, receiving rewards that guide behavior.
- DRL: RL that uses deep neural networks to approximate value or policy functions.
- Distinction from other ML paradigms:
  - Supervised learning requires labelled datasets.
  - Unsupervised learning finds structure in unlabeled data.
  - RL does not require a pre-existing labeled dataset; it learns from interactions with an environment.
- Core RL terminologies (as summarized in Table 1):
  - State, s ∈ S: representation of an environment.
  - Action, a ∈ A: behavior of the RL agent.
  - Reward, R(s,a): stimulus due to action and state.
  - Policy function, π(s): mapping from state to action or distribution over actions (stochastic or deterministic).
  - Value function, Q(s,a): expected cumulative rewards for a state-action pair.
- Objective: find optimal policy π* = argmax_π Q(s,a).

### Brief history (selected milestones)
- Origins in behaviorism and links to psychology and neuroscience; early contributors include Alan Turing, Norbert Wiener, Richard Bellman.
- 1950s–1960s: trail-and-error ideas (Minsky 1954, Farley and Clark 1954); term “reinforcement” introduced by Minsky (1961); Bellman (1957) formalized dynamic programming and the Bellman equation.
- 1970s–1980s: temporal-difference learning and trial-and-error linked to psychology; Klopf (1972); Sutton and Barto (1981); RL combined with control theory and adaptive control.
- Late 1980s: Watkins (1989) developed Q-learning (first modern RL algorithm).
- 1990s–2000s: RL applied to inverted pendulum, backgammon (Tesauro 1995), connections to neuroscience (Schultz et al 1997); combination with function approximation and neural networks.
- Late 2000s onward: RL combined with deep learning → modern deep RL applied broadly in robotics, healthcare, finance, etc.

### Theory and solution structure
- RL solves Markov Decision Processes (MDPs); environment often modeled as a Markov Chain where next state depends only on current state.
- Typical DRL workflow:
  - Initiate memory.
  - Define state and action spaces and specify reward function.
  - Agent-environment interactions produce transitions (state, action, next state, reward).
  - Store transitions in memory and sample from memory to update value or policy functions.
  - Repeat until task solved or terminal state reached.
- Exploration vs Exploitation:
  - Trade-off between exploring new actions and exploiting known good actions.
  - Two common strategies: ε-greedy (with probability ε choose actions randomly; with probability 1−ε choose greedy action) and noise perturbation (a_t = π(s_t) + N_t).
- TD (temporal-difference) updating:
  - Learns from intermediate rewards and bootstraps from next-state value; TD-error central to updates.
  - Two main TD methods: SARSA and Q-learning.

### Main algorithm classes and examples
- Model-based vs Model-free:
  - Model-based (planning) requires environment model (e.g., dynamic programming).
  - Model-free learns from experience without a model (e.g., temporal-difference, policy gradient, actor-critic).
- Value-based vs Policy-based:
  - Value-based learns value function (e.g., Q-learning).
  - Policy-based learns policy parameters directly via gradient ascent on performance J(θ) (e.g., REINFORCE).
- On-policy vs Off-policy:
  - On-policy learns from the current policy.
  - Off-policy can learn from data generated by other policies.
- Prominent DRL algorithms:
  - DQN (Deep Q Networks) — Q-learning combined with ANNs; off-policy; applied to Atari games.
  - DDPG (Deep Deterministic Policy Gradient) — off-policy actor-critic for continuous control.
  - TD3 (Twin Delayed DDPG) — addresses DDPG Q-value overestimation.
  - PPO (Proximal Policy Optimization), TRPO (Trust Region Policy Optimization).
  - A2C (Advantage Actor-Critic) and A3C (Asynchronous Advantage Actor Critic).
- Actor-critic methods:
  - Jointly learn value function and policy; reduce variance of policy gradient estimates but can introduce bias and instability.
- Experience replay:
  - Stores past transitions and replays them to decorrelate samples and improve learning.

### Multiagent RL
- Multiagent problems generalize MDPs to stochastic games (Markov games) with n agents and joint action space A = A1 × A2 × ... × An.
- Challenges:
  - Non-stationary environment for each agent due to interdependent agents.
  - Private information and strategic interactions complicate equilibrium selection.
  - Centralized training vs decentralized training trade-offs; hybrid centralized/decentralized approaches emerging.
- Use cases and studies:
  - Curry et al (2022): multi-agent deep RL in real-business cycle models with 100 consumer-workers, 10 firms, and a government; discover ε-Nash equilibria via meta-game best-response analyses.
  - Johnson et al (2022): agents produce, trade, and consume in a spatial world; emergent production, consumption, and pricing behaviors including arbitrage.
  - Applications to ABMs, mechanism design, and game-theoretic problems cited.

### Recent applications in economics and macroeconomics
- DRL used across domains: AlphaGo, AlphaFold, WaveNet, energy optimization in data centers, personalized recommendations, trading strategies, automated diagnosis and therapy.
- Macro applications split into two main areas:
  1. Using DRL as a solution method to find optimal policy or policy response functions:
     - Hinterlang and Tänzer (2022): DRL to find optimal interest rate reaction functions using ANNs to estimate transition equations from a three-equation New Keynesian model.
     - Castro et al (2022): REINFORCE algorithm for banks’ initial liquidity choice in a real-time gross settlement payments system (daily episodes divided into T intraday periods); agents risk neutral and payment demand profile assumed common knowledge; show RL can learn policies reducing liquidity costs.
     - Curry et al (2022): multi-agent DRL identifies stable ε-Nash equilibria in RBC models; evaluate equilibria by holding other agent-type policies fixed and testing best-response improvements.
  2. Modeling bounded rationality, learning, and transition dynamics:
     - Chen et al (2021): DDPG in a monetary model with multiple equilibria; RL agent locally converges to steady states.
     - Shi (2021a, 2021b): DDPG and DDPG-like studies on consumption-saving behavior and money demand with transaction costs and policy regime changes; show exploration levels affect beliefs, convergence, and welfare; agents with experience of regime shifts adjust better.
     - Hill et al (2021): solve rational expectations models with discrete heterogeneous agents across topics like precautionary savings and ‘epi-macro’ models.
- Other cited studies: Jirnyi and Lepetyuk (2011) on incomplete market with liquidity constraints; Zheng et al (2020) on optimal tax policy; Covarrubias (2022) on oligopolistic competition and monetary transmission; Calvano et al (2020) on potential algorithmic collusion.

### Prospects (selected avenues)
- Forecasting: potential to apply DRL to forecasting macro variables and strategic behaviors; prior ANN applications in energy usage and stock trading cited.
- Macro simulation and ABM integration: simulate many interacting agents to study cooperative/competitive behavior; RL agents could share information to improve welfare; Sert et al (2020) and Song et al (2021) noted.
- Rational expectations and learnability: study local vs global convergence in models with multiple equilibria (Chen et al 2021).
- Transfer learning: reuse models trained on one country/region to improve learning in another; alleviate data hunger of deep models.
- Meta learning: train meta-models that can be fine-tuned quickly for subproblems; Zhang et al. (2022) DRL meta-learning example reduces training time for many submodels.
- Automated hyperparameter optimization and model selection/construction: RL frameworks for automatic model selection and repair (Shang et al., Barriga Rodriguez et al., Laredo et al., Zhu and Yuan).
- Quantum reinforcement learning (QRL): potential future speedups with quantum computing (Dunjko et al., Wu et al., Madsen et al., Zhang & Ni).
- Super integrated policy frameworks and interdisciplinary integration: DRL could enable highly granular, interconnected policy models, integrating natural language processing and social-science inputs for more holistic welfare-focused macro models.
- Combined NLP-DRL approaches: DRL4NLP could ingest textual data (e.g., central bank minutes, news) to enhance macro models.
- Causal deep reinforcement learning (CRL & CDRL): combine RL with causal inference to simulate counterfactual policy scenarios where experiments are infeasible.

### Issues and challenges
- Training costs and scalability:
  - DRL requires long simulation times and high computational power.
  - Scalability problematic for large environments and multiagent settings; ANNs are data- and computation-intensive.
  - Transfer learning may help.
- Reward function design:
  - Critical and often difficult to design; multiagent reward alignment can be problematic.
- Stability and sensitivity:
  - Instabilities during training and sensitivity to hyperparameters (batch size, exploration level, episode length).
- Interpretability and algorithm choice:
  - ANNs are opaque; many algorithms exist and algorithm selection depends on state/action design and reward structure.
- Ergodicity:
  - Non-ergodic processes where some states/actions are unreachable impede learning from all experiences.
- Credit assignment and sparse rewards:
  - Identifying which actions lead to long-term goals is difficult, especially with delayed or sparse rewards.
- Long-term planning:
  - Hard to account for long-term consequences and future uncertainty.
- Bias, fairness, transparency:
  - Potential for bias in learned policies; fairness issues in multiagent settings; opacity undermines transparency.
- Benchmarking and safety:
  - Lack of standardized benchmarks; safety guarantees are difficult in stochastic RL systems—critical in industrial, health, and policy contexts.
- Computational and identifiability concerns:
  - Identifiability of heterogeneous agent models and robustness of learning to environmental changes remain open problems.

### Conclusion (summary points)
- DRL applications in macroeconomics have grown rapidly but remain limited in scope.
- Two primary uses in macroeconomics:
  - As flexible, model-free solution methods for optimal policies and general equilibrium problems.
  - To model bounded rationality, learning dynamics, and convergence properties.
- DRL holds promise across forecasting, simulation, transfer learning, automated model construction, interdisciplinary modeling, NLP integration, causal simulation, and potentially quantum-enhanced methods.
- Significant technical and conceptual challenges persist—computational demands, reward design, stability, interpretability, ergodicity, credit assignment, fairness, transparency, benchmarking, and safety—requiring focused research.

*Source: wpiea2022259-print-pdf - Introduction*

### References

### References

### Foundational texts on reinforcement learning and deep learning
- Goodfellow, I., Bengio, Y. and Courville, A., 2016. Deep learning. MIT press. https://www.deeplearningbook.org  
- Sutton, R.S. and Barto, A.G., 2018. Reinforcement learning: An introduction. MIT press.  
- Mnih, V., Kavukcuoglu, K., Silver, D., Graves, A., Antonoglou, I., Wierstra, D. and Riedmiller, M., 2013. Playing atari with deep reinforcement learning. arXiv preprint arXiv:1312.5602.  
- Lillicrap, T.P., Hunt, J.J., Pritzel, A., Heess, N., Erez, T., Tassa, Y., Silver, D. and Wierstra, D., 2015. Continuous control with deep reinforcement learning. arXiv preprint arXiv:1509.02971.  
- Barto, A. G., & Singh, S. P. (1991). On the computational economics of reinforcement learning. In Connectionist Models (pp. 35-44). Morgan Kaufmann.  

### Surveys, reviews, and challenges in RL and deep RL
- Dulac-Arnold, G., Levine, N., Mankowitz, D.J., Li, J., Paduraru, C., Gowal, S. and Hester, T., 2021. Challenges of real-world reinforcement learning: definitions, benchmarks and analysis. Machine Learning, 110(9), pp.2419-2468.  
- Ding, Z., & Dong, H. (2020). Challenges of reinforcement learning. In Deep Reinforcement Learning (pp. 249-272). Springer, Singapore.  
- Du, W., & Ding, S. (2021). A survey on multi-agent deep reinforcement learning: from the perspective of challenges and applications. Artificial Intelligence Review, 54(5), 3215-3238.  
- Mosavi, A., Faghan, Y., Ghamisi, P., Duan, P., Ardabili, S. F., Salwana, E., & Band, S. S. (2020). Comprehensive review of deep reinforcement learning methods and applications in economics. Mathematics, 8(10), 1640.  
- Nguyen, T.T., Nguyen, N.D. and Nahavandi, S., 2020. Deep reinforcement learning for multiagent systems: A review of challenges, solutions, and applications. IEEE transactions on cybernetics, 50(9), pp.3826-3839.  
- Hernandez-Leal, P., Kartal, B. and Taylor, M.E., 2019. A survey and critique of multiagent deep reinforcement learning. Autonomous Agents and Multi-Agent Systems, 33(6), pp.750-797.  

### Machine learning methodology, model selection, and hyperparameter optimization
- Agrawal, T. (2021). Hyperparameter Optimization in Machine Learning: Make Your Machine Learning and Deep Learning Models More Efficient. New York, NY, USA: Apress.  
- Laredo, D., Qin, Y., Schütze, O., & Sun, J. Q. (2019). Automatic model selection for neural networks. arXiv preprint arXiv:1905.06010.  
- Shang, D., Sun, H., & Zeng, Q. (2020, February). A Reinforcement-Algorithm Framework for Automatic Model Selection. In IOP Conference Series: Earth and Environmental Science (Vol. 440, No. 2, p. 022060). IOP Publishing.  
- Barriga, A., Rutle, A., & Heldal, R. (2018, October). Automatic model repair using reinforcement learning. In MoDELS (Workshops) (pp. 781-786).  

### Multi-agent learning, game theory, and equilibrium concepts
- Fudenberg, D. and Levine, D.K., 2007. An economist's perspective on multi-agent learning. Artificial intelligence, 171(7), pp.378-381.  
- Nowé, A., Vrancx, P. and Hauwere, Y.M.D., 2012. Game theory and multi-agent reinforcement learning. In Reinforcement Learning (pp. 441-470). Springer, Berlin, Heidelberg.  
- Bai, Y., Jin, C., Wang, H. and Xiong, C., 2021. Sample-efficient learning of stackelberg equilibria in general-sum games. Advances in Neural Information Processing Systems, 34, pp.25799-25811.  
- Rajeswaran, A., Mordatch, I., & Kumar, V. (2020). A game theoretic framework for model based reinforcement learning. In International conference on machine learning (pp. 7953-7963). PMLR.  
- Curry, M., Trott, A., Phade, S., Bai, Y., & Zheng, S. (2022). Finding General Equilibria in Many-Agent Economic Simulations Using Deep Reinforcement Learning. arXiv preprint arXiv:2201.01163.  

### Reinforcement learning applications in economics, finance, and macroeconomics
- Charpentier, A., Elie, R., & Remlinger, C. (2021). Reinforcement learning in economics and finance. Computational Economics, 1-38.  
- Chen, M., Joseph, A., Kumhof, M., Pan, X., Shi, R. and Zhou, X., (2021) Deep reinforcement learning in a monetary model. arXiv preprint arXiv:2104.09368.  
- Trott, A., Srinivasa, S., van der Wal, D., Haneuse, S. and Zheng, S., 2021. Building a foundation for data-driven, interpretable, and robust policy design using the ai economist. arXiv preprint arXiv:2108.02904.  
- Zheng, S., Trott, A., Srinivasa, S., Naik, N., Gruesbeck, M., Parkes, D.C. and Socher, R., 2020. The ai economist: Improving equality and productivity with ai-driven tax policies. arXiv preprint arXiv:2004.13332.  
- Hill, E., Bardoscia, M. and Turrell, A., (2021) Solving heterogeneous general equilibrium economic models with deep reinforcement learning. arXiv preprint arXiv:2103.16977.  
- Hinterlang, N., & Tänzer, A. (2021). Optimal monetary policy using reinforcement learning.  
- Li, Y., Ni, P. and Chang, V., 2020. Application of deep reinforcement learning in stock trading strategies and stock forecasting. Computing, 102(6), pp.1305-1322.  
- Pendharkar, P. C., & Cusatis, P. (2018). Trading financial indices with reinforcement learning agents. Expert Systems with Applications, 103, 1-13.  
- Shi, R.A., (2021a) Learning from Zero: How to Make Consumption-Saving Decisions in a Stochastic Environment with an AI Algorithm. arXiv preprint arXiv:2105.10099.  
- Shi, R. A., (2021b) Can an AI agent hit a moving target?. arXiv preprint arXiv:2110.02474.  
- Covarrubias, M. (2022), Dynamic Oligopoly and Monetary Policy: A Deep Reinforcement Learning Approach.  

### Causality, reasoning, and interpretability in RL
- Dasgupta, I., Wang, J., Chiappa, S., Mitrovic, J., Ortega, P., Raposo, D., ... & Kurth-Nelson, Z. (2019). Causal reasoning from meta-reinforcement learning. arXiv preprint arXiv:1901.08162.  
- Gasse, M., Grasset, D., Gaudron, G., & Oudeyer, P. Y. (2021). Causal reinforcement learning using observational and interventional data. arXiv preprint arXiv:2106.14421.  
- Gershman, S. J. (2017). Reinforcement learning and causal models. The Oxford handbook of causal reasoning, 295.  
- Grimbly, St John. “Causal Reinforcement Learning: A Primer | Asking Why.” Asking Why., stjohngrimbly.com, 9 Dec. 2020, https://stjohngrimbly.com/causal-reinforcement-learning/.  

### Quantum machine learning and quantum reinforcement learning
- Biamonte, J., Wittek, P., Pancotti, N., Rebentrost, P., Wiebe, N., & Lloyd, S. (2017). Quantum machine learning. Nature, 549(7671), 195-202.  
- Wittek, P. (2014). Quantum machine learning: what quantum computing means to data mining. Academic Press.  
- Dunjko, V., Taylor, J. M., & Briegel, H. J. (2017, October). Advances in quantum reinforcement learning. In 2017 IEEE International Conference on Systems, Man, and Cybernetics (SMC) (pp. 282-287). IEEE.  
- Wu, S., Jin, S., Wen, D., & Wang, X. (2020). Quantum reinforcement learning in continuous action space. arXiv preprint arXiv:2012.10711.  
- Madsen, L. S., Laudenbach, F., Askarani, M. F., Rortais, F., Vincent, T., Bulmer, J. F., ... & Lavoie, J. (2022). Quantum computational advantage with a programmable photonic processor. Nature, 606(7912), 75-81.  
- Zhang, Y., & Ni, Q. (2020). Recent advances in quantum machine learning. Quantum Engineering, 2(1), e34.  

### Domain-specific and engineering applications of RL
- Liu, T., Tan, Z., Xu, C., Chen, H. and Li, Z., 2020. Study on deep reinforcement learning techniques for building energy consumption forecasting. Energy and Buildings, 208, p.109675.  
- Li, B., Xie, K., Huang, X., Wu, Y., & Xie, S. (2022). Deep Reinforcement Learning based Incentive Mechanism Design for Platoon Autonomous Driving with Social Effect. IEEE Transactions on Vehicular Technology.  
- Zhan, Y., & Zhang, J. (2020, July). An incentive mechanism design for efficient edge learning by deep reinforcement learning approach. In IEEE INFOCOM 2020-IEEE Conference on Computer Communications (pp. 2489-2498). IEEE.  
- Zhu, F., Liao, P., Zhu, X., Yao, J., & Huang, J. (2018, August). Cohesion-driven Online Actor-Critic Reinforcement Learning for mHealth Intervention. In Proceedings of the 2018 ACM International Conference on Bioinformatics, Computational Biology, and Health Informatics (pp. 482-491).  
- Vezhnevets, A.S., Osindero, S., Schaul, T., Heess, N., Jaderberg, M., Silver, D. and Kavukcuoglu, K., 2017, July. Feudal networks for hierarchical reinforcement learning. In International Conference on Machine Learning (pp. 3540-3549). PMLR.  
- Nagabandi, A., Clavera, I., Liu, S., Fearing, R.S., Abbeel, P., Levine, S. and Finn, C., 2018. Learning to adapt in dynamic, real-world environments through meta-reinforcement learning. arXiv preprint arXiv:1803.11347.  
- Zhu, Z., Lin, K., & Zhou, J. (2020). Transfer learning in deep reinforcement learning: A survey. arXiv preprint arXiv:2009.07888.  

### Agent-based modeling, calibration, and socio-economic dynamics
- Sert, E., Bar-Yam, Y. and Morales, A.J., 2020. Segregation dynamics with reinforcement learning and agent based modeling. Scientific reports, 10(1), pp.1-12.  
- Song, B., Xiong, G., Yu, S., Ye, P., Dong, X. and Lv, Y., 2021, July. Calibration of agent-based model using reinforcement learning. In 2021 IEEE 1st International Conference on Digital Twins and Parallel Intelligence (DTPI) (pp. 278-281). IEEE.  
- Dasgupta, I., Wang, J., Chiappa, S., Mitrovic, J., Ortega, P., Raposo, D., ... & Kurth-Nelson, Z. (2019). Causal reasoning from meta-reinforcement learning. arXiv preprint arXiv:1901.08162.  

_Working Paper No. WP/2022/259_

---


_Source: https://www.imf.org/-/media/files/publications/wp/2022/english/wpiea2022259-print-pdf.pdf_
