Approximate Policy Iteration
Simulation
Classification-based Policy Iteration
Lagoudakis and Parr (2003) proposed an classification-based policy iteration that uses classifiers such as SVM and neural networks to approximate the policy.
Lazaric et al. (2016) introduced cost-sensitive loss function weighting each classification mistake by its actual regret.
References
Lagoudakis, Michail G, and Ronald Parr. 2003. Reinforcement Learning as Classification: Leveraging Modern Classifiers. Https://cdn.aaai.org/ICML/2003/ICML03-057.pdf.
Lazaric, Alessandro, Mohammad Ghavamzadeh, and Rémi Munos. 2016. “Analysis of Classification-Based Policy Iteration Algorithms.” Journal of Machine Learning Research 17 (19): 1–30.