Approximate Policy Iteration

Simulation
Author

Ziang Liu

Published

September 1, 2026

Classification-based Policy Iteration

Lagoudakis and Parr (2003) proposed an classification-based policy iteration that uses classifiers such as SVM and neural networks to approximate the policy.

Lazaric et al. (2016) introduced cost-sensitive loss function weighting each classification mistake by its actual regret.

References

Lagoudakis, Michail G, and Ronald Parr. 2003. Reinforcement Learning as Classification: Leveraging Modern Classifiers. Https://cdn.aaai.org/ICML/2003/ICML03-057.pdf.
Lazaric, Alessandro, Mohammad Ghavamzadeh, and Rémi Munos. 2016. “Analysis of Classification-Based Policy Iteration Algorithms.” Journal of Machine Learning Research 17 (19): 1–30.