社区所有版块导航
Python
python开源   Django   Python   DjangoApp   pycharm  
DATA
docker   Elasticsearch  
aigc
aigc   chatgpt  
WEB开发
linux   MongoDB   Redis   DATABASE   NGINX   其他Web框架   web工具   zookeeper   tornado   NoSql   Bootstrap   js   peewee   Git   bottle   IE   MQ   Jquery  
机器学习
机器学习算法  
Python88.com
反馈   公告   社区推广  
产品
短视频  
印度
印度  
Py学习  »  机器学习算法

【干货书】强化学习算法,98页pdf综合讲解人工智能和机器学习

专知 • 5 年前 • 412 次点击  


强化学习是一种学习范式,它关注于如何学习控制一个系统,从而最大化表达一个长期目标的数值性能度量。强化学习与监督学习的区别在于,对于学习者的预测,只向学习者提供部分反馈。此外,预测还可能通过影响被控系统的未来状态而产生长期影响。因此,时间起着特殊的作用。强化学习的目标是开发高效的学习算法,以及了解算法的优点和局限性。强化学习具有广泛的实际应用价值,从人工智能到运筹学或控制工程等领域。在这本书中,我们重点关注那些基于强大的动态规划理论的强化学习算法。我们给出了一个相当全面的学习问题目录,描述了核心思想,关注大量的最新算法,然后讨论了它们的理论性质和局限性。


  1. Preface ix

  2. Acknowledgments xiii

  3. Markov Decision Processes 1

    1. Preliminaries 1

    2. Markov Decision Processes 1

    3. Value functions 6

    4. Dynamic programming algorithms for solving MDPs 10

  4. Value Prediction Problems 11

    1. TD(lambda) with function approximation 22

    2. Gradient temporal difference learning 25

    3. Least-squares methods 27

    4. The choice of the function space 33

    5. Tabular TD(0) 11

    6. Every-visit Monte-Carlo 14

    7. TD(lambda): Unifying Monte-Carlo and TD(0) 16

    1. Temporal difference learning in finite state spaces 11

    2. Algorithms for large state spaces 18

  5. Control 37

    1. Implementing a critic 54

    2. Implementing an actor 56

    3. Q-learning in finite MDPs 47

    4. Q-learning with function approximation 49

    5. Online learning in bandits 38

    6. Active learning in bandits 40

    7. Active learning in Markov Decision Processes 41

    8. Online learning in Markov Decision Processes 42

    1. A catalog of learning problems 37

    2. Closed-loop interactive learning 38

    3. Direct methods 47

    4. Actor-critic methods 52

  6. For Further Exploration 63

    1. Further reading 63

    2. Applications 63

    3. Software 64

  7. Appendix: The Theory of Discounted Markovian Decision Processes 65

    1. A.1 Contractions and Banach’s fixed-point theorem 65

    2. A.2 Application to MDPs 69

  8. Bibliography 73

  9. Author's Biography 89


https://sites.ualberta.ca/~szepesva/rlbook.html



专知便捷查看

便捷下载,请关注专知公众号(点击上方蓝色专知关注)

  • 后台回复“ A98” 可以获取《【干货书】强化学习算法,98页pdf综合讲解人工智能和机器学习》专知下载链接索引

专知,专业可信的人工智能知识分发,让认知协作更快更好!欢迎注册登录专知www.zhuanzhi.ai,获取5000+AI主题干货知识资料!
欢迎微信扫一扫加入专知人工智能知识星球群,获取最新AI专业干货知识教程资料和与专家交流咨询!
点击“阅读原文”,了解使用专知,查看获取5000+AI主题知识资源
Python社区是高质量的Python/Django开发社区
本文地址:http://www.python88.com/topic/107933