跳到主要内容
步芽

【极致中配】斯坦福大学 CS234 强化学习 I 2024春季

斯坦福CS234系统讲授强化学习:从表格型MDP、策略评估到策略搜索、离线RL、探索、多智能体与价值对齐等前沿主题。

难度
难度 4/5需概率、线性代数与机器学习基础,含前沿研究内容
适合人群
有ML基础、想系统学习强化学习的高年级本科生与研究生
前置要求
概率论(马尔可夫决策过程与期望推导必备)、线性代数(函数逼近与价值表示)、机器学习基础(监督学习、神经网络)、Python编程(实现RL算法)
课程规模
16 · 3102播放

主题覆盖

强化学习导论马尔可夫决策过程策略评估Q学习函数逼近策略搜索离线强化学习DPO探索策略多智能体博弈价值对齐

课程大纲(16 讲)

  1. P1 · Lecture 1 Introduction to Reinforcement Learning I 202442 分钟
  2. P2 · Lecture 2 Tabular MDP Planning43 分钟
  3. P3 · Lecture 3 Policy Evaluation46 分钟
  4. P4 · Lecture 4 Q learning and Function Approximation43 分钟
  5. P5 · Lecture 5 Policy Search 139 分钟
  6. P6 · Lecture 6 Policy Search 245 分钟
  7. P7 · Lecture 7 Policy Search 344 分钟
  8. P8 · Lecture 8 Offline RL 141 分钟
  9. P9 · Lecture 9 Guest Lecture on DPO: Rafael Rafailov, Archit Sharma, Eric Mitchell46 分钟
  10. P10 · Lecture 10 Offline RL 345 分钟
  11. P11 · Lecture 11 Exploration 142 分钟
  12. P12 · Lecture 12 Exploration 242 分钟
  13. P13 · Lecture 13 Exploration 340 分钟
  14. P14 · Lecture 14 Multi-Agent Game Playing42 分钟
  15. P15 · Lecture 15 Emma Brunskill & Dan Webber42 分钟
  16. P16 · Lecture 16 Value Alignment40 分钟

本课程卡由 AI 生成,可能存在误差,欢迎反馈。