mirror of
https://github.com/janishutz/eth-summaries.git
synced 2026-09-10 19:15:25 +02:00
[AMR] Add feedback loops
This commit is contained in:
+3
@@ -26,3 +26,6 @@ $\varepsilon$ is prob. to act randomly, $1 - \varepsilon$ is prob. to act on pol
|
||||
Value-based (estimate val or $Q$-func and extract pol., e.g. Q-Learn),
|
||||
Actor-Critic (estim. val or $Q$ of curr. pol., improve pol., e.g. A3C, SAC),
|
||||
Policy-Gradient (diff. expect. reward w.r.t. params of policy network, e.g. REINFORCE)
|
||||
|
||||
\includegraphics[width=0.5\columnwidth]{assets/loop-rl.png}
|
||||
\includegraphics[width=0.5\columnwidth]{assets/loop-drl.png}
|
||||
|
||||
Reference in New Issue
Block a user