{"id":123,"date":"2022-07-21T19:25:21","date_gmt":"2022-07-21T18:25:21","guid":{"rendered":"https:\/\/wp.coventry.domains\/e2edu\/?page_id=123"},"modified":"2022-08-23T16:41:39","modified_gmt":"2022-08-23T15:41:39","slug":"reinforcement-learning","status":"publish","type":"page","link":"https:\/\/wp.coventry.domains\/e2edu\/reinforcement-learning\/","title":{"rendered":"Reinforcement Learning"},"content":{"rendered":"\n<p class=\"wp-block-paragraph\">Reinforcement Learning (RL) is a Machine Learning approach that is inspired by how animals or young children learn from their mistakes (or successes). This approach combines trial and error actions with a reward. A typical natural example of this type of learning would be a young child who touches the door of a hot oven and burns his or her fingers. The pain sensation from getting burnt represents a strong negative reward that will likely cause the child to avoid touching hot oven doors from now on. <\/p>\n\n\n\n<p class=\"wp-block-paragraph\">This type of learning is often applied for agent-based simulations in which an agent needs to develop a strategy for successfully interacting with an environment. The level of success is measured as a reward that an agent receives from the environment after performing an action. This reward can be positive or negative (a punishment) and it can be received instantaneously or in the future. One of the main challenges of RL concerns the correct attribution of rewards to actions when the actions lie far in the past. <\/p>\n\n\n\n<h2 class=\"wp-block-heading\">Agent, Environment, Action, State, Reward<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">A RL system includes the following elements: Environment, Agent, State, Observation, Action, Reward. The environment represents the world within which the agent acts. The world and the agent have at each point in time a particular state.  An agent represents the entity that interacts with the environment and tries to maximise its reward. An agent can observe the environment. Depending on the implementation, such an observation is identical with the state of the environment or it is different. When it is different, then an observation might reveal to the agent only partial information about the environment&#8217;s state. Based on this observation, the agent choses an action. An action is an activity conducted by the agent that might cause  a change in the state of the environment. Once an action has been executed, the environment responds to  it by returning to the agent a reward value and by assuming a new state. <\/p>\n\n\n<div class=\"wp-block-image\">\n<figure class=\"aligncenter size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"908\" height=\"350\" src=\"https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/RL_Agent_Environment_Reward_State.png\" alt=\"\" class=\"wp-image-838\" srcset=\"https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/RL_Agent_Environment_Reward_State.png 908w, https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/RL_Agent_Environment_Reward_State-300x116.png 300w, https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/RL_Agent_Environment_Reward_State-768x296.png 768w, https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/RL_Agent_Environment_Reward_State-788x304.png 788w\" sizes=\"auto, (max-width: 908px) 100vw, 908px\" \/><figcaption>Information passed between an Agent and Its Environment. \u00a9Shweta Bhatt<\/figcaption><\/figure>\n<\/div>\n\n\n<p class=\"wp-block-paragraph\">In the case of the minimalistic and discrete agent and environment depicted below, the state is the position of the agent in a grid world. In this world, there are a total of 11 different states. The action of the agent represents the movement of the agent into a neighbouring grid cell. There are a total of four different actions at the agent&#8217;s disposal. Not in all RL-systems are the states and actions discrete. Other systems possess continuous states or actions or both.<\/p>\n\n\n<div class=\"wp-block-image\">\n<figure class=\"aligncenter size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"338\" height=\"212\" src=\"https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/RL_GridWorld.png\" alt=\"\" class=\"wp-image-842\" srcset=\"https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/RL_GridWorld.png 338w, https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/RL_GridWorld-300x188.png 300w\" sizes=\"auto, (max-width: 338px) 100vw, 338px\" \/><figcaption>A Minimalistic Agent and Environment. The environment is a grid world in which discrete locations cause either a positive (green cell), negative (red cell), or no reward (white cell) when the agent moves into them.  \u00a9Anish Phadnis<\/figcaption><\/figure>\n<\/div>\n\n\n<h2 class=\"wp-block-heading\">Basic Assumptions and Terminology<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">The following section describes some of the basic assumptions, terminology and equations used in reinforcement learning. <\/p>\n\n\n\n<h4 class=\"wp-block-heading\">Reward Hypothesis<\/h4>\n\n\n\n<p class=\"wp-block-paragraph\">The reward hypothesis assumes that the goal of an intelligent agent is to maximise some sort of reward.  The agent&#8217;s intelligence and all its capabilities are understood as being subservient to this goal. <\/p>\n\n\n\n<h4 class=\"wp-block-heading\">Markov Property<\/h4>\n\n\n\n<p class=\"wp-block-paragraph\">The Markov Property assumes that an agent and its environment are memoryless. According to this property, the future state that an agent and environment will transition to depends only on the immediately preceding state and no other states further back in history.  <\/p>\n\n\n\n<h4 class=\"wp-block-heading\">Markov Decision Processes<\/h4>\n\n\n\n<p class=\"wp-block-paragraph\">The change of states due to the actions of an agent can be mathematically framed in terms of Markov Decision Processes (MDP).  A MDP is a stochastic process that proceeds in discrete time steps. This process defines the probability of transitioning from one state to the next when taking an action and the associated reward.<\/p>\n\n\n<div class=\"wp-block-image\">\n<figure class=\"aligncenter size-full is-resized\"><img loading=\"lazy\" decoding=\"async\" src=\"https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/MarkovDecisionProcess.png\" alt=\"\" class=\"wp-image-848\" width=\"559\" height=\"411\" srcset=\"https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/MarkovDecisionProcess.png 751w, https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/MarkovDecisionProcess-300x221.png 300w\" sizes=\"auto, (max-width: 559px) 100vw, 559px\" \/><figcaption>A Markov Decision Process for a Day in the Live of a Student. Circles and rectangles represent states, arrows represent transitions, R stands for rewards, and orange words represent actions. \u00a9David Silver<\/figcaption><\/figure>\n<\/div>\n\n\n<p class=\"wp-block-paragraph\">More information on Markov Decision Processes can be found for example in the article &#8220;<a rel=\"noreferrer noopener\" href=\"https:\/\/levelup.gitconnected.com\/markov-decision-processes-simplified-f5f8d37ab70f\" target=\"_blank\">Markov Decision Processes Simplified<\/a>&#8220;.<\/p>\n\n\n\n<h4 class=\"wp-block-heading\">Model-Free and Model-Based Reinforcement Learning<\/h4>\n\n\n\n<p class=\"wp-block-paragraph\">There exist two principle types of RL algorithms that differ from each other with respect to how they deal with state transition probabilities. A model-based RL algorithm tries to model these transition probabilities while a model-free RL algorithm does not. Therefore, a model-free RL algorithm is fully following a trial-and-error approach while a model-based RL algorithm does this to a lesser degree. All the RL algorithms explained in this article are model-free.<\/p>\n\n\n\n<h4 class=\"wp-block-heading\">Policy<\/h4>\n\n\n\n<p class=\"wp-block-paragraph\">What an agent tries to learn in an RL-scenario is a policy that maximises reward.  A policy defines the probability of an agent to take a certain action when the environment is in a certain state.  The policy that maximizes the total reward is called the\u00a0optimal policy. In mathematical equations, the policy is abbreviated as pi.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Policies can be implemented with any data structure and\/or method that is suitable. For simple RL systems that operate on discrete states and actions, a policy can be implemented for example as a simple lookup table between states and actions. For more complicated RL systems with a large number or continuous states or actions, a policy can be implemented for example with a neural network. The same applies to state and value functions which are described further down.<\/p>\n\n\n\n<h4 class=\"wp-block-heading\">On-Policy and Off-Policy Methods<\/h4>\n\n\n\n<p class=\"wp-block-paragraph\">In on-policy learning, the policy an agent uses to select actions is improved through learning. On-policy methods are called&nbsp;<img loading=\"lazy\" decoding=\"async\" height=\"8\" width=\"7\" src=\"https:\/\/www.baeldung.com\/wp-content\/ql-cache\/quicklatex.com-f1ea683a5e3ac49e12a81be8cd57fe90_l3.svg\" alt=\"\\epsilon\">-greedy policies&nbsp;as they select random actions with an&nbsp;<img loading=\"lazy\" decoding=\"async\" height=\"8\" width=\"7\" src=\"https:\/\/www.baeldung.com\/wp-content\/ql-cache\/quicklatex.com-f1ea683a5e3ac49e12a81be8cd57fe90_l3.svg\" alt=\"\\epsilon\"> probability and follow the optimal action with 1-<img loading=\"lazy\" decoding=\"async\" height=\"8\" width=\"7\" src=\"https:\/\/www.baeldung.com\/wp-content\/ql-cache\/quicklatex.com-f1ea683a5e3ac49e12a81be8cd57fe90_l3.svg\" alt=\"\\epsilon\"> probability.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">In off-policy learning, an optimal policy is discovered independently of the policy that an agent uses to select actions. Accordingly, these methods have two policies, a behaviour policy that is used for exploring and a target policy that is used for improvement. <\/p>\n\n\n\n<p class=\"wp-block-paragraph\">On- and off-policy methods have their advantages and disadvantages. On-policy methods are said to be more stable and off-policy methods are more efficient and provide a better balance between the exploration of new policies and the exploitation of the currently best policies.<\/p>\n\n\n\n<h4 class=\"wp-block-heading\">Discounted Cumulative Reward<\/h4>\n\n\n\n<p class=\"wp-block-paragraph\">In an RL system, not only the reward that an agent receives immediately after performing an action is important. Rather, an agent tries to maximise the accumulation of all rewards it receives from a particular point in time on until the end of what is called an episode (a simulation run). Typically, when summing these rewards, the influence of rewards that lie further in the future is made less strong than more immediate rewards. This is achieved by multiplying the rewards with a discount factor that is taken to the power of time. The discounted cumulative reward is mathematically defined as follows:<\/p>\n\n\n<div class=\"wp-block-image\">\n<figure class=\"aligncenter size-full is-resized\"><img loading=\"lazy\" decoding=\"async\" src=\"https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/DiscountedCumulativeReward.png\" alt=\"\" class=\"wp-image-856\" width=\"548\" height=\"158\" srcset=\"https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/DiscountedCumulativeReward.png 812w, https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/DiscountedCumulativeReward-300x87.png 300w, https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/DiscountedCumulativeReward-768x222.png 768w, https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/DiscountedCumulativeReward-788x228.png 788w, https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/DiscountedCumulativeReward-350x100.png 350w\" sizes=\"auto, (max-width: 548px) 100vw, 548px\" \/><figcaption>Return as Discounted Cumulative Reward. Gt stands for the return at time t, Rt stands for the Reward at time t, gamma is the discount factor. These equations are for an infinitely long episode. \u00a9Wouter van Heeswijk<\/figcaption><\/figure>\n<\/div>\n\n\n<h4 class=\"wp-block-heading\">Value Functions<\/h4>\n\n\n\n<p class=\"wp-block-paragraph\">Value functions specify the return that can be expected by an agent when following its policy. There are two types of value functions, state value function and action value function.<\/p>\n\n\n\n<h4 class=\"wp-block-heading\">State Value Function<\/h4>\n\n\n\n<p class=\"wp-block-paragraph\">The state value function is the expected return starting from the current state and then following a policy. <\/p>\n\n\n<div class=\"wp-block-image\">\n<figure class=\"aligncenter size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"493\" height=\"56\" src=\"https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/StateValueFunction.png\" alt=\"\" class=\"wp-image-858\" srcset=\"https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/StateValueFunction.png 493w, https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/StateValueFunction-300x34.png 300w\" sizes=\"auto, (max-width: 493px) 100vw, 493px\" \/><figcaption>State Value Function. V_pi represents the state value function for state s when following policy pi. E_pi stands for the expectation under policy pi. Gt is the discounted cumulative reward starting from time t and with state s. \u00a9Jordi Torres<\/figcaption><\/figure>\n<\/div>\n\n\n<p class=\"wp-block-paragraph\">If states are discrete, state values are associated to all the states an agent and environment can be in. In case of the simple grid world example from before, the state values might assume the following values after a few training iterations:<\/p>\n\n\n<div class=\"wp-block-image\">\n<figure class=\"aligncenter size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"400\" height=\"301\" src=\"https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/StateValuesOnGrid.png\" alt=\"\" class=\"wp-image-902\" srcset=\"https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/StateValuesOnGrid.png 400w, https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/StateValuesOnGrid-300x226.png 300w\" sizes=\"auto, (max-width: 400px) 100vw, 400px\" \/><figcaption>State Values for a Minimalistic Agent and Environment.  \u00a9Michael Avendi<\/figcaption><\/figure>\n<\/div>\n\n\n<p class=\"wp-block-paragraph\">An optimal state value function can be expressed in terms of the optimal policy:<\/p>\n\n\n<div class=\"wp-block-image\">\n<figure class=\"aligncenter size-large\"><img loading=\"lazy\" decoding=\"async\" width=\"1024\" height=\"69\" src=\"https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/OptimalStateValueFunction-1024x69.png\" alt=\"\" class=\"wp-image-881\" srcset=\"https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/OptimalStateValueFunction-1024x69.png 1024w, https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/OptimalStateValueFunction-300x20.png 300w, https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/OptimalStateValueFunction-768x52.png 768w, https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/OptimalStateValueFunction-1536x104.png 1536w, https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/OptimalStateValueFunction-788x53.png 788w, https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/OptimalStateValueFunction.png 1715w\" sizes=\"auto, (max-width: 1024px) 100vw, 1024px\" \/><figcaption>Optimal State Value Function. Max_pi stands for the optimal policy. \u00a9Jordi Torres<\/figcaption><\/figure>\n<\/div>\n\n\n<h4 class=\"wp-block-heading\">Action Value Function<\/h4>\n\n\n\n<p class=\"wp-block-paragraph\">The action value function is the expected return starting from the current state, then taking an action, and then following a policy.<\/p>\n\n\n<div class=\"wp-block-image\">\n<figure class=\"aligncenter size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"980\" height=\"62\" src=\"https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/ActionValueFunction.png\" alt=\"\" class=\"wp-image-859\" srcset=\"https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/ActionValueFunction.png 980w, https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/ActionValueFunction-300x19.png 300w, https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/ActionValueFunction-768x49.png 768w, https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/ActionValueFunction-788x50.png 788w\" sizes=\"auto, (max-width: 980px) 100vw, 980px\" \/><figcaption>Action Value Function. Q_pi represents the action value function for state s and action a when following policy pi. E_pi stands for the expectation under policy pi. Gt is the discounted cumulative reward starting from time t and with state s and action a. \u00a9Jordi Torres<\/figcaption><\/figure>\n<\/div>\n\n\n<p class=\"wp-block-paragraph\">If both the states and actions are discrete, action values are associated with all state action pairs.  In case of the simple grid world example from before, the action values might assume the following values after a few training iterations:<\/p>\n\n\n<div class=\"wp-block-image\">\n<figure class=\"aligncenter size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"400\" height=\"301\" src=\"https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/ActionValuesOnGrid.png\" alt=\"\" class=\"wp-image-901\" srcset=\"https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/ActionValuesOnGrid.png 400w, https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/ActionValuesOnGrid-300x226.png 300w\" sizes=\"auto, (max-width: 400px) 100vw, 400px\" \/><figcaption>Action Values for a Minimalistic Agent and Environment.  \u00a9Michael Avendi<\/figcaption><\/figure>\n<\/div>\n\n\n<p class=\"wp-block-paragraph\">An optimal action value function can be expressed in terms of the optimal policy:<\/p>\n\n\n<div class=\"wp-block-image\">\n<figure class=\"aligncenter size-large\"><img loading=\"lazy\" decoding=\"async\" width=\"1024\" height=\"51\" src=\"https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/OptimalActionValueFunction-1024x51.png\" alt=\"\" class=\"wp-image-882\" srcset=\"https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/OptimalActionValueFunction-1024x51.png 1024w, https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/OptimalActionValueFunction-300x15.png 300w, https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/OptimalActionValueFunction-768x38.png 768w, https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/OptimalActionValueFunction-1536x76.png 1536w, https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/OptimalActionValueFunction-788x39.png 788w, https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/OptimalActionValueFunction.png 1715w\" sizes=\"auto, (max-width: 1024px) 100vw, 1024px\" \/><figcaption>Optimal Action Value Function. Max_pi stands for the optimal policy. \u00a9Jordi Torres<\/figcaption><\/figure>\n<\/div>\n\n\n<h4 class=\"wp-block-heading\">Bellman Equation<\/h4>\n\n\n\n<p class=\"wp-block-paragraph\">The Bellman equation decomposes a value function into two parts, the immediate reward plus the discounted future rewards. The second part can be formulated recursively.<\/p>\n\n\n<div class=\"wp-block-image\">\n<figure class=\"aligncenter size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"493\" height=\"70\" src=\"https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/BellmanStateValueFunction.png\" alt=\"\" class=\"wp-image-869\" srcset=\"https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/BellmanStateValueFunction.png 493w, https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/BellmanStateValueFunction-300x43.png 300w\" sizes=\"auto, (max-width: 493px) 100vw, 493px\" \/><figcaption>Bellman Equation for State Value Function. This equation multiplies the probabilities of any action an agent might take in state s with the probability of reaching any successor state s&#8217; with the reward associated with state s and action a plus the discounted state value for the successor state. \u00a9Jordi Torres<\/figcaption><\/figure>\n<\/div>\n\n<div class=\"wp-block-image\">\n<figure class=\"aligncenter size-large\"><img loading=\"lazy\" decoding=\"async\" width=\"1024\" height=\"82\" src=\"https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/BellmanActionValueFunction-1024x82.png\" alt=\"\" class=\"wp-image-870\" srcset=\"https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/BellmanActionValueFunction-1024x82.png 1024w, https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/BellmanActionValueFunction-300x24.png 300w, https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/BellmanActionValueFunction-768x62.png 768w, https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/BellmanActionValueFunction-788x63.png 788w, https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/BellmanActionValueFunction.png 1294w\" sizes=\"auto, (max-width: 1024px) 100vw, 1024px\" \/><figcaption>Bellman Equation for Action Value Function. This equation multiplies the probability of reaching any successor state s&#8217; when taking action a with the reward associated with state s and action a plus the discounted action value for any possible successor state and action multiplied by their probabilities. \u00a9Jordi Torres<\/figcaption><\/figure>\n<\/div>\n\n\n<p class=\"wp-block-paragraph\">These two equations can also can also be specified in terms of an optimal policy. In this case, the equations are named Bellman optimality equations. <\/p>\n\n\n<div class=\"wp-block-image\">\n<figure class=\"aligncenter size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"989\" height=\"106\" src=\"https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/BellmanOptimalityStateValueFunction.png\" alt=\"\" class=\"wp-image-886\" srcset=\"https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/BellmanOptimalityStateValueFunction.png 989w, https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/BellmanOptimalityStateValueFunction-300x32.png 300w, https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/BellmanOptimalityStateValueFunction-768x82.png 768w, https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/BellmanOptimalityStateValueFunction-788x84.png 788w\" sizes=\"auto, (max-width: 989px) 100vw, 989px\" \/><figcaption>Bellman Optimality Equation for the State Value Function. Max_a stands for the probability of the action that returns the maximum possible immediate reward. \u00a9Jordi Torres<\/figcaption><\/figure>\n<\/div>\n\n<div class=\"wp-block-image\">\n<figure class=\"aligncenter size-large\"><img loading=\"lazy\" decoding=\"async\" width=\"1024\" height=\"104\" src=\"https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/BellmanOptimalityActionValueFunction-1024x104.png\" alt=\"\" class=\"wp-image-887\" srcset=\"https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/BellmanOptimalityActionValueFunction-1024x104.png 1024w, https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/BellmanOptimalityActionValueFunction-300x31.png 300w, https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/BellmanOptimalityActionValueFunction-768x78.png 768w, https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/BellmanOptimalityActionValueFunction-788x80.png 788w, https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/BellmanOptimalityActionValueFunction.png 1040w\" sizes=\"auto, (max-width: 1024px) 100vw, 1024px\" \/><figcaption>Bellman Optimality Equation for the Action Value Function. Max_a&#8217; stands for the probability of a successor action that returns the maximum possible immediate reward. \u00a9Jordi Torres<\/figcaption><\/figure>\n<\/div>\n\n\n<h4 class=\"wp-block-heading\">TD-Error<\/h4>\n\n\n\n<p class=\"wp-block-paragraph\">TD in TD-Error stands for time difference. The TD-Error is frequently used in RL to improve the state or action value during learning. The TD-Error measures the difference between the currently observed state or action value and the previously assumed state or action value. <\/p>\n\n\n\n<p class=\"wp-block-paragraph\">For the state value, the TD-Error is mathematically formulated as follows:<\/p>\n\n\n<div class=\"wp-block-image\">\n<figure class=\"aligncenter size-full is-resized\"><img loading=\"lazy\" decoding=\"async\" src=\"https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/TDErrorStateValue.jpg\" alt=\"\" class=\"wp-image-919\" width=\"446\" height=\"37\" srcset=\"https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/TDErrorStateValue.jpg 853w, https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/TDErrorStateValue-300x25.jpg 300w, https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/TDErrorStateValue-768x64.jpg 768w, https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/TDErrorStateValue-788x66.jpg 788w\" sizes=\"auto, (max-width: 446px) 100vw, 446px\" \/><figcaption>TD-Error for State Value Function.<\/figcaption><\/figure>\n<\/div>\n\n\n<p class=\"wp-block-paragraph\">For the action value, the TD-Error is mathematically formulated as follows:<\/p>\n\n\n<div class=\"wp-block-image\">\n<figure class=\"aligncenter size-full is-resized\"><img loading=\"lazy\" decoding=\"async\" src=\"https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/TDErrorActionValue.jpg\" alt=\"\" class=\"wp-image-920\" width=\"508\" height=\"34\" srcset=\"https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/TDErrorActionValue.jpg 1004w, https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/TDErrorActionValue-300x20.jpg 300w, https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/TDErrorActionValue-768x51.jpg 768w, https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/TDErrorActionValue-788x53.jpg 788w\" sizes=\"auto, (max-width: 508px) 100vw, 508px\" \/><figcaption>TD-Error for Action Value Function.<\/figcaption><\/figure>\n<\/div>\n\n\n<p class=\"wp-block-paragraph\">Using the TD-Error, the state or action values can be updated during training by adding to the current state or action value the corresponding TD-Error multiplied by the learning rate.<\/p>\n\n\n<div class=\"wp-block-image\">\n<figure class=\"aligncenter size-full is-resized\"><img loading=\"lazy\" decoding=\"async\" src=\"https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/StateValueUpdate.jpg\" alt=\"\" class=\"wp-image-922\" width=\"491\" height=\"33\" srcset=\"https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/StateValueUpdate.jpg 1024w, https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/StateValueUpdate-300x20.jpg 300w, https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/StateValueUpdate-768x51.jpg 768w, https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/StateValueUpdate-788x52.jpg 788w\" sizes=\"auto, (max-width: 491px) 100vw, 491px\" \/><figcaption>Update of the State Value using the TD-Error. Alpha represents the learning rate. <\/figcaption><\/figure>\n<\/div>\n\n\n<figure class=\"wp-block-image size-large\"><img loading=\"lazy\" decoding=\"async\" width=\"1024\" height=\"56\" src=\"https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/ActionValueUpdate-1024x56.jpg\" alt=\"\" class=\"wp-image-923\" srcset=\"https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/ActionValueUpdate-1024x56.jpg 1024w, https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/ActionValueUpdate-300x17.jpg 300w, https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/ActionValueUpdate-768x42.jpg 768w, https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/ActionValueUpdate-788x43.jpg 788w, https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/ActionValueUpdate.jpg 1345w\" sizes=\"auto, (max-width: 1024px) 100vw, 1024px\" \/><figcaption>Update of the Action Value using the TD-Error. Alpha represents the learning rate. <\/figcaption><\/figure>\n\n\n\n<h2 class=\"wp-block-heading\">Example RL-Algorithm: SARSA<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">SARSA is a model-free on-policy method that works best with discrete states and actions. SARSA stands for the tuple S, A, R, S1, A1. S represents the current state, A the currently taken action, R the reward received after taking action A, S1 the next state, and A1 the next action taken. SARSA operates on action values. The pseudocode for training the SARA algorithm is as follows: <\/p>\n\n\n<div class=\"wp-block-image\">\n<figure class=\"aligncenter size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"976\" height=\"379\" src=\"https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/SARSA_Pseudocode.png\" alt=\"\" class=\"wp-image-912\" srcset=\"https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/SARSA_Pseudocode.png 976w, https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/SARSA_Pseudocode-300x116.png 300w, https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/SARSA_Pseudocode-768x298.png 768w, https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/SARSA_Pseudocode-788x306.png 788w\" sizes=\"auto, (max-width: 976px) 100vw, 976px\" \/><figcaption>Pseudocode for SARSA Algorithm. \u00a9Adesh Gautam<\/figcaption><\/figure>\n<\/div>\n\n\n<h2 class=\"wp-block-heading\">Example RL-Algorithm: Q-Learning<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">Q-Learning is a model-free off-policy method that also works best with discrete states and actions. Q-learning seeks to find the best action to take given the current state. Accordingly, the update function for the action value is slightly different from SARSA when it comes to selecting the action for the next state and action. Instead of selecting an action that is suggested by the current policy (as is the case with SARSA), the action chosen is the one whose action value is highest. The pseudocode for training the Q-Learning algorithm is as follows: <\/p>\n\n\n<div class=\"wp-block-image\">\n<figure class=\"aligncenter size-large\"><img loading=\"lazy\" decoding=\"async\" width=\"1024\" height=\"425\" src=\"https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/QLearning_Pseudocode-1024x425.png\" alt=\"\" class=\"wp-image-928\" srcset=\"https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/QLearning_Pseudocode-1024x425.png 1024w, https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/QLearning_Pseudocode-300x124.png 300w, https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/QLearning_Pseudocode-768x318.png 768w, https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/QLearning_Pseudocode-1536x637.png 1536w, https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/QLearning_Pseudocode-788x327.png 788w, https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/QLearning_Pseudocode.png 1647w\" sizes=\"auto, (max-width: 1024px) 100vw, 1024px\" \/><figcaption>Pseudocode for Q-Learning Algorithm. \u00a9Zitao Shen<\/figcaption><\/figure>\n<\/div>\n\n\n<h2 class=\"wp-block-heading\">Policy Gradient Methods<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">Both SARSA and Q-Learning are value-based RL algorithm. Policy Gradient methods differ from value-based RL algorithms in that they operate directly on a policy instead of state or action value functions. Policy-based RL is effective in high dimensional and stochastic continuous action spaces and for learning stochastic policies. Value-based RL excels in sample efficiency and stability. <\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Many of the more sophisticated RL setups employ continuous states and actions. A classical example is the &#8220;MontainCar&#8221; simulation in which a car with a weak motor is located in a valley and needs to learn to accelerate back and forth to reach a mountain top. <\/p>\n\n\n<div class=\"wp-block-image\">\n<figure class=\"aligncenter size-large\"><img loading=\"lazy\" decoding=\"async\" width=\"1024\" height=\"602\" src=\"https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/GridVersusContinuousSpace-1024x602.jpeg\" alt=\"\" class=\"wp-image-978\" srcset=\"https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/GridVersusContinuousSpace-1024x602.jpeg 1024w, https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/GridVersusContinuousSpace-300x176.jpeg 300w, https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/GridVersusContinuousSpace-768x451.jpeg 768w, https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/GridVersusContinuousSpace-1536x903.jpeg 1536w, https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/GridVersusContinuousSpace-788x463.jpeg 788w, https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/GridVersusContinuousSpace.jpeg 1671w\" sizes=\"auto, (max-width: 1024px) 100vw, 1024px\" \/><figcaption> Discrete State and Action Space versus Continuous State and Action Space. \u00a9OpenAI Gym<\/figcaption><\/figure>\n<\/div>\n\n\n<p class=\"wp-block-paragraph\">If the policy is represented by a parameterized function such as a neural network, then training consists of optimising these parameters to maximise the expected cumulative reward.  The objective of maximising the expected cumulative reward can then formulated as follows:<\/p>\n\n\n<div class=\"wp-block-image\">\n<figure class=\"aligncenter size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"744\" height=\"98\" src=\"https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/PolicyGradientObjectiveFunction.png\" alt=\"\" class=\"wp-image-943\" srcset=\"https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/PolicyGradientObjectiveFunction.png 744w, https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/PolicyGradientObjectiveFunction-300x40.png 300w\" sizes=\"auto, (max-width: 744px) 100vw, 744px\" \/><figcaption>Objective Function J. \u03c0\u03b8 stands for the parameterised policy. Tau stands for a trajectory which is the same as an episode. \u00a9Janis Klaise<\/figcaption><\/figure>\n<\/div>\n\n\n<p class=\"wp-block-paragraph\">Policy Gradient methods apply gradient ascent to maximise the objective. The gradient of the objective with respect to the parameters of the policy can be written as follows:<\/p>\n\n\n<div class=\"wp-block-image\">\n<figure class=\"aligncenter size-full is-resized\"><img loading=\"lazy\" decoding=\"async\" src=\"https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/PolicyGradientEquation.png\" alt=\"\" class=\"wp-image-946\" width=\"347\" height=\"76\" srcset=\"https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/PolicyGradientEquation.png 400w, https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/PolicyGradientEquation-300x66.png 300w\" sizes=\"auto, (max-width: 347px) 100vw, 347px\" \/><figcaption>The gradient \u2207 of the Objective Function J. \u00a9Janis Klaise<\/figcaption><\/figure>\n<\/div>\n\n\n<p class=\"wp-block-paragraph\">Updating the policy parameters during training then corresponds to:<\/p>\n\n\n<div class=\"wp-block-image\">\n<figure class=\"aligncenter size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"380\" height=\"226\" src=\"https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/PolicyParameterUpdateEquation.png\" alt=\"\" class=\"wp-image-949\" srcset=\"https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/PolicyParameterUpdateEquation.png 380w, https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/PolicyParameterUpdateEquation-300x178.png 300w\" sizes=\"auto, (max-width: 380px) 100vw, 380px\" \/><figcaption>Updating Policy Parameters using Gradient Ascent. \u00a9Janis Klaise<\/figcaption><\/figure>\n<\/div>\n\n\n<p class=\"wp-block-paragraph\">The gradient can be rewritten using action probabilities of the parameterised policy function and parameterised action value function:<\/p>\n\n\n<div class=\"wp-block-image\">\n<figure class=\"aligncenter size-large\"><img loading=\"lazy\" decoding=\"async\" width=\"1024\" height=\"172\" src=\"https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/PolicyGradientEquation_2-1024x172.png\" alt=\"\" class=\"wp-image-952\" srcset=\"https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/PolicyGradientEquation_2-1024x172.png 1024w, https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/PolicyGradientEquation_2-300x50.png 300w, https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/PolicyGradientEquation_2-768x129.png 768w, https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/PolicyGradientEquation_2-788x132.png 788w, https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/PolicyGradientEquation_2.png 1150w\" sizes=\"auto, (max-width: 1024px) 100vw, 1024px\" \/><figcaption>The gradient \u2207 of the Objective Function J. \u00a9Chris Yoon<\/figcaption><\/figure>\n<\/div>\n\n\n<h2 class=\"wp-block-heading\">Actor-Critique Methods<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">Actor-Critique Methods employ both a policy function and a value function, both of them typically implemented as neural networks. The policy function is called the actor, the value function the critique. The actor decides which action to take, and the critic informs the actor how good the action taken was and how it should be adjusted. The term critique is reminiscent of the critique in <a href=\"https:\/\/wp.coventry.domains\/e2edu\/generative-adversarial-network\/\" target=\"_blank\" rel=\"noreferrer noopener\">Generative Adversarial Networks <\/a>but its role is very different in Actor-Critique methods. In Actor-Critique methods, the actor and critique cooperate with each other and the both improve over time. <\/p>\n\n\n<div class=\"wp-block-image\">\n<figure class=\"aligncenter size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"353\" height=\"350\" src=\"https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/ActorCritique.png\" alt=\"\" class=\"wp-image-957\" srcset=\"https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/ActorCritique.png 353w, https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/ActorCritique-300x297.png 300w, https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/ActorCritique-150x150.png 150w\" sizes=\"auto, (max-width: 353px) 100vw, 353px\" \/><figcaption>Actor-Critic Architecture. \u00a9Richard S. Sutton and Andrew G. Barto<\/figcaption><\/figure>\n<\/div>\n\n\n<p class=\"wp-block-paragraph\">During training, instead of dealing with the expectation over the cumulative rewards of all possible trajectories, the cumulative rewards are only sampled from some trajectories. Furthermore, the action value function is replaced by an advantage function. An advantage function returns how much better (or worse) the reward for the currently chosen action is as compared with the current value function. For this reason, the advantage function is identical with the TD-Error.<\/p>\n\n\n<div class=\"wp-block-image\">\n<figure class=\"aligncenter size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"399\" height=\"40\" src=\"https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/AdvantageFunction.png\" alt=\"\" class=\"wp-image-966\" srcset=\"https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/AdvantageFunction.png 399w, https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/AdvantageFunction-300x30.png 300w\" sizes=\"auto, (max-width: 399px) 100vw, 399px\" \/><figcaption>Advantage Function of Actor-Critic Algorithm. \u00a9Dhanoop Karunakaran<\/figcaption><\/figure>\n<\/div>\n\n<div class=\"wp-block-image\">\n<figure class=\"aligncenter size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"362\" height=\"74\" src=\"https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/PolicyGradientWithAdvantageFunction.png\" alt=\"\" class=\"wp-image-967\" srcset=\"https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/PolicyGradientWithAdvantageFunction.png 362w, https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/PolicyGradientWithAdvantageFunction-300x61.png 300w\" sizes=\"auto, (max-width: 362px) 100vw, 362px\" \/><figcaption>Policy-Gradient of the Actor. \u00a9Dhanoop Karunakaran<\/figcaption><\/figure>\n<\/div>\n\n\n<p class=\"wp-block-paragraph\">The pseudocode for training an Actor-Critique RL system is as follows: <\/p>\n\n\n\n<figure class=\"wp-block-image size-large\"><img loading=\"lazy\" decoding=\"async\" width=\"1024\" height=\"440\" src=\"https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/ActorCritiquePseudoCode-1024x440.png\" alt=\"\" class=\"wp-image-969\" srcset=\"https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/ActorCritiquePseudoCode-1024x440.png 1024w, https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/ActorCritiquePseudoCode-300x129.png 300w, https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/ActorCritiquePseudoCode-768x330.png 768w, https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/ActorCritiquePseudoCode-788x339.png 788w, https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/ActorCritiquePseudoCode.png 1400w\" sizes=\"auto, (max-width: 1024px) 100vw, 1024px\" \/><figcaption>Pseudocode for Actor-Critique RL System. \u00a9Chris Yoon<\/figcaption><\/figure>\n","protected":false},"excerpt":{"rendered":"<p>Reinforcement Learning (RL) is a Machine Learning approach that is inspired by how animals or young children learn from their mistakes (or successes). This approach combines trial and error actions with a reward. A typical natural example of this type of learning would be a young child who touches the door of a hot oven [&hellip;]<\/p>\n","protected":false},"author":2154,"featured_media":0,"parent":0,"menu_order":0,"comment_status":"closed","ping_status":"closed","template":"","meta":{"_monsterinsights_skip_tracking":false,"_monsterinsights_sitenote_active":false,"_monsterinsights_sitenote_note":"","_monsterinsights_sitenote_category":0,"_coblocks_attr":"","_coblocks_dimensions":"","_coblocks_responsive_height":"","_coblocks_accordion_ie_support":"","footnotes":""},"class_list":["post-123","page","type-page","status-publish","hentry"],"_links":{"self":[{"href":"https:\/\/wp.coventry.domains\/e2edu\/wp-json\/wp\/v2\/pages\/123","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/wp.coventry.domains\/e2edu\/wp-json\/wp\/v2\/pages"}],"about":[{"href":"https:\/\/wp.coventry.domains\/e2edu\/wp-json\/wp\/v2\/types\/page"}],"author":[{"embeddable":true,"href":"https:\/\/wp.coventry.domains\/e2edu\/wp-json\/wp\/v2\/users\/2154"}],"replies":[{"embeddable":true,"href":"https:\/\/wp.coventry.domains\/e2edu\/wp-json\/wp\/v2\/comments?post=123"}],"version-history":[{"count":112,"href":"https:\/\/wp.coventry.domains\/e2edu\/wp-json\/wp\/v2\/pages\/123\/revisions"}],"predecessor-version":[{"id":2909,"href":"https:\/\/wp.coventry.domains\/e2edu\/wp-json\/wp\/v2\/pages\/123\/revisions\/2909"}],"wp:attachment":[{"href":"https:\/\/wp.coventry.domains\/e2edu\/wp-json\/wp\/v2\/media?parent=123"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}