Question: For the plot of the total reward as a function of time as in Figure 13.4 (page 594), the minimum and zero crossing are only
For the plot of the total reward as a function of time as in Figure 13.4
(page 594), the minimum and zero crossing are only meaningful statistics when balancing positive and negative rewards is reasonable behavior. Suggest what should replace these statistics when zero reward is not an appropriate definition of reasonable behavior. [Hint: Think about the cases that have only positive reward or only negative reward.]
Step by Step Solution
There are 3 Steps involved in it
Get step-by-step solutions from verified subject matter experts
