How do I define a continuous reward function for RL environment?

I am trying to follow the double integrator example for giving a continuous reward function. When I used the custom template, and defined the reward using the QR cost function, I get an error stating that the reward should be a scalar value. Where can I find the property of reward and change it to accept vector values?

댓글 수: 3

Not sure why you want the reward to be scalar. Typically, rewards are treated as cost functions - they output a scalar value. If you have more than one states, you can turn it into a scalar using e.g. an l2 norm for example/some distance metric.
Yes I did that, thank you, Just to confirm the output of the cost function will always be a scalar value, right? So in the double integrator continuous example there are two states but the output reward at each step is a scalar value, right?

댓글을 달려면 로그인하십시오.

 채택된 답변

Priysha LNU
Priysha LNU 2020년 10월 8일

0 개 추천

Here is an excerpt from the documentation :
To guide the learning process, reinforcement learning uses a scalar reward signal generated from the environment.
For detailed information on defining reward signals, discrete and continous rewards, please refer to this documentation link.

추가 답변 (0개)

카테고리

도움말 센터 및 File Exchange에서 Environments에 대해 자세히 알아보기

제품

릴리스

R2020a

Community Treasure Hunt

Find the treasures in MATLAB Central and discover how the community can help you!

Start Hunting!

Translated by