Oscillation of Episode Q0 during DDPG training

조회 수: 7 (최근 30일)

이전 댓글 표시

Heesu Kim 2021년 4월 6일

0
링크

이 질문에 대한 바로 가기 링크

https://kr.mathworks.com/matlabcentral/answers/794607-oscillation-of-episode-q0-during-ddpg-training

댓글: Heesu Kim 2021년 4월 6일

How do I interpret this kind of Episode Q0 oscillation?

The oscillation shows a pattern like up and down and the range also increases quite regularly.

According to other docs, they're saying the Q0 is supposed to approach actual discounted future reward as long as the critic network is designed properly.

Is this kind of Q0 oscillation just evidence that my critic network is not well-designed?

Is there any solution to work it out?

I'm not sure this question is acceptable to this community because I think it's more or less a theoretical issue.

댓글 수: 1
이전 댓글 -1개 표시이전 댓글 -1개 숨기기

Heesu Kim 2021년 4월 6일

As a side note, I'm using DDPG + LSTM model that RL toolbox provides

댓글을 달려면 로그인하십시오.

이 질문에 답변하려면 로그인하십시오.

답변 (0개)

이 질문에 답변하려면 로그인하십시오.

카테고리

AI, Data Science, and Statistics Deep Learning Toolbox Applications Autonomous and Control Systems Reinforcement Learning

Help Center 및 File Exchange에서 Reinforcement Learning에 대해 자세히 알아보기

제품

Reinforcement Learning Toolbox

릴리스

R2021a

Community Treasure Hunt

Find the treasures in MATLAB Central and discover how the community can help you!

Start Hunting!

Translated by