The gradient of mini batches

Question

MAHSA YOUSEFI 2020년 11월 23일

0
링크

이 질문에 대한 바로 가기 링크

https://kr.mathworks.com/matlabcentral/answers/658543-the-gradient-of-mini-batches

댓글: Mahesh Taparia 2020년 12월 21일

채택된 답변: Mahesh Taparia

MATLAB Online에서 열기

Hi there.

I need your confimation or rejection for this question...

In following code, if the minibatch size is h,

[grad,loss] = dlfeval(@modelGradients,dlnet,dlX_miniBatch,Y_miniBatch);

the grad is the average of gradients of loss over h samples? Does it calculate dradients automatically and at the end with:

grad = 1/h * sum_i=1:h (\nabla loss(y_i,yHat_i)) ??

Following this question, for computing the total loss and geadient (for a full batch), does we should take avarage of losses and averages of gradients (averaging with the number of batches, say 1000 batches each with h size)??

댓글 수: 0
이전 댓글 -2개 표시이전 댓글 -2개 숨기기

댓글을 달려면 로그인하십시오.

이 질문에 답변하려면 로그인하십시오.

Answer 1

Mahesh Taparia 2020년 12월 14일

0
링크

이 답변에 대한 바로 가기 링크

https://kr.mathworks.com/matlabcentral/answers/658543-the-gradient-of-mini-batches#answer_575280

Hi

The function dlfeval evaluate the custom deep learning models. The loss are calculated based on what has been defined in modelGradients function. So if you are calculating the average loss in this function, then it will return the averaged one. For example, consider this modelGradient function, it is calculating the average cross entropy loss, so it will return the average loss. The gradients are calculated with respect to the loss function defined in for the network.

댓글 수: 2
없음 표시없음 숨기기

MAHSA YOUSEFI 2020년 12월 19일

MATLAB Online에서 열기

In the example you mentioned, there is a mistake.

function [gradients, loss] = modelGradients(parameters, dlX, T)
    % Forward data through the model function.
    dlY = model(parameters,dlX);
    % Compute loss.
    loss = crossentropy(dlX,T);
    % Compute gradients.
    gradients = dlgradient(loss,parameters);
end

dlY must be feed to crossentropy!

Mahesh Taparia 2020년 12월 21일

Yeah, crossentropy loss will be calculated between dlY and T. The documentation page will be updated.

댓글을 달려면 로그인하십시오.

The gradient of mini batches

댓글 수: 0
이전 댓글 -2개 표시이전 댓글 -2개 숨기기

채택된 답변

댓글 수: 2
없음 표시없음 숨기기

추가 답변 (0개)

참고 항목

카테고리

태그

제품

Community Treasure Hunt

The gradient of mini batches

댓글 수: 0 이전 댓글 -2개 표시이전 댓글 -2개 숨기기

채택된 답변

댓글 수: 2 없음 표시없음 숨기기

추가 답변 (0개)

참고 항목

카테고리

태그

제품

Community Treasure Hunt

댓글 수: 0
이전 댓글 -2개 표시이전 댓글 -2개 숨기기

댓글 수: 2
없음 표시없음 숨기기