selfAttentionLayer can't process sequence-to-label problem?

Question

xingxingcui 2024년 1월 5일

1
링크

이 질문에 대한 바로 가기 링크

https://kr.mathworks.com/matlabcentral/answers/2066601-selfattentionlayer-can-t-process-sequence-to-label-problem

편집: xingxingcui 2024년 4월 27일

selfAttentionLayer why can't handle the following simple sequence classification problem, already through the flattenLayer into one-dimensional data, on the contrary, lstm specify "outputMode" as "last" will pass.

% Here use simple data, for demonstration purposes only
XTrain = rand(3,200,1000); % dims "CTB"
TTrain = categorical(randi(4,1000,1));
% define my layers
numClasses = numel(categories(TTrain));
layers = [inputLayer(size(XTrain),"CTB");
    flattenLayer;
    selfAttentionLayer(6,48);
    % lstmLayer(20,OutputMode="last"); % use lstmLayer is ok!
    layerNormalizationLayer;
    fullyConnectedLayer(numClasses);
    softmaxLayer];
net = dlnetwork(layers);
% train network
lossFcn = "crossentropy";
options = trainingOptions("adam", ...
    MaxEpochs=1, ...
    InitialLearnRate=0.01,...
    Shuffle="every-epoch", ...
    GradientThreshold=1, ...
    Verbose=true);
netTrained = trainnet(XTrain,TTrain,net,lossFcn,options);
Error using trainnet
Number of observations in predictors (1000) and targets (1) must match. Check that the data and network are consistent.

댓글 수: 0
이전 댓글 -2개 표시이전 댓글 -2개 숨기기

댓글을 달려면 로그인하십시오.

이 질문에 답변하려면 로그인하십시오.

Answer 1

xingxingcui 2024년 1월 7일

0
링크

이 답변에 대한 바로 가기 링크

https://kr.mathworks.com/matlabcentral/answers/2066601-selfattentionlayer-can-t-process-sequence-to-label-problem#answer_1384691

편집: xingxingcui 2024년 4월 27일

MATLAB Online에서 열기

In terms of the output feature map dimensions, there is a time "T" dimension that has to be eliminated in order to match the output dimensions, which can usually be done by indexing1dLayer. So the layers array is added before the fullyConnectedLayer.

% Here use simple data, for demonstration purposes only
XTrain = rand(3,200,1000); % dims "CTB"
TTrain = categorical(randi(4,1000,1));
% define my layers
numClasses = numel(categories(TTrain));
layers = [inputLayer(size(XTrain),"CTB");
    flattenLayer;
    selfAttentionLayer(6,48);
    % lstmLayer(20,OutputMode="last"); % use lstmLayer is ok!
    layerNormalizationLayer;
    
    indexing1dLayer; % Add this!!!
    fullyConnectedLayer(numClasses);
    softmaxLayer];
net = dlnetwork(layers);
% train network
lossFcn = "crossentropy";
options = trainingOptions("adam", ...
    MaxEpochs=1, ...
    InitialLearnRate=0.01,...
    Shuffle="every-epoch", ...
    GradientThreshold=1, ...
    Verbose=true);
netTrained = trainnet(XTrain,TTrain,net,lossFcn,options);
    Iteration    Epoch    TimeElapsed    LearnRate    TrainingLoss
    _________    _____    ___________    _________    ____________
            1        1       00:00:02         0.01          1.5374
            7        1       00:00:06         0.01          1.5272
Training stopped: Max epochs completed

-------------------------Off-topic interlude-------------------------------

I am currently looking for a job in the field of CV algorithm development, based in Shenzhen, Guangdong, China. I would be very grateful if anyone is willing to offer me a job or make a recommendation. My preliminary resume can be found at: https://cuixing158.github.io/about/ . Thank you!

Email: cuixingxing150@gmail.com

댓글 수: 5
이전 댓글 3개 표시이전 댓글 3개 숨기기

DGM 2024년 3월 5일

Posted as a comment-as-flag by chang gao:

Useful answer.

jingwen 2024년 4월 15일

Your answer helps me! Thank you

댓글을 달려면 로그인하십시오.

selfAttentionLayer can't process sequence-to-label problem?

댓글 수: 0
이전 댓글 -2개 표시이전 댓글 -2개 숨기기

채택된 답변

댓글 수: 5
이전 댓글 3개 표시이전 댓글 3개 숨기기

추가 답변 (0개)

참고 항목

카테고리

태그

제품

릴리스

Community Treasure Hunt

selfAttentionLayer can't process sequence-to-label problem?

댓글 수: 0 이전 댓글 -2개 표시이전 댓글 -2개 숨기기

채택된 답변

댓글 수: 5 이전 댓글 3개 표시이전 댓글 3개 숨기기

추가 답변 (0개)

참고 항목

카테고리

태그

제품

릴리스

Community Treasure Hunt

댓글 수: 0
이전 댓글 -2개 표시이전 댓글 -2개 숨기기

댓글 수: 5
이전 댓글 3개 표시이전 댓글 3개 숨기기