How to import big data CSV files

조회 수: 3 (최근 30일)

Abhishek Singh 2019년 5월 25일

0
링크

이 질문에 대한 바로 가기 링크

https://kr.mathworks.com/matlabcentral/answers/463938-how-to-import-big-data-csv-files

댓글: Abhishek Singh 2019년 5월 30일

Hi,

Similar questions have been already asked but I wanted to know if there is an alternative to importing. I have a csv file with 3.5 million rows and 56 columns. At present while importing, I have to select the range I need. For instance I am selecting only 36 columns and barring few rows almost all. I tried textscan() but unable to achieve what I want which is without importing my code should select those rows and columns needed and also my main aim is to save time of importing and since textscan() is pretty fast I was using that. My column values are numerics.

Please let me know if I need to update with something here. I have a 8GB RAM so I think it should be easily feasible.

Thanks in advance!text

댓글 수: 0
이전 댓글 -2개 표시이전 댓글 -2개 숨기기

댓글을 달려면 로그인하십시오.

이 질문에 답변하려면 로그인하십시오.

채택된 답변

Walter Roberson 2019년 5월 25일

1
링크

이 답변에 대한 바로 가기 링크

https://kr.mathworks.com/matlabcentral/answers/463938-how-to-import-big-data-csv-files#answer_376552

편집: per isakson 2019년 5월 25일

MATLAB Online에서 열기

These days, often the most convenient way is to use detectImportOptions, and set the SelectedVariableNames property of that to choose particular columns, and the readtable() -- or as of R2019a, readmatrix() if you are pure numeric.

textscan() is also a possibility. The easiest way to use that might be something like,

numcol = 56;
wanted_cols = [5 17:33 44:47];   %adjust to suit
fmt_cell = repmat('%*s', 1, numcol);
fmt_cell(wanted_cols) = {'%f'};
fmt = [fmt_cell{:}];
data = cell2mat( textscan(fid, fmt, 'Delimiter', ',', 'CollectOutput', 1) );

댓글 수: 17
이전 댓글 15개 표시이전 댓글 15개 숨기기

Abhishek Singh 2019년 5월 30일

MATLAB Online에서 열기

I will show you a little bit of the data here. The first 4 rows below are the rows 34:37 so here I need to reed the column 1 to know that they are characters and then start from 37 since they are the numbers and also only include from column 2:37. Also the next 4 rows are last 4 columns from the data, again I have to read the first column and delete rows which have their first entity as character hence I was deleting last two rows also. I could do all these with the code you have helped me with but for that I need to be aware with the structure of the file. But I somehow want to automate it so that even if I do not know the structure I am just running and getting the numbers in workspace.

 ChannelAttributes  0  0  0 
 
  ChannelAttributesUseable  1     
 
  0  -0.80882  0  -8.67647 
 
  1  -0.88235  0  -8.82353 
  
  61094  -18.0882  0  -24.6324 
 
  61095  -18.0882  0  -24.6324 
 
  ChannelPhysicalInputPort  0  1  2 
 
  ChannelPhysicalInputHardware  1  1  1

And where should I define that I do not need the first 35 rows.

I tried this and it shows me this error "Index exceeds the number of array elements" at the line

opt.SelectedVariableNames = opt.SelectedVariableNames(wanted_cols);

Walter Roberson 2019년 5월 30일

Shrug. You can get a wrong and unusable input quickly, or you can get a correct and usable input more slowly.

The extract you showed is not a csv file. csv files are, by definition, Comma Separated Values, never whitespace separated. RFC4180 specifically says that spaces are to be considered part of a field, not a separator.

It is not difficult to extract from a file only the lines that start with numbers. It is not all that much more difficult to divide a file into blocks of numbers, so that non-number marks the end of a block and the blocks are to be imported separately. But as soon as you start wanting to extract particular information from the non-number sections and associating it with a block of numbers, the task becomes more complicated.

Abhishek Singh 2019년 5월 30일

I am sorry. Importing in matlab showed it in columns and since it was a csv format file I thought it must be like that. But I do understand your point. I think I will stick to the first plan of using textscan() since it was also efficient.

댓글을 달려면 로그인하십시오.

추가 답변 (0개)

이 질문에 답변하려면 로그인하십시오.

카테고리

MATLAB Data Import and Analysis Data Import and Export Standard File Formats Text Files

Help Center 및 File Exchange에서 Text Files에 대해 자세히 알아보기

Community Treasure Hunt

Find the treasures in MATLAB Central and discover how the community can help you!

Start Hunting!

Translated by

How to import big data CSV files

댓글 수: 0
이전 댓글 -2개 표시이전 댓글 -2개 숨기기

채택된 답변

댓글 수: 17
이전 댓글 15개 표시이전 댓글 15개 숨기기

추가 답변 (0개)

참고 항목

카테고리

태그

Community Treasure Hunt

How to import big data CSV files

댓글 수: 0 이전 댓글 -2개 표시이전 댓글 -2개 숨기기

채택된 답변

댓글 수: 17 이전 댓글 15개 표시이전 댓글 15개 숨기기

추가 답변 (0개)

참고 항목

카테고리

태그

Community Treasure Hunt

댓글 수: 0
이전 댓글 -2개 표시이전 댓글 -2개 숨기기

댓글 수: 17
이전 댓글 15개 표시이전 댓글 15개 숨기기