Extracting data using regular expression

Question

Shuvashish Roy on 20 May 2021

0
Link

Direct link to this question

https://se.mathworks.com/matlabcentral/answers/835638-extracting-data-using-regular-expression

Commented: Shuvashish Roy on 21 May 2021

AR_20base_201214_adh.txt

Hi,

I have the attached text file. I want to extract all the columns starting from line 1472(if used notepad) named "Physics", "Time", "dt", "Progress", "Nonlinear Iteration" "Linear Iterations"...."Nodes After Adaption". I don't know how to specify the header names so that only the numeric values after that headers are extracted in a dataframe or matrix format. Thanks a lot for your help.

Input file format:

Unnecessary lines with text

Unnevessary lines with text

................................

many unnecessay lines............

adh_run_func :: tfinal = 12513600.000000

Physics Time dt Progress Nonlinear Iteration Linear Iteration Max Resid Norm ... Nodes After Adaption

HYD_1 11908800 5 0 1 ........ ...65926

HYD_1 11908800 5 0 2 ...... ...65926

............................................................................................. ................................

100% COMPLETE

output file format:

Physics Time dt Progress Nonlinear Iteration Linear Iteration Max Resid Norm ... Nodes After Adaption

HYD_1 11908800 5 0 1 ........ ...65926

HYD_1 11908800 5 0 2 ...... ...65926

............................................................................................. ................................

0 Comments
Show -2 older commentsHide -2 older comments

Sign in to comment.

Sign in to answer this question.

Answer 1

per isakson on 21 May 2021

1
Link

Direct link to this answer

https://se.mathworks.com/matlabcentral/answers/835638-extracting-data-using-regular-expression#answer_704978

Edited: per isakson on 21 May 2021

Open in MATLAB Online

AR_20base_201214_adh.txt

"all the columns [...] named "Physics", "Time", "dt", "Progress", "Nonlinear Iteration" "Linear Iterations"...."Nodes After Adaption" " I understand that as all the columns, none excluded.

There is a choice. Shall we use readtable() or textscan()? I don't think readtable() can handle this file without relying on the critical line numbers, which I hessitate to do. It is however possible to determine the line numbers needed in a separate step and then use readtable(). textscan() is able to parse a 1D character array, which readtabe() is not. Only TMW knows why.

I choose textscan().

%%  Read file
chr = fileread('AR_20base_201214_adh.txt');
%%  Remove meta data
%   Using 'adh_run_func :: tfinal' feels more robust than using the line number 
pos = regexp( chr, '^adh_run_func :: tfinal', 'once', 'lineanchors' );
chr(1:pos-1) = [];  % remove until the first line that begins with 'adh_run_func :: tfinal' 
%%  Remove the summary lines at the end
pos = regexp( chr, '^\d+[\% ]+COMPLETE', 'once', 'lineanchors' );
chr(pos:end) = [];
%%  Get the column headers
txt = regexp( chr, '^Physics.+?$', 'match', 'once', 'lineanchors' );
column_headers = strsplit( txt, '\t' );
%%
cac = textscan( chr, ['%s',repmat('%f',1,numel(column_headers)-1)]  ...
            ,   'Headerlines'   , 2     ...     two remains after meta-data is removed
            ,   'Delimiter'     , '\t'  ...
            ,   'Whitespace'    , ' %'  ...     ignore the %-sign in Progress
            ,   'CollectOutput' , true  );
Physics = cac{1};    
matrix  = cac{2};
whos Physics matrix column_headers
  Name                    Size              Bytes  Class     Attributes

  Physics             13487x1             1537454  cell                
  column_headers          1x17               2026  cell                
  matrix              13487x16            1726336  double              

1 Comment
Show -1 older commentsHide -1 older comments

Shuvashish Roy on 21 May 2021

Per Isakon,

I got your answer.It worked! You are awesome. Thanks a lot both you and Stephen for your valueable times.

Sign in to comment.

Extracting data using regular expression

0 Comments
Show -2 older commentsHide -2 older comments

Accepted Answer

1 Comment
Show -1 older commentsHide -1 older comments

More Answers (0)

See Also

Categories

Tags

Community Treasure Hunt

Extracting data using regular expression

0 Comments Show -2 older commentsHide -2 older comments

Accepted Answer

1 Comment Show -1 older commentsHide -1 older comments

More Answers (0)

See Also

Categories

Tags

Community Treasure Hunt

0 Comments
Show -2 older commentsHide -2 older comments

1 Comment
Show -1 older commentsHide -1 older comments