Alex Rivera | Logout

How to make this sed script faster?

Asked 2009-12-01T19:13:47.777
11

I have inherited this sed script snippet that attempts to remove certain empty spaces:

s/[\s\t]*|/|/g
s/|[\s\t]*/|/g
s/[\s] *$//g
s/^|/null|/g

that operates on a file that is around 1Gb large. This script runs for 2 hours on our unix server. Any ideas how to speed it up?

Notes that the \s stands for a space and \t stands for a tab, the actual script uses the actual space and tab and not those symbols

The input file is a pipe delimited file and is located locally not on the network. The 4 lines are in a file executed with sed -f

Edit
Report

2 Answers

4

My testing indicated that sed can become CPU bound pretty easily on something like this. If you have a multi-core machine you can try spawning off multiple sed processes with a script that looks something like this:

#!/bin/sh
INFILE=data.txt
OUTFILE=fixed.txt
SEDSCRIPT=script.sed
SPLITLIMIT=`wc -l $INFILE | awk '{print $1 / 20}'`

split -d -l $SPLITLIMT $INFILE x_

for chunk in x_??
do
  sed -f $SEDSCRIPT $chunk > $chunk.out &
done

wait 

cat x_??.out >> output.txt

rm -f x_??
rm -f x_??.out
answered 2009-12-02T01:24:56.407
3

Try changing the first two lines to:

s/[ \t]*|[ \t]*/|/g
answered 2009-12-01T20:21:15.967

Your Answer