Alex Rivera | Logout

What is an efficient way to replace list of strings with another list in Unix file?

Asked 2011-08-25T22:47:57.420
16

Suppose I have two lists of strings (list A and list B) with the exact same number of entries, N, in each list, and I want to replace all occurrences of the the nth element of A with the nth element of B in a file in Unix (ideally using Bash scripting).

What's the most efficient way to do this?

An inefficient way would be to make N calls to "sed s/stringA/stringB/g".

Edit
Report

2 Answers

12

This will do it in one pass. It reads listA and listB into awk arrays, then for each line of the linput, it examines each word and if the word is found in listA, the word is replaced by the corresponding word in listB.

awk '
    FILENAME == ARGV[1] { listA[$1] = FNR; next }
    FILENAME == ARGV[2] { listB[FNR] = $1; next }
    {
        for (i = 1; i <= NF; i++) {
            if ($i in listA) {
                $i = listB[listA[$i]]
            }
        }
        print
    }
' listA listB filename > filename.new
mv filename.new filename

I'm assuming the strings in listA do not contain whitespace (awk's default field separator)

answered 2011-08-26T00:46:22.230
6

I needed to do something similar, and I wound up generating sed commands based on a map file:

$ cat file.map
abc => 123
def => 456
ghi => 789

$ cat stuff.txt
abc jdy kdt
kdb def gbk
qng pbf ghi
non non non
try one abc

$ sed `cat file.map | awk '{print "-e s/"$1"/"$3"/"}'`<<<"`cat stuff.txt`"
123 jdy kdt
kdb 456 gbk
qng pbf 789
non non non
try one 123

Make sure your shell supports as many parameters to sed as you have in your map.

answered 2012-12-05T18:59:35.990

Your Answer