Alex Rivera | Logout

Remove all text before colon

Asked 2012-09-06T10:17:50.597
77

I have a file containing a certain number of lines. Each line looks like this:

TF_list_to_test10004/Nus_k0.345_t0.1_e0.1.adj:PKMYT1

I would like to remove all before ":" character in order to retain only PKMYT1 that is a gene name. Since I'm not an expert in regex scripting can anyone help me to do this using Unix (sed or awk) or in R?

Edit
Report

3 Answers

10

Using sed:

sed 's/.*://' < your_input_file > output_file

This will replace anything followed by a colon with nothing, so it'll remove everything up to and including the last colon on each line (because * is greedy by default).

As per Josh O'Brien's comment, if you wanted to only replace up to and including the first colon, do this:

sed "s/[^:]*://"

That will match anything that isn't a colon, followed by one colon, and replace with nothing.

Note that for both of these patterns they'll stop on the first match on each line. If you want to make a replace happen for every match on a line, add the 'g' (global) option to the end of the command.

Also note that on linux (but not on OSX) you can edit a file in-place with -i eg:

sed -i 's/.*://' your_file
answered 2012-09-06T10:26:46.400
6

You can use awk like this:

awk -F: '{print $2}' /your/file
answered 2012-09-06T10:31:50.777
2

If you have GNU coreutils available use cut:

cut -d: -f2 infile
answered 2012-09-06T12:49:01.030

Your Answer