KnowledgeHub
Questions
Tags
Users
Search
Alex Rivera
|
Logout
Edit Question
Title
Body
I have a task of creating a script which takes a huge text file as an input. It then needs to find all words and the number of occurrences and create a new file with each line displaying a unique word and its occurrence. As an example take a file with this content: Lorem ipsum dolor sit amet, consectetur adipisicing elit, sed do eiusmod tempor incididunt ut labore et dolore magna aliqua. Ut enim ad minim veniam, quis nostrud exercitation ullamco laboris nisi ut aliquip ex ea commodo consequat. Duis aute irure dolor in reprehenderit in voluptate velit esse cillum dolore eu fugiat nulla pariatur. Excepteur sint occaecat cupidatat non proident, sunt in culpa qui officia deserunt mollit anim id est laborum. I need to create a file which looks like this: 1 AD 1 ADIPISICING 1 ALIQUA ... 1 ALIQUIP 1 DO 2 DOLOR 2 DOLORE ... For this I wrote a script using tr , sort and uniq : #!/bin/sh INPUT=$1 OUTPUT=$2 if [ -a $INPUT ] then tr '[:space:][\-_?!.;\:]' '\n' < $INPUT | tr -d '[:punct:][:special:][:digit:]' | tr '[:lower:]' '[:upper:]' | sort | uniq -c > $OUTPUT fi What this does is split the words by space as the delimiter. If the word contains -_?!.;: I break them into words again. I remove the punctuations, special characters and digits and convert the entire string to uppercase. Once this is done I sort it and pass it through uniq to get it to the format I want. Now I downloaded the bible in txt format and used it as the input. Timing this I got: scripts|$ time ./text-to-word.sh text.txt b ./text-to-word.sh text.txt b 16.17s user 0.09s system 102% cpu 15.934 total I did the same with a Python script: import re from collections import Counter from itertools import chain import sys fi
Tags (comma-separated)
Save Edits
Cancel