The challenge:
Build an ASCII chart of the most commonly used words in a given text.
The rules:
- Only accept
a-zandA-Z(alphabetic characters) as part of a word. - Ignore casing (
She==shefor our purpose). - Ignore the following words (quite arbitary, I know):
the, and, of, to, a, i, it, in, or, is Clarification: considering
don't: this would be taken as 2 different 'words' in the rangesa-zandA-Z: (donandt).Optionally (it's too late to be formally changing the specifications now) you may choose to drop all single-letter 'words' (this could potentially make for a shortening of the ignore list too).
Parse a given text (read a file specified via command line arguments or piped in; presume us-ascii) and build us a word frequency chart with the following characteristics:
- Display the chart (also see the example below) for the 22 most common words (ordered by descending frequency).
- The bar
widthrepresents the number of occurences (frequency) of the word (proportionally). Append one space and print the word. - Make sure these bars (plus space-word-space) always fit:
bar+[space]+word+[space]should be always <=80characters (make sure you account for possible differing bar and word lengths: e.g.: the second most common word could be a lot longer then the first while not differing so much in frequency). Maximize bar width within these constraints and scale the bars appropriately (according to the frequencies they represent).
An example:
The text for the example <