KnowledgeHub
Questions
Tags
Users
Search
Alex Rivera
|
Logout
Edit Question
Title
Body
I have a file (size = ~1.9 GB) which contains ~220,000,000 (~220 million) words / strings. They have duplication, almost 1 duplicate word every 100 words. In my second program, I want to read the file. I am successful to read the file by lines using BufferedReader. Now to remove duplicates, we can use Set (and it's implementations), but Set has problems, as described following in 3 different scenarios: With default JVM size, Set can contain up to 0.7-0.8 million words, and then OutOfMemoryError. With 512M JVM size, Set can contain up to 5-6 million words, and then OOM error. With 1024M JVM size, Set can contain up to 12-13 million words, and then OOM error. Here after 10 million records addition into Set, operations become extremely slow. For example, addition of next ~4000 records, it took 60 seconds. I have restrictions that I can't increase the JVM size further, and I want to remove duplicate words from the file. Please let me know if you have any idea about any other ways/approaches to remove duplicate words using Java from such a gigantic file. Many Thanks :) Addition of info to question: My words are basically alpha-numeric and they are IDs which are unique in our system. Hence they are not plain English words.
Tags (comma-separated)
Save Edits
Cancel