KnowledgeHub
Questions
Tags
Users
Search
Alex Rivera
|
Logout
Edit Question
Title
Body
I'm trying to read a large text corpus into memory with Java. At some point it hits a wall and just garbage collects interminably. I'd like to know if anyone has experience beating Java's GC into submission with large data sets. I'm reading an 8 GB file of English text, in UTF-8, with one sentence to a line. I want to split() each line on whitespace and store the resulting String arrays in an ArrayList<String[]> for further processing. Here's a simplified program that exhibits the problem: /** Load whitespace-delimited tokens from stdin into memory. */ public class LoadTokens { private static final int INITIAL_SENTENCES = 66000000; public static void main(String[] args) throws IOException { List<String[]> sentences = new ArrayList<String[]>(INITIAL_SENTENCES); BufferedReader stdin = new BufferedReader(new InputStreamReader(System.in)); long numTokens = 0; String line; while ((line = stdin.readLine()) != null) { String[] sentence = line.split("\\s+"); if (sentence.length > 0) { sentences.add(sentence); numTokens += sentence.length; } } System.out.println("Read " + sentences.size() + " sentences, " + numTokens + " tokens."); } } Seems pretty cut-and-dried, right? You'll notice I even pre-size my ArrayList ; I have a little less than 66 million sentences and 1.3 billion tokens. Now if you whip out your Java object sizes reference and your pencil, you'll find that should require about: 66e6 String[] references @ 8 bytes ea = 0.5 GB 66e6 String[] objects @ 32 bytes ea = 2 GB 66e6 char[] objects @ 32 bytes ea = 2 GB 1.3e9 String referen
Tags (comma-separated)
Save Edits
Cancel