Alex Rivera | Logout

How to access directly and efficiently on very large text file?

Asked 2012-04-30T08:40:37.653
20

I have a very large text files (+10GB) which i want to read for some data mining technics. To do that, i use parallel technics with MPI so many processes can access together to the same file.
In fact, i want that each process read N number of lines. Since the file is not structured (same number of fields but each field can contain different number of characters), i'm in the obligation to parse the file and that is not parallel and it takes a lot of time. Is there any way to access directly to a specific number of line withount parsing and counting the lines? Thank you for you help.

Edit
Report

1 Answer

4

No there isn't: until you don't read through your unknown data nobody will know how many new line characters there are. This problem complexity is O(n) thus meaning that at least once you'll have to read the whole file. Then you might want to build an index table where you record where there are new line characters in your file: this can be used by all process and with fseek you can speed up dramatically further access.

answered 2012-04-30T08:47:33.377

Your Answer