KnowledgeHub
Questions
Tags
Users
Search
Alex Rivera
|
Logout
Edit Question
Title
Body
I investigated some time here on StackOverflow to find good algorithms to split strings with multiple delimiters into a vector< string > . I also found some methods: The Boost way: boost::split(vector, string, boost::is_any_of(" \t")); the getline method: std::stringstream ss(string); std::string item; while(std::getline(ss, item, ' ')) { vector.push_back(item); } the tokenize way of Boost: char_separator<char> sep(" \t"); tokenizer<char_separator<char>> tokens(string, sep); BOOST_FOREACH(string t, tokens) { vector.push_back(t); } and the cool STL way: istringstream iss(string); copy(istream_iterator<string>(iss), istream_iterator<string>(), back_inserter<vector<string> >(vector)); and the method of Shadow2531 (see the linked topic). Most of them came from this topic . But they unfortunately don't solve my problem: Boost's split is easy to use but with the big data (about 1.5*10^6 single elements in best cases) and about 10 delimiters I am using it's horrific slow. The getline , STL and Shadow2531's method have the problem that I can only use one single char as delimiter. I need a few more. Boost's tokenize is even more horrific in the aspect of speed. It took 11 seconds with 10 delimiters to split a string into 1.5*10^6 elements. So I don't know what to do: I want to have a really fast string splitting algorithm with multiple delimiters. Is Boost's split the maximum or is there a way to do it faster ?
Tags (comma-separated)
Save Edits
Cancel