KnowledgeHub
Questions
Tags
Users
Search
Alex Rivera
|
Logout
Edit Question
Title
Body
I do a lot of HTML parsing in my line of work. Up until now, I was using the HtmlUnit headless browser for parsing and browser automation. Now, I want to separate both tasks. I want to use a light HTML parser because it takes much time in HTMLUnit to first load a page, then get the source, and then parse it. I want to know which HTML parser can parse HTML efficiently. I need Speed Ease to locate any HtmlElement by its "id" or "name" or "tag type". It would be ok for me if it doesn't clean the dirty HTML code. I don't need to clean any HTML source. I just need the easiest way to move across HtmlElements and harvest data from them.
Tags (comma-separated)
Save Edits
Cancel