Alex Rivera | Logout

Best library to parse HTML with Python 3 and example?

Asked 2010-03-24T02:54:14.433
26

I'm new to Python completely and am using Python 3.1 on Windows (pywin). I need to parse some HTML, to essentially extra values between specific HTML tags and am confused at my array of options, and everything I find is suited for Python 2.x. I've read raves about Beautiful Soup, HTML5Lib and lxml, but I cannot figure out how to install any of these on Windows.

Questions:

  1. What HTML parser do you recommend?
  2. How do I install it? (Be gentle, I'm completely new to Python and remember I'm on Windows)
  3. Do you have a simple example on how to use the recommended library to snag HTML from a specific URL and return the value out of say something like this:

    <div class="foo"><table><tr><td>foo</td></tr></table><a class="link" href='/blahblah'>Link</a></div>

(say we want to return "/blahblah")

Edit
Report

1 Answer

7

If your HTML is well formed, you have many options, such as sax and dom. If it is not well formed you need a fault tolerant parser such as Beautiful soup, element tidy, or lxml's HTML parser. No parser is perfect, when presented with a variety of broken HTML, sometimes I have to try more then one. Lxml and Elementree use a mostly compatible api that is more of a standard than Beautiful soup.

In my opinion, lxml is the best module for working with xml documents, but the ElementTree included with python is still pretty good. In the past I have used Beautiful soup to convert HTML to xml and construct ElementTree for processing the data.

answered 2010-03-24T03:23:11.210

Your Answer