Alex Rivera | Logout

Using python to edit html, but lxml converts nice html entities to strange encoding

Asked 2011-02-02T16:00:12.217
10

I'm trying to use python (with pyquery and lxml) to alter and clean up some html.

Eg. html = "<div><!-- word style><bleep><omgz 1,000 tags><--><p>It&#146;s a spicy meatball!</div>"

The lxml.html.clean function, clean_html(), works well, except that it replaces the nice html entities like

&#146; 

with some unicode string

\xc2\x92

The unicode looks strange in different browsers (firefox and opera using auto encoding, utf8, latin-1, etc), like an empty box. How can I stop lxml converting the entities? How can I get it all in latin-1 encoding? Seems strange that a module built specifically for html would do this.

I can't be sure of which characters are there, so I can't just use

replace("\xc2\x92","&#146;").

I've tried using

clean_html(html).encode('latin-1')

but the unicode persists.

And yes, I'd tell people to stop using word to write html, but then I'd hear the whole

"iz th wayz i liks it u cant mak me chang hitlr".

Edit: a beautifulsoup solution:

from BeautifulSoup import BeautifulSoup, Comment
soup = BeautifulSoup(str(desc[desc_type]))
                    comments = soup.findAll(text=lambda text:isinstance(text, Comment))
                    [comment.extract() for comment in comments]
                    print soup
Edit
Report

1 Answer

12

There are a few things that - if you know them - will lead to the easiest/best solution:

  • clean_html() returns the same type you provide it with: if you give it a string, it will return a string, but if you give it an Element or ElementTree, it will return an Element or ElementTree respectively

  • you can control the way an Element or ElementTree is serialized, by giving encoding options to lxml.html.tostring() method or the tree's write() method (same goes for xml by the way). You can do this with encoding='utf-8' for example.

  • any content that CAN be encoded in that encoding, will be output as an encoded string, any content that cannot will be "escaped" as entities. Using encoding="ascii" will force any non-ascii characters to "nice" entities like you wish.

Put together, this means: first parse the string into an element (or tree if you wish), clean it, and serialize it as needed:

html = lxml.html.fromstring("<div><!-- word style><bleep><omgz 1,000 tags><--><p>It&#146;s a spicy meatball!</div>")
html = clean_html(html)
result = lxml.html.tostring(html, encoding="ascii")

(and a slightly dirtier trick is to use the errors parameter on the encode() method of a unicode string: try encoding a unicode string containing "special" characters with s.encode('ascii', 'xmlcharrefreplace') and see what that does...)

answered 2011-02-03T21:18:44.847

Your Answer