KnowledgeHub
Questions
Tags
Users
Search
Alex Rivera
|
Logout
Edit Question
Title
Body
The following test code reads a file, and using lxml.html generates the leaf nodes of the DOM/Graph for the page. However, I'm also trying to figure out how to get the input from a "string". Using: lxml.html.fromstring(s) doesn't work, as this generates an Element as opposed to an ElementTree . So, I'm trying to figure out how to convert an element to an ElementTree . [my test code] import lxml.html from lxml import etree # trying this to see if needed # to convert from element to elementtree #cmd='cat osu_test.txt' cmd='cat o2.txt' proc=subprocess.Popen(cmd, shell=True,stdout=subprocess.PIPE) s=proc.communicate()[0].strip() # s contains HTML not XML text #doc = lxml.html.parse(s) doc = lxml.html.parse('osu_test.txt') doc1 = lxml.html.fromstring(s) for node in doc.iter(): if len(node) == 0: print "aaa ",node.tag, doc.getpath(node) #print "aaa ",node.tag nt = etree.ElementTree(doc1) <<<<< doesn't work.. so what will?? for node in nt.iter(): if len(node) == 0: print "aaa ",node.tag, doc.getpath(node) #print "aaa ",node.tag UPDATE 1: (parsing html instead of xml) Added the changes suggested by Abbas. got the following errs: doc1 = etree.fromstring(s) File "lxml.etree.pyx", line 2532, in lxml.etree.fromstring (src/lxml/lxml.etree.c:48621) File "parser.pxi", line 1545, in lxml.etree._parseMemoryDocument (src/lxml/lxml.etree.c:72232) File "parser.pxi", line 1424, in lxml.etree._parseDoc (src/lxml/lxml.etree.c:71093) File "parser.pxi", line 938, in lxml.etree._BaseParser._parseDoc (src/lxml/lxml.etree.c:67862) File "parser.pxi", line 539, in lxml.etree._ParserContext._handleParseResultDoc
Tags (comma-separated)
Save Edits
Cancel