Alex Rivera | Logout

how to get specific nodes in xml file with python

Asked 2010-02-09T16:29:55.717
10

im searching for a way to get a specific tags .. from a very big xml document with python dom built in module
for example :

<AssetType longname="characters" shortname="chr" shortnames="chrs">
  <type>
    pub
  </type>
  <type>
    geo
  </type>
  <type>
    rig
  </type>
</AssetType>

<AssetType longname="camera" shortname="cam" shortnames="cams">
  <type>
    cam1
  </type>
  <type>
    cam2
  </type>
  <type>
    cam4
  </type>
</AssetType>

i want to retrieve the value of children of AssetType node who got attribute (longname= "characters" ) to have the result of 'pub','geo','rig'
please put in mind that i have more than 1000 < AssetType> nodes
thanx in advance

Edit
Report

1 Answer

1

Similar to eswald's solution, again stripping whitespace, again loading the document into memory, but returning the three text items at a time

from lxml import etree

data = """<AssetType longname="characters" shortname="chr" shortnames="chrs"
  <type>
    pub
  </type>
  <type>
    geo
  </type>
  <type>
    rig
  </type>
</AssetType>
"""

doc = etree.XML(data)

for asset in doc.xpath('//AssetType[@longname="characters"]'):
  threetypes = [ x.strip() for x in asset.xpath('./type/text()') ]
  print threetypes
answered 2010-02-09T16:56:06.757

Your Answer