Alex Rivera | Logout

Clean namespace handling with dom4j

Asked 2009-09-14T15:48:30.773
9

We are using dom4j 1.6.1, to parse XML comming from somewhere. Sometime, the balise have mention of the namespace ( eg : ) and sometime not ( ). And it's make call of Element.selectSingleNode(String s ) fails.

For now we have 3 solutions, and we are not happy with them

1 - Remove all namespace occurence before doing anything with the xml document

xml = xml .replaceAll("xmlns=\"[^\"]*\"","");
xml = xml .replaceAll("ds:","");
xml = xml .replaceAll("etm:","");
[...] // and so on for each kind of namespace

2 - Remove namespace just before getting a node By calling

Element.remove(Namespace ns)

But it's works only for a node and the first level of child

3 - Clutter the code by

node = rootElement.selectSingleNode(NameWithoutNameSpace)
if ( node == null )
    node = rootElement.selectSingleNode(NameWithNameSpace)

So ... what do you think ? Witch one is the less worse ? Have you other solution to propose ?

Edit
Report

2 Answers

1

Option 1 is dangerous because you can't guarantee the prefixes for a given namespace without pre-parsing the document, and because you can end up with namespace collision. If you're consuming a document and not outputting anything, it might be ok, depending on the source of the doc, but otherwise it just loses too much information.

Option 2 could be applied recursively but its got many of the same problems as option 1.

Option 3 sounds like the best approach, but rather than clutter your code, make a static method that does both checks rather than putting the same if statement throughout your codebase.

The best approach is to get whoever is sending you the bad XML to fix it. Of course this begs the question is it actually broken. Specifically, are you getting XML where the default namespace is defined as X and then a namespace also representing X is given a prefix of 'es'? If this is the case then the XML is well formed and you just need code that is agnostic about the prefix, but still uses a qualified name to fetch the element. I'm not familiar enough with Dom4j to know if creating a Namespace with a null prefix will cause it to match all elements with a matching URI or only those with no prefix, but its worth experimenting with.

answered 2009-09-14T16:25:36.280
0

As Abhishek, I needed to strip the namespace from XML to simplify XPath queries in system testing scripts. (the XML is first XSD validated)

Here are the problems I faced:

  1. I needed to process deeply structured XML that had a tendency of blowing up the stack.
  2. On most complex XML, for a reason I didn't investigate fully, stripping all the namespaces only worked in reliably when traversing the DOM tree depth first. So that excluded the visitor, or getting the list of nodes with document.selectNodes("//*")

I ended up with the following (not the most elegant, but if that can help solving somebody's problem ...):

public static String normaliseXml(final String message) {
    org.dom4j.Document document;
    document = DocumentHelper.parseText(message);

    Queue stack = new LinkedList();

    Object current = document.getRootElement();

    while (current != null) {
        if (current instanceof Element) {
            Element element = (Element) current;

            Iterator iterator = element.elementIterator();

            if (iterator.hasNext()) {
                stack.offer(element);
                current = iterator;
            } else {
                stripNamespace(element);

                current = stack.poll();
            }
        } else {
            Iterator iterator = (Iterator) current;

            if (iterator.hasNext()) {
                stack.offer(iterator);
                current = iterator.next();
            } else {
                current = stack.poll();

                if (current instanceof Element) {
                    stripNamespace((Element) current);

                    current = stack.poll();
                }
            }
        }
    }

    return document.asXML();
}

private static void stripNamespace(Element element) {
    QName name = new QName(element.getName(), Namespace.NO_NAMESPACE, element.getName());
    element.setQName(name);

    for (Obj
answered 2013-03-23T01:54:00.807

Your Answer