Alex Rivera | Logout

What are best practices for designing XML schemas?

Asked 2008-10-23T21:12:16.833
44

As an amateur software developer (I'm still in academia) I've written a few schemas for XML documents. I routinely run into design flubs that cause ugly-looking XML documents because I'm not entirely certain what the semantics of XML exactly are.

My assumptions:

<property> value </property>

property = value

<property attribute="attval"> value </property>

A property with a special descriptor, the attribute.

<parent>
  <child> value </child>
</parent>

The parent has a characteristic "child" which has the value "value."

<tag />

"Tag" is a flag or it directly translates to text. I'm not sure on this one.

<parent>
  <child />
</parent>

"child" describes "parent." "child" is a flag or boolean. I'm not sure on this one, either.

Ambiguity arises if you want to do something like representing cartesian coordinates:

<coordinate x="0" y="1" />

<coordinate> 0,1 </coordinate>

<coordinate> <x> 0 </x> <y> 1 </y> </coordinate>

Which one of these options is most correct? I would lean towards the third based upon my current conception of XML schema design, but I really don't know.

What are some resources that succinctly describe how to effectively design xml schemas?

Edit
Report

3 Answers

25

One general (but important!) recommendation is never to store multiple logical pieces of data in a single node (be it a text node or an attribute node). Otherwise, you end up needing your own parsing logic on top of the XML parsing logic you normally get for free from your framework.

So in your coordinate example, <coordinate x="0" y="1" /> and <coordinate> <x>0</x> <y>1</y> </coordinate> are both reasonable to me.

But <coordinate> 0,1 </coordinate> isn’t very good, because it’s storing two logical pieces of data (the X-coordinate and the Y-coordinate) in a single XML node—forcing the consumer to parse the data outside of their XML parser. And while splitting a string by a comma is pretty simple, there are still some ambiguities like what happens if there's an extra comma at the end.

answered 2008-10-23T21:52:31.073
1

There's nothing inherently wrong with using an element or sub-element for every value you'd like to represent.

The main consideration is that sometimes it's cleaner to use an attribute. Since an element can only have one attribute of a given name, you're stuck with a 1:1 cardinality. If you're representing the data as a child element, you can use whatever cardinality you'd like (or be open to extending it later).

Rob Wells' response above is right: it depends on the relationships you're trying to model.

Any time there's clearly never going to be anything but a 1:1 relationship, an attribute may be cleaner.

answered 2008-10-23T21:40:40.490
0

I guess, it depends on how complex or simple the structure is.
I will make x and y as attribute, unless x and y have their own details

You can look at HTML or any other form of markup, which is used to define things (XAML in case of WPF, MXML in case of flash) to understand, why something is chosen as attribute as against a child node)

if x and y are not to be repeated, they can be attributes.

Lets say co-ordinates has multiple x and y, I guess xml doesnt allow multiple attributes with same name for a node. In that case, you will have to use child nodes.

answered 2008-10-23T21:40:24.063

Your Answer