Alex Rivera | Logout

What data type is recommended for ID columns?

Asked 2009-05-31T14:12:50.010
18

I realize this question is very likely to have been asked before, but I've searched around a little among questions on StackOverflow, and I didn't really find an answer to mine, so here goes. If you find a duplicate, please link to it.

For some reason I prefer to use Guids (uniqueidentifier in MsSql) for my primary key fields, but I really don't know why this would be better. In many of tutorials I've walked myself through lately an automatically incremented int has been used. I can see pro's and cons with both:

  • A Guid is always of the same size and length, and there is no reason to worry about running out of them, whereas there is a limit to how many records you could have before you'd run out of numbers that fit in an int.
  • int is (at least in C#) a nullable type, which opens for a couple of shortcuts when querying for data.
  • And int is easier to read.
  • I bet you could come up with at least a couple of more things here.

So, as simple as the title says it: What is the recommended data type for ID (primary key) columns in a database?

EDIT: After recieving a couple of short answer, I must also add this follow-up question. Without it, your answer is neither compelling nor educating... ;) Why do you think so, and what are the cons of the other option that make you not choose that instead?

Edit
Report

2 Answers

7

Popular databases allow for larger autoincrement fields for years now, so it's much less of an issue.

As for what to use, it's always a choice. One is not clearly better than the other, they have different characteristics and each is good in different scenarios. I have used both over time, and the next schema I work with I'll consider both.

Pros for GUID:

  • Should be unique across computers.
  • Random, unmemorable goo means people are likely to use this only for its intended purpose of an opaque identifier.

Pros for autoincrement:

  • Human understandable.
  • Sequential assignment means you can use a clustered index and impact performance.
  • Suitable for data partitioning.
answered 2009-05-31T14:24:31.890
0

If the database is distributed, where you could get records from other databases, the primary key needs to be unique within a table across all the databases. GUID solves this issue, albeit at the cost of space. A combination of autoincrement and namespace would be a good tradeoff.

It would be nice if databases could provide inbuild support for autoincrements with "prefixes". So in one database, I get IDs like X1,X2,X3 ... and so on whereas in the other database it could be Y1,Y2,Y3 ... and so on.

answered 2009-05-31T14:47:38.770

Your Answer