Alex Rivera | Logout

Case-insensitive storage and unicode compatibility

Asked 2011-08-09T03:21:33.937
18

After I heard of someone at my work using String.toLowerCase() to store case-insensitive codes in a database for searchability, I had an epic fail moment thinking about the number of ways that it can go wrong:

  • Turkey test (in particular changing locales on the running computer)
  • Unicode version upgrades - I mean, who knows about this stuff? If I upgrade to Java 7, I have to reindex my data if I'm being case-insensitive?

What technologies are affected by Unicode versions?

Do I need to worry about Oracle or SQL Server (or other vendors) changing their unicode versions and resulting in one of my locales not resulting in the same lower or upper character conversion?

How do I manage this? I'm tempted by the "simplicity" of ensuring I use the database conversion, but when there's an upgrade it'll be the same sort of issue.

Edit
Report

1 Answer

0

I think the most long term solution is to

  • record the current default locale and technology stack version (in my case Java version) into configuration
  • if it's changed (since last start up, or running for locale - depending on how it's loaded by said technology stack), then lock the store and re-index all affected data sets.

Obviously, this needs to occur at the primary interface level; if I'm doing these changes in java, I better hope that it's my only data interface mechanism (e.g. that other techs are not querying the underlying table store)

answered 2011-08-09T03:25:27.373

Your Answer