Alex Rivera | Logout

Is UTF8 injective mapping?

Asked 2011-11-13T20:53:03.027
9

We write a C++ application and need to know this:

Is UTF8 text encoding an injective mapping from bytes to characters, meaning that every single character (letter...) is encoded in only one way? So, e.g. letter 'Ž' cannot be encoded as, say, both 3231 and 32119.

Edit
Report

1 Answer

0

Yes, sort of. If used properly, each unicode code point should only be encoded one way in UTF-8, but that's partly because of the requirement that only the shortest applicable UTF-8 byte sequence should be used for any character.

The method used to encode the characters, however, could encode many characters more than one way if not for this requirement -- and though not proper, there are some cases where this is done.

For example, 'Z' could be encoded as 0x5a or {0xc1, 0x9a} (among others) though the only 0x5a is considered correct because it is the shortest sequence.

answered 2011-11-13T21:10:23.993

Your Answer