I searched Java's internal representation for String, but I've got two materials which look reliable but inconsistent.
One is:
http://www.codeguru.com/cpp/misc/misc/multi-lingualsupport/article.php/c10451
and it says:
Java uses UTF-16 for the internal text representation and supports a non-standard modification of UTF-8 for string serialization.
The other is:
and it says:
Tcl also uses the same modified UTF-8[25] as Java for internal representation of Unicode data, but uses strict CESU-8 for external data.
Modified UTF-8? Or UTF-16? Which one is correct? And how many bytes does Java use for a char in memory?
Please let me know which one is correct and how many bytes it uses.