KnowledgeHub
Questions
Tags
Users
Search
Alex Rivera
|
Logout
Edit Question
Title
Body
Before anyone will tells me to RTFM, I must say - I have digged through: Why does modern Perl avoid UTF-8 by default? Checklist for going the Unicode way with Perl How to match string with diacritic in perl? How to make "use My::defaults" with modern perl & utf8 defaults? and many others (like perluniintro and others) - but - sure , missed something So, the basic code: use 5.014; #getting 'unicode_strings' feature use uni::perl; #turning on many utf8 things use Unicode::Normalize qw(NFD NFC); use warnings; while(<>) { chomp; my $data = NFD($_); say "OK" if utf8::is_utf8($data); } At this point, from the utf8 encoded STDIN I got a correct unicode string in $data , e.g. "\w" will match multibyte [\p{Alphabetic}\p{Decimal_Number}\p{Letter_Number}] (maybe something more). That's ok and works. AFAIK $data does not contain utf8, but a string in perl's internal Unicode format. Now the questions: HOW can I ensure (test it), that any $other_data contains valid Unicode string? For what purpose is the utf8::is_utf8($data)? The whole utf8 pragma is a mystery for me. I understand that the use utf8; is only for the purpose of
Tags (comma-separated)
Save Edits
Cancel