10
Before anyone will tells me to RTFM, I must say - I have digged through:
- Why does modern Perl avoid UTF-8 by default?
- Checklist for going the Unicode way with Perl
- How to match string with diacritic in perl?
- How to make "use My::defaults" with modern perl & utf8 defaults?
- and many others (like perluniintro and others) - but - sure, missed something
So, the basic code:
use 5.014; #getting 'unicode_strings' feature
use uni::perl; #turning on many utf8 things
use Unicode::Normalize qw(NFD NFC);
use warnings;
while(<>) {
chomp;
my $data = NFD($_);
say "OK" if utf8::is_utf8($data);
}
At this point, from the utf8 encoded STDIN I got a correct unicode string in $data, e.g. "\w" will match multibyte [\p{Alphabetic}\p{Decimal_Number}\p{Letter_Number}] (maybe something more). That's ok and works.
AFAIK $data does not contain utf8, but a string in perl's internal Unicode format.
Now the questions:
- HOW can I ensure (test it), that any
$other_datacontains valid Unicode string? - For what purpose is the utf8::is_utf8($data)? The whole utf8 pragma is a mystery for me.
I understand that the use utf8; is only for the purpose of