Before anyone will tells me to RTFM, I must say - I have digged through:

So, the basic code:

use 5.014;           #getting 'unicode_strings' feature
use uni::perl;       #turning on many utf8 things
use Unicode::Normalize  qw(NFD NFC);
use warnings;
while(<>) {
    chomp;
    my $data = NFD($_);
    say "OK" if utf8::is_utf8($data);
}

At this point, from the utf8 encoded STDIN I got a correct unicode string in $data, e.g. "\w" will match multibyte [\p{Alphabetic}\p{Decimal_Number}\p{Letter_Number}] (maybe something more). That's ok and works.

AFAIK $data does not contain utf8, but a string in perl's internal Unicode format.

Now the questions:

  • HOW can I ensure (test it), that any $other_data contains valid Unicode string?
  • For what purpose is the utf8::is_utf8($data)? The whole utf8 pragma is a mystery for me.

I understand that the use utf8; is only for the purpose of

Edit
Report