Alex Rivera | Logout

Ruby String.encode still gives "invalid byte sequence in UTF-8"

Asked 2012-05-05T21:31:11.567
8

In IRB, I'm trying the following:

1.9.3p194 :001 > foo = "\xBF".encode("utf-8", :invalid => :replace, :undef => :replace)
 => "\xBF" 
1.9.3p194 :002 > foo.match /foo/
ArgumentError: invalid byte sequence in UTF-8
from (irb):2:in `match'

Any ideas what's going wrong?

Edit
Report

1 Answer

2

This is fixed if you read the source text file in using an explicit code page:

File.open( 'thefile.txt', 'r:iso8859-1' )
answered 2013-03-19T18:41:41.047

Your Answer