KnowledgeHub
Questions
Tags
Users
Search
Alex Rivera
|
Logout
Edit Question
Title
Body
I'm trying to log a UTF-8 encoded string to a file using Python's logging package. As a toy example: import logging def logging_test(): handler = logging.FileHandler("/home/ted/logfile.txt", "w", encoding = "UTF-8") formatter = logging.Formatter("%(message)s") handler.setFormatter(formatter) root_logger = logging.getLogger() root_logger.addHandler(handler) root_logger.setLevel(logging.INFO) # This is an o with a hat on it. byte_string = '\xc3\xb4' unicode_string = unicode("\xc3\xb4", "utf-8") print "printed unicode object: %s" % unicode_string # Explode root_logger.info(unicode_string) if __name__ == "__main__": logging_test() This explodes with UnicodeDecodeError on the logging.info() call. At a lower level, Python's logging package is using the codecs package to open the log file, passing in the "UTF-8" argument as the encoding. That's all well and good, but it's trying to write byte strings to the file instead of unicode objects, which explodes. Essentially, Python is doing this: file_handler.write(unicode_string.encode("UTF-8")) When it should be doing this: file_handler.write(unicode_string) Is this a bug in Python, or am I taking crazy pills? FWIW, this is a stock Python 2.6 installation.
Tags (comma-separated)
Save Edits
Cancel