When writing a parser in a parser combinator library like Haskell's Parsec, you usually have 2 choices:

  • Write a lexer to split your String input into tokens, then perform parsing on [Token]
  • Directly write parser combinators on String

The first method often seems to make sense given that many parsing inputs can be understood as tokens separated by whitespace.

In other places, I have seen people recommend against tokenizing (or scanning or lexing, how some call it), with simplicity being quoted as the main reason.

What are general trade-offs between lexing and not doing it?

Edit
Report