Alex Rivera | Logout

Why are people using regexp for email and other complex validation?

Asked 2008-10-17T11:50:17.110
13

There are a number of email regexp questions popping up here, and I'm honestly baffled why people are using these insanely obtuse matching expressions rather than a very simple parser that splits the email up into the name and domain tokens, and then validates those against the valid characters allowed for name (there's no further check that can be done on this portion) and the valid characters for the domain (and I suppose you could add checking for all the world's TLDs, and then another level of second level domains for countries with such (ie, com.uk)).

The real problem is that the tlds and slds keep changing (contrary to popular belief), so you have to keep updating the regexp if you plan on doing all this high level checking whenever the root name servers send down a change.

Why not have a module that simply validates domains, which pulls from a database, or flat file, and optionally checks DNS for matching records?

I'm being serious here, why is everyone so keen on inventing the perfect regexp for this? It doesn't seem to be a suitable solution to the problem...

Convince me that it's not only possible to do in regexp (and satisfy everyone) but that it's a better solution than a custom parser/validator.

-Adam

Edit
Report

3 Answers

3

People do it because in most languages it is way easier to write regexp than to write and use a parser in your code (or so it seems, at least).

If you decide to eschew regexes, you will have to either write parsers by hand, or you resort to external tools (like yacc) for lexer/parser generation. This is way more complex than single-line regex match.

One need to have a library that makes it easy to write parsers directly in the language X (where 'X' is C, C++, C#, Java) to be able to build custom parsers with the same ease as regular expression matchers.

Such libraries originated in the functional land (Haskell and ML), but nowadays "parser combinators libraries" exist for Java, C++, C#, Scala and other mainstream languages.

answered 2008-10-17T12:43:06.257
3

People use regexes for email addresses, HTML, XML, etc. because:

  1. It looks like they should work and they often do work for the obvious cases.
  2. They "know" regular expressions. When all you have is a hammer all your problems look like nails.
  3. Writing a parser is harder (or seems harder) than writing a regular expression. In particular, writing a parser is harder than writing a regex that handles the obvious cases in #1.
  4. They don't understand the full complexity of the task.
  5. They don't understand the limitations of regular expressions.
  6. They start with a regex that handles the obvious cases and then try to extend it to handle others. They get locked into one approach.
  7. They aren't aware that there's (probably) a library available to do the work for them.
answered 2008-10-17T13:15:04.960
3

and then validates those against the valid characters allowed for name (there's no further check that can be done on this portion)

This is not true. For example, "ben..doom@gmail.com" contains only valid characters in the name section, but is not valid.

In languages that do not have libraries for email validation, I generally use regex becasue

  1. I know regex, and find it easy to use
  2. I have many friends who know regex, and I can collaborate with
  3. It's fast for me to code, and me-time is more expensive than processor-time for most applications
  4. For the majority of email addresses, it works.

I'm sure many built-in libraries do use your approach, and if you want to cover all the possibilities, it does get ridiculous. However, so does your parser. The formal spec for email addresses is absurdly complex. So, we use a regex that gets close enough.

answered 2008-10-17T13:51:34.870

Your Answer