KnowledgeHub
Questions
Tags
Users
Search
Alex Rivera
|
Logout
Edit Question
Title
Body
I've taken the Liberal URL Regex from Daring Fireball , merged it with some of Alan Storm improvements and hacked my way into fixing some bugs like support for IDN chars inside parentheses. This is what I've: /(?:[\w-]+:\/\/?|www[.])[^\s()<>]+(?:(?:\([^\s()<>]*\)[^\s()<>]*)+|[^[:punct:]\s]|\/)/ However I've encountered a bug that I'm not being able to solve: 'www.dsd(sd)sdsd.com' // can also be the valid 'www.dsd.com/whatever(whatever)' The above URL is being recognized as www.dsd(sd)sdsd.com' (or www.dsd.com/whatever(whatever)' ) instead of www.dsd(sd)sdsd.com (or www.dsd.com/whatever(whatever) ). This only seems to happen when the URL has parentheses, since the following URL: 'www.sampleurl.com' Is correctly being recognized as www.sampleurl.com . I think the [^[:punct:]\s]|\/ part of the regex is not being executed when the URL has parentheses , I've been trying for some time but I can't seem to find a solution. Can anyone help me? For commodity, I've set up a Rubular permalink with the regex and some test data (the last URL fails). I think that Gruber's regex was a little rushed, for instance it doesn't match URL's like: http://en.wikipedia.org/wiki/Something_(Special)_For_You I'm even more impressed by seeing that both Gruber and Alan
Tags (comma-separated)
Save Edits
Cancel