I'm having trouble understanding the finer details of negative lookahead regular expressions. After reading Regex lookahead, lookbehind and atomic groups, I thought I had a good summary of negative lookaheads when I found this description:
(?!REGEX_1)REGEX_2Match only if
REGEX_1does not match; after checkingREGEX_1, the search forREGEX_2starts at the same position.
Hoping I understood the algorithm, I cooked up a two-sentence test insult; I wanted to find the sentence without a certain word. Specifically...
Insult: 'Yomama is ugly. And, she smells like a wet dog.'
Requirements:
- Test 1: Return a sentence without 'ugly'.
- Test 2: Return a sentence without 'looks'.
- Test 3: Return a sentence without 'smells'.
I assigned the test words to $arg, and I used (?:(?![A-Z].*?$arg.*?\.))([A-Z].*?\.) to implement the test.
(?![A-Z].*?$arg.*?\.)is a negative lookahead to reject a sentence with the test word([A-Z].*?\.)matches at least one sentence.
The critical piece seems to be in understanding where the regex engine starts matching after processing the negative lookahead.
Expected Results:
- Test 1 ($arg = "ugly"): "And, she smells like a wet dog."
- Test 2 ($arg = "looks"): "Yomama is ugly."
- Test 3 ($arg = "smells"): "Yomama is ugly."
Actual Results:
- Test 1 ($arg = "ugly"): "And, she smells like a wet dog." (Success)
- Test 2 ($arg