Stax
Tools

Regex in 8 Steps: Patterns Every Developer Actually Uses

Learn regex patterns step by step, from literal characters to lookaheads, with 12 real-world patterns to copy into production and the mistakes that break them.

Harshil
Harshil
··6 min read
Regex in 8 Steps: Patterns Every Developer Actually Uses

Regex is the tool most developers avoid until they need it, then Google hurriedly, and finally paste something that half-works. The patterns aren't arbitrary — each piece of syntax solves a specific matching problem. Learn the building blocks in order and the rest becomes readable.

What you'll have by the end: a working mental model of regex syntax and 12 patterns you can use directly in code. Already comfortable with the basics and just need a lookup table? Jump straight to the regex cheat sheet.


What you need to know first

A regular expression is a pattern that a regex engine matches against a string. Most languages ship a built-in engine (re in Python, RegExp in JavaScript, java.util.regex in Java) with largely compatible syntax. The differences (lookahead support, Unicode handling, POSIX vs PCRE) become relevant at the advanced level — the core syntax in this guide works in every major language.

Test every pattern at the Stax Regex Tester as you read — real-time matching makes the concepts click faster than reading alone.


Step 1: Literal characters — the foundation

The simplest regex matches exact text.

hello

Matches the string hello wherever it appears: "say hello there" ✅, "hello world" ✅, "helo" ❌.

Regex is case-sensitive by default. Hello does not match hello unless you use the case-insensitive flag (/i in JavaScript, re.IGNORECASE in Python).

Characters that need escaping: . * + ? ^ $ { } [ ] | ( ) \

These have special meaning in regex. To match a literal dot in a URL, use \. not .. The backslash is the escape character.

stax\.tools    # matches "stax.tools" exactly
stax.tools     # matches "staxXtools" too (. matches any character)

Step 2: Character classes — match one of several characters

Square brackets define a set of characters to match at a single position.

[aeiou]        # matches any single vowel
[a-z]          # matches any lowercase letter (range)
[A-Z]          # matches any uppercase letter
[0-9]          # matches any digit
[a-zA-Z0-9]   # matches any alphanumeric character
[^aeiou]       # ^ inside brackets = NOT — matches any non-vowel

Shorthand classes — these are common enough to have their own syntax:

Shorthand Equivalent Meaning
\d [0-9] Any digit
\D [^0-9] Any non-digit
\w [a-zA-Z0-9_] Word character (letter, digit, underscore)
\W [^a-zA-Z0-9_] Non-word character
\s [ \t\n\r\f] Any whitespace
\S [^ \t\n\r\f] Any non-whitespace
. (everything except \n) Any character except newline

Step 3: Anchors — control where the match can occur

Without anchors, a pattern matches anywhere in the string. Anchors constrain the position.

^hello          # matches "hello" only at the start of the string
hello$          # matches "hello" only at the end of the string
^hello$         # matches ONLY the string "hello" — nothing before or after
\bword\b        # \b = word boundary — matches "word" but not "password"

\b (word boundary) is one of the most useful anchors. \bcat\b matches "cat" and "the cat sat" but not "catch" or "concatenate". It matches the position between a \w character and a \W character (or start/end of string).


Step 4: Quantifiers — control how many times to match

Quantifiers apply to the character or group immediately before them.

a?             # 0 or 1 'a' (optional)
a*             # 0 or more 'a's
a+             # 1 or more 'a's
a{3}           # exactly 3 'a's
a{2,4}         # between 2 and 4 'a's
a{2,}          # 2 or more 'a's

Greedy vs lazy: Quantifiers are greedy by default — they match as much as possible. Add ? to make them lazy (match as little as possible).

<.+>           # greedy: on "<b>bold</b>", matches "<b>bold</b>" (the whole thing)
<.+?>          # lazy: on "<b>bold</b>", matches "<b>" then "</b>" separately

This distinction matters enormously when parsing HTML-like strings or any nested structures.


Step 5: Alternation and grouping — OR logic and reusable units

The pipe | means OR. Parentheses () group expressions together.

cat|dog        # matches "cat" or "dog"
(cat|dog)s     # matches "cats" or "dogs"
(ab)+          # matches "ab", "abab", "ababab", etc.

Groups also capture the matched text, making it available for extraction or replacement.

(\d{4})-(\d{2})-(\d{2})

On "2026-06-23", this captures three groups: 2026, 06, 23. In Python, re.match(pattern, s).groups() returns ('2026', '06', '23'). In JavaScript, s.match(pattern) returns an array where index 1, 2, 3 hold the captured groups.

Non-capturing groups ((?:...)) group without capturing — useful when you need alternation but don't need to extract the group:

(?:cat|dog)s   # matches "cats" or "dogs" but doesn't capture "cat"/"dog"

Step 6: Lookaheads and lookbehinds — match without consuming

Lookarounds assert that something exists ahead of or behind the current position without including it in the match.

\d+(?= dollars)    # positive lookahead: digits followed by " dollars"
                   # "100 dollars" → matches "100", not "100 dollars"

\d+(?! dollars)    # negative lookahead: digits NOT followed by " dollars"

(?<=\$)\d+         # positive lookbehind: digits preceded by "$"
                   # "$100" → matches "100"

(?<!\$)\d+         # negative lookbehind: digits NOT preceded by "$"

Lookaheads are supported in virtually all modern regex engines. Lookbehinds are supported in Python, .NET, and PCRE but have limitations in JavaScript (variable-length lookbehinds require Node 16+).


Step 7: Flags — modify how the engine behaves

Flags change global matching behavior. They're passed outside the regex pattern.

Flag JavaScript Python Effect
Case insensitive /pattern/i re.IGNORECASE [a-z] also matches [A-Z]
Global (find all) /pattern/g re.findall() Don't stop at first match
Multiline /pattern/m re.MULTILINE ^ and $ match start/end of each line
Dotall /pattern/s re.DOTALL . matches newline too
Verbose N/A re.VERBOSE Allow whitespace and comments in pattern

The combination of g and m is common in text processing: find all occurrences of a pattern, across multiple lines, without stopping at the first match.


Step 8: Build and test before shipping

Regex bugs are invisible until production. A pattern that looks correct can match inputs you didn't expect (false positives) or miss edge cases you didn't test (false negatives).

Standard validation workflow:

  1. Write the pattern
  2. Test with expected-to-match inputs
  3. Test with expected-to-NOT-match inputs
  4. Test edge cases: empty string, maximum length, Unicode characters, leading/trailing spaces

Use the Stax Regex Tester or the Stax Regex Cheat Sheet to validate before committing.


12 patterns to copy directly

# Email (pragmatic, not RFC-complete)
^[a-zA-Z0-9._%+\-]+@[a-zA-Z0-9.\-]+\.[a-zA-Z]{2,}$

# Indian mobile number
^[6-9]\d{9}$

# URL (http and https)
https?:\/\/[^\s/$.?#].[^\s]*

# IPv4 address
^((25[0-5]|2[0-4]\d|[01]?\d\d?)\.){3}(25[0-5]|2[0-4]\d|[01]?\d\d?)$

# Date (YYYY-MM-DD)
^\d{4}-(0[1-9]|1[0-2])-(0[1-9]|[12]\d|3[01])$

# Indian PAN number
^[A-Z]{5}[0-9]{4}[A-Z]{1}$

# Indian IFSC code
^[A-Z]{4}0[A-Z0-9]{6}$

# Hex colour code
^#([A-Fa-f0-9]{6}|[A-Fa-f0-9]{3})$

# Strong password (8+ chars, upper, lower, digit, special)
^(?=.*[a-z])(?=.*[A-Z])(?=.*\d)(?=.*[@$!%*?&])[A-Za-z\d@$!%*?&]{8,}$

# Slug (URL-safe lowercase with hyphens)
^[a-z0-9]+(?:-[a-z0-9]+)*$

# UUID v4
^[0-9a-f]{8}-[0-9a-f]{4}-4[0-9a-f]{3}-[89ab][0-9a-f]{3}-[0-9a-f]{12}$

# Extract numbers from a string
\d+(?:\.\d+)?

The UUID v4 pattern above checks structure only — for how UUIDs compare to newer alternatives like ULID and Nanoid, see our UUID vs ULID vs Nanoid guide.


Common mistakes

Using . when you mean a literal dot. google.com in regex matches "googleXcom". Use google\.com.

Anchoring the email regex wrong. [a-zA-Z0-9._%+\-]+@[a-zA-Z0-9.\-]+\.[a-zA-Z]{2,} without ^...$ anchors will accept "invalid@@@@email@example.com" because it finds a valid-looking substring inside. Always anchor validation patterns.

Catastrophic backtracking. Nested quantifiers like (a+)+ can cause exponential matching time on non-matching inputs. A string of 30 'a's followed by a character that breaks the match can take seconds. This is a ReDoS (Regular Expression Denial of Service) vulnerability in web applications. Avoid nested quantifiers on unbounded character classes.

Forgetting the g flag in JavaScript. Without g, str.replace(/pattern/, replacement) only replaces the first match. With g, it replaces all. String.prototype.replaceAll() is available in Node 15+ and modern browsers as an alternative.


By Harshil Shah, developer and founder at Stax Tools.

Sources & methodology

  1. PCRE2 Documentation — Perl-Compatible Regular Expressions syntax reference
  2. MDN Web Docs — Regular expressions in JavaScript
  3. Python re module documentation — Regular expression operations
  4. OWASP — ReDoS (Regular expression Denial of Service)