Regex Tester

Test regular expressions live with match highlighting and capture groups.

Pattern
/ /

Flags: g global · i case-insensitive · m multiline ^/$ · s dot matches newline · u unicode · y sticky

Test string
Matches
Capture groups

How a regex engine actually matches

Most people learn regex as a bag of tokens. It becomes far easier to debug once you understand that an engine walks your string from left to right, trying to match the pattern starting at position 0, then position 1, and so on. At each position it consumes characters one pattern element at a time. When a match succeeds it stops (leftmost, earliest); when an element fails it backtracks — returns to the last choice point and tries the next alternative.

Backtracking explains nearly every "why does my regex do that" moment, and it is also the source of the performance cliff described below.

Greedy vs. lazy: the most common bug

Quantifiers are greedy by default — they consume as much as they can, then backtrack only if the rest of the pattern demands it. Adding ? makes them lazy: consume as little as possible.

Text:    <h1>Title</h1> and <p>Body</p>

<.*>     greedy  ->  <h1>Title</h1> and <p>Body</p>   (one huge match)
<.*?>    lazy    ->  <h1>   then   </h1>   then   <p>   …

Neither is "correct" in the abstract — they answer different questions. If you want "the contents of the first tag", the better answer is not a lazy quantifier but a negated character class, which cannot overrun the closing delimiter:

<[^>]*>   ->  <h1>, </h1>, <p>, </p>   (correct, and fast)

Catastrophic backtracking

This is worth knowing because it causes real production outages. A pattern with nested quantifiers — (a+)+ — can require exponential time to fail on a near-match. A 30-character input can hang a server for minutes. This is ReDoS (regex denial of service) and it is a genuine vulnerability class with its own CVE category.

(a+)+$        # catastrophic on "aaaaaaaaaaaaaaaaaaaaaaaaaaaaX"
^(a|a)*$      # same problem, expressed differently
(a|aa)+       # ambiguous alternation inside a quantifier

The defence: avoid nested quantifiers, prefer bounded or negated classes, and never run a user-supplied regex against user-supplied input on a request thread. If you must accept patterns from users, use a linear-time engine (RE2) or run matching with a timeout.

Anchors: ^ and $ are not "start of string"

By default in many engines, ^ means start of string and $ means end of string. With the m (multiline) flag they mean start and end of line. This distinction matters constantly when validating input:

/^\d+$/       matches "123"        rejects "123abc"   ✓ correct for validation
/^\d+$/m      matches "123" in "abc\n123\ndef"          — probably not what you wanted
/\A\d+\z/     (PCRE/Ruby) true start/end of string, ignores multiline

When validating user input, always anchor both ends. An unanchored pattern such as \d+ "validates" the string abc123def because it finds digits somewhere inside it.

Capture groups vs. non-capturing groups

Parentheses do double duty: they group, and they capture. If you only need grouping, use (?:…) — it is faster and keeps your capture indices clean.

(https?|ftp)://([^/]+)(/.*)?
  group 1 = scheme, group 2 = host, group 3 = path

(?:(https?|ftp)://)?([^/]+)
  group 1 = scheme (undefined if absent), group 2 = host

Named groups improve readability dramatically for anything non-trivial:

/(?<year>\d{4})-(?<month>\d{2})-(?<day>\d{2})/
const { year, month, day } = m.groups;

Practical patterns that are actually correct

Email validation is the classic trap — a fully RFC-5322-compliant regex is thousands of characters long and nobody uses it. The pragmatic answer is a loose check followed by sending a confirmation email:

\b[A-Za-z0-9._%+-]+@[A-Za-z0-9.-]+\.[A-Za-z]{2,}\b

Other patterns worth having:

UTC ISO-8601 timestamp
^\d{4}-\d{2}-\d{2}T\d{2}:\d{2}:\d{2}(?:\.\d+)?(?:Z|[+-]\d{2}:\d{2})$

Semantic version
^v?(0|[1-9]\d*)\.(0|[1-9]\d*)\.(0|[1-9]\d*)(?:-[\w.-]+)?(?:\+[\w.-]+)?$

IPv4 (validates range, not just shape)
^(?:(?:25[0-5]|2[0-4]\d|1\d\d|[1-9]?\d)\.){3}(?:25[0-5]|2[0-4]\d|1\d\d|[1-9]?\d)$

Slug
^[a-z0-9]+(?:-[a-z0-9]+)*$

When not to use a regex

Two cases come up repeatedly. Parsing HTML or XML — a regex cannot handle nesting; use a real parser (DOM, BeautifulSoup, cheerio). Parsing URLs or email addresses — use the standard library for your language; the edge cases are already handled for you. Regex is superb for extracting, validating and rewriting flat text, and poor at anything with recursive structure.

Frequently asked questions

Quantifiers are greedy — * and + consume as much as possible. Add ? to make them lazy (*?, +?), or better, use a negated character class such as [^>]* so the match cannot overrun its delimiter.

Modern browsers support lookbehind (?<=…) and (?<!…) in JavaScript, so it works here. Note that Safari added support later than Chrome and Firefox — if you are targeting older Safari, check compatibility before shipping.

Almost certainly catastrophic backtracking from nested quantifiers such as (a+)+. Restructure using bounded or negated character classes. This tool caps iterations to protect your browser, but the pattern will be just as slow in production.

m makes ^ and $ match at line boundaries instead of string boundaries. s makes . match newlines, which it otherwise does not.

Use a loose pattern like the one above, then confirm by sending an email. Full RFC 5322 compliance is impractical in a regex, and even a perfect pattern cannot tell you whether the mailbox exists.