Greedy vs Lazy: The Core Difference
Regular expression quantifiers (*, +, ?, {n,m}) are greedy by default: they consume as many characters as possible, then backtrack if the rest of the pattern cannot match. Adding a ? after any quantifier makes it lazy (also called non-greedy or reluctant): it consumes as few characters as possible, then extends one character at a time until the full pattern can succeed. This single character — the trailing ? — flips the matching strategy and is behind one of the most common regex misunderstandings on Stack Overflow.
The Lazy Quantifier Family
* zero or more, greedy X* greedy X*? lazy
+ one or more, greedy X+ greedy X+? lazy
? zero or one, greedy X? greedy X?? lazy
{n,m} between n and m, greedy X{2,5} greedy X{2,5}? lazyThe Canonical Example: HTML Tags
// Input:
// "Hello <b>bold</b> and <i>italic</i> text"
/<.*>/g // GREEDY — one giant match
// => ["<b>bold</b> and <i>italic</i>"]
/<.*?>/g // LAZY — four small matches
// => ["<b>", "</b>", "<i>", "</i>"]
/<[^>]+>/g // NEGATED CLASS — same four matches, no backtracking
// => ["<b>", "</b>", "<i>", "</i>"]The lazy pattern <.*?> gives correct results, but the engine still backtracks (one character per advance). A negated character class <[^>]+>is faster and more intention-revealing — it says "match everything except >until you see one".
Quoted Strings
// Input: 'He said "hello" and "goodbye".'
/".*"/g // GREEDY — one match: "hello" and "goodbye"
/".*?"/g // LAZY — two matches: "hello" and "goodbye"
/"[^"]*"/g // NEGATED CLASS — same two matches, fastestIf strings can contain escaped quotes (\"), you need a more careful pattern: /"(?:\\.|[^"\\])*"/. Lazy quantifiers can't handle escapes correctly on their own.
Language-Specific Usage
JavaScript
// Extract tags from HTML
const html = '<p>Hello <b>bold</b> world</p>';
html.match(/<[^>]+>/g);
// => ["<p>", "<b>", "</b>", "</p>"]
// Pull content between two delimiters
const log = 'START payload-A END and START payload-B END';
log.match(/START (.*?) END/g);
// => ["START payload-A END", "START payload-B END"]
// With a capture group, pull just the inside
[...log.matchAll(/START (.*?) END/g)].map(m => m[1]);
// => ["payload-A", "payload-B"]Python
import re
text = 'Hello <b>bold</b> and <i>italic</i>'
re.findall(r'<[^>]+>', text) # fastest
# => ['<b>', '</b>', '<i>', '</i>']
re.findall(r'<.*?>', text) # lazy
# => ['<b>', '</b>', '<i>', '</i>']
# Block comments across multiple lines (re.DOTALL = s flag)
code = "/* first\n line */ x = 1 /* second */"
re.findall(r'/\*.*?\*/', code, re.DOTALL)
# => ['/* first\n line */', '/* second */']PHP
$html = '<p>Hello <b>bold</b> world</p>';
preg_match_all('/<[^>]+>/', $html, $m);
// $m[0] => ["<p>", "<b>", "</b>", "</p>"]
// Lazy variant
preg_match_all('/<.*?>/', $html, $m);
// Same result; slightly more backtracking
// With s flag for multiline block comments
preg_match_all('/\/\*.*?\*\//s', $code, $m);Common Pitfalls
Lazy is not the same as "correct"
Lazy quantifiers tell the engine "try shortest first". They do not stop at structurally meaningful boundaries. ".*?" matching JSON strings will still accept escaped quotes as string content and break on them. For anything beyond trivial parsing, use a proper parser rather than regex.
Dot does not match newlines by default
In JavaScript, Python and PCRE, . matches any character except newline unless you enable the s flag (re.DOTALL in Python). So /<.*?>/ will not match a tag that spans multiple lines until you also pass s.
Lazy with alternation is often the real culprit
/(cat|dog).*?/ is almost never what you want — the .*? can match the empty string. Lazy quantifiers need a specific stop condition after them (like a literal character or anchor) to be meaningful.
Lazy inside a loop is still linear
Beginners sometimes think "lazy = O(n)", but a lazy quantifier can still backtrack repeatedly if the pattern after it is complex. The real performance win is negated character classes, which have no backtracking at all.
Possessive quantifiers (where supported)
Some engines (Java, PCRE, Ruby Onigmo) support possessive quantifiers like .*+ that are even more aggressive than greedy — they commit without backtracking. These are unavailable in JavaScript and Python. If you need possessive behaviour in those languages, wrap the sub-pattern in an atomic group where supported, or use a negated character class.
Lazy-Quantifier Cheatsheet
| Greedy | Lazy | Meaning |
|---|---|---|
| * | *? | zero or more |
| + | +? | one or more |
| ? | ?? | zero or one |
| {n} | {n}? | exactly n (no effect on fixed count) |
| {n,m} | {n,m}? | between n and m, prefer n |
Testing Your Lazy Regex
Use the live Regex Tester above with these inputs:
Test string:
"Hello <b>bold</b> and <i>italic</i> text. He said \"hi\"."
Patterns to compare:
/<.*>/g — one giant greedy match
/<.*?>/g — four small lazy matches
/<[^>]+>/g — four fastest negated-class matches
/".*?"/g — the single quoted string
/[A-Z].*?\./g — "Hello ... text."Performance Rule of Thumb
Negated character class > lazy quantifier > greedy with backtracking. For HTML tags prefer <[^>]+> over <.*?>. For quoted strings prefer "[^"]*". Reserve .*? for cases where you genuinely cannot enumerate what the stop character is — e.g. /<(.*?)>/ when the content can contain any character including > inside quotes (in which case regex is probably the wrong tool anyway).