What a Word Boundary Actually Is
A word boundary is one of the most misunderstood features in regular expressions. It is not a character. It is a zero-width assertion — a check that a certain condition holds at a position in the string, without consuming any characters. \b succeeds whenever one side of the position is a word character ([A-Za-z0-9_]) and the other side is a non-word character or the edge of the string. That's the whole rule, and almost every surprising behaviour of \b comes from forgetting one of its two sides.
Finding Whole Words
The most common use case: find a word standing alone, not as part of a longer word.
/\bcat\b/g
// Test string:
// "The cat sat. Cats scattered. A wildcat appeared. Concatenate."
// Matches: 1 — the standalone "cat"
// Rejected: "Cats" (word char after), "scattered" (word chars both sides),
// "wildcat" (word char before), "Concatenate" (both sides)This is the pattern behind every "match whole words only" checkbox in text editors, IDE search panels, and grep -w.
\\b vs \\B
\Bis the inverse — it matches at any position that is NOT a word boundary. That means both sides of the position are the same "type" (both word characters or both non-word characters).
/\Bcat\B/g
// Matches "cat" only when surrounded by other word characters.
// "concatenate" — match ✓
// "wildcats" — match ✓
// "The cat" — no match (word boundary on left)
// "cats" — no match (word boundary on left)Use \B to find a sequence insidea larger token — handy for finding substrings that are embedded in identifiers but not standalone, e.g. finding every function whose internal name contains "proxy" without flagging the exact string "proxy".
Language-Specific Usage
JavaScript
// Literal regex — backslash does NOT need escaping in literal form
const WHOLE_WORD = /\bERROR\b/g;
"ERROR: timeout. ERRORS=5. preERRORstate".match(WHOLE_WORD);
// => ["ERROR"] (only the standalone one)
// String form (new RegExp) — must double-escape the backslash
const r = new RegExp("\\berror\\b", "gi");
r.test("Error occurred"); // true
r.test("errors occurred"); // false
// Unicode-aware boundary (names with accents)
const WORD = /(?<![\p{L}\p{N}_])José(?![\p{L}\p{N}_])/u;
WORD.test("Meet José today"); // true
WORD.test("Josépha arrived"); // falsePython
import re
# \b in a RAW string works directly
re.findall(r'\bcat\b', 'The cat sat. Cats scattered. wildcat.')
# => ['cat']
# Python 3 re is Unicode-aware by default
re.findall(r'\bJosé\b', 'José met João')
# => ['José'] (works out of the box in Py3)
# re.ASCII flag reverts to ASCII-only word chars
re.findall(r'\bJosé\b', 'José met João', re.ASCII)
# => [] (é becomes a non-word character)PHP
// Without /u, PHP PCRE uses ASCII word chars
preg_match_all('/\bcat\b/i', 'The Cat SAT. CATS!', $m);
// $m[0] => ["Cat"] (just the standalone one)
// With /u, PCRE treats \w as ASCII only — \b still ASCII
// For Unicode-aware boundaries use lookbehind/lookahead with \p{L}
preg_match_all('/(?<![\p{L}\p{N}_])José(?![\p{L}\p{N}_])/u',
'Meet José today', $m);Common Pitfalls
Hyphenated words trip up \\b
The hyphen is a non-word character, so /\bstate\b/treats "state-of-the-art" as four separate tokens and matches "state". If you want hyphenated phrases to be treated as single tokens, replace \b with explicit lookaround over an extended word class:
/(?<![\w-])state(?![\w-])/
// "state-of-the-art" — no match
// "the state is good" — matchUnicode letters are non-word in most engines
In JavaScript (without a workaround), Perl in default mode, and Go's regexp package, \w is strictly [A-Za-z0-9_]. That means \bfires between "s" and "é" in "José", giving you surprise matches on partial names. Python 3 and PCRE/u flip this to Unicode by default. Always check your engine before relying on \b for international text.
Apostrophes break word boundaries too
/\bdon\b/matches "don" inside "don't". To find contractions as single tokens, extend the word set to include apostrophes: /(?<![\w'])don't(?![\w'])/.
Case sensitivity
\b is case-insensitive by nature (both sides check against the same word-char set). The i flag affects the letters you match, not the boundary itself. So /\bcat\b/imatches "Cat", "CAT", "cAt" as whole words equally.
\\b at the start or end of a string
String edges count as non-word positions. /\bhello/matches "hello" at the start of a line because the position before "h" qualifies as a boundary. Similarly /world\b/ matches at the end.
Word-Boundary Cheatsheet
| Goal | Pattern | Example match |
|---|---|---|
| Whole word | /\bcat\b/ | "cat" alone |
| Inside a word | /\Bcat\B/ | "concatenate" |
| Start of word | /\bcat/ | "cat", "cats" |
| End of word | /cat\b/ | "cat", "wildcat" |
| Hyphen-aware whole token | /(?<![\w-])cat(?![\w-])/ | no match on "cat-5" |
| Unicode whole word | /(?<![\p{L}])José(?![\p{L}])/u | accented names |
Testing Your Word Boundary Regex
Use the live Regex Tester above with this canonical test string:
"The cat sat on the mat. Cats scattered. A wildcat appeared.
Concatenate strings carefully. State-of-the-art tooling."- Pattern:
/\bcat\b/gi— exactly one match - Pattern:
/\Bcat\B/gi— inside "Concatenate" - Pattern:
/\bstate\b/gi— matches inside "State-of-the-art" because hyphens are non-word characters - Pattern:
/(?<![\w-])state(?![\w-])/gi— correctly skips the hyphenated form
Performance Notes
\bis extraordinarily fast because it's zero-width — it inspects a single character on each side without allocating anything. Even in loops over millions of lines, word boundaries do not become a bottleneck. The usual performance killers in regex (catastrophic backtracking) require quantifier interaction; \b has no quantifier of its own and never causes backtracking.