Base64 at a Glance
Base64 maps every 3 input bytes (24 bits) to 4 output characters (6 bits each), using an alphabet of 64 characters. The standard alphabet (RFC 4648 §4) is A-Z a-z 0-9 + /. Output length is always a multiple of 4 — if input is not a multiple of 3 bytes, the output is padded with 1 or 2 = characters so the length still divides evenly.
# Encoding math
# input bytes → output chars
# 3 → 4 (no padding)
# 2 → 3 + "=" (one pad)
# 1 → 2 + "==" (two pads)
# Examples
"ABC" → "QUJD" (3 bytes → 4 chars, no pad)
"AB" → "QUI=" (2 bytes → 3 chars + "=")
"A" → "QQ==" (1 byte → 2 chars + "==")
"ABCD" → "QUJDRA==" (4 = 3 + 1 → "QUJD" + "RA==")The Strict RFC 4648 Validator
This is the pattern most code libraries actually use. It enforces the three rules simultaneously: character set, length multiple of 4, and padding position.
const BASE64 = /^(?:[A-Za-z0-9+\/]{4})*(?:[A-Za-z0-9+\/]{2}==|[A-Za-z0-9+\/]{3}=)?$/;
BASE64.test(""); // true — empty is valid base64
BASE64.test("QUJD"); // true — "ABC"
BASE64.test("QUI="); // true — "AB"
BASE64.test("QQ=="); // true — "A"
BASE64.test("YWJjZA=="); // true — "abcd"
BASE64.test("Hello world"); // false — space + comma + case mix OK but length not /4
BASE64.test("QQ="); // false — single = not at correct position
BASE64.test("==QQ"); // false — padding must be at the end
BASE64.test("A"); // false — length 1 (not a multiple of 4)base64url — The JWT & URL Variant
RFC 4648 §5 defines base64url. It uses - instead of + and _ instead of / so the output is safe to use in URL path, query, and filename. Padding is optional in base64url and is usually omitted by producers (JWT libraries, OAuth state parameters).
// Strict base64url with optional padding
const B64URL_STRICT = /^(?:[A-Za-z0-9_-]{4})*(?:[A-Za-z0-9_-]{2}(==)?|[A-Za-z0-9_-]{3}=?)?$/;
// Lenient — just the character set, any length
const B64URL_LOOSE = /^[A-Za-z0-9_-]+={0,2}$/;
// JWT segment (3 of them separated by dots) — never has padding
const JWT_SEGMENT = /^[A-Za-z0-9_-]+$/;
// Full JWT (3 base64url segments)
const JWT = /^[A-Za-z0-9_-]+\.[A-Za-z0-9_-]+\.[A-Za-z0-9_-]*$/;
JWT.test("eyJhbGciOiJIUzI1NiJ9.eyJzdWIiOiIxIn0.sig"); // trueDetecting Base64 Inside Free Text
When scanning logs, source code or JSON for leaked tokens, the problem flips: find something that looks like base64 in the middle of plain text. Short base64 strings collide with English words, so you must gate on length.
// Minimum 40 chars (≈ 30 bytes of data) — low false-positive rate
const DETECT = /\b[A-Za-z0-9+\/]{40,}={0,2}\b/g;
const log = `2026-10-07 user=alice token=aGVsbG8gd29ybGQgaG93IGFyZSB5b3UgdG9kYXkgZnJpZW5kcw== refresh=xxx`;
log.match(DETECT);
// → ["aGVsbG8gd29ybGQgaG93IGFyZSB5b3UgdG9kYXkgZnJpZW5kcw=="]
// Even stricter — require padding (long encoded binary almost always ends in =)
const DETECT_PADDED = /\b[A-Za-z0-9+\/]{20,}=\b/g;
// Also check base64url (common in Authorization headers)
const DETECT_URL = /\b[A-Za-z0-9_-]{20,}\b/g;Matching Data URIs
Data URIs (RFC 2397) look like data:image/png;base64,iVBORw0KGgo.... They have a MIME type, an optional charset, an optional ;base64 marker, and the payload.
// Full data URI with optional charset and optional ;base64
const DATA_URI = /^data:([\w+.-]+\/[\w+.-]+)(;charset=[\w-]+)?(;base64)?,([A-Za-z0-9+\/=]*)$/;
// Image data URIs only (common for inline avatars, favicons)
const IMG_DATA_URI = /^data:image\/(png|jpeg|jpg|gif|webp|svg\+xml);base64,([A-Za-z0-9+\/=]+)$/i;
const m = IMG_DATA_URI.exec("data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAAEAAAABCAQAAAC1HAwCAAAAC0lEQVR42mNk+A8AAQUBAScY42YAAAAASUVORK5CYII=");
// m[1] = "png"
// m[2] = "iVBORw0K...CYII="
// Extract just the payload (strip the prefix) for decoding:
const payload = m ? m[2] : null;Language-Specific Usage
JavaScript / Node
const BASE64 = /^(?:[A-Za-z0-9+\/]{4})*(?:[A-Za-z0-9+\/]{2}==|[A-Za-z0-9+\/]{3}=)?$/;
function isBase64(s) {
if (typeof s !== 'string' || s.length === 0) return false;
return BASE64.test(s);
}
// Node: double-check by round-trip
function isTrueBase64(s) {
if (!isBase64(s)) return false;
const decoded = Buffer.from(s, 'base64');
const reencoded = decoded.toString('base64');
return reencoded === s; // catches cases where Buffer silently fixes bad input
}
// Browser: use atob with try/catch
function isValidBase64Browser(s) {
if (!BASE64.test(s)) return false;
try { atob(s); return true; } catch { return false; }
}Python
import re, base64, binascii
BASE64 = re.compile(r'^(?:[A-Za-z0-9+/]{4})*(?:[A-Za-z0-9+/]{2}==|[A-Za-z0-9+/]{3}=)?$')
BASE64URL = re.compile(r'^(?:[A-Za-z0-9_-]{4})*(?:[A-Za-z0-9_-]{2,3}=?)?$')
def is_base64(s: str) -> bool:
return bool(BASE64.match(s))
def is_true_base64(s: str) -> bool:
if not is_base64(s):
return False
try:
base64.b64decode(s, validate=True)
return True
except binascii.Error:
return False
# Detect inside text:
DETECT = re.compile(r'\b[A-Za-z0-9+/]{40,}={0,2}\b')
for m in DETECT.finditer(log_line):
print("maybe base64:", m.group())Java
import java.util.regex.Pattern;
import java.util.Base64;
static final Pattern BASE64 = Pattern.compile(
"^(?:[A-Za-z0-9+/]{4})*(?:[A-Za-z0-9+/]{2}==|[A-Za-z0-9+/]{3}=)?$"
);
public static boolean looksLikeBase64(String s) {
return s != null && BASE64.matcher(s).matches();
}
public static boolean isTrueBase64(String s) {
if (!looksLikeBase64(s)) return false;
try { Base64.getDecoder().decode(s); return true; }
catch (IllegalArgumentException e) { return false; }
}Common Pitfalls
Whitespace in base64 output
MIME base64 (RFC 2045) inserts a newline every 76 characters. OpenSSL does the same. If your input has line breaks, strip them before applying the strict regex: s.replace(/[\s]/g, ''). Alternatively, allow whitespace in the character set: /^[A-Za-z0-9+\/\s]+=2$/ — but this breaks the length-multiple-of-4 constraint, so prefer stripping first.
Empty strings
The strict regex matches an empty string because (?:X{4})* allows zero repetitions and the trailing group is optional. If you want to reject empty strings add (?=.) at the start: /^(?=.)(?:...)/ or check s.length > 0 in code.
Length not a multiple of 4
Standard base64 strings mustbe a multiple of 4 characters when padded. Decoders like Node's Buffer and Python's b64decode(validate=False) silently add implicit padding, hiding bugs. The strict regex catches this. For base64url without padding, allow length ≡ 2 or 3 (mod 4) as valid partial chunks.
Mixing standard and URL-safe characters
A string containing both / and - is invalid in both standards. Pick one before validation, and normalize if your input mixes them: s.replace(/-/g, '+').replace(/_/g, '/').
Base64 Regex Cheatsheet
| Goal | Pattern | Notes |
|---|---|---|
| Strict base64 | /^(?:[A-Za-z0-9+\/]{4})*(?:[A-Za-z0-9+\/]{2}==|[A-Za-z0-9+\/]{3}=)?$/ | RFC 4648 §4 |
| base64url (optional pad) | /^[A-Za-z0-9_-]+={0,2}$/ | JWT, OAuth |
| Detect in text | /\b[A-Za-z0-9+\/]{40,}={0,2}\b/g | heuristic only |
| Image data URI | /^data:image\/(png|jpeg|gif|webp|svg\+xml);base64,([A-Za-z0-9+\/=]+)$/i | 1=mime, 2=payload |
| JWT segment | /^[A-Za-z0-9_-]+$/ | no padding |
Testing Your Base64 Regex
Open the live Regex Tester and paste this canonical test block:
# VALID — should ALL match /^(?:[A-Za-z0-9+\/]{4})*(?:[A-Za-z0-9+\/]{2}==|[A-Za-z0-9+\/]{3}=)?$/
QUJD # "ABC" (no padding)
QUI= # "AB" (one pad)
QQ== # "A" (two pads)
SGVsbG8gd29ybGQ= # "Hello world"
YWJjZGVmZ2hpams= # "abcdefghijk"
aHR0cHM6Ly9wcm9tcHRzcGFjZS5pbg== # a URL
# INVALID — should ALL fail
Q # length 1
QQ= # single = at wrong position
== # padding only
QQ==extra # extra chars after padding
Hello world! # space + !
QUJD\n # newline (strip first)Regex Is Shape, Decoding Is Truth
Any random mix of base64 characters with the right length and padding will pass the regex. Only decoding tells you whether the result is a valid UTF-8 string, a valid image header, a valid protobuf, or garbage. Treat the regex as a cheap rejection filter — a wall against obvious junk — and let your language's base64 decoder be the final arbiter. For security- sensitive values (JWTs, cookies, OAuth state), you must decode and parse further anyway.