PHPRegex Rule Reference
This reference documents every lint rule that PHPRegex produces. It serves as the authoritative guide for understanding what PHPRegex checks and how to fix issues. Validation error codes live in Diagnostics: Error Codes, and the optimizer’s rewrite suggestions in the API reference.
Table of Contents
| Section | Description |
|---|---|
| Validation Layers | How PHPRegex analyzes patterns |
| Flags | Flag-related diagnostics |
| Anchors | Anchor positioning issues |
| Quantifiers | Quantifier-related patterns |
| Groups | Group-related diagnostics |
| Lookarounds | Lookahead contradictions |
| Alternation | Alternation patterns |
| Backreferences | Reference diagnostics |
| Character Classes | Character class issues |
| Escapes | Escape sequence problems |
| Literals | Literal text that reads badly |
| Bytes Without /u | Multibyte text read as bytes |
| Inline Flags | Inline flag diagnostics |
| Pattern Complexity | When a pattern is too complex |
| ReDoS Security | Catastrophic backtracking |
| Advanced Syntax | Pointers to the tutorial |
| Compatibility | PHP and PCRE2 support |
| Diagnostics Catalog | Error code reference |
| Quick Reference Table | One line per rule |
Validation Layers
PHPRegex validates patterns through four layers — parse, semantic, runtime (optional), lint — each catching different types of issues. The diagram, one example per layer and the category each refused pattern carries live in Diagnostics: Validation Layers. The layers below this one are the last of the four: the lint rules this page documents.
PCRE2 Compatibility Contract
PHPRegex targets PHP’s PCRE2 engine (preg_*). Key behaviors that may surprise users:
| Behavior | Description | Example |
|---|---|---|
| Forward references | Backreference before capture compiles, but the group is unset and the match will fail until it has captured | /\1(a)/ compiles, but \1 fails |
| Branch reset | (?\|...) changes capture numbering |
\2 can be invalid in branches |
\g{0} |
Invalid in PHP | Use \g<0> or (?R) |
| Lookbehind | Must have bounded length | (?<=a+) is invalid |
Flags
Useless Flag ‘s’ (DOTALL)
Identifier: regex.lint.flag.useless.s
When it triggers: The pattern sets the s (DotAll) modifier but contains no unescaped dot outside a character class. The s flag only affects .: an escaped \. and a . inside [...] are literal dots, which it leaves alone.
Message: Flag 's' is useless: the pattern contains no unescaped dot outside a character class.
Visual Explanation:
/^\d+\.\d+$/s
-> no unescaped dot outside a class
-> s has no effect
Example:
// Warning: DotAll does nothing because there is no dot
preg_match('/^user_id:\d+$/s', $input);
// Preferred: Remove unnecessary flag
preg_match('/^user_id:\d+$/', $input);
Fix: Drop the flag or introduce a dot if intended.
Read more:
Useless Flag ‘m’ (Multiline)
Identifier: regex.lint.flag.useless.m
When it triggers: The pattern sets the m (multiline) modifier but contains no start/end anchors (^ or $). \A, \z and \Z are not read by m, so they do not count.
Message: Flag 'm' is useless: the pattern contains no ^ or $ anchor.
Visual Explanation:
/search_term/m
-> no ^ or $ anchors
-> m has no effect
Example:
// Warning: Multiline mode is unused because there are no anchors
preg_match('/search_term/m', $text);
// Preferred: Remove unnecessary flag
preg_match('/search_term/', $text);
Fix: Remove the flag or add anchors for per-line matching.
Read more:
Useless Flag ‘i’ (Caseless)
Identifier: regex.lint.flag.useless.i
When it triggers: The pattern sets the i (case-insensitive) modifier but contains no case-sensitive characters.
Message: Flag 'i' is useless: the pattern contains no case-sensitive characters.
Example:
// Warning: No letters to justify case-insensitive matching
preg_match('/^\d{4}-\d{2}-\d{2}$/i', $date);
// Preferred: Remove unnecessary flag
preg_match('/^\d{4}-\d{2}-\d{2}$/', $date);
Fix: Drop the flag when matching digits/symbols.
Useless Flag ‘D’ (Dollar End Only)
Identifier: regex.lint.flag.useless.D
When it triggers: The pattern sets the D modifier, which only stops a $ from matching before a final newline, but no $ is read without m: there is no $ at all (\Z and a $ in a class are not read by D), or every $ is under m, set as a modifier or inline, which overrides D.
Message: Flag 'D' is useless: the pattern contains no $ anchor. or Flag 'D' is useless: every $ anchor is read under m, which overrides D.
Example:
// Warning: no $ for D to read
preg_match('/^[a-z]+\z/D', $input);
// Warning: m overrides D on every $
preg_match('/^[a-z]+$/mD', $input);
// Preferred: drop the flag
preg_match('/^[a-z]+\z/', $input);
Fix: Drop the flag, or write \z and drop it.
Useless Flag ‘x’ (Extended)
Identifier: regex.lint.flag.useless.x
When it triggers: The pattern sets the x modifier, which only drops whitespace and # comments, but holds no whitespace and no # comment. A # in a class, escaped, or in a (?#...) group opens no comment. Whitespace is looked for in the pattern as written: a space in a class or after a backslash, which x keeps, silences the rule too, and so does a no-break space, which x drops under a UTF-8 LC_CTYPE locale. A pattern holding both a # and a NUL byte is left alone: under (*NUL) a NUL ends a comment.
Message: Flag 'x' is useless: the pattern contains no whitespace and no # comment.
Example:
// Warning: nothing for x to drop
preg_match('/^\d{4}-\d{2}$/x', $input);
// Preferred
preg_match('/^\d{4}-\d{2}$/', $input);
Fix: Drop the flag, or lay the pattern out with spaces and comments.
Anchors
Anchor Conflicts
Identifier: regex.lint.anchor.impossible.start, regex.lint.anchor.impossible.end, regex.lint.anchor.impossible.boundary
When it triggers:
^appears after consuming tokens (without themflag, including an inline(?m)scope) — cannot match start$appears before consuming tokens that cannot continue a line end — asserts end too early.$and\Zstill match before the subject’s final newline (and a multiline$before any newline), so a tail such as$\nis fine; under theDmodifier$is strict like\z(unlessmis set, which disablesD), and\zalways is\bsits between two word characters or two non-word characters, or\Bbetween a word and a non-word character (regex.lint.anchor.impossible.boundary):/a\bb/and/a\B!/never match. Both neighbours must be items that read exactly one character (a letter, an escape, a class, a dot,\d…); an optional one may be absent, so/a?\bb/is not reported. What a word character is follows the pattern’s mode: under/uPHP turns UCP on, soéis one and/é\bx/uis reported, while in byte mode the last byte oféis not, and/é\bx/matcheséx. An ASCII character is a word character or not under every option, and is decided as such; any other atom is put to the automata, which the rule asks at most eight questions about one pattern, staying silent for the rest of it past them. The rule stays silent in a pattern that sets an ASCII option ((?a),(?aD), …) or(?xx), or setsrbeside a(?^): the flags it reads do not say which of them is in force.
Visual Explanation:
Valid:
/^abc/ -> anchor before content
/abc$/ -> anchor after content
Invalid:
/a^bc/ -> ^ after consuming 'a'
/$abc/ -> $ before consuming anything
/a\bb/ -> no boundary between two letters
Fix: Move anchors to the correct position.
Anchor Precedence in Alternation
Identifier: regex.lint.anchor.alternationPrecedence
When it triggers: An anchor binds tighter than |: /^a|b|c$/ is ^a, or b anywhere,
or c$. The rule speaks of intent and never says the pattern is wrong. It fires when the
first alternative of the pattern starts with ^, \A or \G (past the option settings,
verbs and comments before it), or the last one ends with $, \z or \Z, and some
alternative has an anchor on neither side. The anchor may sit inside a group that changes
nothing about where the alternative matches ((?:b$), (b$), (?>b$), a one-alternative
(?|b$)) or be all a positive lookaround asserts (b(?=$)). An alternation whose every
alternative is anchored on one side is left alone, as the trim idiom /^\s+|\s+$/ is, and so
is an alternative that is the anchor alone (/a|$/: “or the end”). As in SonarPHP, an anchor anywhere
but the start of the first alternative and the end of the last shows the anchors are placed
alternative by alternative, and the rule stays silent (/^ +| +$|,/, /^a|b|^c/), as it does
under the A modifier, which anchors every alternative at the start. An alternative of verbs
and comments only, as (*FAIL), is not counted, and an empty one (/^a|/) is
regex.lint.alternation.empty’s. A verb that ends the match
attempt on backtracking, accepts at once or fails ((*COMMIT), (*PRUNE), (*SKIP),
(*ACCEPT), (*FAIL)), in an alternative that reads something, leaves the rule silent: it can
keep the engine from the other alternatives (/(*COMMIT)^a|b/ and /^(*COMMIT)a|b/ never
match b), and a failing verb would carry into the grouped form, which then matches nothing.
A (*SKIP:n) that no (*MARK:n) names is ignored by the engine, and by the rule. The tip
keeps the options before the anchor ((?i)^(?:a|b)), keeps text
quoted with \Q...\E quoted (^(?:a|\Qb)\E) for /^a|\Qb)\E/), and leaves out the comments
ending each alternative, so that it compiles under /x. A group holding the anchor alone
leaves with it (^(?:a|b) for /(?:^)a|b/), but a capturing group stays, keeping its number,
and a group setting m moves as written ((?m:^)(?:a|b)).
Example:
// WARNING: "b" matches anywhere
preg_match('/^a|b|c$/', 'xbx'); // 1
// PREFERRED, if the anchors are meant for every alternative
preg_match('/^(?:a|b|c)$/', 'xbx'); // 0
Fix: Group the alternatives, as the tip shows, when the anchors are meant for all of them.
Quantifiers
Nested Quantifiers (ReDoS Risk)
Identifier: regex.lint.quantifier.nested
When it triggers: A variable quantifier wraps another variable quantifier, creating catastrophic backtracking potential.
It is not reported when the nesting cannot blow up:
- Nothing at all follows the outer loop:
/(a+)+/and/x(\d+\.?)+/end the pattern, so the first way the engine finds is the match. This holds when the outer minimum is 0 or 1, and when no subroutine call runs the loop again. Anything after the loop, an anchor or an assertion included, keeps the issue:/(a+)+$/and/(?:a+)+(?=b)/are reported. - A short run before the end: an outer bound of at most 2 around a greedy run of one literal character, followed only by
$,\zor\Z, without/m, as/(a+){1,2}$/. Any other shape keeps the issue: a bound of 3 or more (/(a+){0,3}$/), a class (/([ab]+){1,2}$/), a lazy run (/(a+?){1,2}$/),/m, or a lookahead after the loop. - A separator splits each iteration: the inner loop is followed right away by an item it cannot match, and every other item of the iteration matches one way only. In the unrolled loop
/"[^"\\]*(?:\\.[^"\\]*)*"/, each iteration ends with[^"\\]*, and the\\that opens the next iteration is what[^"\\]refuses, so there is one way to split the input. Under/ithe two sides are compared case-folded, so/(?:a+,)+$/iis still exempt, but a loop unrolled this way is never exempt:/"[^"\\]*(?:\\.[^"\\]*)*"/iis reported.
With the ReDoS analysis on (--redos, or checks.redos.enabled in regex.json), a pattern the analysis proves linear drops the issue: the proof outranks the heuristic. That holds for a pattern with inline options too, an option setting such as (?s) or a scoped group such as (?i-r:…), which the proof reads as PCRE does. The issue is also left out for a pattern listed in the ignored patterns or found trivially safe.
Visual Explanation:
Pattern: /(a+)+b/
Input: "aaaaa!"
Inner (a+) can match 1..n and the outer + repeats 1..n
Result: many backtracking paths
Example:
// VULNERABLE: Nested quantifiers can cause ReDoS
preg_match('/(a+)+b/', $input);
// SAFER: Use atomic groups
preg_match('/(?>a+)+b/', $input);
// SAFER: Use possessive quantifier
preg_match('/(a++)+b/', $input);
Fix: Refactor to be deterministic, or use atomic groups/possessive quantifiers — but verify the rewrite still matches everything you need: when the inner part is ambiguous ((?:ab|a)+b), an atomic inner group removes the backtracking between iterations and can change the language.
Read more:
Dot-Star in Quantifier
Identifier: regex.lint.dotstar.nested
When it triggers: An unbounded quantifier wraps a dot-star, which can cause extreme backtracking.
Without /s (or an inline (?s)), a dot cannot cross a newline, so /(?:.*\n)+x/ is not reported: each iteration ends at the next newline and the input splits one way. With /s or (?s), the same pattern is reported. Unlike nested quantifiers, a loop at the very end of the pattern is still reported, as /(?:.*)+/ in the example below. With the ReDoS analysis on, a pattern the analysis proves linear drops the issue, inline options such as (?s) or (?s:…) included (see Nested Quantifiers).
Example:
// RISKY: .* in repetition can backtrack heavily
preg_match('/(?:.*)+/', $input);
// SAFER: Make the dot-star atomic or possessive
preg_match('/(?>.*)+/', $input);
preg_match('/.*+/', $input); // If no outer repetition is needed
// BETTER: Use negated character class
preg_match('/[^"]*/', $input); // For double-quoted strings
Fix: Make it atomic/possessive or replace .* with a specific class — and verify the rewrite still matches everything you need.
Useless Quantifier
Identifier: regex.lint.quantifier.useless
When it triggers: A quantifier that matches exactly once (e.g., {1} or {1,1}).
Example:
// WARNING: {1} does not change the match
preg_match('/a{1}/', $input);
// PREFERRED: Remove the quantifier
preg_match('/a/', $input);
Fix: Remove the {1} quantifier.
Zero Quantifier
Identifier: regex.lint.quantifier.zero
When it triggers: A quantifier with a maximum of zero (e.g., {0} or {0,0}), which makes the element disappear.
Example:
// WARNING: {0} always matches zero occurrences
preg_match('/ab{0}c/', $input);
// PREFERRED: Remove the quantified element
preg_match('/ac/', $input);
Fix: Remove the quantified element or replace it with an empty pattern.
Optimal Quantifier Concatenation
Identifier: regex.lint.quantifier.concatenation
When it triggers: Two adjacent quantified tokens can be simplified because one character set is a subset of the other.
The rule speaks only when it knows the subset holds: a token that may take a character above
ASCII fits only a dot, a negated class of ASCII members or, without u, \D, \W and \S
(/^\h{1,3}[\t ]+\z/ stays silent, \h takes "\xA0"), and under i a negated class must
hold both cases of its letters (/^A?[^a]*\z/i stays silent).
Example:
// WARNING: \d is a subset of \w
preg_match('/\d+\w+/', $input);
// PREFERRED: Keep the superset quantifier unbounded
preg_match('/\d\w+/', $input);
// WARNING: \d* before \w* can be dropped — \w covers it
preg_match('/\d*\w*/', $input);
// PREFERRED: Drop the whole quantified term
preg_match('/\w*/', $input);
An atom guarded by a negative lookahead, as in (?:.(?!x))*, does not match every character of its class: .(?!x) refuses a character followed by x. Such a pair is reported only when the rewrite keeps every guard true, so /^(?:.(?!x))*.{1,3}$/ is not.
Fix: Tighten the smaller quantifier to its minimum, or — when it can already match zero times — drop the whole quantified term.
Lazy Quantifier at the End
Identifier: regex.lint.quantifier.lazyEnd
When it triggers: A lazy quantifier, or a greedy one under /U, has nothing after it
in the pattern. The match ends as soon as it may, so the quantifier matches its minimum.
The same holds when only items that may match nothing follow it, and none of them holds an
anchor, a lookaround or another test that can fail: in /a+?b*/ on aab, a+? takes one
a, b* matches nothing after it, and the match is a. A $, a \b or a lookahead after
the suffix lets the quantifier take more, and the rule stays silent.
Under /U the message says where the laziness comes from:
Quantifier "+" is lazy under the U flag and ends the pattern, so it always matches its minimum.
A quantifier inside a group that a subroutine call runs again ((?1), (?&name)) is not
reported: where the call stands, something may follow it. In /(?1)x(a+?)/, the call takes
aaa from aaaxa. A recursive pattern ((?R), (?0)) is not checked by this rule at all.
Example:
// WARNING: +? stops after one digit
preg_match('/id-\d+?/', 'id-123', $m); // $m[0] is "id-1"
// WARNING: under /U, + is lazy: each space is replaced on its own
preg_replace('/\s+/U', ' ', "a b"); // "a b"
// PREFERRED
preg_match('/id-\d+/', 'id-123', $m); // "id-123"
preg_match('/id-\d+?$/', 'id-123', $m); // "id-123": something follows
Fix: Make the quantifier greedy, write its minimum, or anchor what must follow it.
Repeat That Can Match Empty
Identifier: regex.lint.quantifier.emptyRepeat
When it triggers: An unbounded quantifier (*, +, {n,}) repeats an item that can
match the empty string: (a*)*, (?:a|b?)+, (?:a{0,2})+. PCRE ends the loop on an empty
iteration, which matches nothing more. When every way the item matches the empty string goes
through a capturing group, that last, empty iteration sets the capture to the empty string,
and the message says so; in (?:(a)|b?)* the empty way goes around the group, and $m[1]
keeps its a, and a capture inside a lookaround keeps what the lookaround read:
(?:x|(?=(a)))* on xa captures a. An item that fails rather than match the empty string is sound:
(?:a|(*FAIL))* and (?:a|(?!))* are not reported. Where another rule already reports the
same repeat (an empty alternative, a quantified lookaround, a nested quantifier), this one
stays silent, as long as the configuration turns that other rule on: (?:a*)*b with
quantifier.nested off is reported by this rule.
Example:
// WARNING: the capture ends empty
preg_match('/(a*)*/', 'aaa', $m); // $m is ["aaa", ""]
// PREFERRED
preg_match('/(a*)/', 'aaa', $m); // $m is ["aaa", "aaa"]
Fix: Make the repeated item read at least one character, or drop the outer quantifier.
The tip gives a rewrite only where one matches the same text: a* for (?:a*)+, (?:a?)*
or (?:a{0,2})+. Elsewhere, as for (?:a|b?)+ (which is (?:a|b)*) or (?:a{0})+ (which
reads nothing), the issue comes without a tip.
Impossible Possessive Quantifier
Identifier: regex.lint.quantifier.possessiveImpossible
When it triggers: A possessive repeat (a*+), or a greedy one inside an atomic group
((?>a*)), takes every character the atom right after it could read, and never gives one
back: /a*+a/, /\d++5/ and /a*+A/i never match through there. Decided on atoms that
read exactly one character, under the flags in force at each, the automata comparing their
sets: /a*+A/ matches A, /\w++é/ matches aé in byte mode, /\w++é/u never does. A
bounded repeat stops at its bound (/^a{0,3}+a$/ matches aaaa) and an atom that may match
nothing (/a*+a?/) can step aside: neither is reported. Past the eighth repeat it asks the
automata about in a pattern, the rule stays silent for the rest of it. The rule stays silent in a pattern that sets an ASCII option ((?a), (?aD), …) or (?xx), or sets r beside a (?^): the flags it reads do not say which of them is in force.
Example:
// WARNING: the repeat takes every digit, the "5" included
preg_match('/\d++5/', '12345'); // 0
// PREFERRED
preg_match('/\d+5/', '12345'); // 1
Fix: Make the repeat greedy, or exclude from its set what must follow it.
Lazy Quantifier Before a Delimiter
Identifier: regex.lint.quantifier.lazyToClass (off by default)
When it triggers: A lazy dot, .*? or .+?, is followed by one closing character and
nothing that can fail comes after it: /".*?"/ backtracks one character at a time where
/"[^"\n]*"/ reads the run at once, and the two write the same $matches on every subject.
The tip spells the class: [^"\n]* without s, [^"]* under s. The automata prove the
two patterns equivalent before the rule speaks; it stays silent when something follows the
closing character (in /".*?"x/ on "a"b"x the lazy dot crosses the middle quote, the class
cannot), and under a newline convention other than \n ((*CRLF), (*ANYCRLF), (*CR)…),
where the dot stops at another character, and past the eighth lazy dot it asks the automata
about in a pattern. A perf rule: turn it on with
"quantifier.lazyToClass": true under checks.lint.rules.
Example:
// INFO: backtracks at each character
preg_match('/<.*?>/', '<a><b>', $m); // "<a>"
// PREFERRED: same match, read at once
preg_match('/<[^>\n]*>/', '<a><b>', $m); // "<a>"
Fix: Write the negated class the tip gives.
Useless Lazy Quantifier
Identifier: regex.lint.quantifier.uselessLazy (off by default)
When it triggers: A lazy quantifier writes the same $matches as its greedy form on every
subject: /(a+?)b/ and /(a+)b/ stop at the same b, since the run cannot take one, and a
fixed count such as a{3}? has nothing to choose. The automata prove the two patterns
equivalent before the rule speaks, from any offset (a pattern with a start anchor is asked
again with the anchor failing). It stays silent where the lazy marker matters
(/(a+?)a/ on aaa captures a, the greedy form aa), under U, where +? is the greedy
one, where nothing that can fail follows a variable count (quantifier.lazyEnd reports that),
in a pattern that may match the empty string (after an empty match, preg_match_all(),
preg_replace() and preg_split() try again for a non-empty one, where the two forms part:
/(\d+?,?)??/ finds 1 and 2 in 12, the greedy form 12) or that sets a
(*LIMIT_MATCH=…), (*LIMIT_DEPTH=…) or (*LIMIT_HEAP=…), which the lazy form, backtracking
more, may reach first, and past the eighth question it asks the automata about in a pattern. A style rule: turn it
on with "quantifier.uselessLazy": true under checks.lint.rules.
Example:
// INFO: lazy for nothing
preg_match('/(\d+?)\D/', 'ab12x', $m); // "12x", "12"
// PREFERRED: the same $matches
preg_match('/(\d+)\D/', 'ab12x', $m); // "12x", "12"
Fix: Drop the ?.
Repeat at the Edge of a Lookaround
Identifier: regex.lint.lookaround.edgeQuantifier (off by default)
When it triggers: A lookahead only checks that its body can match, so the last item of
the body repeated past its minimum changes nothing: (?=a{2,6}) asserts what (?=a{2})
asserts, (?=ab*) what (?=a) asserts. The same holds for the first item of a lookbehind,
(?<=a{2,6}). The automata prove that the two bodies hold at the same positions before the
rule speaks. It stays silent on a body holding a capture ((?=(a{2,6})) keeps the whole run
in $1), a reference, a call, a verb, a callout or \K, on a lookbehind whose body tests a
position (\z, $, \b, a lookahead…: a lookbehind of variable length reads the end of the
subject at its own position, one of fixed length the real end, and (?<=b{1,2}\z)a matches
ba where (?<=b\z)a does not), in a non-atomic lookahead
((*napla:...)), on a possessive repeat, which the automata do not read, and past the eighth
question it asks them about in a pattern. A perf rule: turn it on with
"lookaround.edgeQuantifier": true under checks.lint.rules.
Example:
// INFO: the x past the first is never checked
preg_match('/\d+(?=px{1,3})/', '12pxx', $m); // "12"
// PREFERRED
preg_match('/\d+(?=px)/', '12pxx', $m); // "12"
Fix: Write the minimum count the tip gives.
Quantified Assertion
Identifier: regex.lint.quantifier.assertion
When it triggers: A lookaround carries a quantifier. With a minimum of zero PCRE tries
the rest of the pattern with and without the assertion, so it constrains nothing; with a
minimum of one or more, PCRE checks it once. A lookaround that captures is left alone:
(?=(\w+))? still sets its group when it holds.
Example:
// WARNING: the lookahead may be skipped
preg_match('/(?=\d)?\w/', 'a'); // 1
// PREFERRED
preg_match('/(?=\d)\w/', 'a'); // 0
Fix: Remove the quantifier, or the assertion.
Groups
Redundant Non-Capturing Group
Identifier: regex.lint.group.redundant
When it triggers: A non-capturing group wraps a single atomic token without changing precedence.
An empty group is regex.lint.group.empty’s, and a group that keeps an escape apart from
the digit after it is not redundant: (a)\1(?:0) matches aa0, (a)\10 does not, and
(?:\N){U+41} reads {U+41} as text where \N{U+41} is the code point under /u, and does
not compile without it. Neither is a group that keeps braces from reading as a quantifier, a{(?:2)} (it matches a{2},
a{2} matches aa), nor a group a quantifier repeats around anything but one character:
without the group, (?:\Qab\E)+ repeats the b alone, (?:é)+ without /u the last byte
of the letter, and (?:^)+ does not compile. (?:a)+ is reported: a+ is the same.
Example:
// WARNING: Unnecessary group
preg_match('/(?:foo)/', $input);
// PREFERRED: Remove the group
preg_match('/foo/', $input);
Fix: Remove the unnecessary group.
Empty Group
Identifier: regex.lint.group.empty
When it triggers: A non-capturing or atomic group is empty, (?:) or (?>): it matches
the empty string and changes nothing. An empty capturing group () is a placeholder that
numbers a group and is left alone, as are empty lookarounds, which have rules of their own.
So is a (?:) that keeps an escape apart from the digit after it, the form the library’s own
printer writes: (a)\1(?:)0 is not (a)\10, and \01(?:)2 is not the newline \012; one
that keeps braces from a count, a{(?:)2}, a{1,(?:)2} or a{(?:),2} ({,2} is a
quantifier since PCRE2 10.43); and one a quantifier repeats: a(?:)? matches a alone,
a? the empty string too. An unbounded repeat of an empty group, (?>)+, is
regex.lint.quantifier.emptyRepeat’s.
Example:
// WARNING: the group does nothing
preg_match('/a(?:)b/', 'ab'); // 1
// PREFERRED
preg_match('/ab/', 'ab'); // 1
Fix: Remove the group.
Capture That Never Captures Text
Identifier: regex.lint.group.alwaysEmptyCapture
When it triggers: A capturing group is empty or unset wherever the pattern matches: in
/a+(a*)/ the greedy a+ takes every a, so $1 is always ""; in /^\w+(\d*)$/ \w+
takes the digits. The automata prove it: the pattern with the group’s body removed writes the
same $matches on every subject, from offset 0 and, for a pattern with a start anchor, again
with the anchor failing, as it does where preg_match_all() goes on (/^.|.(c)/s fills $1
there). It stays silent on a group with nothing in it (() is a marker), on a group in a branch
reset ((?|...)), whose number another alternative may fill, in a pattern that may match the empty string (preg_match_all() tries again after an
empty match, where another path may fill the group), where quantifier.lazyEnd or
quantifier.zero already says why the group stays empty (/^L_(.*?)/, /(b){0}/), on a group
a reference or a call reads, which the automata do not follow, and past the eighth question it
asks them about in a pattern.
Example:
// WARNING: $1 is always ""
preg_match('/^(\w+)(\d*)$/', 'abc123', $m); // ["abc123", "abc123", ""]
// PREFERRED
preg_match('/^([a-z]+)(\d*)$/', 'abc123', $m); // ["abc123", "abc", "123"]
Fix: Remove the group, or narrow what precedes it so the group can read text.
Quantified Capturing Group
Identifier: regex.lint.group.quantifiedCapture
When it triggers: A repeatable quantifier (*, +, {n,}…) sits on a capturing group: each iteration overwrites the capture, and only the last one survives. An unnamed group is reported at info severity, a named one at warning.
Example:
// INFO: Quantified capturing group "(...)" with "+":
// only the last iteration's capture is retained.
preg_match('/(a)+/', 'aaa', $m); // $m is ["aaa", "a"]
// PREFERRED: the repetition does not need to capture
preg_match('/(?:a)+/', 'aaa', $m);
Fix: Use a non-capturing group (?:...) for the repetition and capture the whole match, or restructure the pattern.
Lookarounds
Impossible Lookaround
Identifier: regex.lint.lookaround.impossible
When it triggers: A lookahead contradicts what the pattern reads after it: (?=a)b asks
for an a where the pattern reads a b, (?!a)a forbids the a it reads. The automata
decide it: what follows the lookahead runs through the enclosing groups and quantifiers
(/(?:x(?=a))+b/ is reported: another x or the b comes next, never an a), read for a
few items. A contradiction in one alternative is reported though the pattern still matches
through another (/(?:x(?=a)|y)b/). The rule stays silent where it cannot read: a
backreference, a lookbehind or \K after the lookahead, a group a subroutine call runs from
elsewhere, a pattern past the automata’s work cap or past the eighth lookahead it asks them
about, and in UTF mode a word boundary, \w or a
Unicode property near the lookahead. A repeated lookahead is
regex.lint.quantifier.assertion’s, and an empty negative lookahead (?!) fails on purpose. The rule stays silent in a pattern that sets an ASCII option ((?a), (?aD), …) or (?xx), or sets r beside a (?^): the flags it reads do not say which of them is in force.
Example:
// WARNING: a digit is required where a letter is read
preg_match('/(?=\d)[a-z]+/', 'a1'); // 0
// WARNING: the "a" is forbidden, then read
preg_match('/(?!a)a/', 'aa'); // 0
Fix: Fix the lookahead or what follows it, or remove the alternative that cannot match.
Alternation
Duplicate Alternation Branches
Identifier: regex.lint.alternation.duplicateDisjunction
When it triggers: The same alternative appears more than once.
Example:
// WARNING: Duplicate branch
preg_match('/(a|a)/', $input);
// PREFERRED: Use a single literal
preg_match('/a/', $input);
Fix: Remove duplicates or use a character class.
Empty Alternatives
Identifier: regex.lint.alternation.empty
When it triggers: An alternation contains an empty branch (e.g., trailing | or ||).
Example:
// WARNING: Empty alternative
preg_match('/a|/', $input);
// PREFERRED: Use a quantifier
preg_match('/a?/', $input);
Fix: Replace the empty alternative with a quantifier or an explicit empty group if intentional.
Overlapping Alternation Branches
Identifier: regex.lint.alternation.overlap
When it triggers: One literal alternative is a prefix of another.
Visual Explanation:
Pattern: /(a|aa)+b/
Input: "aaaaab"
The engine can split the a's as: a+a+a+a+a or aa+a+a+a, ...
Result: overlapping paths trigger heavy backtracking
Example:
// VULNERABLE: Overlapping branches in repetition
preg_match('/(a|aa)+b/', $input);
// SAFER: Use atomic groups
preg_match('/(?>a|aa)+b/', $input);
// SIMPLER: Just use a+
preg_match('/a+b/', $input); // Often equivalent
Fix: Use atomic groups or simplify the pattern.
Dot or Newline Alternation
Identifier: regex.lint.alternation.dotNewline
When it triggers: An alternation reads (.|\n), the anti-pattern for “any character including newlines”. The dot already matches every character but a newline, so the branch adds nothing but a slower, two-way choice.
Example (real regex lint output):
→ /(.|\n)/
WARN Alternation (.|\n) is an anti-pattern for matching any character including newlines.
↳ Use the "s" (PCRE_DOTALL) flag to make "." match newlines, or use [\s\S] instead of (.|\n).
Fix: Use the s flag so the dot crosses newlines, or write [\s\S].
Overlapping Character Sets
Identifier: regex.lint.overlap.charset
When it triggers: Alternation branches have overlapping character sets, and the alternation is repeated by an unbounded quantifier: a character both branches accept can be taken by either, on every iteration. An alternation matched once, such as /[a-c]|[b-d]/, is not reported, nor is one inside a lookbehind ((?<!http:|https:)): PCRE runs a lookaround atomically, so the loop around it never comes back to try the other branch. A loop at the very end of the pattern is still reported. With the ReDoS analysis on, a pattern the analysis proves linear drops the issue, inline options such as (?s) or (?s:…) included; like the two rules above, the issue is also left out for a pattern listed in the ignored patterns or found trivially safe.
Example:
// WARNING: Overlapping character classes inside a repetition
preg_match('/(?:[a-c]|[b-d])+$/', $input);
// SAFER: Use an atomic group to avoid backtracking
preg_match('/(?>[a-c]|[b-d])+$/', $input);
// IF EQUIVALENT: Merge ranges
preg_match('/[a-d]+$/', $input);
Backreferences
Useless Backreferences
Identifier: regex.lint.backref.useless
When it triggers: A backreference is used before its capturing group can be set or when the group is guaranteed to be empty.
A backreference inside a group that a subroutine call runs again is not reported: the call may run it after the group has captured, as in /^(?:(a)|(\1))(?2)$/. A recursive pattern ((?R), (?0)) is not checked by this rule.
Example:
// WARNING: Backreference appears before the group closes
preg_match('/\1(a)/', $input);
// WARNING: Capturing group is always empty
preg_match('/(\b)a\1/', $input);
// PREFERRED: Move the backreference after the group
preg_match('/(a)\1/', $input);
Fix: Move the backreference after the capturing group or remove it if it adds no constraint.
Undefined Backreference
Identifier: regex.lint.backref.undefined
When it triggers: A numbered or named backreference points to a group the pattern does not have: \2 with a single group, \k<name> with no group of that name. A \NN of two digits or more that names no group and starts with an octal digit is an octal escape instead (\11 is a tab), and is not reported. Validation refuses the same defect first (regex.backref.missing_group, regex.backref.missing_named_group), so regex lint reports the pattern as invalid; the lint id is what the linter reports on a parsed tree.
Example:
// FAIL (validation): Backreference to non-existent group: \2.
preg_match('/(a)\2/', $input);
// PREFERRED: refer to a group that exists
preg_match('/(a)\1/', $input);
Fix: Number or name the group the reference points to, or drop the reference.
Character Classes
Redundant Character Class Elements
Identifier: regex.lint.charclass.redundant
When it triggers: A character class contains redundant elements or overlapping ranges.
Example:
// WARNING: Redundant elements detected in character class.
// ↳ Redundant elements: range 'a'-'z' (overlaps 'a'-'z')
preg_match('/[a-zA-Za-z]/', $input);
// PREFERRED: Remove duplicates
preg_match('/[a-zA-Z]/', $input);
// WARNING: Redundant elements detected in character class.
// ↳ Redundant elements: range 'c'-'d' (covered by range 'a'-'f')
preg_match('/[a-fc-d]/', $input);
// PREFERRED: Use clean ranges
preg_match('/[a-f]/', $input);
// WARNING: Redundant elements detected in character class.
// ↳ Redundant elements: ranges 'a'-'m' and 'k'-'z' overlap: merge them into 'a'-'z'
preg_match('/[a-mk-z]/', $input);
// PREFERRED: One range
preg_match('/[a-z]/', $input);
The hint (↳) names what to change: a repeated range is named with the one it repeats, a range another one covers is named with the range that covers it, and two ranges that only partly overlap, which are both needed as written, come with the range to merge them into.
Fix: Remove duplicates or merge ranges.
Duplicate Character Class Elements
Identifier: regex.lint.charclass.duplicateChars
When it triggers: A character class contains an element that is fully covered by another element (e.g., \d plus 0-9).
Example:
// WARNING: \d already includes 0-9
preg_match('/[\d0-9]/', $input);
// WARNING: \w already includes A-Z
preg_match('/[\wA-Z]/', $input);
// PREFERRED: Remove the redundant element
preg_match('/[\d]/', $input);
preg_match('/[\w]/', $input);
Fix: Remove the redundant character class element.
Useless Character Range
Identifier: regex.lint.range.useless
When it triggers: A range spans only one or two characters (e.g., [a-a] or [a-b]) and can be written explicitly.
Example:
// WARNING: Range only spans one character
preg_match('/[a-a]/', $input);
// WARNING: Range only spans two characters
preg_match('/[a-b]/', $input);
// PREFERRED: List characters explicitly
preg_match('/[a]/', $input);
preg_match('/[ab]/', $input);
Fix: Replace the range with literal characters inside the class.
Suspicious ASCII Ranges
Identifier: regex.lint.charclass.suspiciousRange
When it triggers: A character range spans the ASCII gap between Z and a (e.g. [A-z]), which includes [ \ ] ^ _ .
Example:
// WARNING: Includes non-letters between Z and a
preg_match('/[A-z]/', $input);
// PREFERRED: Use two ranges
preg_match('/[A-Za-z]/', $input);
Fix: Split into [A-Z] and [a-z], or combine as [A-Za-z].
Alternation-like Character Classes
Identifier: regex.lint.charclass.suspiciousPipe
When it triggers: A character class contains | alongside many letters, suggesting an alternation typo.
Example:
// WARNING: | is literal inside []
preg_match('/[error|failure]/', $input);
// PREFERRED: Use alternation
preg_match('/(error|failure)/', $input);
Fix: Replace the class with an alternation group when you intend multi-character words.
Literal Metacharacter in a Character Class
Identifier: regex.lint.charclass.literalMetachar
When it triggers: A small character class holds a shorthand (\w, \d, \s or their negations) next to an unescaped *, + or ?. Inside [...] these are literal characters, so [\w*] matches a word character or a *: the quantifier was likely meant outside the class. It is not reported when the metacharacter is written with a backslash, as \* in [\w\*], which says the literal is meant, in a negated class such as [^\s+], or in a class that lists three elements or more besides it, such as [\w+.-], which builds a set on purpose.
Example:
// WARNING: "*" is a literal character inside a character class, not a quantifier
preg_match('/^[\w*]$/', '*'); // 1: the class matches a "*"
// PREFERRED: Quantify the shorthand
preg_match('/^\w*$/', '*'); // 0
// PREFERRED: Escape the literal with a backslash when it is meant
preg_match('/^[\w\*]$/', '*'); // 1, not reported
Fix: Move the quantifier out of the class, or escape the character with a backslash when the literal is meant.
Backreference Written Inside a Character Class
Identifier: regex.lint.charclass.backrefAsOctal
When it triggers: A class holds \1-\9 while a group of that number exists. Inside a class the escape is octal, not a backreference: [\1] is the byte \x01. Without a group of that number the escape is plain octal and nothing is reported.
Example (real regex lint output):
→ /(a)[\1]/
WARN Suspicious \1 inside character class: this is octal (\x01), not a backreference to group 1.
↳ Backreferences do not work inside character classes. If you intended to match the same text as group 1, move the check outside the class.
Fix: Match the same text as the group outside the class ((a)\1), or write the byte you mean ([\x01]).
Single-Character Class
Identifier: regex.lint.charclass.single (off by default)
When it triggers: A class holds one character, [a], which reads the same without the
class. A style rule: turn it on with "charclass.single": true under checks.lint.rules.
A class holding one metacharacter ([.], [*], and under x, [#] or any white space x
skips: a space, tab, line feed, vertical tab, form feed, carriage return or next line, and
under /u U+200E, U+200F, U+2028 and U+2029) is the clearer escape and is left alone, as are a negated class [^a], the delimiter [\/], an escape a
class reads otherwise ([\b] is a backspace, [\1] an octal escape where \1 is a
reference), a multibyte character without /u, whose bytes [é] reads one at a time, and a
character that would join the text around the class: \1[0] is not \10, a{[2]} not
a{2}. The message quotes the class as written, [\Qa\E]. [aa] is
regex.lint.charclass.redundant’s.
Example:
// INFO: the class adds nothing
preg_match('/x[a]y/', 'xay'); // 1
// PREFERRED
preg_match('/xay/', 'xay'); // 1
Fix: Write the character alone.
Escapes
Suspicious Escapes
Identifier: regex.lint.escape.suspicious
When it triggers: An escape names a value PCRE cannot read: a \x{...} code point past U+10FFFF, an octal escape past \377 in byte mode (\777 without /u), or a \N{name} no character answers to (the name lookup needs intl). \d and the other shorthands are never suspicious, and \8 is not an escape question at all: PCRE reads it as a backreference, and validation refuses it as regex.backref.missing_group.
The same defects are refused at validation first — regex.unicode.out_of_range, regex.octal.out_of_range, regex.escape.unsupported — so regex lint reports the pattern as invalid rather than printing a warning. The lint id is what the linter reports on a parsed tree.
Example (real regex lint output):
→ /\x{110000}/
FAIL Invalid Unicode codepoint "\x{110000}" (out of range).
→ /\777/
FAIL Invalid legacy octal codepoint "\777" (out of range).
Fix: Write a value in range: \x{10FFFF} under /u, \777 under /u, or \377 in byte mode.
Literals
Multiple Spaces
Identifier: regex.lint.literal.multipleSpaces (off by default)
When it triggers: Two or more literal spaces in a row are hard to count; ` {n} says how
many. A style rule: turn it on with “literal.multipleSpaces”: true under
checks.lint.rules. Only spaces written bare are counted: under x or xx, or after
(?x), they are not literal, and quoted (\Q \E), escaped or class spaces are not bare. A
quantifier on the last space enters the count: the tip for /a +/ is {2,}`.
Example:
// INFO: how many spaces?
preg_match('/a b/', 'a b'); // 1
// PREFERRED
preg_match('/a {3}b/', 'a b'); // 1
Fix: Write the space once with its count.
Bytes Without /u
Without /u, PCRE reads the pattern and the subject as bytes. A character written in
UTF-8 that takes several bytes is several items to it, and three places get it wrong: a
multibyte character in a class, a quantifier after one, and a Unicode property such as
\p{L} (regex.lint.unicode.propertyWithoutU), which then covers only the first 256
code points.
The three rules report at error severity: the pattern compiles, but does not do what it
says, so regex lint exits with 1. Turn one off under checks.lint.rules in regex.json
("unicode.multibyteInClassWithoutU": false) when bytes are really meant. Two further
rules watch byte mode: the braced hex escape PCRE2 refuses outright, and — off by
default — the ASCII-only shorthands.
Multibyte Character in a Class
Identifier: regex.lint.unicode.multibyteInClassWithoutU
When it triggers: A character class holds a character of two bytes or more, and the
pattern has neither /u nor (*UTF). The class holds each byte on its own.
Example:
// ERROR: [é] is the class of the bytes \xC3 and \xA9
preg_match('/[é]/', 'à'); // 1: "à" starts with \xC3 too
// PREFERRED: read code points
preg_match('/[é]/u', 'à'); // 0
Fix: Add /u. If bytes are meant, write them as \x escapes so the class says so.
Quantifier After a Multibyte Character
Identifier: regex.lint.unicode.quantifiedMultibyteWithoutU
When it triggers: A quantifier follows a character of two bytes or more, and the
pattern has neither /u nor (*UTF). The quantifier repeats the last byte only.
Example:
// ERROR: + repeats \xA9, the last byte of "é"
preg_match('/^é+$/', 'éé'); // 0
preg_match('/^é+$/', "é\xA9"); // 1
// PREFERRED
preg_match('/^é+$/u', 'éé'); // 1
preg_match('/^(?:é)+$/', 'éé'); // 1, still byte mode
Fix: Add /u, or group the character so the quantifier takes all of it.
Unicode Property Without /u
Identifier: regex.lint.unicode.propertyWithoutU
When it triggers: A Unicode property (\p{…} or \P{…}) is used, and the pattern has
neither /u nor (*UTF). The property then reads one byte at a time, as a code point
below 256, so it covers only the first 256 code points.
Example:
// ERROR: "é" is the bytes \xC3 \xA9, and \xA9 is no letter
preg_match('/^\p{L}+$/', 'é'); // 0
// PREFERRED
preg_match('/^\p{L}+$/u', 'é'); // 1
Fix: Add /u.
Braced Hex Escape Without /u
Identifier: regex.lint.unicode.bracedHexWithoutU
When it triggers: A braced escape \x{...} names a code point above U+FF, and the pattern has neither /u nor (*UTF). Unlike the three rules above, the pattern does not even compile in byte mode — PCRE2 refuses it — so validation reports it first, and the lint id names the same defect on a parsed tree, at error severity.
Example:
// FAIL (validation): Invalid code point "\x{100}":
// without the "u" flag, a character is at most \xFF.
preg_match('/\x{100}/', $input);
// PREFERRED
preg_match('/\x{100}/u', $input);
Fix: Add /u.
Unicode Shorthand Without /u
Identifier: regex.lint.unicode.shorthandWithoutU (off by default)
When it triggers: A \w, \d, \s or one of their negations runs without /u: each matches ASCII only, so é is no \w match. A style rule: turn it on with "unicode.shorthandWithoutU": true under checks.lint.rules, or --enable-rule=unicode.shorthandWithoutU on the command line.
Example (real regex lint output):
→ /\w+/
INFO Shorthand "\w" matches only ASCII without /u flag.
↳ Add /u flag for Unicode support, or use \p{L} for letters.
preg_match('/^\w+$/', 'é'); // 0: é is no ASCII word character
preg_match('/^\w+$/u', 'é'); // 1
Fix: Add /u, or write the property you mean (\p{L}).
Inline Flags
Inline Flag Redundant
Identifier: regex.lint.flag.redundant
When it triggers: An inline flag sets/unsets a modifier already in the desired state.
Example:
// WARNING: Redundant inline flag
preg_match('/(?i)foo/i', $input); // Global i already set
// PREFERRED: Remove redundancy
preg_match('/foo/i', $input);
Inline Flag Override
Identifier: regex.lint.flag.override
When it triggers: An inline flag explicitly unsets a global modifier.
Example:
// WARNING: Unset global flag
preg_match('/(?-i:foo)/i', $input);
// CONSIDER: Scope the flag instead
preg_match('/(?i:foo)bar/', $input);
Pattern Complexity
Complexity Threshold
Identifier: regex.lint.complexity
When it triggers: The analysis’ complexity score for the pattern — the same number Regex::validate() returns as complexityScore — reaches the warning threshold, 50 by default.
Message: Pattern is complex (score: 74). — here for /((a|b)+c?){3}(x|y)*[0-9]{2,4}(p|q|r)+\s*z{5,}/:
use PHPRegex\Toolkit\Regex;
$result = Regex::create()->validate('/((a|b)+c?){3}(x|y)*[0-9]{2,4}(p|q|r)+\s*z{5,}/');
echo $result->complexityScore; // 74
regex lint and the PHPStan rule leave this one out of their report: read the score itself. The Symfony route analysis reports it, with its own threshold under php_regex.analysis.warning_threshold.
Fix: Split the pattern, or factor its repeated parts out.
Security (ReDoS)
Catastrophic Backtracking
Identifier: regex.redos in PHPStan, whatever the severity; the message names it
When it triggers: The ReDoS analyzer proves how one match attempt grows with the input, by following PCRE’s backtracking order, and hands back the input that triggers it. Patterns outside that model (backreferences, conditionals, recursion, an analysis over budget) are judged by heuristics, and the report says which of the two decided. The message prefix names the class: Exponential backtracking (ReDoS), Polynomial backtracking (ReDoS) or Potential backtracking (ReDoS).
Risk Levels:
| Level | Proven class | Action Required |
|---|---|---|
critical |
exponential | Refactor immediately |
high |
polynomial, degree 3 or more | Consider refactoring |
medium |
polynomial, degree 2 (quadratic) | Monitor and plan fix |
low |
heuristic finding only | Accept with logging |
See the ReDoS guide for the guarantee and its limits.
Example:
// VULNERABLE: Exponential backtracking
preg_match('/(a+)+$/', $input); // CRITICAL
// SAFER: Atomic group
preg_match('/(?>a+)+$/', $input); // SAFE
// SAFER: Possessive quantifier
preg_match('/(a++)+$/', $input); // SAFE
Fix: Make the ambiguous part atomic or possessive, or refactor.
Read more:
Quadratic Search
Identifier: regex.lint.redos.search; regex.redos.search in PHPStan, with the message Quadratic search (ReDoS): <pattern>
When it triggers: One match attempt is proven linear, but the search is not anchored: preg_match() starts an attempt at each position of the subject, and on a run of characters each attempt reads to the end of the run before it fails. On the run repeated n times then a breaking character, PCRE2’s interpreter takes a number of steps quadratic in n; preg_match_all(), preg_replace() and preg_split() retry the same way. An anchored alternative does not protect the others: the trim regex /^\s+|\s+$/ is quadratic on "!" . " " x n . "!". pcre.backtrack_limit does not stop it: the limit counts each attempt apart, and trips only when one attempt exceeds it. The JIT may avoid it for some patterns, not for all (see the ReDoS guide).
It runs under the ReDoS check, with no switch of its own. Its severity is that of a proven quadratic attempt, medium, so the default high threshold hides it: --redos-threshold=medium shows it. It is a warning in every mode, and --disable-rule=regex.lint.redos.search or "redos.search": false in the checks.lint.rules of regex.json turns it off.
Example:
// Quadratic search: on " " x n . "!" each attempt reads the rest of the run
preg_match('/\s+$/', $input);
// Linear, and the same answer: the characters \s matches without /u
rtrim($input, " \t\n\v\f\r") !== $input;
The two agree on the characters \s matches without /u in the default C locale. PHP builds PCRE2’s character tables from LC_CTYPE when a script sets another locale, which may change what \s matches; under /u, \s also matches Unicode spaces such as U+00A0 and U+2028, which this rtrim() keeps.
Fix: Anchor the pattern when every match starts at a known place (^, \A, \G or the A modifier), or bound the length of the run, or do the work without a regex (rtrim() for trailing whitespace).
Advanced Syntax
The fixes above lean on three constructs this reference does not teach from scratch:
possessive quantifiers (a++), atomic groups ((?>a+)) and lookarounds
((?=…), (?<=…)). The tutorial covers each in depth —
Quantifiers and Greediness and
Performance and ReDoS for the first two,
Lookarounds and Assertions for the third, and
Backreferences, Subroutines, and Recursion
for the recursive (?R) calls some fixes keep working.
Compatibility & Limitations
Supported PHP Versions
- PHP 8.2 and above
- Uses readonly classes and enum features
- Needs the
mbstringextension.intlis optional: it only lets a lint message name the character a\N{name}escape spells, an escape PCRE refuses anyway; every verdict and every other lint result is the same without it - Reads digits, letters and spaces in ASCII, as PCRE does: the locale of the process changes no verdict
Target Engine
- PCRE2 via PHP’s
preg_*functions
Known Differences from PCRE1
| Feature | PCRE2 | Note |
|---|---|---|
\p{...} Unicode |
Supported | |
Branch reset (?\|...) |
Supported |
Meaning That Changes Across PHP Versions
Identifier: regex.lint.compat.meaningChanges
When it triggers: regex lint judges a project over the PHP range of its composer.json.
A pattern every version of the range accepts may still be parsed into another meaning by a
later one: /a{,3}/ is the text a{,3} under PCRE2 before 10.43 (PHP 8.2 and 8.3) and a
zero to three times from 10.43 (PHP 8.4); /a{ 2 }/ likewise. The warning names the first
version that reads the pattern otherwise, under target in the JSON report. Changes that keep
the parse and move the engine’s semantics are not covered. "compat.meaningChanges": false in
checks.lint.rules turns it off.
Message: From PHP 8.4 (PCRE2 10.44) the pattern parses differently: the same text means something else there.
Fix: Write the part that differs in a form every version reads alike: a{0,3} for the
quantifier, a\{,3} for the text.
Diagnostics Catalog
Every code an exception or a failed validation carries is listed, with its meaning, in Diagnostics: Error Codes. The lint rule ids below are a separate vocabulary: they name advice, not a refused pattern.
Quick Reference Table
Every lint rule reports at warning severity unless the table says otherwise. The rules
SonarPHP users know are mapped to these ids in SonarPHP Regex Rules. A rule of
error severity fails regex lint (exit code 1); a warning is printed and leaves the code
at 0; an info, style included, is printed under an INFO badge and leaves the code at 0.
| Category | Rule ID | Severity | Quick Fix |
|---|---|---|---|
| Flags | regex.lint.flag.useless.s, .m, .i, .D, .x |
warning | Remove the unused flag |
| Anchors | regex.lint.anchor.impossible.start, .end, .boundary |
warning | Move the anchor |
| Anchors | regex.lint.anchor.alternationPrecedence |
warning | Group the alternatives |
| Quantifiers | regex.lint.quantifier.nested, regex.lint.dotstar.nested |
warning | Use atomic groups |
| Quantifiers | regex.lint.quantifier.useless, .zero, .concatenation, .lazyEnd, .assertion |
warning | Simplify the quantifier |
| Quantifiers | regex.lint.quantifier.emptyRepeat, .possessiveImpossible |
warning | Fix the repeated item |
| Quantifiers | regex.lint.quantifier.lazyToClass (off by default) |
perf | Use a negated class |
| Quantifiers | regex.lint.quantifier.uselessLazy (off by default) |
style | Drop the ? |
| Lookarounds | regex.lint.lookaround.edgeQuantifier (off by default) |
perf | Keep the minimum count |
| Groups | regex.lint.group.redundant, .empty |
warning | Remove the group |
| Groups | regex.lint.group.alwaysEmptyCapture |
warning | Narrow what precedes the group |
| Lookarounds | regex.lint.lookaround.impossible |
warning | Fix the lookahead |
| Groups | regex.lint.group.quantifiedCapture |
info for an unnamed group, warning for a named one | Repeat a non-capturing group |
| Alternation | regex.lint.alternation.duplicateDisjunction, .empty, .overlap, .dotNewline |
warning | Simplify or use atomic |
| Alternation | regex.lint.overlap.charset |
warning | Use atomic or merge the sets |
| Backrefs | regex.lint.backref.useless, .undefined |
warning | Move or remove the backreference |
| Character | regex.lint.charclass.redundant, .duplicateChars, .suspiciousRange, .suspiciousPipe, .literalMetachar, .backrefAsOctal |
warning | Clean up the class |
| Character | regex.lint.charclass.single (off by default) |
style | Write the character alone |
| Ranges | regex.lint.range.useless |
warning | Replace with literals |
| Escapes | regex.lint.escape.suspicious |
warning | Fix the escape sequence |
| Literals | regex.lint.literal.multipleSpaces (off by default) |
style | Count the spaces |
| Unicode | regex.lint.unicode.multibyteInClassWithoutU, .quantifiedMultibyteWithoutU, .propertyWithoutU, .bracedHexWithoutU |
error | Add the /u flag |
| Unicode | regex.lint.unicode.shorthandWithoutU (off by default) |
style | Add the /u flag |
| Inline | regex.lint.flag.redundant, .override |
warning | Remove or scope the inline flag |
| Complexity | regex.lint.complexity |
warning | Split the pattern |
| ReDoS | regex.lint.redos (regex.redos in PHPStan) |
warning; error when --redos-mode=confirmed reproduces a verdict at high or above or proves one it cannot replay |
Use possessive quantifiers |
| Targets | regex.lint.compat.meaningChanges |
warning | Write it alike for every PHP |
| ReDoS | regex.lint.redos.search (regex.redos.search in PHPStan) |
warning in every mode; ReDoS severity medium, shown from --redos-threshold=medium |
Anchor the pattern or bound the run |
| Sources | regex.lint.source.unreadable |
error | A source file could not be read, so its patterns were not linted — fix what the message names: raise memory_limit, or exclude the file |