FAQ and Glossary
Short answers to common questions plus quick definitions of core terms used throughout PHPRegex documentation.
Frequently Asked Questions
General Questions
Does PHPRegex execute regexes?
No. PHPRegex parses and analyzes patterns statically. It never actually runs the regex against input. Runtime validation is optional and uses a safe compile check with preg_match().
use PHPRegex\Toolkit\Regex;
// Static analysis - does NOT execute
$analysis = Regex::create()->redos('/(a+)+b/');
echo $analysis->severity->value; // 'critical' - without running it!
// Runtime validation (optional, uses PCRE)
$regex = Regex::create(['runtime_pcre_validation' => true]);
$result = $regex->validate('/test/');
Is this PCRE2-only?
Yes. PHPRegex targets PHP’s preg_* engine, which uses PCRE2. Patterns are validated against PCRE2 semantics.
// PCRE2-specific features work
preg_match('/\p{L}/u', $text); // Unicode properties
// PCRE1 patterns may not work
// preg_match('/\g{0}/', $text); // Invalid in PCRE2
Does this guarantee ReDoS safety?
Within limits. For the patterns its backtracking model covers, PHPRegex proves the cost of one match attempt: safe (proven) means no input makes that attempt backtrack beyond a linear number of steps. The guarantee is per match attempt (the retries of an unanchored search and preg_match_all() are not counted), about the pattern as analysed (see abstractions), and for the PCRE2 release it ran with. Patterns with backreferences, conditionals or recursion are judged by heuristics, and say so. The ReDoS guide lists every limit.
$analysis = Regex::create()->redos('/(a+)+b/');
echo $analysis->headline(); // 'Exponential backtracking (proven)'
echo $analysis->witness->render(); // '"a" x n . "!b"': the input that triggers it
var_dump(Regex::create()->redos('/a+b/')->isProvenSafe()); // bool(true)
// You should still:
// 1. Check preg_* results for false
// 2. Validate input length
// 3. Prefer atomic groups and possessive quantifiers
What is tolerant parsing?
Tolerant parsing returns a partial AST plus errors, allowing tools to continue even when patterns are partially invalid.
use PHPRegex\Toolkit\Regex;
// Strict parsing - throws on error
$ast = Regex::create()->parse('/[broken/'); // Throws LexerException (extends RegexException)
// Tolerant parsing - returns partial AST
$result = Regex::create()->parseTolerant('/[broken/');
echo $result->ast instanceof \PHPRegex\Parser\Node\RegexNode; // true (partial)
echo count($result->errors); // 1
Can I use this in CI?
Yes. PHPRegex is designed for CI/CD integration.
# CLI linting
vendor/bin/regex lint src/ --format=json > regex-issues.json
# Fail on any error-severity issue
if [ "$(jq '[.results[] | .issues[] | select(.severity == "error")] | length' regex-issues.json)" -eq 0 ]; then
echo "No error-severity regex issues found"
else
echo "Error-severity regex issues found!"
exit 1
fi
# GitHub Actions example
- name: Run PHPRegex
run: vendor/bin/regex lint src/ --format=github
The exit code fails the job: 0 clean, 1 errors found, 2 unusable configuration or command line. See
the CLI guide for the report formats and the exit codes.
Why an AST?
The Abstract Syntax Tree provides:
| Benefit | Description |
|---|---|
| Precision | Exact error locations, not just “somewhere in pattern” |
| Analysis | Detect complex issues like ReDoS, proven on a model of PCRE |
| Transformation | Refactor patterns safely without string hacking |
| Tooling | Support IDEs, linters, formatters |
// String-based tools can only guess:
preg_match('/test/', $pattern); // What if 'test' is escaped?
// AST knows the structure:
$ast = Regex::create()->parse('/test/');
$sequence = $ast->pattern; // Exact structure known
Usage Questions
How do I check if a pattern is safe from ReDoS?
use PHPRegex\Toolkit\Regex;
$analysis = Regex::create()->redos('/(a+)+b/');
echo $analysis->severity->value; // 'critical' ('safe', 'low', 'medium', 'unknown', 'high', 'critical')
echo $analysis->headline(); // 'Exponential backtracking (proven)'
echo $analysis->confidenceLevel()->value; // 'medium' ('high' once replayed on PCRE)
echo $analysis->recommendations[0]; // Suggested fix
$analysis->isProvenSafe(); // false: true only for 'safe (proven)'
How do I optimize a pattern?
use PHPRegex\Toolkit\Regex;
$result = Regex::create()->optimize('/[0-9]+/');
echo $result->original; // '/[0-9]+/'
echo $result->optimized; // '/\d+/'
var_export($result->changes);
// array (
// 0 => 'Optimized pattern.',
// )
How do I explain a pattern to users?
use PHPRegex\Toolkit\Regex;
$explanation = Regex::create()->explain('/\d{3}-\d{4}/');
echo $explanation;
/*
Regex matches
Character Type: A digit: [0-9] (exactly 3 times)
'-'
Character Type: A digit: [0-9] (exactly 4 times)
*/
How do I generate a matching sample?
use PHPRegex\Toolkit\Regex;
$sample = Regex::create()->generate('/[A-Z][a-z]{3,5}\d{2}/');
echo $sample; // e.g., "Word12"
Technical Questions
What’s the difference between validate() and parse()?
| Method | Returns | On Error |
|---|---|---|
parse() |
RegexNode (AST) | Throws exception |
validate() |
ValidationResult | Returns result with isValid = false |
use PHPRegex\Toolkit\Regex;
// parse() - throws
try {
$ast = Regex::create()->parse('/[broken/');
} catch (\PHPRegex\Parser\Exception\ExceptionInterface $e) {
echo "Parse failed: {$e->getMessage()}";
}
// validate() - returns result
$result = Regex::create()->validate('/[broken/');
echo $result->isValid ? 'Valid' : "Invalid: {$result->error}";
How does caching work?
use PHPRegex\Toolkit\Regex;
// Default: the latest 1024 trees in memory, nothing on disk
$regex = Regex::create();
// Custom cache location
$regex = Regex::create(['cache' => '/my/app/cache']);
// Disable cache
$regex = Regex::create(['cache' => null]);
// Long-running processes should clear cache
$regex->clearCaches();
What PHP versions are supported?
- PHP 8.2 and above
- Uses modern PHP features (readonly classes, enums, etc.)
// Requires PHP 8.2+
$regex = Regex::create([
'php_version' => '8.2', // Target version
]);
Glossary
| Term | Definition |
|---|---|
| AST | Abstract Syntax Tree - structured representation of the regex |
| Node | Single element in the AST (literal, group, quantifier, etc.) |
| Visitor | Algorithm that traverses the AST (compile, explain, lint) |
| PCRE2 | Perl Compatible Regular Expressions - the engine PHP uses |
| ReDoS | Regular Expression Denial of Service - catastrophic backtracking |
| Backtracking | Engine behavior that retries alternative paths on failure |
| Lookaround | Zero-width assertion like (?=...) or (?<=...) |
| Atomic group | (?>...) - prevents backtracking inside the group |
| Possessive quantifier | *+, ++, {m,n}+ - no backtracking |
| Branch reset | (?\|...) - resets capture numbering per branch |
| Subroutine | (?1) or (?&name) - reuses a group definition |
| Lexer | Tokenizes the pattern string into tokens |
| Parser | Builds the AST from tokens |
| Tokenizer | Same as Lexer |
| Delimiter | Character marking pattern boundaries (e.g., / in /pattern/) |
| Flag | Modifier like i (case-insensitive) or s (dotall) |
| Quantifier | *, +, ?, {m,n} - specifies repetition |
| Greedy | Default quantifier behavior - matches as much as possible |
| Lazy | *?, +?, ?? - matches as little as possible |
| Capturing group | (...) - captures matched text |
| Non-capturing group | (?:...) - groups without capture |
| Named group | (?<name>...) - captures with a name |
| Backreference | \1, \k<name> - refers to previous capture |
| Escape sequence | \d, \w, \x{...} - special character representation |
| Character class | [...] - matches one character from a set |
| Negated class | [^...] - matches any character NOT in the set |
| Shorthand class | \d, \w, \s - common character classes |
| Anchor | ^, $, \A, \z - matches position, not characters |
| Assertion | Zero-width check like \b or (?=...) |
| Word boundary | \b - transition between word and non-word characters |
| Unicode property | \p{L}, \p{N} - characters matching Unicode properties |
Pattern Quick Reference
The regex basics — from . and \d to possessive quantifiers — are taught by the
tutorial, which starts at the atoms of a pattern.