AST Node Reference
This reference documents every node type in the PHPRegex AST. Nodes are the building blocks that represent parsed regex patterns. Understanding nodes is essential for building custom visitors, debugging parsing issues, or extending the library.
How to Read This Reference
Each node includes:
- Purpose — What the node represents in a pattern
- Fields — Public properties you can access
- Example — A PHP code snippet showing how the node is created
- Common Errors — Pitfalls to avoid when working with this node type
If you are new to regex, start with the Tutorial first.
Core Structure Nodes
These nodes form the backbone of every parsed pattern.
RegexNode
Purpose: The root node that wraps the entire pattern, including delimiter, pattern body, and flags.
Fields:
| Field | Type | Description |
|---|---|---|
delimiter |
string | The delimiter character (e.g., /, #, ~) |
pattern |
NodeInterface | The parsed pattern body |
flags |
string | All flags combined (e.g., imsxu) |
Example:
use PHPRegex\Toolkit\Regex;
$ast = Regex::create()->parse('/foo/i');
echo $ast->delimiter; // '/'
echo $ast->flags; // 'i'
echo $ast->pattern; // SequenceNode instance
Common Errors:
// WRONG: Trying to change the pattern directly
$ast->flags = 'g'; // Flags are immutable
// RIGHT: Create a new RegexNode
$newFlags = $ast->flags . 's';
SequenceNode
Purpose: An ordered list of nodes that must match in sequence from left to right.
Fields:
| Field | Type | Description |
|——-|——|————-|
| children | array | Array of child nodes in order |
Example:
use PHPRegex\Toolkit\Regex;
$ast = Regex::create()->parse('/foo/');
$sequence = $ast->pattern;
echo count($sequence->children); // 3
echo $sequence->children[0]->value; // 'f'
echo $sequence->children[1]->value; // 'o'
echo $sequence->children[2]->value; // 'o'
Common Errors:
// WRONG: Modifying children array directly
$sequence->children[] = new LiteralNode('x');
// RIGHT: Create a new SequenceNode
$newChildren = array_merge($sequence->children, [new LiteralNode('x')]);
$newSequence = new SequenceNode($newChildren, $sequence->startPosition, $sequence->endPosition);
AlternationNode
Purpose: Branches separated by the | operator. Matches if any alternative matches.
Fields:
| Field | Type | Description |
|---|---|---|
alternatives |
array | The alternative nodes: a multi-item branch is a SequenceNode, a single atom stands bare |
Example:
use PHPRegex\Toolkit\Regex;
$ast = Regex::create()->parse('/foo|bar|baz/');
$alternation = $ast->pattern;
foreach ($alternation->alternatives as $index => $alt) {
$text = '';
foreach ($alt->children as $child) {
$text .= $child->value;
}
echo "Alternative $index: $text\n";
}
// Output:
// Alternative 0: foo
// Alternative 1: bar
// Alternative 2: baz
Common Errors:
// WRONG: Assuming every alternative is a SequenceNode
// Pattern: /a|b/ holds two bare LiteralNodes, one per alternative;
// only a multi-item branch such as "foo" in /foo|bar/ is a SequenceNode.
foreach ($alternation->alternatives as $alt) {
echo $alt->children[0]->value; // fails on /a|b/: a LiteralNode has no children
}
// RIGHT: Compile an alternative when you need its text
$printer = new \PHPRegex\Parser\Printer\PatternPrinter();
foreach ($alternation->alternatives as $alt) {
echo $alt->accept($printer), "\n"; // 'a', 'b' for /a|b/
}
GroupNode
Purpose: Groups multiple nodes together. Can be capturing, non-capturing, lookaround, atomic, or inline flags.
Fields:
| Field | Type | Description |
|---|---|---|
child |
NodeInterface | The grouped content |
type |
GroupType | The type of group |
name |
string|null | Group name for named groups; for a substring scan, its spelling (scs or scan_substring) |
flags |
string|null | Inline flags (e.g., i in (?i:foo)) |
scannedGroups |
list<string> | For a substring scan, the groups it matches, as written (1, -1, <name>, 'name') |
Group Types:
| Case | Pattern | Description |
|---|---|---|
GroupType::Capturing |
(foo) |
Captures matched text |
GroupType::NonCapturing |
(?:foo) |
Groups without capture |
GroupType::Named |
(?<name>foo) |
Captures with name |
GroupType::LookaheadPositive |
(?=foo) |
Lookahead (matches position) |
GroupType::LookaheadNegative |
(?!foo) |
Negative lookahead |
GroupType::LookbehindPositive |
(?<=foo) |
Lookbehind |
GroupType::LookbehindNegative |
(?<!foo) |
Negative lookbehind |
GroupType::InlineFlags |
(?i:foo) |
Inline flag modification |
GroupType::Atomic |
(?>foo) |
Atomic group (no backtracking) |
GroupType::BranchReset |
(?\|foo\|bar) |
Same group numbers in branches |
GroupType::ScanSubstring |
(*scs:(1)foo) |
Matches its body against what the listed groups captured (PCRE2 10.45) |
Example:
use PHPRegex\Toolkit\Regex;
use PHPRegex\Parser\Node\GroupType;
$ast = Regex::create()->parse('/(?<year>\d{4})-(?<month>\d{2})/');
$group = $ast->pattern->children[0]; // First child is GroupNode
echo $group->type === GroupType::Named; // true
echo $group->name; // 'year'
echo $group->child->quantifier; // '{4}' — the group holds a QuantifierNode over \d
echo $group->child->node->value; // 'd' — the CharTypeNode being quantified
Common Errors:
// WRONG: Assuming lookarounds consume characters
// Pattern: /(?=foo)bar/ matches "bar" at position where "foo" follows
// It does NOT match "foobar"!
$lookahead = $ast->pattern->children[0];
echo $lookahead->type === GroupType::LookaheadPositive; // true
// RIGHT: Lookarounds are zero-width assertions
// They match a POSITION, not characters
QuantifierNode
Purpose: Repeats a node a specified number of times.
Fields:
| Field | Type | Description |
|---|---|---|
node |
NodeInterface | The node being quantified |
quantifier |
string | The quantifier token, without its greed suffix: +, *, {2,4} |
type |
QuantifierType | Greedy, lazy, or possessive — the greed suffix |
The repetition bounds are not fields. QuantifierBounds parses them from the
token — the same way the parser does, for {n}, {n,}, {n,m}, {,m} and
the * / + / ? shorthands alike.
Quantifier Types:
| Case | Pattern | Behavior |
|---|---|---|
QuantifierType::Greedy |
+, *, {m,n} |
Matches as much as possible, then backtracks |
QuantifierType::Lazy |
+?, *?, {m,n}? |
Matches as little as possible |
QuantifierType::Possessive |
++, *+, {m,n}+ |
Matches as much as possible, no backtracking |
Example:
use PHPRegex\Toolkit\Regex;
use PHPRegex\Parser\Node\QuantifierBounds;
use PHPRegex\Parser\Node\QuantifierType;
$ast = Regex::create()->parse('/a{2,4}?/'); // Lazy quantifier
$quantifier = $ast->pattern;
echo $quantifier->quantifier; // '{2,4}' — the lazy '?' lives in the type
echo $quantifier->type === QuantifierType::Lazy; // true
$bounds = QuantifierBounds::parse($quantifier->quantifier);
echo $bounds->min; // 2
echo $bounds->max; // 4
var_dump($bounds->max); // int(4); null when unbounded, as in '{2,}'
Common Errors:
// WRONG: Confusing quantifier with literal
// Pattern: /a+/ contains a QUANTIFIER, not a plus literal
$literal = $ast->pattern;
echo $literal instanceof QuantifierNode; // true
// RIGHT: Access the quantified node
echo $quantifier->node instanceof LiteralNode; // true
echo $quantifier->node->value; // 'a'
Literal and Character Nodes
LiteralNode
Purpose: One literal character — unescaped, or an escaped form of one (\., \-): value holds the character itself.
Fields:
| Field | Type | Description |
|---|---|---|
value |
string | The character itself, one long |
Example:
use PHPRegex\Toolkit\Regex;
$ast = Regex::create()->parse('/hello/');
$literal = $ast->pattern->children[0];
echo $literal->value; // 'h'
Common Errors:
// WRONG: Expecting one LiteralNode per run of text
// The parser emits one LiteralNode per character: /hello/ holds five
// ('h', 'e', 'l', 'l', 'o'), and /he l lo/ holds seven, spaces included.
echo count($ast->pattern->children); // 5 for /hello/
// RIGHT: Concatenate the values when you need the text
$text = '';
foreach ($ast->pattern->children as $child) {
$text .= $child->value;
}
echo $text; // 'hello'
CharLiteralNode
Purpose: Represents a single escaped character with a specific representation (Unicode, octal, hex, etc.).
Fields:
| Field | Type | Description |
|---|---|---|
originalRepresentation |
string | The original escape sequence |
codePoint |
int | The Unicode code point value |
type |
CharLiteralType | The type of escape |
CharLiteralType Values:
| Value | Example | Description |
|---|---|---|
Unicode |
\x{1F600} |
Unicode code point escape (\x{...}, \u{...}, \uFFFF, \xFF) |
UnicodeNamed |
\N{LATIN SMALL LETTER A} |
Named Unicode character |
Octal |
\o{141} |
Octal representation |
OctalLegacy |
\141, \012 |
Legacy octal: up to 3 octal digits, \0 first or a number past the groups opened before it (\101 before group 101 exists is A, \1000 is @ then 0) |
Example:
use PHPRegex\Toolkit\Regex;
$ast = Regex::create()->parse('/\x{1F600}/'); // Grinning face emoji
$char = $ast->pattern;
echo $char->codePoint; // 128512
echo $char->originalRepresentation; // '\x{1F600}'
ControlCharNode
Purpose: A control-character escape \cX: the letter after \c, as written, with the control code point it produces.
Fields:
| Field | Type | Description |
|---|---|---|
char |
string | The letter after \c, as written |
codePoint |
int | The control character it stands for |
Example:
use PHPRegex\Toolkit\Regex;
$ast = Regex::create()->parse('/\cM/');
$ctrl = $ast->pattern;
echo $ctrl->char; // 'M'
echo $ctrl->codePoint; // 13 — \cM is CR
CharTypeNode
Purpose: Character type escapes like \d, \w, \s, and their negations.
Fields:
| Field | Type | Description |
|---|---|---|
value |
string | The type character (d, D, w, W, s, S) |
Type Reference:
| Value | Meaning | Negation |
|---|---|---|
\d |
Digits [0-9] | \D |
\w |
Word characters [a-zA-Z0-9_] | \W |
\s |
Whitespace | \S |
Example:
use PHPRegex\Toolkit\Regex;
$ast = Regex::create()->parse('/\d+/');
$digitClass = $ast->pattern;
echo $digitClass->value; // 'd'
DotNode
Purpose: The dot token . which matches any character except newlines (unless s flag is set).
Example:
use PHPRegex\Toolkit\Regex;
$ast = Regex::create()->parse('/./');
$dot = $ast->pattern;
echo $dot instanceof \PHPRegex\Parser\Node\DotNode; // true
Common Errors:
// WRONG: Assuming . matches newlines
// Without /s flag, . does NOT match \n
preg_match('/./', "\n", $matches); // Match: no
// With /s flag, . matches newlines
preg_match('/./s', "\n", $matches); // Match: yes
AnchorNode
Purpose: Start (^) and end ($) anchors, and string anchors like \A, \z.
Fields:
| Field | Type | Description |
|---|---|---|
value |
string | The anchor character |
Anchor Reference:
| Anchor | Meaning |
|---|---|
^ |
Start of string (or line with m flag) |
$ |
End of string (or line with m flag) |
\A |
Absolute start of string |
\z |
Absolute end of string |
\Z |
End of string, or before trailing newline |
Example:
use PHPRegex\Toolkit\Regex;
$ast = Regex::create()->parse('/^foo$/');
$startAnchor = $ast->pattern->children[0]; // the sequence starts with '^'
$endAnchor = $ast->pattern->children[4]; // after 'f', 'o', 'o'
echo $startAnchor->value; // '^'
echo $endAnchor->value; // '$'
AssertionNode
Purpose: Zero-width assertions that are not anchors, like word boundaries (\b, \B).
Fields:
| Field | Type | Description |
|---|---|---|
value |
string | The assertion character |
Assertion Reference:
| Assertion | Meaning |
|---|---|
\b |
Word boundary (transition between \w and \W) |
\B |
Not a word boundary |
(?=...) |
Positive lookahead (GroupNode) |
(?!...) |
Negative lookahead (GroupNode) |
(?<=...) |
Positive lookbehind (GroupNode) |
(?<!...) |
Negative lookbehind (GroupNode) |
(?*...), (*napla:...) |
Non-atomic positive lookahead: a GroupType::LookaheadPositive GroupNode whose flags is * |
(?<*...), (*naplb:...) |
Non-atomic positive lookbehind: a GroupType::LookbehindPositive GroupNode whose flags is * |
Example:
use PHPRegex\Toolkit\Regex;
$ast = Regex::create()->parse('/\bword\b/');
$assertion = $ast->pattern->children[0];
echo $assertion->value; // 'b'
Common Errors:
// WRONG: Confusing \b with [a-zA-Z_]
// \b is a POSITION assertion, not a character class
preg_match('/\b/', 'word', $matches); // Match: yes (at position 0)
preg_match('/[a-z]/', 'word', $matches); // Match: yes ('w')
// They are different!
Character Class Nodes
CharClassNode
Purpose: Character classes [...] including negated classes [^...]. PHP has no class intersection or subtraction: inside a class, && is two & members and -- a range through -, so [a&&b] matches & and [a--b] is refused as a range out of order.
Fields:
| Field | Type | Description |
|---|---|---|
expression |
NodeInterface | The class content (ranges, characters, escapes) |
isNegated |
bool | True for [^...], false for [...] |
Example:
use PHPRegex\Toolkit\Regex;
$ast = Regex::create()->parse('/[a-z]/');
$class = $ast->pattern;
echo $class->isNegated; // false
echo $class->expression instanceof \PHPRegex\Parser\Node\RangeNode; // true
Common Errors:
// WRONG: [A-z] includes characters between Z and a
// In ASCII: [, \, ], ^, _, `
preg_match('/[A-z]/', '_', $matches); // Match: yes (includes _)
// RIGHT: Use [A-Za-z]
preg_match('/[A-Za-z]/', '_', $matches); // Match: no (no underscore)
RangeNode
Purpose: A range within a character class, like a-z or 0-9.
Fields:
| Field | Type | Description |
|---|---|---|
start |
NodeInterface | The start of range (usually CharLiteralNode) |
end |
NodeInterface | The end of range (usually CharLiteralNode) |
Example:
use PHPRegex\Toolkit\Regex;
$ast = Regex::create()->parse('/[a-z]/');
$range = $ast->pattern->expression;
echo $range->start->value; // 'a'
echo $range->end->value; // 'z'
PosixClassNode
Purpose: POSIX character classes inside character classes, like [[:alpha:]].
Fields:
| Field | Type | Description |
|---|---|---|
class |
string | The POSIX class name |
POSIX Classes:
| Class | Meaning |
|---|---|
[:alpha:] |
Letters |
[:digit:] |
Digits |
[:alnum:] |
Letters and digits |
[:space:] |
Whitespace |
[:upper:] |
Uppercase letters |
[:lower:] |
Lowercase letters |
[:xdigit:] |
Hexadecimal digits |
Example:
use PHPRegex\Toolkit\Regex;
$ast = Regex::create()->parse('/[[:digit:]]/');
$posix = $ast->pattern->expression; // the class content, under the CharClassNode
echo $posix->class; // 'digit'
UnicodePropNode
Purpose: Unicode property escapes like \p{L} (letters) or \P{Lu} (non-uppercase letters).
Fields:
| Field | Type | Description |
|---|---|---|
prop |
string | The property specifier, as written ({L}, L) |
hasBraces |
bool | True for \p{L}, false for \pL |
Common Unicode Properties:
| Property | Meaning |
|---|---|
\p{L} |
Any letter |
\p{Lu} |
Uppercase letter |
\p{Ll} |
Lowercase letter |
\p{N} |
Any number |
\p{P} |
Any punctuation |
\p{S} |
Any symbol |
\p{Z} |
Any separator |
\p{Sc} |
Currency symbol |
Example:
use PHPRegex\Toolkit\Regex;
$ast = Regex::create()->parse('/^\p{L}+$/u'); // Unicode letters only
$prop = $ast->pattern->children[1]->node; // the \p{L} under its '+' quantifier
echo $prop->prop; // '{L}' — kept as written; bare '\pL' would give 'L'
echo $prop->hasBraces; // true
ExtendedCharClassNode
Purpose: A Perl extended class, (?[ \p{L} - [aeiou] ]): a set expression
over classes, read for a target on PCRE2 10.45 or later. It matches one
character, like a character class.
Fields:
| Field | Type | Description |
|---|---|---|
expression |
NodeInterface | An operand, or a ClassSetOperationNode tree |
text |
string | The class as written, (?[ to ]); empty for a node built by hand |
Operands are the nodes a class holds: a CharClassNode, a PosixClassNode
([:alpha:]), a CharTypeNode (\d), a UnicodePropNode (\p{L}) or a
character (\n, \x61, \!). An escape is read as inside a class, so \b
is a backspace and \1 an octal escape; a nested class is read as under xx,
where spaces and tabs are skipped. A plain character is no operand.
ClassSetOperationNode
Purpose: One set operation inside an extended class.
Fields:
| Field | Type | Description |
|---|---|---|
operator |
ClassSetOperator | The operation |
left |
NodeInterface|null | The left operand; null for a complement |
right |
NodeInterface | The right operand |
symbol |
string | The operator as written (+ or \| for a union) |
Operators, ! binding tightest, then &, then the others left to right:
| Operator | Syntax |
|---|---|
ClassSetOperator::Complement |
!a |
ClassSetOperator::Intersection |
a & b |
ClassSetOperator::Union |
a + b, a \| b |
ClassSetOperator::Difference |
a - b |
ClassSetOperator::SymmetricDifference |
a ^ b |
Example:
use PHPRegex\Toolkit\Regex;
$ast = Regex::create(['pcre_version' => '10.45'])->parse('/(?[ \p{L} - [aeiou] ])/u');
$difference = $ast->pattern->expression;
echo $difference->operator->name; // 'Difference'
Group and Reference Nodes
BackrefNode
Purpose: Backreference to a previously captured group, like \1 or \k<name>.
Fields:
| Field | Type | Description |
|---|---|---|
ref |
string | The group number or name, as written (\1, \k<name>) |
Example:
use PHPRegex\Toolkit\Regex;
$ast = Regex::create()->parse('/(\w+)\1/'); // Match doubled word
$backref = $ast->pattern->children[1];
echo $backref->ref; // '\1' — the reference keeps its original spelling
The same node serves as a conditional’s group test, where ref holds the bare
number: in /(?(1)b|c)/ the condition is a BackrefNode whose ref is '1'.
Common Errors:
// WRONG: Backreference before capture
// Pattern: /\1(\w)/ is invalid - no group 1 yet
preg_match('/\1(\w)/', 'a', $matches); // Error or no match
// RIGHT: Reference after capture
preg_match('/(\w)\1/', 'aa', $matches); // Match: yes
ConditionalNode
Purpose: Conditional pattern (?(condition)yes|no) that matches different things based on a condition.
Fields:
| Field | Type | Description |
|---|---|---|
condition |
NodeInterface | The condition (usually BackrefNode) |
yes |
NodeInterface | Pattern if condition is true |
no |
NodeInterface | Pattern if condition is false (optional) |
A group test, (?(1)...) or (?(<n>)...), is a BackrefNode; a recursion
test, (?(R)...), (?(R2)...) or (?(R&n)...), is a SubroutineNode. As in
PCRE2, (?(R)...) and (?(R2)...) test a group instead when one is named R
or R2, before or after the condition: /(?<R2>a)(?(R2)b|c)/ holds a
BackrefNode whose ref is 'R2'.
Example:
use PHPRegex\Toolkit\Regex;
$ast = Regex::create()->parse('/(a)?(?(1)b|c)/'); // If 'a' captured, expect 'b'; else expect 'c'
$conditional = $ast->pattern->children[1]; // after the quantified (a)?
echo $conditional->condition instanceof \PHPRegex\Parser\Node\BackrefNode; // true
echo $conditional->condition->ref; // '1'
echo $conditional->yes->value; // 'b'
VersionConditionNode
Purpose: The PCRE2 version condition, (?(VERSION=10.44)...) or (?(VERSION>=10.44)...): the engine tests its own version before choosing a branch.
Fields:
| Field | Type | Description |
|---|---|---|
operator |
string | = or >= |
version |
string | The version compared against |
Example:
use PHPRegex\Toolkit\Regex;
$ast = Regex::create()->parse('/(?(VERSION>=10.44)a|b)/');
$condition = $ast->pattern->condition; // the ConditionalNode's condition
echo $condition->operator; // '>='
echo $condition->version; // '10.44'
SubroutineNode
Purpose: Subroutine call to reuse a capture group pattern, like (?1) or (?&name).
Fields:
| Field | Type | Description |
|---|---|---|
reference |
string | The group number or name |
syntax |
string | The call marker: '' for (?1), '&' for (?&name), 'g' for \g{...}, 'P>' for (?P>name) |
Example:
use PHPRegex\Toolkit\Regex;
$ast = Regex::create()->parse('/(?<paren>\((?:[^()]++|(?&paren))*\))/');
// Match balanced parentheses using recursion
$subroutine = $ast->pattern // the (?<paren>...) group
->child // its sequence: \( ... \)
->children[1] // the quantified (?:...)*
->node // the non-capturing group
->child // the alternation [^()]++ | (?&paren)
->alternatives[1]; // the (?&paren) call
echo $subroutine->reference; // 'paren'
echo $subroutine->syntax; // '&', the (?&name) call marker
DefineNode
Purpose: (?DEFINE...) block that defines subpatterns without matching them.
Fields:
| Field | Type | Description |
|---|---|---|
content |
NodeInterface | The defined subpatterns |
Example:
use PHPRegex\Toolkit\Regex;
$ast = Regex::create()->parse('/(?(DEFINE)(?<digit>\d+)(?<number>\g<digit>))/');
$define = $ast->pattern;
echo $define->content instanceof \PHPRegex\Parser\Node\SequenceNode; // true
Advanced Nodes
PcreVerbNode
Purpose: PCRE verbs like (*ACCEPT), (*FAIL), (*SKIP) that control matching behavior.
Fields:
| Field | Type | Description |
|---|---|---|
verb |
string | The verb name |
PCRE Verbs:
| Verb | Meaning |
|---|---|
(*ACCEPT) |
Force match success |
(*FAIL) |
Force match failure |
(*SKIP) |
Skip and remember position |
(*COMMIT) |
No backtracking on failure |
(*PRUNE) |
Prune backtracking stack |
(*THEN) |
Jump to alternation branch |
Example:
use PHPRegex\Toolkit\Regex;
$ast = Regex::create()->parse('/foo(*FAIL)bar/');
$verb = $ast->pattern->children[3]; // after 'f', 'o', 'o'
echo $verb->verb; // 'FAIL'
LimitMatchNode
Purpose: (*LIMIT_MATCH=...) verb that sets a protective limit against runaway backtracking.
Fields:
| Field | Type | Description |
|---|---|---|
limit |
int | The match limit |
Example:
use PHPRegex\Toolkit\Regex;
$ast = Regex::create()->parse('/(*LIMIT_MATCH=1000)foo/');
$limit = $ast->pattern->children[0];
echo $limit->limit; // 1000
ScriptRunNode
Purpose: (*script_run:...), short (*sr:...), which matches its body only when every character it matches belongs to one script. (*atomic_script_run:...), short (*asr:...), is the same with an atomic body: (*sr:(?>...)).
Fields:
| Field | Type | Description |
|---|---|---|
script |
string | The body as written |
content |
NodeInterface|null | The parsed body; its positions count from its start |
atomic |
bool | Whether the body is atomic ((*asr:...)) |
Example:
use PHPRegex\Toolkit\Regex;
$ast = Regex::create()->parse('/(*asr:\d+)/');
$run = $ast->pattern;
echo $run->script; // '\d+'
var_dump($run->atomic); // bool(true)
KeepNode
Purpose: The \K keep assertion: everything matched before it stays out of the reported match, which starts at this position instead. It is zero-width, and does not change the length range of a match (see LengthRangeCalculator).
Fields: none of its own — every node carries startPosition and endPosition.
Example:
use PHPRegex\Toolkit\Regex;
$ast = Regex::create()->parse('/foo\Kbar/');
$keep = $ast->pattern->children[3]; // after 'f', 'o', 'o'
var_dump($keep instanceof \PHPRegex\Parser\Node\KeepNode); // bool(true)
CommentNode
Purpose: A comment in the pattern: a (?#...) group, or a # ... line under the x flag.
Fields:
| Field | Type | Description |
|---|---|---|
comment |
string | The comment text, as written |
extended |
bool | True for a # ... line under x, false for (?#...) |
Example:
use PHPRegex\Toolkit\Regex;
$ast = Regex::create()->parse('/a(?# a note)b/');
$comment = $ast->pattern->children[1];
echo $comment->comment; // ' a note'
var_dump($comment->extended); // bool(false)
$ast = Regex::create()->parse('/a # trailing note/x');
$comment = $ast->pattern->children[1];
echo $comment->comment; // '# trailing note'
var_dump($comment->extended); // bool(true)
CalloutNode
Purpose: A callout — (?C), (?C1) or (?C"name") — where PCRE2 hands control to the embedder’s callback mid-match.
Fields:
| Field | Type | Description |
|---|---|---|
identifier |
int|string|null | The callout number or name, null for (?C) |
isStringIdentifier |
bool | Whether the identifier was a quoted name |
Example:
use PHPRegex\Toolkit\Regex;
$ast = Regex::create()->parse('/a(?C1)b(?C"tag")/');
$numbered = $ast->pattern->children[1];
var_dump($numbered->identifier); // int(1)
var_dump($numbered->isStringIdentifier); // bool(false)
$named = $ast->pattern->children[3];
var_dump($named->identifier); // string(3) "tag"
var_dump($named->isStringIdentifier); // bool(true)
Supporting Types
NodeInterface
Purpose: The base interface that all nodes implement. Defines the visitor pattern entry point.
interface NodeInterface
{
public function accept(NodeVisitorInterface $visitor): mixed;
}
AbstractNode
Purpose: Base class for nodes with position tracking. All major nodes extend this.
Fields:
| Field | Type | Description |
|---|---|---|
startPosition |
int | Offset in pattern where node begins |
endPosition |
int | Offset in pattern where node ends |
Position Reference:
Positions are byte offsets into the pattern body — the text between the
delimiters, flags excluded — and they are 0-based. endPosition is exclusive:
one past the node’s last byte.
use PHPRegex\Toolkit\Regex;
$ast = Regex::create()->parse('/foo/');
$literal = $ast->pattern->children[0];
echo $literal->startPosition; // 0 — the body 'foo' starts at 0
echo $literal->endPosition; // 1 — one past the node's single byte
echo $ast->pattern->children[2]->endPosition; // 3, the body's length
Quick Reference Table
| Node | Purpose | Key Fields |
|---|---|---|
RegexNode |
Root of AST | delimiter, pattern, flags |
SequenceNode |
Ordered list | children[] |
AlternationNode |
Branches | alternatives[] |
GroupNode |
Grouping | child, type, name, flags |
QuantifierNode |
Repetition | node, quantifier, type |
LiteralNode |
Literal char | value |
CharLiteralNode |
Escaped char | codePoint, type |
ControlCharNode |
\cX escape |
char, codePoint |
CharTypeNode |
Type escape | value |
DotNode |
Any char | (none) |
AnchorNode |
Position anchor | value |
AssertionNode |
Assertion | value |
KeepNode |
\K keep |
(none) |
CharClassNode |
Character set | expression, isNegated |
RangeNode |
Range in class | start, end |
PosixClassNode |
[:alpha:] etc. |
class |
UnicodePropNode |
\p{...} |
prop, hasBraces |
ExtendedCharClassNode |
Extended class | expression, text |
ClassSetOperationNode |
Set operation | operator, left, right |
BackrefNode |
Backreference | ref |
ConditionalNode |
Conditional | condition, yes, no |
VersionConditionNode |
VERSION test | operator, version |
SubroutineNode |
Group call | reference, syntax |
DefineNode |
(?DEFINE...) |
content |
CommentNode |
Comment | comment, extended |
CalloutNode |
(?C...) |
identifier, isStringIdentifier |
PcreVerbNode |
(*VERB) |
verb |
LimitMatchNode |
(*LIMIT_MATCH=) |
limit |
ScriptRunNode |
(*sr:...) |
script, content, atomic |
Summary
Understanding nodes is essential for working with the AST directly. Key takeaways:
- Nodes are immutable — create new instances to transform
- Positions are important — 0-based byte offsets into the pattern body, preserved for diagnostics
- Literals are per character — one
LiteralNodeper character, escaped or not - GroupNode is versatile — handles many group types
- QuantifierNode has types — greedy, lazy, possessive; bounds come from
QuantifierBounds - Character classes are complex — can contain ranges, operations, POSIX classes