@lanexio/parser-core
This page documents @lanexio/parser-core, the foundational package that defines the shared buffer protocol, tree data structures, and parser host for all Lanexio Parser packages.
- Version: Stable
- Module name:
parser-core - Package:
@lanexio/parser-core - Import path:
@lanexio/parser-core - Layer: 1 (Core)
- Runtime: Universal (browser, server, edge worker, test runner)
- Module format: ESM
- Stability: Stable
- Primary use case: Build and traverse zero-copy flat ASTs, register grammars, use core primitives.
Layer contract
Section titled “Layer contract”When to use this module
Section titled “When to use this module”- You need to create, traverse, or modify a
LexTreedirectly. - You want to use the grammar registry to manage multiple grammars.
- You need to build a custom grammar or parser integration.
- You need the incremental editing API (applyEdit, reparse) or streaming parse (createParseStream).
- You want the shared buffer protocol constants for value-level interop.
Module boundary
Section titled “Module boundary”| Boundary | Description |
|---|---|
| Inputs | Uint8Array (source bytes for parse), or preconstructed buffers for tree manipulation |
| Outputs | Uint8Array (edited source), LexTree, LexNode, LexCursor |
| Side effects | None |
| Determinism | Yes (same bytes and same protocol version produce same tree) |
| External dependencies | None |
| Never-throw guarantee | Yes for parse() |
| Security surface | None (no HTML output, no I/O) |
Installation
Section titled “Installation”-
Install the package.
Terminal window pnpm add @lanexio/parser-coreTerminal window npm install @lanexio/parser-coreTerminal window yarn add @lanexio/parser-core -
Import the named export.
import { parse, LexTree, LexNode, LexCursor } from '@lanexio/parser-core';
Peer dependencies
Section titled “Peer dependencies”This package has no peer dependencies. It is the foundation that all other packages depend on.
Basic Usage
Section titled “Basic Usage”Headless usage (pure TypeScript)
Section titled “Headless usage (pure TypeScript)”import { parse, LexNode, LexCursor, LexTree } from '@lanexio/parser-core';
const encoder = new TextEncoder();const bytes = encoder.encode('toy input');
const tree: LexTree = parse(bytes);console.log(tree.nodeCount);console.log(tree.root);Traverse a tree with cursor
Section titled “Traverse a tree with cursor”const tree = parse(encoder.encode('example'));const cursor = tree.cursor();
// Preorder DFS: descend into the first child when possible, otherwise move to// the next sibling or climb back toward the root.walk: for (;;) { const node = cursor.current; console.log('kind:', node.kind, 'range:', node.range); if (cursor.gotoFirstChild()) continue; for (;;) { if (cursor.gotoNextSibling()) continue walk; if (!cursor.gotoParent()) break walk; }}Exports
Section titled “Exports”| Export | Type | Description |
|---|---|---|
parse | (source: Uint8Array) => LexTree | Parse toy-grammar bytes into a flat AST. Never throws. |
LexTree | class | Root handle for a zero-copy flat AST backed by a single ArrayBuffer. |
LexNode | class | Reference to a single 16-byte node within a LexTree. |
LexCursor | class | Preorder DFS cursor over a LexTree, backed by a node index stack. |
PROTOCOL_VERSION | number | Current shared buffer protocol version (value: 1). |
LEX_TREE_MAGIC | number | Magic number identifying a valid tree header. |
LEX_TREE_VERSION | number | Current tree layout version. |
createLexErrorTree | (source: Uint8Array, diagnostics?: readonly string[]) => LexTree | Create a tree with a single root LexError node. |
applyEdit | (source: Uint8Array, edit: LexEdit) => Uint8Array | Apply a byte-range edit to a source buffer and return the edited bytes. |
reparse | (tree: LexTree, edit: LexEdit, oracle?: ReuseOracle, options?: ReparseOptions) => LexTree | Incrementally reparse after an edit; falls back to a full reparse without an oracle. Never throws. |
reparseWithStats | (tree: LexTree, edit: LexEdit, oracle: ReuseOracle, options?: ReparseOptions) => { tree: LexTree; stats: ReparseStats } | Reparse and return subtree-reuse statistics. |
computeDirtyRegion | (tree: LexTree, edit: LexEdit, oracle: ReuseOracle) => DirtyRegion | Compute the dirty re-parse window for an edit using a reuse oracle. |
createParseStream | (grammar: string, parseFn: (bytes: Uint8Array) => LexTree) => LexParseStream | Create a streaming push parser bound to a grammar label and a never-throwing parse function. |
grammarRegistry | object | Global registry for grammar registrations. |
embedGuests | (hostTree: LexTree, options?: EmbedGuestsOptions) => LexTree | Embed guest-language subtrees into host tree. |
embedRegistry | object | Registry for guest-language embeddings. |
graftSubtree | (hostTree: LexTree, parentIndex: number, insertBeforeIndex: number, guestRecords: Uint32Array, metadata?: LexTreeMetadata) => LexTree | Splice packed guest node records into a host tree under a parent node. |
findParentIndex | (records: Uint32Array, nodeCount: number, index: number) => number | null | Find the parent node index of a child in a packed records array, or null. |
fixSubtreeSizes | (records: Uint32Array, startIndex: number, delta: number) => void | Add delta to subtree_size of every ancestor of startIndex, in place. |
rebaseRecords | (records: Uint32Array, delta: number) => Uint32Array | Offset all byte ranges in a node record array. |
assertLossless | (source: Uint8Array, leaves: ReadonlyArray<LeafRange>) => Uint8Array | Assert leaf ranges tile the source byte-for-byte; returns the concatenated leaf bytes. |
emitSource | (tree: LexTree) => Uint8Array | Reconstruct the source bytes from the tree. |
extractLeafRanges | (tree: LexTree) => LeafRange[] | Extract leaf-level source byte ranges. |
countFalseHasError | (tree: LexTree, knownErrorKinds: ReadonlySet<number>) => number | Count non-Error leaf nodes whose flags collide with the reserved error bit. |
countInvertedRanges | (tree: LexTree) => { count: number; firstIndex: number | null } | Count nodes with start > end ranges and report the first offender. |
assertWellFormed | (tree: LexTree) => string[] | Validate all flat-AST invariants; returns violation descriptions (empty when well-formed). |
LexEditError | class | Error thrown on invalid edits. |
LexEditErrorCode | const object | Error code constants for edit failures. |
LosslessError | class | Error thrown when lossless integrity check fails. |
NODE_STRIDE | number | Stride (in Uint32 slots) per node record. |
SLOT_KIND | number | Slot index for node kind. |
SLOT_FLAGS | number | Slot index for node flags. |
SLOT_FIELD | number | Slot index for node field id. |
SLOT_START | number | Slot index for byte range start. |
SLOT_END | number | Slot index for byte range end. |
SLOT_SIZE | number | Slot index for subtree size. |
createTree | (source: Uint8Array, nodes: ReadonlyArray<PendingNode>, metadata?: LexTreeMetadata) => LexTree | Create a tree from pending node descriptions. |
createTreeFromUint32Array | (nodeData: Uint32Array, nodeCount: number, source: Uint8Array, metadata?: LexTreeMetadata) => LexTree | Create a tree from packed Uint32Array node records (NODE_STRIDE slots per node). |
growNodeData | (nd: Uint32Array, needed: number) => Uint32Array | Grow the node data buffer. |
writeTreeHeader | (records: Uint32Array, nodeCount: number, offset?: number) => void | Write the standard tree header into a Uint32Array at the given word offset. |
createPendingNode | (kind: number, flags: number, fieldId: number, rangeStart: number, rangeEnd: number) => PendingNode | Create a pending leaf node for tree construction. |
LANEXIO_PARSER_CORE_PACKAGE_NAME | string | Stable npm package name constant. |
LexToyKind | const object | Kind constants for the built-in toy grammar. |
LexToyField | const object | Field constants for the built-in toy grammar. |
toValue | (target: LexTree | LexNode, options?: ToValueOptions) => LexValueResult | Project a tree or node subtree into a JSON-expressible host value. Never throws. |
registerValueExtractor | (language: string, extractor: LexValueExtractor) => void | Register a value extractor under a case-insensitive language name. |
valueExtractorRegistry | ValueExtractorRegistry | Singleton extractor registry (case-insensitive, last registration wins). |
findSubtreeError | (node: LexNode) => LexRange | null | Return the first error byte range inside a node’s subtree, or null. |
rebuildTreeWithMetadata | (original: LexTree, extra: Partial<LexTreeMetadata>) => LexTree | Return a new LexTree view over the same buffer with merged metadata. |
walk | (tree: LexTree, visitors: Visitor, opts?: WalkOptions) => void | Preorder/postorder visitor walk over a flat AST. Returning false from enter prunes the subtree. |
transform | (tree: LexTree, visitor: TransformVisitor) => TransformResult | Rewrite a document in one ascending, non-overlapping pass over a copy of the source. Never mutates the tree. |
TransformSkipReason | const object | Stable machine-readable skip reasons (Overlap, Descending, InvalidRange). |
createPipeline | (inputs: ReadonlyArray<GrammarRegistration | Plugin>) => Pipeline | Compose grammar registrations and plugins into one parse/transform pipeline over the shared registry. |
assertDomWellFormed | (tree: LexTree) => string[] | Assert the ADR-0016 DOM-shape flat-AST invariants; returns violation strings (empty when well-formed). |
repairContainerRanges | (records: Uint32Array, nodeCount: number) => void | Extend container node ranges to enclose all children, in place. |
ValueExtractorRegistry | class | Case-insensitive value-extractor registry (register / byLanguage). |
LEX_TOY_FIELD_NAMES_BY_ID | Readonly<Record<number, string>> | Field-id to field-name map for the built-in toy grammar. |
LEX_NODE_HAS_ERROR | number | Node flag (bit 0) set when a node or its subtree contains a parse error (value: 1). |
GRAMMAR_FLAG_NODE_ERROR | number | Companion flag (bit 6) marking intentional error propagation (value: 64). |
NODE_RECORD_SIZE | number | Size in bytes of one flat AST node record (value: 16). |
TREE_HEADER_SIZE | number | Size in bytes of the tree header (value: 16). |
NODE_KIND_OFFSET | number | Node byte offset of the kind field (value: 0). |
NODE_FLAGS_OFFSET | number | Node byte offset of the flags field (value: 2). |
NODE_FIELD_ID_OFFSET | number | Node byte offset of the field id field (value: 3). |
NODE_RANGE_START_OFFSET | number | Node byte offset of the range start field (value: 4). |
NODE_RANGE_END_OFFSET | number | Node byte offset of the range end field (value: 8). |
NODE_SUBTREE_SIZE_OFFSET | number | Node byte offset of the subtree size field (value: 12). |
decodeUtf8CodePoint | (bytes: Uint8Array, offset: number) => { codePoint: number; width: number } | Decode one UTF-8 code point at offset. |
isLegalXmlChar | (codePoint: number) => boolean | True when a code point is a legal XML 1.0 character. |
isNameStartCp | (codePoint: number) => boolean | True when a code point is an XML 1.0 NameStartChar. |
isNameCharCp | (codePoint: number) => boolean | True when a code point is an XML 1.0 NameChar. |
isAsciiNameStart | (byte: number) => boolean | Fast ASCII-only NameStartChar check. |
isAsciiNameChar | (byte: number) => boolean | Fast ASCII-only NameChar check. |
isWs | (byte: number) => boolean | True for space, tab, LF, CR. |
isDecDigit | (byte: number) => boolean | True for 0-9. |
isHexDigit | (byte: number) => boolean | True for 0-9, a-f, A-F. |
Options
Section titled “Options”The core module accepts options through LexTreeOptions when constructing trees programmatically. LexTreeOptions has a single optional field, metadata, which carries up to three optional sub-fields.
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
metadata | LexTreeMetadata | No | {} | Optional grammar field-name map, resolved language name, and diagnostics. |
metadata.fieldNamesById | Readonly<Record<number, string>> | No | undefined | Grammar-specific field-id to field-name map. |
metadata.language | string | No | undefined | Resolved language name (set by the unified parse() dispatcher). |
metadata.diagnostics | readonly string[] | No | undefined | Actionable fix messages surfaced on error trees. |
Return shape
Section titled “Return shape”The primary return value from parse() is LexTree:
| Property | Type | Description |
|---|---|---|
root | LexNode | Root node of the flat AST. |
nodeCount | number | Total nodes in the tree. |
source | Uint8Array | Original parsed bytes. |
cursor() | () => LexCursor | Create a new preorder DFS cursor starting at root. |
Exported types
Section titled “Exported types”| Type | Purpose | Notes |
|---|---|---|
LexTreeOptions | Options for programmatic tree construction | Single optional metadata field; see Options |
LexTreeMetadata | Immutable tree metadata passed at construction | { fieldNamesById?, language?, diagnostics? }, all optional |
LexRange | readonly [start: number, end: number] half-open byte offset tuple | Used across all grammar APIs |
LexEdit | Structural edit descriptor | Passed to applyEdit |
DirtyRegion | Byte range affected by an edit | Returned by computeDirtyRegion |
ReparseOptions | Options for incremental reparse | Passed to reparse |
ReparseStats | Performance statistics from reparse | Returned by reparseWithStats |
ReuseOracle | Oracle for incremental node reuse | Used by grammar incremental modules |
LexParseStream | Streaming push parser interface | Created by createParseStream |
LexTreeValidationCode | Validation result code | Produced by integrity checks |
LexTreeValidationError | Error from tree validation | Thrown on invalid tree structure |
LeafRange | Leaf-level source range | Used by extractLeafRanges |
LosslessDiagnostics | Diagnostic info from lossless check | Returned by lossless operations |
PendingNode | Node pending insertion | Used during tree construction |
LanexioParserPureGrammar | Pure-TS grammar descriptor interface | Used by grammar packages to register themselves |
GrammarRegistration | Full grammar registration descriptor | Used with grammarRegistry |
EmbedGuestsOptions | Options for guest embedding | Passed to embedGuests |
EmbedRule | Guest embedding rule | Used in embed registry |
LexValue | JSON-expressible host value | Union of string, number, boolean, null, object, array |
LexValueObject | Projected object value | { readonly [key: string]: LexValue } |
LexValueArray | Projected array value | readonly LexValue[] |
LexValueResult | Result of one toValue call | { value: LexValue | undefined; diagnostics: readonly LexValueDiagnostic[] } |
LexValueDiagnostic | Structured value-extraction failure | { code: LexValueDiagnosticCode; rangeStart; rangeEnd; message } |
LexValueDiagnosticCode | Stable machine-readable failure code | "has_error", "unresolved_language", "unsafe_integer", "overflow", "duplicate_key", "missing_anchor", "circular_alias", "max_depth", "excess_field", "empty_document" |
LexValueContext | Recursion context shared with an extractor | { maxDepth: number } |
LexValueExtractor | Projection contract implemented by grammar packs | (node: LexNode, context: LexValueContext) => LexValueResult |
ToValueOptions | Options for toValue | { language?, extractor?, maxDepth? } |
Visitor | Callbacks run over every node by walk | { enter?: (node) => void | boolean; leave?: (node) => void }; enter returning false prunes |
WalkOptions | Options for walk | { cursor?: boolean } selects cursor-driven navigation |
TransformVisitor | Callback for transform | (node: LexNode) => string | Uint8Array | null | undefined; returning text replaces the node’s range |
Edit | One applied replacement | { range: LexRange; replacement: Uint8Array } |
SkippedEdit | One skipped candidate edit | { range: LexRange; reason: string } with a stable machine-readable reason |
TransformResult | Result of one transform call | { source: Uint8Array; applied: ReadonlyArray<Edit>; skipped: ReadonlyArray<SkippedEdit> } |
TransformSkipReasonType | Union of skip-reason strings | "overlaps-applied-edit", "descending-edit", "invalid-range" |
Plugin | A pipeline plugin | { name: string; onParse?: (tree, ctx) => LexTree | null | undefined; transform?: (tree) => TransformResult | null | undefined } |
Pipeline | A pipeline returned by createPipeline | parse / transform / walk / register / grammars() |
PipelineContext | Context handed to an onParse hook | { language: string }, the resolved canonical language |
PipelineParseOptions | Options for Pipeline.parse | { language?: string; filename?: string } |
LexToyKindType | Type alias for LexToyKind | Same keys as the LexToyKind const |
LexToyFieldType | Type alias for LexToyField | Same keys as the LexToyField const |
LexTreeValidationCodeType | Type alias for LexTreeValidationCode | Same codes as the LexTreeValidationCode const |
Configuration and Extension
Section titled “Configuration and Extension”Grammar registry
Section titled “Grammar registry”The grammarRegistry manages grammar registrations globally. Register a grammar to make it available for auto-detection and unified parsing.
import { grammarRegistry } from '@lanexio/parser-core';import { htmlRegistration } from '@lanexio/parser-grammar-html';
grammarRegistry.register(htmlRegistration);Embedding guests
Section titled “Embedding guests”embedGuests inserts guest-language subtrees (for example, a CSS block inside an HTML style element) into a host tree. The embedRegistry stores the embedding rules.
Streaming parse
Section titled “Streaming parse”createParseStream(grammar, parseFn) creates a LexParseStream that accepts chunked input via push(), finalizes with end(), and yields one tree per push through its async iterator:
import { createParseStream, parse } from '@lanexio/parser-core';
const encoder = new TextEncoder();const stream = createParseStream('@lanexio/parser-core', (bytes) => parse(bytes));stream.push(encoder.encode('chunk1 '));stream.push(encoder.encode('chunk2'));stream.end();
for await (const tree of stream) { console.log(tree.nodeCount);}Value extraction
Section titled “Value extraction”toValue() projects a parse tree (or any node subtree) into a JSON-expressible host value without walking the AST by hand. It is a pure read of the frozen flat AST, never throws, and resolves a grammar-specific extractor from the tree’s metadata.language (or an explicit language / extractor option). Extractors are registered by @lanexio/parser/all in one place; grammar packs do not self-register. Each pack exports its extractor and a <grammar>ToValue convenience function that presets it, so direct grammar-pack users need no registration step.
import { toValue } from '@lanexio/parser';import { parse } from '@lanexio/parser/all';
const tree = parse('{"a": 1, "b": [true, null]}', { language: 'json' });const { value, diagnostics } = toValue(tree);
console.log(value); // { a: 1, b: [true, null] }console.log(diagnostics); // []@lanexio/parser/all registers every grammar and extractor but does not export toValue, hence the split import.
See the value-extraction reference for the LexValue model, the diagnostic codes, and the per-grammar coercion rules.
Plugin and visitor API
Section titled “Plugin and visitor API”The plugin API is the grammar-agnostic composition surface of Lanexio™ Parser:
walk visits a flat AST, transform rewrites its source safely, and
createPipeline composes grammars with plugin hooks. All three live in this
package, which keeps its “depends on nothing” property: the pipeline accepts
grammars and callbacks as arguments and never imports a grammar pack. See the
Plugins guide for the lifecycle walkthrough; the
contracts below are the reference.
walk(tree, visitors, opts?) runs enter in preorder and leave in
postorder over every node, with node.range (half-open byte span) and
node.text available on each visit. enter returning false prunes the
subtree: nothing below the pruned node is visited and no leave fires for the
pruned node itself. The walk is iterative and allocation-light; it never
materializes the parent side table and never mutates the tree or its source.
Passing { cursor: true } selects cursor-driven navigation with identical
observable behavior.
import { walk } from '@lanexio/parser';
const leafRanges: Array<[number, number]> = [];walk(tree, { enter(node) { if (node.subtreeSize === 1) { leafRanges.push([node.range[0], node.range[1]]); return false; // a leaf has no descendants; prune the empty descent } return undefined; }});transform and the mutation-consistency contract
Section titled “transform and the mutation-consistency contract”transform(tree, visitor) rewrites a document without touching the frozen
flat buffer. The visitor is called once per node in preorder; returning a
string (UTF-8 encoded), a Uint8Array, or an empty string (a deletion)
replaces the node’s source range, while null/undefined leaves the node
alone. Edits are applied in one ascending, non-overlapping pass over a copy of
the source, and the result reports every edit:
source: a freshUint8Arraywith each applied replacement spliced in, never a view over the tree’s source.applied: the edits that were applied, in ascending range order.skipped: candidate edits that were not applied, each with a stable reason:overlaps-applied-edit(the range starts inside an already-applied edit, for example a container edit followed by an edit on a node inside it),descending-edit(out of document order), orinvalid-range.
The original tree is never mutated: tree.source, tree.buffer, and every
node range stay byte-for-byte identical. The transform produces a new source
deterministically and requires a re-parse to obtain a new tree; it never
rewrites subtree_size or rebuilds the buffer. The visitor callback is caller
code, so its errors propagate to the transform caller.
createPipeline
Section titled “createPipeline”createPipeline(inputs) takes grammar registrations and plugins, in the order
their hooks run, and returns a pipeline:
parse(source, { language?, filename? })resolves the grammar through the shared registry by language or filename extension, runsonParsehooks in plugin order, thentransformhooks in plugin order (re-parsing between hooks whenever a hook changes the source). Unresolvable dispatch returns an error tree with diagnostics, and a throwing plugin is recorded as a structured pipeline diagnostic on the returned last-good tree, so the parse path never throws.transform(tree)runs the plugin transform hooks over an existing tree; without hooks it returns an identity result over a copy of the source. As caller code, direct calls propagate plugin errors.walk(tree, visitors),register(reg), andgrammars()round out the surface. Registrations go into the same sharedgrammarRegistrythat the unifiedparse()andregister()use.
Accessibility
Section titled “Accessibility”Accessibility requirements
Section titled “Accessibility requirements”- No direct accessibility surface. The core module produces flat AST data structures. Consuming code is responsible for rendering output with appropriate semantics.
Accessibility checklist
Section titled “Accessibility checklist”| Concern | Status |
|---|---|
| Generated output semantics | Not applicable (core data structures only) |
| ARIA attributes in serialized output | Not applicable |
| Semantic element round-trip | Not applicable |
Security
Section titled “Security”Security considerations
Section titled “Security considerations”- No direct security surface. The core module provides data structures and parsing infrastructure. All parse functions carry the never-throw guarantee.
- The toy grammar (
parse()) is not intended for untrusted input validation — it is a minimal reference implementation.
| Threat | Mitigation | Status |
|---|---|---|
| Malformed input byte sequence | Panic-free guarantee: all inputs accepted, errors produce LexError AST nodes | Implemented |
Companion packages
Section titled “Companion packages”| Package | Relationship | Layer | Notes |
|---|---|---|---|
@lanexio/parser-grammar-html | Consumes | 2 | HTML grammar built on core primitives |
@lanexio/parser-grammar-markdown | Consumes | 2 | Markdown grammar built on core primitives |
@lanexio/parser-grammar-json | Consumes | 2 | JSON grammar built on core primitives |
@lanexio/parser-query | Consumes | 3 | Query engine operates on core LexTree |
@lanexio/parser | Re-exports | 6 | Entry point re-exporting core API |
Changelog
Section titled “Changelog”| Version | Date | Status | Notable changes |
|---|---|---|---|
1.0.0 | 2026-05-29 | Current | Initial stable release. Apache-2.0. |
Migration notes
Section titled “Migration notes”- None. This is the initial stable release.