Skip to content

@lanexio/parser-core

This page documents @lanexio/parser-core, the foundational package that defines the shared buffer protocol, tree data structures, and parser host for all Lanexio Parser packages.

  • Version: Stable
  • Module name: parser-core
  • Package: @lanexio/parser-core
  • Import path: @lanexio/parser-core
  • Layer: 1 (Core)
  • Runtime: Universal (browser, server, edge worker, test runner)
  • Module format: ESM
  • Stability: Stable
  • Primary use case: Build and traverse zero-copy flat ASTs, register grammars, use core primitives.
  • You need to create, traverse, or modify a LexTree directly.
  • You want to use the grammar registry to manage multiple grammars.
  • You need to build a custom grammar or parser integration.
  • You need the incremental editing API (applyEdit, reparse) or streaming parse (createParseStream).
  • You want the shared buffer protocol constants for value-level interop.
BoundaryDescription
InputsUint8Array (source bytes for parse), or preconstructed buffers for tree manipulation
OutputsUint8Array (edited source), LexTree, LexNode, LexCursor
Side effectsNone
DeterminismYes (same bytes and same protocol version produce same tree)
External dependenciesNone
Never-throw guaranteeYes for parse()
Security surfaceNone (no HTML output, no I/O)
  1. Install the package.

    Terminal window
    pnpm add @lanexio/parser-core
  2. Import the named export.

    import { parse, LexTree, LexNode, LexCursor } from '@lanexio/parser-core';

This package has no peer dependencies. It is the foundation that all other packages depend on.

import { parse, LexNode, LexCursor, LexTree } from '@lanexio/parser-core';
const encoder = new TextEncoder();
const bytes = encoder.encode('toy input');
const tree: LexTree = parse(bytes);
console.log(tree.nodeCount);
console.log(tree.root);
const tree = parse(encoder.encode('example'));
const cursor = tree.cursor();
// Preorder DFS: descend into the first child when possible, otherwise move to
// the next sibling or climb back toward the root.
walk: for (;;) {
const node = cursor.current;
console.log('kind:', node.kind, 'range:', node.range);
if (cursor.gotoFirstChild()) continue;
for (;;) {
if (cursor.gotoNextSibling()) continue walk;
if (!cursor.gotoParent()) break walk;
}
}
ExportTypeDescription
parse(source: Uint8Array) => LexTreeParse toy-grammar bytes into a flat AST. Never throws.
LexTreeclassRoot handle for a zero-copy flat AST backed by a single ArrayBuffer.
LexNodeclassReference to a single 16-byte node within a LexTree.
LexCursorclassPreorder DFS cursor over a LexTree, backed by a node index stack.
PROTOCOL_VERSIONnumberCurrent shared buffer protocol version (value: 1).
LEX_TREE_MAGICnumberMagic number identifying a valid tree header.
LEX_TREE_VERSIONnumberCurrent tree layout version.
createLexErrorTree(source: Uint8Array, diagnostics?: readonly string[]) => LexTreeCreate a tree with a single root LexError node.
applyEdit(source: Uint8Array, edit: LexEdit) => Uint8ArrayApply a byte-range edit to a source buffer and return the edited bytes.
reparse(tree: LexTree, edit: LexEdit, oracle?: ReuseOracle, options?: ReparseOptions) => LexTreeIncrementally reparse after an edit; falls back to a full reparse without an oracle. Never throws.
reparseWithStats(tree: LexTree, edit: LexEdit, oracle: ReuseOracle, options?: ReparseOptions) => { tree: LexTree; stats: ReparseStats }Reparse and return subtree-reuse statistics.
computeDirtyRegion(tree: LexTree, edit: LexEdit, oracle: ReuseOracle) => DirtyRegionCompute the dirty re-parse window for an edit using a reuse oracle.
createParseStream(grammar: string, parseFn: (bytes: Uint8Array) => LexTree) => LexParseStreamCreate a streaming push parser bound to a grammar label and a never-throwing parse function.
grammarRegistryobjectGlobal registry for grammar registrations.
embedGuests(hostTree: LexTree, options?: EmbedGuestsOptions) => LexTreeEmbed guest-language subtrees into host tree.
embedRegistryobjectRegistry for guest-language embeddings.
graftSubtree(hostTree: LexTree, parentIndex: number, insertBeforeIndex: number, guestRecords: Uint32Array, metadata?: LexTreeMetadata) => LexTreeSplice packed guest node records into a host tree under a parent node.
findParentIndex(records: Uint32Array, nodeCount: number, index: number) => number | nullFind the parent node index of a child in a packed records array, or null.
fixSubtreeSizes(records: Uint32Array, startIndex: number, delta: number) => voidAdd delta to subtree_size of every ancestor of startIndex, in place.
rebaseRecords(records: Uint32Array, delta: number) => Uint32ArrayOffset all byte ranges in a node record array.
assertLossless(source: Uint8Array, leaves: ReadonlyArray<LeafRange>) => Uint8ArrayAssert leaf ranges tile the source byte-for-byte; returns the concatenated leaf bytes.
emitSource(tree: LexTree) => Uint8ArrayReconstruct the source bytes from the tree.
extractLeafRanges(tree: LexTree) => LeafRange[]Extract leaf-level source byte ranges.
countFalseHasError(tree: LexTree, knownErrorKinds: ReadonlySet<number>) => numberCount non-Error leaf nodes whose flags collide with the reserved error bit.
countInvertedRanges(tree: LexTree) => { count: number; firstIndex: number | null }Count nodes with start > end ranges and report the first offender.
assertWellFormed(tree: LexTree) => string[]Validate all flat-AST invariants; returns violation descriptions (empty when well-formed).
LexEditErrorclassError thrown on invalid edits.
LexEditErrorCodeconst objectError code constants for edit failures.
LosslessErrorclassError thrown when lossless integrity check fails.
NODE_STRIDEnumberStride (in Uint32 slots) per node record.
SLOT_KINDnumberSlot index for node kind.
SLOT_FLAGSnumberSlot index for node flags.
SLOT_FIELDnumberSlot index for node field id.
SLOT_STARTnumberSlot index for byte range start.
SLOT_ENDnumberSlot index for byte range end.
SLOT_SIZEnumberSlot index for subtree size.
createTree(source: Uint8Array, nodes: ReadonlyArray<PendingNode>, metadata?: LexTreeMetadata) => LexTreeCreate a tree from pending node descriptions.
createTreeFromUint32Array(nodeData: Uint32Array, nodeCount: number, source: Uint8Array, metadata?: LexTreeMetadata) => LexTreeCreate a tree from packed Uint32Array node records (NODE_STRIDE slots per node).
growNodeData(nd: Uint32Array, needed: number) => Uint32ArrayGrow the node data buffer.
writeTreeHeader(records: Uint32Array, nodeCount: number, offset?: number) => voidWrite the standard tree header into a Uint32Array at the given word offset.
createPendingNode(kind: number, flags: number, fieldId: number, rangeStart: number, rangeEnd: number) => PendingNodeCreate a pending leaf node for tree construction.
LANEXIO_PARSER_CORE_PACKAGE_NAMEstringStable npm package name constant.
LexToyKindconst objectKind constants for the built-in toy grammar.
LexToyFieldconst objectField constants for the built-in toy grammar.
toValue(target: LexTree | LexNode, options?: ToValueOptions) => LexValueResultProject a tree or node subtree into a JSON-expressible host value. Never throws.
registerValueExtractor(language: string, extractor: LexValueExtractor) => voidRegister a value extractor under a case-insensitive language name.
valueExtractorRegistryValueExtractorRegistrySingleton extractor registry (case-insensitive, last registration wins).
findSubtreeError(node: LexNode) => LexRange | nullReturn the first error byte range inside a node’s subtree, or null.
rebuildTreeWithMetadata(original: LexTree, extra: Partial<LexTreeMetadata>) => LexTreeReturn a new LexTree view over the same buffer with merged metadata.
walk(tree: LexTree, visitors: Visitor, opts?: WalkOptions) => voidPreorder/postorder visitor walk over a flat AST. Returning false from enter prunes the subtree.
transform(tree: LexTree, visitor: TransformVisitor) => TransformResultRewrite a document in one ascending, non-overlapping pass over a copy of the source. Never mutates the tree.
TransformSkipReasonconst objectStable machine-readable skip reasons (Overlap, Descending, InvalidRange).
createPipeline(inputs: ReadonlyArray<GrammarRegistration | Plugin>) => PipelineCompose grammar registrations and plugins into one parse/transform pipeline over the shared registry.
assertDomWellFormed(tree: LexTree) => string[]Assert the ADR-0016 DOM-shape flat-AST invariants; returns violation strings (empty when well-formed).
repairContainerRanges(records: Uint32Array, nodeCount: number) => voidExtend container node ranges to enclose all children, in place.
ValueExtractorRegistryclassCase-insensitive value-extractor registry (register / byLanguage).
LEX_TOY_FIELD_NAMES_BY_IDReadonly<Record<number, string>>Field-id to field-name map for the built-in toy grammar.
LEX_NODE_HAS_ERRORnumberNode flag (bit 0) set when a node or its subtree contains a parse error (value: 1).
GRAMMAR_FLAG_NODE_ERRORnumberCompanion flag (bit 6) marking intentional error propagation (value: 64).
NODE_RECORD_SIZEnumberSize in bytes of one flat AST node record (value: 16).
TREE_HEADER_SIZEnumberSize in bytes of the tree header (value: 16).
NODE_KIND_OFFSETnumberNode byte offset of the kind field (value: 0).
NODE_FLAGS_OFFSETnumberNode byte offset of the flags field (value: 2).
NODE_FIELD_ID_OFFSETnumberNode byte offset of the field id field (value: 3).
NODE_RANGE_START_OFFSETnumberNode byte offset of the range start field (value: 4).
NODE_RANGE_END_OFFSETnumberNode byte offset of the range end field (value: 8).
NODE_SUBTREE_SIZE_OFFSETnumberNode byte offset of the subtree size field (value: 12).
decodeUtf8CodePoint(bytes: Uint8Array, offset: number) => { codePoint: number; width: number }Decode one UTF-8 code point at offset.
isLegalXmlChar(codePoint: number) => booleanTrue when a code point is a legal XML 1.0 character.
isNameStartCp(codePoint: number) => booleanTrue when a code point is an XML 1.0 NameStartChar.
isNameCharCp(codePoint: number) => booleanTrue when a code point is an XML 1.0 NameChar.
isAsciiNameStart(byte: number) => booleanFast ASCII-only NameStartChar check.
isAsciiNameChar(byte: number) => booleanFast ASCII-only NameChar check.
isWs(byte: number) => booleanTrue for space, tab, LF, CR.
isDecDigit(byte: number) => booleanTrue for 0-9.
isHexDigit(byte: number) => booleanTrue for 0-9, a-f, A-F.

The core module accepts options through LexTreeOptions when constructing trees programmatically. LexTreeOptions has a single optional field, metadata, which carries up to three optional sub-fields.

FieldTypeRequiredDefaultDescription
metadataLexTreeMetadataNo{}Optional grammar field-name map, resolved language name, and diagnostics.
metadata.fieldNamesByIdReadonly<Record<number, string>>NoundefinedGrammar-specific field-id to field-name map.
metadata.languagestringNoundefinedResolved language name (set by the unified parse() dispatcher).
metadata.diagnosticsreadonly string[]NoundefinedActionable fix messages surfaced on error trees.

The primary return value from parse() is LexTree:

PropertyTypeDescription
rootLexNodeRoot node of the flat AST.
nodeCountnumberTotal nodes in the tree.
sourceUint8ArrayOriginal parsed bytes.
cursor()() => LexCursorCreate a new preorder DFS cursor starting at root.
TypePurposeNotes
LexTreeOptionsOptions for programmatic tree constructionSingle optional metadata field; see Options
LexTreeMetadataImmutable tree metadata passed at construction{ fieldNamesById?, language?, diagnostics? }, all optional
LexRangereadonly [start: number, end: number] half-open byte offset tupleUsed across all grammar APIs
LexEditStructural edit descriptorPassed to applyEdit
DirtyRegionByte range affected by an editReturned by computeDirtyRegion
ReparseOptionsOptions for incremental reparsePassed to reparse
ReparseStatsPerformance statistics from reparseReturned by reparseWithStats
ReuseOracleOracle for incremental node reuseUsed by grammar incremental modules
LexParseStreamStreaming push parser interfaceCreated by createParseStream
LexTreeValidationCodeValidation result codeProduced by integrity checks
LexTreeValidationErrorError from tree validationThrown on invalid tree structure
LeafRangeLeaf-level source rangeUsed by extractLeafRanges
LosslessDiagnosticsDiagnostic info from lossless checkReturned by lossless operations
PendingNodeNode pending insertionUsed during tree construction
LanexioParserPureGrammarPure-TS grammar descriptor interfaceUsed by grammar packages to register themselves
GrammarRegistrationFull grammar registration descriptorUsed with grammarRegistry
EmbedGuestsOptionsOptions for guest embeddingPassed to embedGuests
EmbedRuleGuest embedding ruleUsed in embed registry
LexValueJSON-expressible host valueUnion of string, number, boolean, null, object, array
LexValueObjectProjected object value{ readonly [key: string]: LexValue }
LexValueArrayProjected array valuereadonly LexValue[]
LexValueResultResult of one toValue call{ value: LexValue | undefined; diagnostics: readonly LexValueDiagnostic[] }
LexValueDiagnosticStructured value-extraction failure{ code: LexValueDiagnosticCode; rangeStart; rangeEnd; message }
LexValueDiagnosticCodeStable machine-readable failure code"has_error", "unresolved_language", "unsafe_integer", "overflow", "duplicate_key", "missing_anchor", "circular_alias", "max_depth", "excess_field", "empty_document"
LexValueContextRecursion context shared with an extractor{ maxDepth: number }
LexValueExtractorProjection contract implemented by grammar packs(node: LexNode, context: LexValueContext) => LexValueResult
ToValueOptionsOptions for toValue{ language?, extractor?, maxDepth? }
VisitorCallbacks run over every node by walk{ enter?: (node) => void | boolean; leave?: (node) => void }; enter returning false prunes
WalkOptionsOptions for walk{ cursor?: boolean } selects cursor-driven navigation
TransformVisitorCallback for transform(node: LexNode) => string | Uint8Array | null | undefined; returning text replaces the node’s range
EditOne applied replacement{ range: LexRange; replacement: Uint8Array }
SkippedEditOne skipped candidate edit{ range: LexRange; reason: string } with a stable machine-readable reason
TransformResultResult of one transform call{ source: Uint8Array; applied: ReadonlyArray<Edit>; skipped: ReadonlyArray<SkippedEdit> }
TransformSkipReasonTypeUnion of skip-reason strings"overlaps-applied-edit", "descending-edit", "invalid-range"
PluginA pipeline plugin{ name: string; onParse?: (tree, ctx) => LexTree | null | undefined; transform?: (tree) => TransformResult | null | undefined }
PipelineA pipeline returned by createPipelineparse / transform / walk / register / grammars()
PipelineContextContext handed to an onParse hook{ language: string }, the resolved canonical language
PipelineParseOptionsOptions for Pipeline.parse{ language?: string; filename?: string }
LexToyKindTypeType alias for LexToyKindSame keys as the LexToyKind const
LexToyFieldTypeType alias for LexToyFieldSame keys as the LexToyField const
LexTreeValidationCodeTypeType alias for LexTreeValidationCodeSame codes as the LexTreeValidationCode const

The grammarRegistry manages grammar registrations globally. Register a grammar to make it available for auto-detection and unified parsing.

import { grammarRegistry } from '@lanexio/parser-core';
import { htmlRegistration } from '@lanexio/parser-grammar-html';
grammarRegistry.register(htmlRegistration);

embedGuests inserts guest-language subtrees (for example, a CSS block inside an HTML style element) into a host tree. The embedRegistry stores the embedding rules.

createParseStream(grammar, parseFn) creates a LexParseStream that accepts chunked input via push(), finalizes with end(), and yields one tree per push through its async iterator:

import { createParseStream, parse } from '@lanexio/parser-core';
const encoder = new TextEncoder();
const stream = createParseStream('@lanexio/parser-core', (bytes) => parse(bytes));
stream.push(encoder.encode('chunk1 '));
stream.push(encoder.encode('chunk2'));
stream.end();
for await (const tree of stream) {
console.log(tree.nodeCount);
}

toValue() projects a parse tree (or any node subtree) into a JSON-expressible host value without walking the AST by hand. It is a pure read of the frozen flat AST, never throws, and resolves a grammar-specific extractor from the tree’s metadata.language (or an explicit language / extractor option). Extractors are registered by @lanexio/parser/all in one place; grammar packs do not self-register. Each pack exports its extractor and a <grammar>ToValue convenience function that presets it, so direct grammar-pack users need no registration step.

import { toValue } from '@lanexio/parser';
import { parse } from '@lanexio/parser/all';
const tree = parse('{"a": 1, "b": [true, null]}', { language: 'json' });
const { value, diagnostics } = toValue(tree);
console.log(value); // { a: 1, b: [true, null] }
console.log(diagnostics); // []

@lanexio/parser/all registers every grammar and extractor but does not export toValue, hence the split import.

See the value-extraction reference for the LexValue model, the diagnostic codes, and the per-grammar coercion rules.

The plugin API is the grammar-agnostic composition surface of Lanexio™ Parser: walk visits a flat AST, transform rewrites its source safely, and createPipeline composes grammars with plugin hooks. All three live in this package, which keeps its “depends on nothing” property: the pipeline accepts grammars and callbacks as arguments and never imports a grammar pack. See the Plugins guide for the lifecycle walkthrough; the contracts below are the reference.

walk(tree, visitors, opts?) runs enter in preorder and leave in postorder over every node, with node.range (half-open byte span) and node.text available on each visit. enter returning false prunes the subtree: nothing below the pruned node is visited and no leave fires for the pruned node itself. The walk is iterative and allocation-light; it never materializes the parent side table and never mutates the tree or its source. Passing { cursor: true } selects cursor-driven navigation with identical observable behavior.

import { walk } from '@lanexio/parser';
const leafRanges: Array<[number, number]> = [];
walk(tree, {
enter(node) {
if (node.subtreeSize === 1) {
leafRanges.push([node.range[0], node.range[1]]);
return false; // a leaf has no descendants; prune the empty descent
}
return undefined;
}
});

transform and the mutation-consistency contract

Section titled “transform and the mutation-consistency contract”

transform(tree, visitor) rewrites a document without touching the frozen flat buffer. The visitor is called once per node in preorder; returning a string (UTF-8 encoded), a Uint8Array, or an empty string (a deletion) replaces the node’s source range, while null/undefined leaves the node alone. Edits are applied in one ascending, non-overlapping pass over a copy of the source, and the result reports every edit:

  • source: a fresh Uint8Array with each applied replacement spliced in, never a view over the tree’s source.
  • applied: the edits that were applied, in ascending range order.
  • skipped: candidate edits that were not applied, each with a stable reason: overlaps-applied-edit (the range starts inside an already-applied edit, for example a container edit followed by an edit on a node inside it), descending-edit (out of document order), or invalid-range.

The original tree is never mutated: tree.source, tree.buffer, and every node range stay byte-for-byte identical. The transform produces a new source deterministically and requires a re-parse to obtain a new tree; it never rewrites subtree_size or rebuilds the buffer. The visitor callback is caller code, so its errors propagate to the transform caller.

createPipeline(inputs) takes grammar registrations and plugins, in the order their hooks run, and returns a pipeline:

  • parse(source, { language?, filename? }) resolves the grammar through the shared registry by language or filename extension, runs onParse hooks in plugin order, then transform hooks in plugin order (re-parsing between hooks whenever a hook changes the source). Unresolvable dispatch returns an error tree with diagnostics, and a throwing plugin is recorded as a structured pipeline diagnostic on the returned last-good tree, so the parse path never throws.
  • transform(tree) runs the plugin transform hooks over an existing tree; without hooks it returns an identity result over a copy of the source. As caller code, direct calls propagate plugin errors.
  • walk(tree, visitors), register(reg), and grammars() round out the surface. Registrations go into the same shared grammarRegistry that the unified parse() and register() use.
  • No direct accessibility surface. The core module produces flat AST data structures. Consuming code is responsible for rendering output with appropriate semantics.
ConcernStatus
Generated output semanticsNot applicable (core data structures only)
ARIA attributes in serialized outputNot applicable
Semantic element round-tripNot applicable
  • No direct security surface. The core module provides data structures and parsing infrastructure. All parse functions carry the never-throw guarantee.
  • The toy grammar (parse()) is not intended for untrusted input validation — it is a minimal reference implementation.
ThreatMitigationStatus
Malformed input byte sequencePanic-free guarantee: all inputs accepted, errors produce LexError AST nodesImplemented
PackageRelationshipLayerNotes
@lanexio/parser-grammar-htmlConsumes2HTML grammar built on core primitives
@lanexio/parser-grammar-markdownConsumes2Markdown grammar built on core primitives
@lanexio/parser-grammar-jsonConsumes2JSON grammar built on core primitives
@lanexio/parser-queryConsumes3Query engine operates on core LexTree
@lanexio/parserRe-exports6Entry point re-exporting core API
VersionDateStatusNotable changes
1.0.02026-05-29CurrentInitial stable release. Apache-2.0.
  • None. This is the initial stable release.