DOM Adapter
The DOM adapter is a first-party facade of @lanexio/parser-grammar-html that
projects the flat HTML AST of the Lanexio™ Parser into DOM-shaped views. It
lives at packages/parser-grammar-html/src/dom.ts, ships from the package root
as parseDom and toDom, and supersedes the deferral recorded in ADR 0047
(see ADR 0049).
The adapter is a read layer: element views read the live LexTree, and
setAttribute / classList mutate only a per-instance override map. The flat
buffer, tree.source, and any serialized output are never touched. Views are
value-like: two views over the same tree node are structurally equal but not
reference-equal, so compare by value (tagName, getAttribute, data), never
with ===.
The ecosystem example examples/ecosystem/dom-adapter.ts used to contain a
hand-rolled projection. It is now a thin re-export of the package facade, kept
so downstream examples import a stable surface.
Install and import
Section titled “Install and import”parseDom and toDom are exports of @lanexio/parser-grammar-html, so no
extra dependency is needed:
import { parseDom, parseHtml, toDom } from '@lanexio/parser-grammar-html';
const bytes = new TextEncoder().encode('<p>Hello <b>world</b></p>');
// Parse bytes and get a DOM-shaped facade in one step.const doc = parseDom(bytes);
// Or wrap a tree you already parsed (document or fragment).const sameDoc = toDom(parseHtml(bytes));Both entry points never throw. Malformed input produces LexError nodes in the
tree (per the parser’s never-throw integrity) and the facade projects only the
node kinds it understands. A LexTree from a non-HTML grammar degrades to an
empty document rather than throwing.
Node shapes
Section titled “Node shapes”Every projected node carries a nodeType discriminator matching the DOM
constants:
| Node type | nodeType | Shape |
|---|---|---|
DomDocument | 9 | children, doctype, querySelector, querySelectorAll |
DomElement | 1 | tagName, getAttribute, setAttribute, attributes, children, firstChild, parentElement, textContent, outerHTML, id, matches, querySelector, querySelectorAll, classList |
DomText | 3 | data (decoded text) |
DomComment | 8 | data (contents without <!-- and -->) |
DomDoctype | 10 | name (lowercased) |
Element tagName is lowercased for HTML elements and preserves source casing
for SVG/MathML (foreign) elements, matching parse5: foreignObject,
linearGradient. Implied scaffolding (a synthesized html/head/body) is
not materialized: a
top-level element whose source had no <html> wrapper reports
parentElement === null, matching the flat AST the parser actually builds.
const doc = parseDom(new TextEncoder().encode('<!DOCTYPE html><p>Hi</p>'));
console.log(doc.doctype?.name); // "html"const p = doc.querySelector('p');console.log(p?.nodeType, p?.tagName); // 1 "p"console.log(p?.textContent); // "Hi"console.log(p?.children[0].data); // "Hi" (nodeType 3)Attributes and text are decoded
Section titled “Attributes and text are decoded”Attribute names match case-insensitively on HTML elements (the DOM contract)
and case-sensitively on SVG/MathML elements. Values are entity-decoded:
getAttribute('title') on title="a&b" returns a&b. A boolean
attribute returns ''. Child text runs are decoded with the element’s content
context, mirroring the grammar’s textContent: normal HTML decodes character
references, raw-text elements (script, style, …) stay verbatim.
const doc = parseDom(new TextEncoder().encode('<p title="a&b">a & b</p>'));const p = doc.querySelector('p');console.log(p?.getAttribute('title')); // "a&b"console.log(p?.textContent); // "a & b"Read layer with decorator override
Section titled “Read layer with decorator override”setAttribute and classList do not mutate the tree. They write to the
facade instance’s override map, which layers on top of the read-only base
attributes. Fresh queries build fresh views with empty overrides, so any
mutation stays local to the view you hold:
import { parseHtml, parseDom, serializeHtml, toDom } from '@lanexio/parser-grammar-html';
const tree = parseHtml(new TextEncoder().encode('<p class="a">x</p>'));const doc = toDom(tree);const p = doc.querySelector('p');
p.classList.add('b'); // this view onlyp.setAttribute('data-k', '1');console.log(p.getAttribute('class')); // "a b"
const fresh = doc.querySelector('p');console.log(fresh.getAttribute('class')); // "a" (base unchanged)console.log(serializeHtml(tree)); // byte-for-byte identicalclassList follows the DOM ordered-set rules: tokens split on ASCII
whitespace, duplicate tokens collapse, and the serialized value joins with a
single space. One deliberate divergence: an empty token or a token containing
whitespace is inert instead of throwing, matching the never-throw posture.
attributes lists base attributes in source order first, then any attributes
added through the view’s override map, with overridden base values already
reflected.
Selector subset
Section titled “Selector subset”querySelector, querySelectorAll, and matches support a documented,
dependency-free subset:
| Form | Example | Notes |
|---|---|---|
| Tag | p | Case-insensitive on HTML elements; case-sensitive on SVG/MathML elements |
| Class | .item | Token-based, case-sensitive |
| Id | #menu | Exact match |
| Tag + class | p.item | Compound on one element |
| Descendant | div p | Space-separated chains |
| Comma list | h1, p.item | Union of the groups |
Unsupported syntax (child >, :pseudo, [attr]) never throws: the affected
group is treated as never-matching. Matching walks ancestor chains above the
query scope, so el.querySelectorAll('div .x') can match descendants of el
whose div ancestor sits above el itself, and element-scoped queries never
return the scope root.
const doc = parseDom(new TextEncoder().encode( '<ul id="menu"><li class="item">one</li><li class="item">two</li></ul>'));console.log(doc.querySelectorAll('li.item').length); // 2console.log(doc.querySelector('#menu li.item').textContent); // "one"outerHTML
Section titled “outerHTML”outerHTML delegates to serializeHtml for element subtrees inside an HTML
document or fragment tree. When the wrapped subtree is not HTML (for example a
hand-built tree from another grammar), it falls back to a byte-exact slice of
node.range over tree.source instead of throwing.
What the facade deliberately is not
Section titled “What the facade deliberately is not”The adapter is a projection surface, not a full browser DOM:
- no implied-element synthesis beyond what the flat AST stores;
- no live DOM identity, events, or mutation of the underlying tree;
- no full Selectors engine (see the subset above);
- no
ownerDocument, namespaces, or attribute nodes.
For a lossless, editable handle over the same source, keep using LexTree and
LexNode directly.
Run the example and its tests
Section titled “Run the example and its tests”From the repo root:
pnpm vitest run examples/ecosystem/dom-adapter.test.tsThe example test asserts the documented facade shapes: nodeType
discriminators, decoded attributes, the override map, the selector subset,
outerHTML, and malformed-input recovery.
Related
Section titled “Related”- parser-grammar-html reference:
parseHtml,serializeHtml, and the other grammar exports - GraphQL validation: the sibling seam on the flat AST
- Plugins: the plugin-registry seam
- Flat AST: the 16-byte node layout the facade reads
Design context: ADR 0049 (HTML DOM adapter) records the market-pull trigger and the home of this facade inside the grammar pack, superseding the deferral in ADR 0047.