Skip to content

DOM Adapter

The DOM adapter is a first-party facade of @lanexio/parser-grammar-html that projects the flat HTML AST of the Lanexio™ Parser into DOM-shaped views. It lives at packages/parser-grammar-html/src/dom.ts, ships from the package root as parseDom and toDom, and supersedes the deferral recorded in ADR 0047 (see ADR 0049).

The adapter is a read layer: element views read the live LexTree, and setAttribute / classList mutate only a per-instance override map. The flat buffer, tree.source, and any serialized output are never touched. Views are value-like: two views over the same tree node are structurally equal but not reference-equal, so compare by value (tagName, getAttribute, data), never with ===.

The ecosystem example examples/ecosystem/dom-adapter.ts used to contain a hand-rolled projection. It is now a thin re-export of the package facade, kept so downstream examples import a stable surface.

parseDom and toDom are exports of @lanexio/parser-grammar-html, so no extra dependency is needed:

import { parseDom, parseHtml, toDom } from '@lanexio/parser-grammar-html';
const bytes = new TextEncoder().encode('<p>Hello <b>world</b></p>');
// Parse bytes and get a DOM-shaped facade in one step.
const doc = parseDom(bytes);
// Or wrap a tree you already parsed (document or fragment).
const sameDoc = toDom(parseHtml(bytes));

Both entry points never throw. Malformed input produces LexError nodes in the tree (per the parser’s never-throw integrity) and the facade projects only the node kinds it understands. A LexTree from a non-HTML grammar degrades to an empty document rather than throwing.

Every projected node carries a nodeType discriminator matching the DOM constants:

Node typenodeTypeShape
DomDocument9children, doctype, querySelector, querySelectorAll
DomElement1tagName, getAttribute, setAttribute, attributes, children, firstChild, parentElement, textContent, outerHTML, id, matches, querySelector, querySelectorAll, classList
DomText3data (decoded text)
DomComment8data (contents without <!-- and -->)
DomDoctype10name (lowercased)

Element tagName is lowercased for HTML elements and preserves source casing for SVG/MathML (foreign) elements, matching parse5: foreignObject, linearGradient. Implied scaffolding (a synthesized html/head/body) is not materialized: a top-level element whose source had no <html> wrapper reports parentElement === null, matching the flat AST the parser actually builds.

const doc = parseDom(new TextEncoder().encode('<!DOCTYPE html><p>Hi</p>'));
console.log(doc.doctype?.name); // "html"
const p = doc.querySelector('p');
console.log(p?.nodeType, p?.tagName); // 1 "p"
console.log(p?.textContent); // "Hi"
console.log(p?.children[0].data); // "Hi" (nodeType 3)

Attribute names match case-insensitively on HTML elements (the DOM contract) and case-sensitively on SVG/MathML elements. Values are entity-decoded: getAttribute('title') on title="a&amp;b" returns a&b. A boolean attribute returns ''. Child text runs are decoded with the element’s content context, mirroring the grammar’s textContent: normal HTML decodes character references, raw-text elements (script, style, …) stay verbatim.

const doc = parseDom(new TextEncoder().encode('<p title="a&amp;b">a &amp; b</p>'));
const p = doc.querySelector('p');
console.log(p?.getAttribute('title')); // "a&b"
console.log(p?.textContent); // "a & b"

setAttribute and classList do not mutate the tree. They write to the facade instance’s override map, which layers on top of the read-only base attributes. Fresh queries build fresh views with empty overrides, so any mutation stays local to the view you hold:

import { parseHtml, parseDom, serializeHtml, toDom } from '@lanexio/parser-grammar-html';
const tree = parseHtml(new TextEncoder().encode('<p class="a">x</p>'));
const doc = toDom(tree);
const p = doc.querySelector('p');
p.classList.add('b'); // this view only
p.setAttribute('data-k', '1');
console.log(p.getAttribute('class')); // "a b"
const fresh = doc.querySelector('p');
console.log(fresh.getAttribute('class')); // "a" (base unchanged)
console.log(serializeHtml(tree)); // byte-for-byte identical

classList follows the DOM ordered-set rules: tokens split on ASCII whitespace, duplicate tokens collapse, and the serialized value joins with a single space. One deliberate divergence: an empty token or a token containing whitespace is inert instead of throwing, matching the never-throw posture.

attributes lists base attributes in source order first, then any attributes added through the view’s override map, with overridden base values already reflected.

querySelector, querySelectorAll, and matches support a documented, dependency-free subset:

FormExampleNotes
TagpCase-insensitive on HTML elements; case-sensitive on SVG/MathML elements
Class.itemToken-based, case-sensitive
Id#menuExact match
Tag + classp.itemCompound on one element
Descendantdiv pSpace-separated chains
Comma listh1, p.itemUnion of the groups

Unsupported syntax (child >, :pseudo, [attr]) never throws: the affected group is treated as never-matching. Matching walks ancestor chains above the query scope, so el.querySelectorAll('div .x') can match descendants of el whose div ancestor sits above el itself, and element-scoped queries never return the scope root.

const doc = parseDom(new TextEncoder().encode(
'<ul id="menu"><li class="item">one</li><li class="item">two</li></ul>'
));
console.log(doc.querySelectorAll('li.item').length); // 2
console.log(doc.querySelector('#menu li.item').textContent); // "one"

outerHTML delegates to serializeHtml for element subtrees inside an HTML document or fragment tree. When the wrapped subtree is not HTML (for example a hand-built tree from another grammar), it falls back to a byte-exact slice of node.range over tree.source instead of throwing.

The adapter is a projection surface, not a full browser DOM:

  • no implied-element synthesis beyond what the flat AST stores;
  • no live DOM identity, events, or mutation of the underlying tree;
  • no full Selectors engine (see the subset above);
  • no ownerDocument, namespaces, or attribute nodes.

For a lossless, editable handle over the same source, keep using LexTree and LexNode directly.

From the repo root:

Terminal window
pnpm vitest run examples/ecosystem/dom-adapter.test.ts

The example test asserts the documented facade shapes: nodeType discriminators, decoded attributes, the override map, the selector subset, outerHTML, and malformed-input recovery.

Design context: ADR 0049 (HTML DOM adapter) records the market-pull trigger and the home of this facade inside the grammar pack, superseding the deferral in ADR 0047.