Parsing XML
Lanexio™ Parser implements XML 1.0 5th Edition with optional DTD validation. The parser handles multi-byte encodings, external entities, and the full W3C XML conformance suite.
Scope. The XML grammar targets XML 1.0 5th Edition only. XML 1.1 remains out of scope: documents that require XML 1.1-only features are not supported by this grammar.
Encodings. Input is UTF-8 bytes by default. A UTF-16 or UTF-32 byte order mark (BOM) is autodetected and transcoded losslessly, so node ranges stay in the original byte space. A declared encoding that contradicts the BOM is a fatal condition that surfaces as an error node, per XML 1.0 5th Ed. section 4.3.3.
Quick start
Section titled “Quick start”-
Install the package.
Terminal window pnpm add @lanexio/parser-grammar-xmlTerminal window npm install @lanexio/parser-grammar-xmlTerminal window yarn add @lanexio/parser-grammar-xml -
Parse an XML document.
import { parseXml } from '@lanexio/parser-grammar-xml';const encoder = new TextEncoder();const tree = parseXml(encoder.encode(`<?xml version="1.0" encoding="UTF-8"?><catalog><book id="bk101"><title>XML Guide</title><price>44.95</price></book></catalog>`));
DTD validation
Section titled “DTD validation”DTD validation is opt-in. The parser never performs I/O: an external subset is
only ever read through a caller-supplied resolveExternal callback. Return the
subset bytes to read it, or null to leave the reference unresolved (the
callback is the only mechanism for obtaining external DTD content).
import { parseXml } from '@lanexio/parser-grammar-xml';
// Application-owned DTD content. The parser never performs I/O: the resolver// is the only channel for external subset bytes.const catalogDtd = new TextEncoder().encode( '<!ELEMENT catalog (book+)><!ELEMENT book (title, price)>',);
// systemId / publicId are the quoted identifiers (quotes stripped); baseUri is// the document base URI for resolving relative system ids. Returning the bytes// marks the subset as read; null leaves the DOCTYPE reference unresolved.const resolver = ( publicId: string | null, systemId: string | null, baseUri: string,): Uint8Array | null => { const key = systemId ?? publicId ?? ''; return baseUri + key === 'https://lanexio.com/docs/catalog/catalog.dtd' ? catalogDtd : null;};
const tree = parseXml(documentBytes, { resolveExternal: resolver, baseUri: 'https://lanexio.com/docs/catalog/',});With no resolver, a <!DOCTYPE ... SYSTEM ...> document parses without reading
the external subset. A non-validating processor is permitted to skip an external
subset, so the document stays clean by default.
To run the full DTD validating pass (VC Root Element Type, element content
models, attribute constraints, entity-name checks), pass validate: true:
const tree = parseXml(documentBytes, { validate: true });An external-DTD document with no resolver then flags, because no declaration for
the root element is available. validate is forwarded through the registered
grammar and the pack’s own parse entry point, so
parse(bytes, { grammarOptions: { validate: true } }) behaves identically to the
direct parseXml(bytes, { validate: true }) call.
Strict mode
Section titled “Strict mode”The default (mode: "lenient") recovers from well-formedness and validity
problems without throwing. Strict mode keeps the never-throw guarantee but
surfaces spec-level rejections on the tree:
- Namespace-prefix well-formedness runs unconditionally. An unbound
element or attribute prefix (no matching
xmlnsdeclaration) flags the element withXML_FLAG_NAMESPACE_WF_ERROR, even when the document carries no namespace declarations at all. - Internal-subset DTD validation runs when the document carries an internal
DTD subset (VC Root Element Type, element content models, attribute
constraints). The parser never performs I/O: external DTDs are only ever
read through a caller-supplied
resolveExternalcallback. - An external-DTD reference surfaces as an error when nothing read the
subset. A DOCTYPE that references an external subset (SYSTEM or PUBLIC)
which no resolver read flags the DocType node as an error node and the
Document root with
LEX_NODE_HAS_ERROR | GRAMMAR_FLAG_NODE_ERROR. This is the strict contract “WF + reject external-DTD / undefined-entity”: a subset the parser never read cannot be certified under strict. A resolver that read the subset keeps the reference clean. Lenient mode never surfaces it.
The parse profiles compose validate and strict:
| Profile | What runs | Unresolved external-DTD document |
|---|---|---|
| default (lenient) | WF-only recovery | parses clean |
validate: true | full DTD validating pass | flagged (no root declaration available) |
mode: "strict" / strict: true | namespace-prefix WF + internal-subset DTD validation + external-DTD surface | flagged on the DocType + Document root |
validate: true + strict | full DTD validating pass on top of the strict surface | flagged |
import { parseXml } from '@lanexio/parser-grammar-xml';
const encoder = new TextEncoder();
// Unbound element prefix: clean in lenient, flagged in strictconst lenient = parseXml(encoder.encode('<a:root/>'));console.log(lenient.root.hasError); // false
const strict = parseXml(encoder.encode('<a:root/>'), { mode: 'strict' });console.log(strict.root.hasError); // true
// The { strict: true } alias is the same option (ADR 0044)const strictAlias = parseXml(encoder.encode('<a:root/>'), { strict: true });console.log(strictAlias.root.hasError); // true
// Bound prefix stays clean in strict modeconst bound = parseXml( encoder.encode('<a:root xmlns:a="urn:test"/>'), { mode: 'strict' },);console.log(bound.root.hasError); // false
// External-DTD reference: clean in lenient, flagged in strict when no// resolver read the subsetconst external = encoder.encode('<!DOCTYPE root SYSTEM "x.dtd"><root/>');console.log(parseXml(external).root.hasError); // falseconsole.log(parseXml(external, { mode: 'strict' }).root.hasError); // truePrecedence is mode > strict > default lenient, so an explicit mode always
wins over a strict boolean, and strict: false is the explicit lenient
spelling. The unified parse() entry point forwards the same options through
grammarOptions.mode or the top-level strict key, so parse(src, { language: "xml", grammarOptions: { mode: "strict" } }) and parse(src, { language: "xml", strict: true }) both behave identically to parseXml(bytes, { mode: "strict" }). Strict mode never throws; a rejected construct is a flag on the tree, never
an exception. The validate: true option composes with strict and stays
available under leniency on its own. See
Leniency and strict mode for the full policy.
Inspecting the tree
Section titled “Inspecting the tree”Elements, attributes, text nodes, comments, PIs, CDATA sections, and doctype declarations are all represented in the flat AST. The tree preserves the document’s hierarchical structure.
const cursor = tree.cursor();cursor.gotoFirstChild(); // XmlDeclaration or DocumentElement
while (cursor.gotoNextSibling()) { console.log(cursor.current.kind);}