Skip to content

Parsing XML

Lanexio™ Parser implements XML 1.0 5th Edition with optional DTD validation. The parser handles multi-byte encodings, external entities, and the full W3C XML conformance suite.

Scope. The XML grammar targets XML 1.0 5th Edition only. XML 1.1 remains out of scope: documents that require XML 1.1-only features are not supported by this grammar.

Encodings. Input is UTF-8 bytes by default. A UTF-16 or UTF-32 byte order mark (BOM) is autodetected and transcoded losslessly, so node ranges stay in the original byte space. A declared encoding that contradicts the BOM is a fatal condition that surfaces as an error node, per XML 1.0 5th Ed. section 4.3.3.

  1. Install the package.

    Terminal window
    pnpm add @lanexio/parser-grammar-xml
  2. Parse an XML document.

    import { parseXml } from '@lanexio/parser-grammar-xml';
    const encoder = new TextEncoder();
    const tree = parseXml(encoder.encode(`
    <?xml version="1.0" encoding="UTF-8"?>
    <catalog>
    <book id="bk101">
    <title>XML Guide</title>
    <price>44.95</price>
    </book>
    </catalog>
    `));

DTD validation is opt-in. The parser never performs I/O: an external subset is only ever read through a caller-supplied resolveExternal callback. Return the subset bytes to read it, or null to leave the reference unresolved (the callback is the only mechanism for obtaining external DTD content).

import { parseXml } from '@lanexio/parser-grammar-xml';
// Application-owned DTD content. The parser never performs I/O: the resolver
// is the only channel for external subset bytes.
const catalogDtd = new TextEncoder().encode(
'<!ELEMENT catalog (book+)><!ELEMENT book (title, price)>',
);
// systemId / publicId are the quoted identifiers (quotes stripped); baseUri is
// the document base URI for resolving relative system ids. Returning the bytes
// marks the subset as read; null leaves the DOCTYPE reference unresolved.
const resolver = (
publicId: string | null,
systemId: string | null,
baseUri: string,
): Uint8Array | null => {
const key = systemId ?? publicId ?? '';
return baseUri + key === 'https://lanexio.com/docs/catalog/catalog.dtd'
? catalogDtd
: null;
};
const tree = parseXml(documentBytes, {
resolveExternal: resolver,
baseUri: 'https://lanexio.com/docs/catalog/',
});

With no resolver, a <!DOCTYPE ... SYSTEM ...> document parses without reading the external subset. A non-validating processor is permitted to skip an external subset, so the document stays clean by default.

To run the full DTD validating pass (VC Root Element Type, element content models, attribute constraints, entity-name checks), pass validate: true:

const tree = parseXml(documentBytes, { validate: true });

An external-DTD document with no resolver then flags, because no declaration for the root element is available. validate is forwarded through the registered grammar and the pack’s own parse entry point, so parse(bytes, { grammarOptions: { validate: true } }) behaves identically to the direct parseXml(bytes, { validate: true }) call.

The default (mode: "lenient") recovers from well-formedness and validity problems without throwing. Strict mode keeps the never-throw guarantee but surfaces spec-level rejections on the tree:

  • Namespace-prefix well-formedness runs unconditionally. An unbound element or attribute prefix (no matching xmlns declaration) flags the element with XML_FLAG_NAMESPACE_WF_ERROR, even when the document carries no namespace declarations at all.
  • Internal-subset DTD validation runs when the document carries an internal DTD subset (VC Root Element Type, element content models, attribute constraints). The parser never performs I/O: external DTDs are only ever read through a caller-supplied resolveExternal callback.
  • An external-DTD reference surfaces as an error when nothing read the subset. A DOCTYPE that references an external subset (SYSTEM or PUBLIC) which no resolver read flags the DocType node as an error node and the Document root with LEX_NODE_HAS_ERROR | GRAMMAR_FLAG_NODE_ERROR. This is the strict contract “WF + reject external-DTD / undefined-entity”: a subset the parser never read cannot be certified under strict. A resolver that read the subset keeps the reference clean. Lenient mode never surfaces it.

The parse profiles compose validate and strict:

ProfileWhat runsUnresolved external-DTD document
default (lenient)WF-only recoveryparses clean
validate: truefull DTD validating passflagged (no root declaration available)
mode: "strict" / strict: truenamespace-prefix WF + internal-subset DTD validation + external-DTD surfaceflagged on the DocType + Document root
validate: true + strictfull DTD validating pass on top of the strict surfaceflagged
import { parseXml } from '@lanexio/parser-grammar-xml';
const encoder = new TextEncoder();
// Unbound element prefix: clean in lenient, flagged in strict
const lenient = parseXml(encoder.encode('<a:root/>'));
console.log(lenient.root.hasError); // false
const strict = parseXml(encoder.encode('<a:root/>'), { mode: 'strict' });
console.log(strict.root.hasError); // true
// The { strict: true } alias is the same option (ADR 0044)
const strictAlias = parseXml(encoder.encode('<a:root/>'), { strict: true });
console.log(strictAlias.root.hasError); // true
// Bound prefix stays clean in strict mode
const bound = parseXml(
encoder.encode('<a:root xmlns:a="urn:test"/>'),
{ mode: 'strict' },
);
console.log(bound.root.hasError); // false
// External-DTD reference: clean in lenient, flagged in strict when no
// resolver read the subset
const external = encoder.encode('<!DOCTYPE root SYSTEM "x.dtd"><root/>');
console.log(parseXml(external).root.hasError); // false
console.log(parseXml(external, { mode: 'strict' }).root.hasError); // true

Precedence is mode > strict > default lenient, so an explicit mode always wins over a strict boolean, and strict: false is the explicit lenient spelling. The unified parse() entry point forwards the same options through grammarOptions.mode or the top-level strict key, so parse(src, { language: "xml", grammarOptions: { mode: "strict" } }) and parse(src, { language: "xml", strict: true }) both behave identically to parseXml(bytes, { mode: "strict" }). Strict mode never throws; a rejected construct is a flag on the tree, never an exception. The validate: true option composes with strict and stays available under leniency on its own. See Leniency and strict mode for the full policy.

Elements, attributes, text nodes, comments, PIs, CDATA sections, and doctype declarations are all represented in the flat AST. The tree preserves the document’s hierarchical structure.

const cursor = tree.cursor();
cursor.gotoFirstChild(); // XmlDeclaration or DocumentElement
while (cursor.gotoNextSibling()) {
console.log(cursor.current.kind);
}