Skip to content

XML Conformance

Lanexio™ Parser implements XML 1.0 5th Edition with optional DTD validation. The conformance suite draws from the W3C XML Conformance Test Suite and the libxml2 test corpus, all under test_files/xml.

Experimental (1.0). @lanexio/parser-grammar-xml is marked experimental for the 1.0 release (versioned 0.1.0) and its API may change without a major version bump.

pnpm conformance:suite measures the XML grammar over the classified test_files/xml corpus (5,859 valid-classified and 191 invalid-classified files) and prints three rates from the same parse:

  • Raw classification conformance (83.3%): both directions over every classified file, (clean valid + flagged invalid) / (valid + invalid).
  • Valid-set conformance (85.8%): clean valid inputs over all valid-classified inputs, 5,027 of 5,859.
  • Valid-subset contract (100.0%): excluding the documented artifacts, 5,227 of 5,227 contract files parse clean. The contract gate is >= 95, with REAL-DEFECT 0.

Every residual flag falls into a documented artifact bucket (not-wf/invalid fixtures, xml-1.1, errata, libxml2 intentional errors, non-UTF-8 encodings, c14n, sun corpus drivers, parameter-entity dependence, and xmlconf invalid cases correctly flagged).

node scripts/conformance-classify.mjs (pnpm conformance:xml) is the package-level gate; its valid-subset contract matches the 100.0% above, and its strict-profile invalid-to-clean gate is <= 5 (external-DTD-only survivors of ADR 0041).

The strict profile (mode: "strict") adds the DTD-validity contract; residual strict flags are validity-coverage limitations (external DTDs are never resolved), not well-formedness defects.

pnpm conformance:suite measured this family over the test_files/xml corpus on 2026-10-06 at Lanexio™ Parser version 1.0.0:

familyvalidinvalidrawcontract-subsetREAL-DEFECT
xml585919183.3%100.0% (5227/5227)0

Over-acceptance (invalid-to-clean) disclosure

Section titled “Over-acceptance (invalid-to-clean) disclosure”

The valid-subset number above measures one direction only: valid-to-error. conformance-classify.mjs counts a file as conforming when the default lenient path parses it clean, so a not-well-formed file the parser accepts (the invalid-to-clean, or over-acceptance, class) is invisible to it. The 100.0% valid-subset number therefore claims nothing about over-acceptance.

The enforcement surface for the not-well-formed direction is the curated manifest ledger at corpus/xmlts/manifest.json, gated in CI by pnpm test:curated-suites. Two known classes are relevant here:

  • Fixed (PR3): declaration contradicting the BOM on the transcode path. A BOM-marked UTF-16/UTF-32 entity whose XML declaration encoding contradicts the BOM-detected encoding is a fatal error per XML 1.0 section 4.3.3. eduni/misc/008.xml (UTF-16BE BOM + encoding='utf-8') is flagged in both parse modes and its frozen ledger expectation (reject) holds again.
  • Known pre-existing gap (documented, not fixed): the UTF-8 path. The UTF-8 path never runs the declaration-vs-bytes check, so a UTF-8 BOM plus a non-UTF-8 declaration is accepted. eduni/misc/007.xml (UTF-8 BOM + encoding='iso-8859-1') is ledgered expect-fail (ENCODING_LENIENCY), so it is a documented carve-out rather than an uncovered divergence; it predates PR3 and stays a disclosed gap.

The parser autodetects a leading UTF-8, UTF-16LE/BE, or UTF-32LE/BE Byte Order Mark per XML 1.0 section 4.3.3. UTF-16 and UTF-32 entities are transcoded to UTF-8 on a rare path and every node range is rewritten back to the original byte offsets, so the flat AST, emitSource, and assertLossless round-trip the original bytes byte-exactly with no loss. A truncated or malformed UTF-16/UTF-32 entity produces a LexError tree and never throws.

On the transcode path the other half of section 4.3.3 is enforced: an entity whose XML declaration encoding contradicts the BOM-derived encoding is a fatal error, so the declaration node and the Document root carry the error flags in both parse modes. The UTF-8 path does not run that check, which is the documented eduni/misc/007.xml carve-out above. See ADR 0046 (docs/decisions/0046-xml-conformance-encoding.md).

The declared ASCII-compatible encodings Shift_JIS, EUC-JP, and ISO-2022-JP remain out of scope: entities that declare them (with no supported BOM) stay on the documented error path. LexNode.text and the value extractors decode the tree source as UTF-8, so for a UTF-16/UTF-32 tree they mojibake non-ASCII text; consumers that need decoded text from a non-UTF-8 document must transcode themselves. AST ranges and emitSource are byte-exact regardless.

The W3C XML Conformance Test Suite (xmlts20130923.zip) is vendored-and-pinned into corpus/xmlts/xmlconf/ with a manifest-ledger at corpus/xmlts/manifest.json, and gated in CI by pnpm test:curated-suites (see ADR 0034). A well-formedness parser (no DTD validation pass) gates unambiguous expectations on the apply subset, which is currently not-wf-only: 316 not-wf (must reject) tests. Valid-accept coverage is deferred behind DTD/entity support — every valid catalog test (including all 654 entity-free ones) carries a <!DOCTYPE and is scope DTD_ENTITY, so it is recorded but not asserted today.

  • Apply: 316 not-wf (must reject) tests gated by parseXml(bytes) default options.
  • ok: 290 — 290/316 meet their well-formedness expectation today.
  • ledger: 26 — 26 apply tests are ledgered expect-fail (currently fail the naive expectation; frozen, with entries removed as the parser improves): 14 namespace well-formedness gaps (NAMESPACE_WF), 7 structural WF gaps (WF_GAP), 3 XML 1.0 name-character leniency gaps (XML10_NAME_LENIENCY), 2 encoding-declaration gaps (ENCODING_LENIENCY).
  • scopeExcluded: 2270 — recorded, not asserted: DTD_ENTITY (requires DTD / external entity resolution or internal-DTD-subset parsing, 2068 tests, including all 812 valid catalog tests, which each carry a <!DOCTYPE) and WF_ONLY_SCOPE (expectation is DTD-validity-dependent: the invalid and error categories, 202 tests).

The ledger shrinks as DTD/entity support and strict-mode surfaces improve; an entry is added only to record a documented carve-out (ADR 0046). The apply set cannot silently shrink (the gate asserts the frozen counts).

parseXml accepts an optional ExternalResolver for DTD fetching. When provided, the parser validates DTD entities, notations, and element declarations. When absent — the default — it parses the document without validation.

parseXml accepts a mode option ("lenient" or "strict"). The default is "lenient", which keeps the recovery behavior of every previous release.

In "strict" mode the parse path still never throws, but two spec-level checks run in addition to the lenient surface:

  • Namespace prefix well-formedness, unconditionally. An element prefix that is not bound to any namespace URI flags the element node, and an attribute-name prefix that is not bound flags the attribute-name node, each with XML_FLAG_NAMESPACE_WF_ERROR, even in documents that declare no xmlns at all. The attribute check is strict-only; lenient never checks attribute prefixes. The reserved xml and xmlns prefixes stay exempt.
  • Internal-subset DTD validation. When the document carries an internal DTD subset (<!DOCTYPE ... [ ... ]>), the validating checks run: VC Root Element Type, element content models, and attribute constraints. External DTDs are never resolved in strict mode and no network or filesystem I/O happens.

When any strict check fails, the Document root carries LEX_NODE_HAS_ERROR | GRAMMAR_FLAG_NODE_ERROR and tree.root.hasError reads true. validate: true remains the separate, opt-in full DTD validation.

AxisResult
never-throw0 throws across 6,112 corpus files, both parse modes
well-formedcorpus WF = 0
losslesscorpus lossless = 0, byte-exact round-trip (UTF-8, UTF-16LE/BE, UTF-32LE/BE)
conformancevalid-subset 5,227/5,227 (100%), REAL-DEFECT 0, strict invalid-to-clean 1