XML Conformance
Lanexio™ Parser implements XML 1.0 5th Edition with optional DTD validation. The conformance suite draws from the W3C XML Conformance Test Suite and the libxml2 test corpus, all under test_files/xml.
Experimental (1.0). @lanexio/parser-grammar-xml is marked experimental for the 1.0 release (versioned 0.1.0) and its API may change without a major version bump.
Test results
Section titled “Test results”pnpm conformance:suite measures the XML grammar over the classified test_files/xml corpus (5,859 valid-classified and 191 invalid-classified files) and prints three rates from the same parse:
- Raw classification conformance (83.3%): both directions over every classified file,
(clean valid + flagged invalid) / (valid + invalid). - Valid-set conformance (85.8%): clean valid inputs over all valid-classified inputs, 5,027 of 5,859.
- Valid-subset contract (100.0%): excluding the documented artifacts, 5,227 of 5,227 contract files parse clean. The contract gate is
>= 95, withREAL-DEFECT 0.
Every residual flag falls into a documented artifact bucket (not-wf/invalid fixtures, xml-1.1, errata, libxml2 intentional errors, non-UTF-8 encodings, c14n, sun corpus drivers, parameter-entity dependence, and xmlconf invalid cases correctly flagged).
node scripts/conformance-classify.mjs (pnpm conformance:xml) is the package-level gate; its valid-subset contract matches the 100.0% above, and its strict-profile invalid-to-clean gate is <= 5 (external-DTD-only survivors of ADR 0041).
The strict profile (mode: "strict") adds the DTD-validity contract; residual strict flags are validity-coverage limitations (external DTDs are never resolved), not well-formedness defects.
Corpus conformance
Section titled “Corpus conformance”pnpm conformance:suite measured this family over the test_files/xml corpus on 2026-10-06 at Lanexio™ Parser version 1.0.0:
| family | valid | invalid | raw | contract-subset | REAL-DEFECT |
|---|---|---|---|---|---|
| xml | 5859 | 191 | 83.3% | 100.0% (5227/5227) | 0 |
Over-acceptance (invalid-to-clean) disclosure
Section titled “Over-acceptance (invalid-to-clean) disclosure”The valid-subset number above measures one direction only: valid-to-error.
conformance-classify.mjs counts a file as conforming when the default lenient
path parses it clean, so a not-well-formed file the parser accepts (the
invalid-to-clean, or over-acceptance, class) is invisible to it. The 100.0%
valid-subset number therefore claims nothing about over-acceptance.
The enforcement surface for the not-well-formed direction is the curated
manifest ledger at corpus/xmlts/manifest.json, gated in CI by
pnpm test:curated-suites. Two known classes are relevant here:
- Fixed (PR3): declaration contradicting the BOM on the transcode path.
A BOM-marked UTF-16/UTF-32 entity whose XML declaration
encodingcontradicts the BOM-detected encoding is a fatal error per XML 1.0 section 4.3.3.eduni/misc/008.xml(UTF-16BE BOM +encoding='utf-8') is flagged in both parse modes and its frozen ledger expectation (reject) holds again. - Known pre-existing gap (documented, not fixed): the UTF-8 path. The
UTF-8 path never runs the declaration-vs-bytes check, so a UTF-8 BOM plus a
non-UTF-8 declaration is accepted.
eduni/misc/007.xml(UTF-8 BOM +encoding='iso-8859-1') is ledgeredexpect-fail(ENCODING_LENIENCY), so it is a documented carve-out rather than an uncovered divergence; it predates PR3 and stays a disclosed gap.
Encoding contract
Section titled “Encoding contract”The parser autodetects a leading UTF-8, UTF-16LE/BE, or UTF-32LE/BE Byte Order
Mark per XML 1.0 section 4.3.3.
UTF-16 and UTF-32 entities are transcoded to UTF-8 on a rare path and every
node range is rewritten back to the original byte offsets, so the flat AST,
emitSource, and assertLossless round-trip the original bytes byte-exactly
with no loss. A truncated or malformed UTF-16/UTF-32 entity produces a
LexError tree and never throws.
On the transcode path the other half of section 4.3.3 is enforced: an entity
whose XML declaration encoding contradicts the BOM-derived encoding is a
fatal error, so the declaration node and the Document root carry the error
flags in both parse modes. The UTF-8 path does not run that check, which is
the documented eduni/misc/007.xml carve-out above. See ADR 0046
(docs/decisions/0046-xml-conformance-encoding.md).
The declared ASCII-compatible encodings Shift_JIS, EUC-JP, and ISO-2022-JP
remain out of scope: entities that declare them (with no supported BOM) stay
on the documented error path. LexNode.text and the value extractors decode
the tree source as UTF-8, so for a UTF-16/UTF-32 tree they mojibake non-ASCII
text; consumers that need decoded text from a non-UTF-8 document must transcode
themselves. AST ranges and emitSource are byte-exact regardless.
Curated xmlts gate (CI)
Section titled “Curated xmlts gate (CI)”The W3C XML Conformance Test Suite (xmlts20130923.zip) is vendored-and-pinned
into corpus/xmlts/xmlconf/ with a manifest-ledger at
corpus/xmlts/manifest.json, and gated in CI by pnpm test:curated-suites (see
ADR 0034). A well-formedness parser (no DTD validation pass) gates unambiguous
expectations on the apply subset, which is currently not-wf-only: 316
not-wf (must reject) tests. Valid-accept coverage is deferred behind
DTD/entity support — every valid catalog test (including all 654 entity-free
ones) carries a <!DOCTYPE and is scope DTD_ENTITY, so it is recorded but
not asserted today.
- Apply: 316
not-wf(must reject) tests gated byparseXml(bytes)default options. - ok: 290 — 290/316 meet their well-formedness expectation today.
- ledger: 26 — 26 apply tests are ledgered
expect-fail(currently fail the naive expectation; frozen, with entries removed as the parser improves): 14 namespace well-formedness gaps (NAMESPACE_WF), 7 structural WF gaps (WF_GAP), 3 XML 1.0 name-character leniency gaps (XML10_NAME_LENIENCY), 2 encoding-declaration gaps (ENCODING_LENIENCY). - scopeExcluded: 2270 — recorded, not asserted:
DTD_ENTITY(requires DTD / external entity resolution or internal-DTD-subset parsing, 2068 tests, including all 812validcatalog tests, which each carry a<!DOCTYPE) andWF_ONLY_SCOPE(expectation is DTD-validity-dependent: theinvalidanderrorcategories, 202 tests).
The ledger shrinks as DTD/entity support and strict-mode surfaces improve; an
entry is added only to record a documented carve-out (ADR 0046). The apply
set cannot silently shrink (the gate asserts the frozen counts).
DTD validation
Section titled “DTD validation”parseXml accepts an optional ExternalResolver for DTD fetching. When provided, the parser validates DTD entities, notations, and element declarations. When absent — the default — it parses the document without validation.
Strict mode
Section titled “Strict mode”parseXml accepts a mode option ("lenient" or "strict"). The default is "lenient", which keeps the recovery behavior of every previous release.
In "strict" mode the parse path still never throws, but two spec-level checks run in addition to the lenient surface:
- Namespace prefix well-formedness, unconditionally. An element prefix that is not bound to any namespace URI flags the element node, and an attribute-name prefix that is not bound flags the attribute-name node, each with
XML_FLAG_NAMESPACE_WF_ERROR, even in documents that declare noxmlnsat all. The attribute check is strict-only; lenient never checks attribute prefixes. The reservedxmlandxmlnsprefixes stay exempt. - Internal-subset DTD validation. When the document carries an internal DTD subset (
<!DOCTYPE ... [ ... ]>), the validating checks run: VC Root Element Type, element content models, and attribute constraints. External DTDs are never resolved in strict mode and no network or filesystem I/O happens.
When any strict check fails, the Document root carries LEX_NODE_HAS_ERROR | GRAMMAR_FLAG_NODE_ERROR and tree.root.hasError reads true. validate: true remains the separate, opt-in full DTD validation.
Four-axis bar
Section titled “Four-axis bar”| Axis | Result |
|---|---|
| never-throw | 0 throws across 6,112 corpus files, both parse modes |
| well-formed | corpus WF = 0 |
| lossless | corpus lossless = 0, byte-exact round-trip (UTF-8, UTF-16LE/BE, UTF-32LE/BE) |
| conformance | valid-subset 5,227/5,227 (100%), REAL-DEFECT 0, strict invalid-to-clean 1 |