Skip to content

SQL Conformance

Lanexio™ Parser implements SQL:2016 core as the always-on base and four named dialects (PostgreSQL, MySQL, SQLite, and T-SQL) as a layered scope: lexer-level recognition extensions (quoting, comment styles, operator atoms, identifiers) plus a curated set of statement containers. Full per-vendor grammar coverage is out of scope for the v0.x line. ADR 0035 records this scope decision.

Experimental (1.0). @lanexio/parser-grammar-sql is marked experimental for the 1.0 release (versioned 0.0.1) and its API may change without a major version bump.

The SQL grammar targets the curated intersection defined in ADR 0035:

  • ANSI SQL:2016 core is always on and is the default dialect.
  • parseSql(bytes, { dialect }) and the unified parse(src, { language: "sql", grammarOptions: { dialect } }) select one of 'ansi', 'postgres', 'mysql', 'sqlite', or 'tsql'.
  • The unified parse(src, { language: "tsql" }) (also mssql / sqlserver) routes to the SQL grammar with the tsql dialect forced (ADR 0045).
  • Each named dialect adds lexer-level rules: backtick identifiers and # comments (MySQL), :: casts plus dollar-quoted and E'' strings (PostgreSQL), bracket identifiers (T-SQL), and the shared comment styles.

The parser does not claim full per-vendor grammar coverage and does not target “28 dialects”. Full vendor grammar remains out of scope (ADR 0035, ADR 0041); the placeholder dialect matrix is closed in ADR 0041 and documented below.

Dialect support is a lexer-level recognition surface plus a curated statement container set; it is not a per-vendor parser.

In scope (v0.x line):

  • ANSI SQL:2016 core as the always-on base and default.
  • The five-dialect selector (ansi, postgres, mysql, sqlite, tsql) with lexer-level rules: quoting, comment styles, operator atoms, and dialect identifiers.
  • The placeholder marker matrix (ADR 0041).
  • The tsql-only client-batch surface: GO and sqlcmd directives, the clientBatch option (ADR 0045, ADR 0048).
  • Curated postgres and T-SQL DDL statement-body containers and the T-SQL statement forms enumerated in ADR 0045.
  • metadata.dialect stamping on every SQL tree.
  • Explicit dialect selection only; dialects are never auto-detected from input.
  • The dialect-language aliases tsql, mssql, and sqlserver on the unified parse().

Out of scope (v0.x line):

  • Full per-vendor grammar and parser semantics: type resolution, operator-precedence semantics that change tree meaning, and vendor value semantics.
  • PL/pgSQL and other procedural bodies as dialect grammars.
  • Statement-level window-function and CTE+UNION dialect coverage, the documented postgres-regress residual.
  • Optional SQL:2016 feature semantics.
  • Anything beyond the five named dialects.

Prepared-statement bind markers are emitted as SqlKind.Parameter. The per-dialect acceptance matrix (ADR 0041 amendment, 2026-08-28):

Markeransipostgresmysqlsqlitetsql
?acceptreject (JSON operator)acceptacceptaccept (ODBC)
$1acceptacceptrejectrejectmoney literal
:nameacceptacceptacceptacceptaccept
@p1rejectrejectacceptrejectaccept

? under mysql is the native prepared-statement marker (bound via mysql_stmt_bind_param); $1 under ansi is a deliberate leniency extension for tooling that emits postgres-style markers and parses with the default dialect. metadata.dialect still reports the active dialect, so tooling can detect a $1-under-ansi mismatch. Under tsql, ? is accepted as an ODBC-style parameter placeholder, and $1 is not a marker at all: it lexes as a money literal (ADR 0045).

pnpm conformance:sql runs scripts/conformance-classify.mjs, which parses the whole test_files/sql corpus (3795 files) with per-family dialect selection and reports two SQL gates:

GateValue
Whole-corpus valid-to-error< 190 (conform > 95% of 3795)
Contract-subset conform %>= 95
SQL REAL-DEFECT0

Artifact buckets, documented and rule-based, separate client-script content from statements: sqlglot semantic fixtures (by family), psql meta-command batches (a line whose first non-whitespace bytes are \ + letter), and #!-header client scripts. Artifact files are parsed for never-throw and stay in the whole-corpus count, but drop out of the contract-subset denominator.

Measured on this tree: whole-corpus valid-to-error is 177 (< 190) and the contract-subset conform is 98.4% (3589/3646). The unified pnpm conformance:suite command additionally reports raw classification conformance 93.1% over the full classified corpus; see the corpus block below.

The postgres-regress statement-level residual (window functions, PL/pgSQL bodies, and other statement-level constructs) is disclosed in the coverage-gap bucket per ADR 0035 and ADR 0041; it is never absorbed into the conform number.

pnpm conformance:suite measured this family over the test_files/sql corpus on 2026-10-06 at Lanexio™ Parser version 1.0.0:

familyvalidinvalidrawcontract-subsetREAL-DEFECT
sql3791493.1%98.4% (3589/3646)0

The raw figure counts both directions over every classified file; the contract-subset is the existing per-dialect SQL contract-subset (the same predicate pnpm conformance:sql gates), so the published contract figure cannot diverge from the package gate.

  • sqlite’s ?NNN and $name forms are not emitted as markers; only bare ?, :name, and @name are.
  • Markers are accepted in any position: the grammar is a flat token scanner with no clause-context validation, the same leniency class as SELECT * FROM 1.
  • ? under postgres is a JSON operator by design, not a marker; a bare ? with no operand reverts to the marker-rejection path.

SQL conformance is gated by the committed fixtures in corpus/sql/ and the package test suite:

GateWhat it proves
corpus/sql/ fixturesCommitted ANSI-core statements parse with root.hasError === false, well-formed and lossless.
dialect.test.tsThe per-dialect lexer surface parses under its dialect and is rejected under the ANSI default; metadata.dialect is stamped.
never-throwEvery parse path returns a LexTree; malformed input yields SqlKind.Error nodes, never exceptions.

The exhaustive harness (pnpm test:exhaustive:sql) runs a corrected-gate entry for SQL: byte-tiled assertWellFormed, byte-lossless emitSource, and metadata-stamped (metadata.language === "sql") in all six modes. The pass-rate bar is deferred; the GRAMMAR_GATES entry is metadata-stamp-only this run.

Every parse tree carries metadata.language === "sql", and the selected dialect is stamped on metadata.dialect (defaults to "ansi").