Skip to content

@lanexio/parser-grammar-sql

This page documents @lanexio/parser-grammar-sql, the SQL grammar package targeting ANSI SQL:2016 core plus PostgreSQL, MySQL, SQLite, and T-SQL dialects. Dialect support is layered: a lexer-level recognition surface (quoting, comment styles, operator atoms, identifiers) plus a curated set of statement containers, per the ADR 0035 scope boundary. This package is at version 0.0.1 — the API may change before the 1.0.0 stable release.

  • Version: Beta
  • Module name: parser-grammar-sql
  • Package: @lanexio/parser-grammar-sql
  • Import path: @lanexio/parser-grammar-sql
  • Layer: 2 (Grammar)
  • Runtime: Universal
  • Module format: ESM
  • Stability: Beta (v0.0.1)
  • Primary use case: Parse SQL statements into a flat AST.
  • You need to parse SQL statements into a traversable AST, with an ANSI SQL:2016 core plus optional PostgreSQL, MySQL, SQLite, or T-SQL layered dialect support (lexer-level recognition extensions and curated statement containers).
  • You are evaluating the SQL grammar for pre-1.0 use cases.
BoundaryDescription
InputsUint8Array (source bytes)
OutputsLexTree (flat AST)
Side effectsNone
DeterminismYes
External dependencies@lanexio/parser-core
Never-throw guaranteeYes
Security surfaceNone (parse only, no output generation)
  1. Install the package.

    Terminal window
    pnpm add @lanexio/parser-grammar-sql
  2. Import the named export.

    import { parseSql } from '@lanexio/parser-grammar-sql';
RequirementRequiredLayerNotes
@lanexio/parser-coreYes1^1.0.0
import { parseSql } from '@lanexio/parser-grammar-sql';
const encoder = new TextEncoder();
const bytes = encoder.encode('SELECT id, name FROM users WHERE active = true');
const tree = parseSql(bytes);
console.log(tree.nodeCount);
ExportTypeDescription
parseSql(bytes: Uint8Array, options?: ParseSqlOptions) => LexTreeParse SQL. Never throws.
SqlKindconst objectNumeric kind IDs for all SQL node types.
SqlFieldconst objectNumeric field IDs for SQL slots.
SQL_FIELD_NAMES_BY_IDreadonly string[]Field name lookup by numeric field ID.
SQL_KIND_NAMES_BY_IDreadonly Record<number, string>Kind-name lookup by numeric ID.
LANEXIO_PARSER_GRAMMAR_SQL_PACKAGE_NAME"@lanexio/parser-grammar-sql"Stable npm package name constant for the SQL grammar package.
sqlGrammarLanexioParserPureGrammarGrammar descriptor for use with parser-pure.
sqlRegistrationGrammarRegistrationRegistration for the unified grammar registry.
SQL_DIALECTSreadonly SqlDialect[]The five dialects the grammar targets.
pickDialect(options?: Record<string, unknown> | null) => SqlDialectReads .grammarOptions.dialect first, then a top-level dialect; defaults to "ansi" for absent or unrecognized values.

parseSql accepts three optional options:

OptionTypeDefaultDescription
dialectSqlDialect"ansi"Selects lexer-level dialect rules.
clientBatchClientBatchMode"accept"How the tsql dialect surfaces a batch-boundary GO separator and a sqlcmd :directive line: "accept", "split", or "reject". tsql-only.
strictbooleanfalseWhen true and clientBatch is absent, behaves as clientBatch: "reject". An explicit clientBatch wins.

An absent or unrecognized dialect value falls back to "ansi" deterministically. The selected dialect is stamped on tree.metadata.dialect.

The clientBatch mode is tsql-specific. On the other four dialects GO stays a plain keyword and : stays a bind-marker context in every mode. "reject" emits an Error node over each boundary GO or :directive and sets hasError while the parse continues, so later statements still materialize. "split" emits SqlKind.BatchBoundary (GO) and SqlKind.ClientDirective (:directive) leaf nodes covering their exact byte spans, which a loader can use to slice the file into batches. All three modes keep the flat AST a lossless leaf partition and never throw.

The registry-facing parse(src, { language: "sql", ... }) and sqlGrammar.parse(src, options) accept the dialect in either shape: .grammarOptions.dialect (the shape the unified parser forwards) or a top-level dialect key. pickDialect reads .grammarOptions.dialect first and a top-level dialect second. Both of these select the postgres dialect and parse SELECT * FROM t WHERE id = $1 clean (and, since the ADR 0041 leniency extension, the default ansi dialect accepts $1 too):

parse(bytes, { language: "sql", grammarOptions: { dialect: "postgres" } });
parse(bytes, { language: "sql", dialect: "postgres" });
PropertyTypeDescription
rootLexNodeRoot node of the SQL statement.
nodeCountnumberTotal nodes in the tree.
sourceUint8ArrayOriginal parsed bytes.
TypePurposeNotes
SqlKindTypeUnion of all SqlKind valuesType-safe kind reference
SqlFieldTypeUnion of all SqlField valuesType-safe field reference
SqlDialect"ansi" | "postgres" | "mysql" | "sqlite" | "tsql"The five dialects the grammar targets
ClientBatchMode"accept" | "reject" | "split"Client-boundary handling for tsql (GO / sqlcmd directives)
ParseSqlOptions{ dialect?: SqlDialect; clientBatch?: ClientBatchMode; strict?: boolean }Options for parseSql

Lanexio™ Parser’s SQL grammar targets the curated intersection in ADR 0035: ANSI SQL:2016 core as the always-on base, plus dialect variants for PostgreSQL, MySQL, SQLite, and T-SQL. For PostgreSQL and T-SQL the surface is curated statement containers plus lexer-level rules: the postgres DDL statement-body containers and the T-SQL statement forms from ADR 0045, layered over the lexer atoms (MySQL backticks and # comments, SQLite none, T-SQL bracket identifiers) in the subsections below. Dialect selection is explicit, never auto-detected from input: pass { dialect } or accept the "ansi" default.

Dialect support is a lexer-level recognition surface plus a curated statement container set; it is not a per-vendor parser.

In scope (v0.x line):

  • ANSI SQL:2016 core as the always-on base and default.
  • The five-dialect selector (ansi, postgres, mysql, sqlite, tsql) with lexer-level rules: quoting, comment styles, operator atoms, and dialect identifiers.
  • The placeholder marker matrix (ADR 0041).
  • The tsql-only client-batch surface: GO and sqlcmd directives, the clientBatch option (ADR 0045, ADR 0048).
  • Curated postgres and T-SQL DDL statement-body containers and the T-SQL statement forms enumerated in ADR 0045.
  • metadata.dialect stamping on every SQL tree.
  • Explicit dialect selection only; dialects are never auto-detected from input.
  • The dialect-language aliases tsql, mssql, and sqlserver on the unified parse().

Out of scope (v0.x line):

  • Full per-vendor grammar and parser semantics: type resolution, operator-precedence semantics that change tree meaning, and vendor value semantics.
  • PL/pgSQL and other procedural bodies as dialect grammars.
  • Statement-level window-function and CTE+UNION dialect coverage, the documented postgres-regress residual.
  • Optional SQL:2016 feature semantics.
  • Anything beyond the five named dialects.

PostgreSQL lexer atoms (postgres dialect only)

Section titled “PostgreSQL lexer atoms (postgres dialect only)”
  • :: casts, dollar-quoted and E'' strings.
  • Array literals array[1,2] / ARRAY[1,2], subscripts a[1], and type suffixes text[] (LeftBracket/RightBracket kinds). Bracket atoms are postgres-only; under the ansi default they still error.
  • JSON operators ->, ->>, #>, #>>, @>, <@, ?, ?|, ?& (JsonOp). The bare ? is a JSON operator under postgres only when followed by an operand; with no operand it reverts to the marker-rejection path.
  • Regex operators ~, ~*, !~, !~* (RegexOp).
  • Trailing-colon casts '1':int (ColonCast), after a string or numeric literal followed by a name-start byte.
  • DDL statement-body containers: CREATE/DROP/ALTER followed by a recognized body head (sequence type role user aggregate operator function access statistics schema rule temp unique) opens a Statement container that closes at ;. An unsupported body construct emits a ranged Error node rather than a silent root-only EOF flag.

Prepared-statement placeholders (SqlKind.Parameter)

Section titled “Prepared-statement placeholders (SqlKind.Parameter)”
MarkerDialectsNotes
$1postgres, ansiansi acceptance is a leniency extension (ADR 0041 decision 4)
?ansi, mysql, sqlitemysql ? is the native prepared-statement marker; postgres ? is a JSON operator when followed by an operand
:nameallaccepted everywhere, never diagnosed
@p1tsql, mysql

A well-formed marker is accepted in any position within an enabled dialect (the grammar is a flat token scanner with no clause-level context validation). A wrong-dialect marker emits a ranged Error node and appends a hint to tree.diagnostics, for example SQL parameter marker `$1` requires the postgres or ansi dialect (active: mysql) or SQL parameter marker `?` requires the ansi, mysql, or sqlite dialect (active: postgres). tree.diagnostics is an optional LexTreeMetadata field populated only by the SQL grammar; treat it as possibly undefined. Known limits: sqlite’s ?NNN and $name forms are not emitted as markers, and ? under postgres stays a JSON operator by design.

T-SQL client-boundary atoms (tsql dialect, clientBatch: "split")

Section titled “T-SQL client-boundary atoms (tsql dialect, clientBatch: "split")”

A batch-boundary GO separator and a sqlcmd :directive line at statement head are client-boundary markers, not T-SQL statements (ADR 0045). Under clientBatch: "split" the grammar surfaces them as first-class leaf atoms, each covering its exact byte span so the AST stays a lossless leaf partition:

  • SqlKind.BatchBoundary (0x133D): a GO that appears at a batch boundary (the statement-head latch is set after a ;, another GO, or the start of the file). Mid-statement GO stays a plain keyword.
  • SqlKind.ClientDirective (0x133E): a sqlcmd directive line such as :setvar DatabaseName TestDb, :r .\include.sql, or :connect server.

The numeric constants are hand-written (the SQL grammar has no Zig grammar and no codegen gate), so a kind typo is caught only by the kind pins in clientbatch.test.ts and the SQL_KIND_NAMES_BY_ID reverse-lookup.

Corpus mapping: test_files/tsql/manifest.json records the client_batches fixtures as accept-client (valid only under sqlcmd or SSMS preprocessing) and invalid/core_tsql/270 / 271 as reject-core. Under clientBatch: "reject" those GO separators and directives carry hasError, the recover fixtures keep a later statement node, and the conformance gate runs over all 155 cases with no tolerated or skipped classifications.

  • No direct accessibility surface. The SQL parser produces flat AST data structures.
ConcernStatus
Generated output semanticsNot applicable (query format)
ARIA attributes in serialized outputNot applicable
Semantic element round-tripNot applicable
  • parseSql never throws on any byte sequence.
ThreatMitigationStatus
Malformed input byte sequencePanic-free guarantee: all inputs accepted, errors produce LexError AST nodesImplemented
PackageRelationshipLayerNotes
@lanexio/parser-coreRequires1Provides LexTree, LexNode, LexCursor
@lanexio/parserConsumes6Unified entry point

This package is at v0.0.1 and is in active development. The API surface is minimal and may change significantly before the 1.0.0 stable release. The package is published to npm for early evaluation. Use with caution in production.

Planned additions for v1.0.0 include:

  • Grammar-level per-vendor coverage beyond the current lexer-level recognition surface and curated statement containers (ADR 0035 keeps this out of scope for the v0.x line)
  • Statement-level coverage for window functions, CTE+UNION, and PL/pgSQL bodies, currently a documented coverage gap
VersionDateStatusNotable changes
0.0.12026-08-28CurrentPostgres lexer atoms (arrays, JSON/regex operators, colon casts), DDL statement-body containers, top-level dialect option, placeholder rejection diagnostics.
0.0.12026-05-29PriorInitial beta release. Apache-2.0.
  • This is a beta release. No migration guarantees until v1.0.0.