@lanexio/parser-grammar-sql
This page documents @lanexio/parser-grammar-sql, the SQL grammar package targeting ANSI SQL:2016 core plus PostgreSQL, MySQL, SQLite, and T-SQL dialects. Dialect support is layered: a lexer-level recognition surface (quoting, comment styles, operator atoms, identifiers) plus a curated set of statement containers, per the ADR 0035 scope boundary. This package is at version 0.0.1 — the API may change before the 1.0.0 stable release.
- Version: Beta
- Module name:
parser-grammar-sql - Package:
@lanexio/parser-grammar-sql - Import path:
@lanexio/parser-grammar-sql - Layer: 2 (Grammar)
- Runtime: Universal
- Module format: ESM
- Stability: Beta (v0.0.1)
- Primary use case: Parse SQL statements into a flat AST.
Layer contract
Section titled “Layer contract”When to use this module
Section titled “When to use this module”- You need to parse SQL statements into a traversable AST, with an ANSI SQL:2016 core plus optional PostgreSQL, MySQL, SQLite, or T-SQL layered dialect support (lexer-level recognition extensions and curated statement containers).
- You are evaluating the SQL grammar for pre-1.0 use cases.
Module boundary
Section titled “Module boundary”| Boundary | Description |
|---|---|
| Inputs | Uint8Array (source bytes) |
| Outputs | LexTree (flat AST) |
| Side effects | None |
| Determinism | Yes |
| External dependencies | @lanexio/parser-core |
| Never-throw guarantee | Yes |
| Security surface | None (parse only, no output generation) |
Installation
Section titled “Installation”-
Install the package.
Terminal window pnpm add @lanexio/parser-grammar-sqlTerminal window npm install @lanexio/parser-grammar-sqlTerminal window yarn add @lanexio/parser-grammar-sql -
Import the named export.
import { parseSql } from '@lanexio/parser-grammar-sql';
Peer dependencies
Section titled “Peer dependencies”| Requirement | Required | Layer | Notes |
|---|---|---|---|
@lanexio/parser-core | Yes | 1 | ^1.0.0 |
Basic Usage
Section titled “Basic Usage”import { parseSql } from '@lanexio/parser-grammar-sql';
const encoder = new TextEncoder();const bytes = encoder.encode('SELECT id, name FROM users WHERE active = true');
const tree = parseSql(bytes);
console.log(tree.nodeCount);Exports
Section titled “Exports”| Export | Type | Description |
|---|---|---|
parseSql | (bytes: Uint8Array, options?: ParseSqlOptions) => LexTree | Parse SQL. Never throws. |
SqlKind | const object | Numeric kind IDs for all SQL node types. |
SqlField | const object | Numeric field IDs for SQL slots. |
SQL_FIELD_NAMES_BY_ID | readonly string[] | Field name lookup by numeric field ID. |
SQL_KIND_NAMES_BY_ID | readonly Record<number, string> | Kind-name lookup by numeric ID. |
LANEXIO_PARSER_GRAMMAR_SQL_PACKAGE_NAME | "@lanexio/parser-grammar-sql" | Stable npm package name constant for the SQL grammar package. |
sqlGrammar | LanexioParserPureGrammar | Grammar descriptor for use with parser-pure. |
sqlRegistration | GrammarRegistration | Registration for the unified grammar registry. |
SQL_DIALECTS | readonly SqlDialect[] | The five dialects the grammar targets. |
pickDialect | (options?: Record<string, unknown> | null) => SqlDialect | Reads .grammarOptions.dialect first, then a top-level dialect; defaults to "ansi" for absent or unrecognized values. |
Options
Section titled “Options”parseSql accepts three optional options:
| Option | Type | Default | Description |
|---|---|---|---|
dialect | SqlDialect | "ansi" | Selects lexer-level dialect rules. |
clientBatch | ClientBatchMode | "accept" | How the tsql dialect surfaces a batch-boundary GO separator and a sqlcmd :directive line: "accept", "split", or "reject". tsql-only. |
strict | boolean | false | When true and clientBatch is absent, behaves as clientBatch: "reject". An explicit clientBatch wins. |
An absent or unrecognized dialect value falls back to "ansi" deterministically. The selected dialect is stamped on tree.metadata.dialect.
The clientBatch mode is tsql-specific. On the other four dialects GO stays a
plain keyword and : stays a bind-marker context in every mode. "reject"
emits an Error node over each boundary GO or :directive and sets
hasError while the parse continues, so later statements still materialize.
"split" emits SqlKind.BatchBoundary (GO) and SqlKind.ClientDirective
(:directive) leaf nodes covering their exact byte spans, which a loader can
use to slice the file into batches. All three modes keep the flat AST a lossless
leaf partition and never throw.
Unified parse option shapes
Section titled “Unified parse option shapes”The registry-facing parse(src, { language: "sql", ... }) and
sqlGrammar.parse(src, options) accept the dialect in either shape:
.grammarOptions.dialect (the shape the unified parser forwards) or a top-level
dialect key. pickDialect reads .grammarOptions.dialect first and a
top-level dialect second. Both of these select the postgres dialect and parse
SELECT * FROM t WHERE id = $1 clean (and, since the ADR 0041 leniency
extension, the default ansi dialect accepts $1 too):
parse(bytes, { language: "sql", grammarOptions: { dialect: "postgres" } });parse(bytes, { language: "sql", dialect: "postgres" });Return shape
Section titled “Return shape”| Property | Type | Description |
|---|---|---|
root | LexNode | Root node of the SQL statement. |
nodeCount | number | Total nodes in the tree. |
source | Uint8Array | Original parsed bytes. |
Exported types
Section titled “Exported types”| Type | Purpose | Notes |
|---|---|---|
SqlKindType | Union of all SqlKind values | Type-safe kind reference |
SqlFieldType | Union of all SqlField values | Type-safe field reference |
SqlDialect | "ansi" | "postgres" | "mysql" | "sqlite" | "tsql" | The five dialects the grammar targets |
ClientBatchMode | "accept" | "reject" | "split" | Client-boundary handling for tsql (GO / sqlcmd directives) |
ParseSqlOptions | { dialect?: SqlDialect; clientBatch?: ClientBatchMode; strict?: boolean } | Options for parseSql |
Dialect scope
Section titled “Dialect scope”Lanexio™ Parser’s SQL grammar targets the curated intersection in ADR 0035:
ANSI SQL:2016 core as the always-on base, plus dialect variants for
PostgreSQL, MySQL, SQLite, and T-SQL. For PostgreSQL and T-SQL the surface is
curated statement containers plus lexer-level rules: the postgres DDL
statement-body containers and the T-SQL statement forms from ADR 0045, layered
over the lexer atoms (MySQL backticks and # comments, SQLite none, T-SQL
bracket identifiers) in the subsections below. Dialect selection is explicit,
never auto-detected from input: pass { dialect } or accept the "ansi"
default.
Scope boundary
Section titled “Scope boundary”Dialect support is a lexer-level recognition surface plus a curated statement container set; it is not a per-vendor parser.
In scope (v0.x line):
- ANSI SQL:2016 core as the always-on base and default.
- The five-dialect selector (
ansi,postgres,mysql,sqlite,tsql) with lexer-level rules: quoting, comment styles, operator atoms, and dialect identifiers. - The placeholder marker matrix (ADR 0041).
- The tsql-only client-batch surface: GO and sqlcmd directives, the
clientBatchoption (ADR 0045, ADR 0048). - Curated postgres and T-SQL DDL statement-body containers and the T-SQL statement forms enumerated in ADR 0045.
metadata.dialectstamping on every SQL tree.- Explicit dialect selection only; dialects are never auto-detected from input.
- The dialect-language aliases
tsql,mssql, andsqlserveron the unifiedparse().
Out of scope (v0.x line):
- Full per-vendor grammar and parser semantics: type resolution, operator-precedence semantics that change tree meaning, and vendor value semantics.
- PL/pgSQL and other procedural bodies as dialect grammars.
- Statement-level window-function and CTE+UNION dialect coverage, the documented postgres-regress residual.
- Optional SQL:2016 feature semantics.
- Anything beyond the five named dialects.
PostgreSQL lexer atoms (postgres dialect only)
Section titled “PostgreSQL lexer atoms (postgres dialect only)”::casts, dollar-quoted andE''strings.- Array literals
array[1,2]/ARRAY[1,2], subscriptsa[1], and type suffixestext[](LeftBracket/RightBracketkinds). Bracket atoms are postgres-only; under the ansi default they still error. - JSON operators
->,->>,#>,#>>,@>,<@,?,?|,?&(JsonOp). The bare?is a JSON operator under postgres only when followed by an operand; with no operand it reverts to the marker-rejection path. - Regex operators
~,~*,!~,!~*(RegexOp). - Trailing-colon casts
'1':int(ColonCast), after a string or numeric literal followed by a name-start byte. - DDL statement-body containers:
CREATE/DROP/ALTERfollowed by a recognized body head (sequence type role user aggregate operator function access statistics schema rule temp unique) opens aStatementcontainer that closes at;. An unsupported body construct emits a rangedErrornode rather than a silent root-only EOF flag.
Prepared-statement placeholders (SqlKind.Parameter)
Section titled “Prepared-statement placeholders (SqlKind.Parameter)”| Marker | Dialects | Notes |
|---|---|---|
$1 | postgres, ansi | ansi acceptance is a leniency extension (ADR 0041 decision 4) |
? | ansi, mysql, sqlite | mysql ? is the native prepared-statement marker; postgres ? is a JSON operator when followed by an operand |
:name | all | accepted everywhere, never diagnosed |
@p1 | tsql, mysql |
A well-formed marker is accepted in any position within an enabled dialect (the
grammar is a flat token scanner with no clause-level context validation). A
wrong-dialect marker emits a ranged Error node and appends a hint to
tree.diagnostics, for example SQL parameter marker `$1` requires the postgres or ansi dialect (active: mysql) or SQL parameter marker `?` requires the ansi, mysql, or sqlite dialect (active: postgres).
tree.diagnostics is an optional LexTreeMetadata field populated only by the
SQL grammar; treat it as possibly undefined. Known limits: sqlite’s ?NNN and
$name forms are not emitted as markers, and ? under postgres stays a JSON
operator by design.
T-SQL client-boundary atoms (tsql dialect, clientBatch: "split")
Section titled “T-SQL client-boundary atoms (tsql dialect, clientBatch: "split")”A batch-boundary GO separator and a sqlcmd :directive line at statement head
are client-boundary markers, not T-SQL statements (ADR 0045). Under
clientBatch: "split" the grammar surfaces them as first-class leaf atoms, each
covering its exact byte span so the AST stays a lossless leaf partition:
SqlKind.BatchBoundary(0x133D): aGOthat appears at a batch boundary (the statement-head latch is set after a;, anotherGO, or the start of the file). Mid-statementGOstays a plain keyword.SqlKind.ClientDirective(0x133E): a sqlcmd directive line such as:setvar DatabaseName TestDb,:r .\include.sql, or:connect server.
The numeric constants are hand-written (the SQL grammar has no Zig grammar and
no codegen gate), so a kind typo is caught only by the kind pins in
clientbatch.test.ts and the SQL_KIND_NAMES_BY_ID reverse-lookup.
Corpus mapping: test_files/tsql/manifest.json records the client_batches
fixtures as accept-client (valid only under sqlcmd or SSMS preprocessing) and
invalid/core_tsql/270 / 271 as reject-core. Under clientBatch: "reject"
those GO separators and directives carry hasError, the recover fixtures keep
a later statement node, and the conformance gate runs over all 155 cases with no
tolerated or skipped classifications.
Accessibility
Section titled “Accessibility”Accessibility requirements
Section titled “Accessibility requirements”- No direct accessibility surface. The SQL parser produces flat AST data structures.
Accessibility checklist
Section titled “Accessibility checklist”| Concern | Status |
|---|---|
| Generated output semantics | Not applicable (query format) |
| ARIA attributes in serialized output | Not applicable |
| Semantic element round-trip | Not applicable |
Security
Section titled “Security”Security considerations
Section titled “Security considerations”parseSqlnever throws on any byte sequence.
| Threat | Mitigation | Status |
|---|---|---|
| Malformed input byte sequence | Panic-free guarantee: all inputs accepted, errors produce LexError AST nodes | Implemented |
Companion packages
Section titled “Companion packages”| Package | Relationship | Layer | Notes |
|---|---|---|---|
@lanexio/parser-core | Requires | 1 | Provides LexTree, LexNode, LexCursor |
@lanexio/parser | Consumes | 6 | Unified entry point |
Beta status
Section titled “Beta status”This package is at v0.0.1 and is in active development. The API surface is minimal and may change significantly before the 1.0.0 stable release. The package is published to npm for early evaluation. Use with caution in production.
Planned additions for v1.0.0 include:
- Grammar-level per-vendor coverage beyond the current lexer-level recognition surface and curated statement containers (ADR 0035 keeps this out of scope for the v0.x line)
- Statement-level coverage for window functions, CTE+UNION, and PL/pgSQL bodies, currently a documented coverage gap
Changelog
Section titled “Changelog”| Version | Date | Status | Notable changes |
|---|---|---|---|
0.0.1 | 2026-08-28 | Current | Postgres lexer atoms (arrays, JSON/regex operators, colon casts), DDL statement-body containers, top-level dialect option, placeholder rejection diagnostics. |
0.0.1 | 2026-05-29 | Prior | Initial beta release. Apache-2.0. |
Migration notes
Section titled “Migration notes”- This is a beta release. No migration guarantees until v1.0.0.