This page documents @lanexio/parser-grammar-csv, the CSV grammar package for parsing RFC 4180 compliant CSV and TSV (tab-separated values) with configurable delimiter, quoting, and header detection.
Version:Stable
Module name:parser-grammar-csv
Package:@lanexio/parser-grammar-csv
Import path:@lanexio/parser-grammar-csv
Layer: 2 (Grammar)
Runtime: Universal
Module format: ESM
Stability: Stable
Primary use case: Parse CSV and TSV data into a flat AST.
csvToValue() returns string[][] by default. When the first record carries CSV_FLAG_HEADER (the header parse option, on by default), the first record is consumed as the header and the result is Record<string, string>[]. Quoted fields strip the surrounding quotes and unescape "" to "; TSV cells are raw. Duplicate header names last-win with a duplicate_key diagnostic, rows wider than the header drop the excess fields with an excess_field diagnostic, and shorter rows omit the missing keys. Empty rows are preserved as empty arrays.
Skip lines that are entirely empty (zero-length or whitespace-only).
comments
number | readonly number[]
No
undefined
Byte value(s) that start a comment line. Lines starting with any of these bytes are skipped.
mode
"lenient" | "strict"
No
"lenient"
Lenient recovery (default) or strict flagging (ADR 0033). Forwarded to parseCsv.
strict
boolean
No
undefined
Convenience alias for mode: "strict" (ADR 0044). An explicit mode wins. Forwarded to parseCsv.
rfc4180LineEndings
boolean
No
false
Enforce CRLF-only record terminators per RFC 4180 §2.1. A bare LF or bare CR terminator is flagged on its Newline node. Forwarded to parseCsv. Independent of strict.
whitespaceBeforeQuote
boolean
No
false
Reject a field whose opening quote is preceded by whitespace (RFC 4180 §2.4). TSV disables quoting, so this is a no-op for TSV and is accepted only for API symmetry. Forwarded to parseCsv.
Every tree returned by parseCsv or parseTsv carries tree.metadata.language —
"csv" on parseCsv results and "tsv" on parseTsv results, including the
never-throw error-tree path. The stamp matches the unified parse() dispatcher and
the HTML/YAML/CSS grammar entry points.
Lanexio™ Parser keeps the default parse lenient: mode: "lenient" recovers from
malformed input without throwing, producing CsvKind.Error leaves. Under
mode: "strict" (or the strict: true alias), spec-level rejections are
surfaced as flags, never exceptions:
A record that fails to parse (unclosed quote, bad quote position) carries
LEX_NODE_HAS_ERROR | GRAMMAR_FLAG_NODE_ERROR on its Record node.
A record whose field count differs from the reference row is flagged the
same way. The reference is the header row when header: true (the
default), otherwise the first record.
On a column-count mismatch the tree’s metadata.diagnostics gains a
per-row entry: CSV record <n>: expected <h> fields, found <f> (TSV trees
use the TSV record prefix). The over-wide record’s surplus Field nodes
also carry LEX_NODE_HAS_ERROR, identifying the divergent column by byte
range.
When any record is rejected, the Document root carries
LEX_NODE_HAS_ERROR | GRAMMAR_FLAG_NODE_ERROR and tree.root.hasError
reads true.
Two independent toggles tighten RFC 4180 conformance without switching to
strict mode. Both default to false, keeping the default lenient, and both
work with or without strict:
Option
Effect
rfc4180LineEndings: true
Enforce CRLF-only record terminators (RFC 4180 §2.1). A bare LF or bare CR terminator is flagged on its Newline node. Newlines inside quoted fields and blank lines skipped by skipEmptyLines are unaffected.
whitespaceBeforeQuote: true
Reject a field whose opening quote is preceded by whitespace (RFC 4180 §2.4). The field is flagged on its Field node.
Under strict: true a toggle violation additionally flags the affected record
and the Document root, so tree.root.hasError reads true.