Parsing CSV and TSV
Lanexio™ Parser handles CSV and TSV through a single package with configurable parsing. Both formats use the same never-throw parse path and produce zero-copy flat ASTs.
Quick start
Section titled “Quick start”-
Install the package.
Terminal window pnpm add @lanexio/parser-grammar-csvTerminal window npm install @lanexio/parser-grammar-csvTerminal window yarn add @lanexio/parser-grammar-csv -
Parse a CSV file.
import { parseCsv } from '@lanexio/parser-grammar-csv';const encoder = new TextEncoder();const tree = parseCsv(encoder.encode('Name,Age,City\nAlice,30,NYC\nBob,25,LA')); -
Parse with options.
import { parseCsv } from '@lanexio/parser-grammar-csv';const tree = parseCsv(encoder.encode('Name|Age|City\nAlice|30|NYC'), {delimiter: 0x7c, // '|' as a byte; the delimiter is a byte, not a stringheader: true,});
TSV files
Section titled “TSV files”import { parseTsv } from '@lanexio/parser-grammar-csv';
const tree = parseTsv(encoder.encode('Name\tAge\tCity\nAlice\t30\tNYC'));Options
Section titled “Options”| Option | Type | Default | Description |
|---|---|---|---|
delimiter | number | 0x2C (,) | Field delimiter byte. The delimiter is a byte, not a string |
quote | number | 0x22 (") | Quote character byte |
header | boolean | true | Whether the first record is a header row |
skipEmptyLines | boolean | false | Whether to skip empty lines |
mode | "lenient" | "strict" | "lenient" | Lenient recovery (default) or strict flagging |
strict | boolean | undefined | Convenience alias for mode: "strict" (ADR 0044). true picks strict, false is the explicit lenient spelling; an explicit mode wins |
rfc4180LineEndings | boolean | false | Enforce CRLF-only record terminators per RFC 4180 §2.1. A bare LF or bare CR terminator is flagged as an error. Independent toggle, works with or without strict |
whitespaceBeforeQuote | boolean | false | Reject a field whose opening quote is preceded by whitespace (RFC 4180 §2.4). Flags the field as an error. Independent toggle, works with or without strict |
Strict mode
Section titled “Strict mode”The default (mode: "lenient") keeps the historical recovery behavior:
malformed records produce error leaves but the parser never throws. Strict
mode flags spec-level rejections on the tree instead of only the offending
field: an unclosed quote and a record whose field count differs from the
reference row (the header row when header is true, otherwise the first
record) both set hasError on the record and the document root.
import { parseCsv } from '@lanexio/parser-grammar-csv';
const encoder = new TextEncoder();
// Unclosed quote: clean in lenient, flagged in strictconst lenient = parseCsv(encoder.encode('a,b\n"unclosed,1\n2,3'));console.log(lenient.root.hasError); // false
const strict = parseCsv(encoder.encode('a,b\n"unclosed,1\n2,3'), { mode: 'strict' });console.log(strict.root.hasError); // true
// Column-count mismatch: clean in lenient, flagged in strictconst wide = parseCsv(encoder.encode('a,b,c\n1,2\n3,4,5\n6,7'), { mode: 'strict' });console.log(wide.root.hasError); // trueThe { strict: true } alias (ADR 0044) is the same option under the
ecosystem-standard spelling: parseCsv(bytes, { strict: true }) produces the
same flags as parseCsv(bytes, { mode: "strict" }), and strict: false is the
explicit lenient spelling. Precedence is mode > strict > default lenient,
so an explicit mode always wins over a strict boolean. The unified
parse() entry point forwards the same options through grammarOptions.mode
or the top-level strict key, so parse(src, { language: "csv", grammarOptions: { mode: "strict" } }) and parse(src, { language: "csv", strict: true }) both behave identically to parseCsv(bytes, { mode: "strict" }). Strict mode never throws; a rejected record is a flag on the tree, never
an exception. Well-formed input parses clean in both modes. See
Leniency and strict mode for the full policy.
Per-row diagnostics
Section titled “Per-row diagnostics”On a column-count mismatch in strict mode, the tree’s metadata.diagnostics
gains a per-row entry that names the record and the expected-vs-found counts:
const tree = parseCsv(encoder.encode('a,b,c\n1,2\n3,4,5\n6,7'), { mode: 'strict' });console.log(tree.diagnostics);// ["CSV record 2: expected 3 fields, found 2", "CSV record 4: expected 3 fields, found 2"]The surplus Field nodes of an over-wide record also carry
LEX_NODE_HAS_ERROR, so the AST identifies the divergent column by byte
range. TSV diagnostics use the same shape with a TSV record prefix.
RFC toggles
Section titled “RFC toggles”Two independent toggles tighten RFC 4180 conformance without switching the
whole parse to strict mode. Both default to false, keeping the default
lenient, and both work with or without strict:
rfc4180LineEndingsenforces CRLF-only record terminators (RFC 4180 §2.1). A bare LF or bare CR terminator is flagged as an error on itsNewlinenode. Newlines inside quoted fields and blank lines skipped byskipEmptyLinesare unaffected.whitespaceBeforeQuoterejects a field whose opening quote is preceded by whitespace (RFC 4180 §2.4). The field is flagged as an error on itsFieldnode.
// A bare-LF record terminator is flagged when the toggle is on.const crlf = parseCsv(encoder.encode('a,b\r\n1,2\r\n'), { rfc4180LineEndings: true });console.log(crlf.root.hasError); // false
const lf = parseCsv(encoder.encode('a,b\n1,2'), { rfc4180LineEndings: true });// the bare-LF Newline node carries LEX_NODE_HAS_ERROR
// A field whose opening quote is preceded by whitespace is flagged.const ws = parseCsv(encoder.encode('a,b\n1, "two"\n'), { whitespaceBeforeQuote: true });// the ' "two"' Field node carries LEX_NODE_HAS_ERRORUnder strict: true a toggle violation also flags the affected record and the
Document root, so tree.root.hasError reads true.
Conformance corpus
Section titled “Conformance corpus”Strict mode is validated against the conformance corpus under
test_files/csv/:
- RFC 4180 clauses (header row, record separation, quoting, escaping) via the conformance suite.
csv/invalid/(unclosed-quote.csv,bad-quote-position.csv) carrieshasErrorunderstrict: true.edge-cases/(includingmismatched-columns.csv) carrieshasErrorunderstrict: truewhere a record is over- or under-wide.
See CSV conformance for the full four-axis results.
Inspecting the tree
Section titled “Inspecting the tree”Each record is a CsvRecord node containing CsvField children. Fields in the header row carry CSV_FLAG_HEADER; quoted fields carry CSV_FLAG_QUOTED.
const cursor = tree.cursor();cursor.gotoFirstChild(); // CsvRecord (first row)cursor.gotoFirstChild(); // CsvFieldconsole.log(cursor.current.kind); // CsvField kind idconsole.log(cursor.current.flags); // CSV_FLAG_HEADER | CSV_FLAG_QUOTED