Skip to content

Parsing CSV and TSV

Lanexio™ Parser handles CSV and TSV through a single package with configurable parsing. Both formats use the same never-throw parse path and produce zero-copy flat ASTs.

  1. Install the package.

    Terminal window
    pnpm add @lanexio/parser-grammar-csv
  2. Parse a CSV file.

    import { parseCsv } from '@lanexio/parser-grammar-csv';
    const encoder = new TextEncoder();
    const tree = parseCsv(encoder.encode('Name,Age,City\nAlice,30,NYC\nBob,25,LA'));
  3. Parse with options.

    import { parseCsv } from '@lanexio/parser-grammar-csv';
    const tree = parseCsv(encoder.encode('Name|Age|City\nAlice|30|NYC'), {
    delimiter: 0x7c, // '|' as a byte; the delimiter is a byte, not a string
    header: true,
    });
import { parseTsv } from '@lanexio/parser-grammar-csv';
const tree = parseTsv(encoder.encode('Name\tAge\tCity\nAlice\t30\tNYC'));
OptionTypeDefaultDescription
delimiternumber0x2C (,)Field delimiter byte. The delimiter is a byte, not a string
quotenumber0x22 (")Quote character byte
headerbooleantrueWhether the first record is a header row
skipEmptyLinesbooleanfalseWhether to skip empty lines
mode"lenient" | "strict""lenient"Lenient recovery (default) or strict flagging
strictbooleanundefinedConvenience alias for mode: "strict" (ADR 0044). true picks strict, false is the explicit lenient spelling; an explicit mode wins
rfc4180LineEndingsbooleanfalseEnforce CRLF-only record terminators per RFC 4180 §2.1. A bare LF or bare CR terminator is flagged as an error. Independent toggle, works with or without strict
whitespaceBeforeQuotebooleanfalseReject a field whose opening quote is preceded by whitespace (RFC 4180 §2.4). Flags the field as an error. Independent toggle, works with or without strict

The default (mode: "lenient") keeps the historical recovery behavior: malformed records produce error leaves but the parser never throws. Strict mode flags spec-level rejections on the tree instead of only the offending field: an unclosed quote and a record whose field count differs from the reference row (the header row when header is true, otherwise the first record) both set hasError on the record and the document root.

import { parseCsv } from '@lanexio/parser-grammar-csv';
const encoder = new TextEncoder();
// Unclosed quote: clean in lenient, flagged in strict
const lenient = parseCsv(encoder.encode('a,b\n"unclosed,1\n2,3'));
console.log(lenient.root.hasError); // false
const strict = parseCsv(encoder.encode('a,b\n"unclosed,1\n2,3'), { mode: 'strict' });
console.log(strict.root.hasError); // true
// Column-count mismatch: clean in lenient, flagged in strict
const wide = parseCsv(encoder.encode('a,b,c\n1,2\n3,4,5\n6,7'), { mode: 'strict' });
console.log(wide.root.hasError); // true

The { strict: true } alias (ADR 0044) is the same option under the ecosystem-standard spelling: parseCsv(bytes, { strict: true }) produces the same flags as parseCsv(bytes, { mode: "strict" }), and strict: false is the explicit lenient spelling. Precedence is mode > strict > default lenient, so an explicit mode always wins over a strict boolean. The unified parse() entry point forwards the same options through grammarOptions.mode or the top-level strict key, so parse(src, { language: "csv", grammarOptions: { mode: "strict" } }) and parse(src, { language: "csv", strict: true }) both behave identically to parseCsv(bytes, { mode: "strict" }). Strict mode never throws; a rejected record is a flag on the tree, never an exception. Well-formed input parses clean in both modes. See Leniency and strict mode for the full policy.

On a column-count mismatch in strict mode, the tree’s metadata.diagnostics gains a per-row entry that names the record and the expected-vs-found counts:

const tree = parseCsv(encoder.encode('a,b,c\n1,2\n3,4,5\n6,7'), { mode: 'strict' });
console.log(tree.diagnostics);
// ["CSV record 2: expected 3 fields, found 2", "CSV record 4: expected 3 fields, found 2"]

The surplus Field nodes of an over-wide record also carry LEX_NODE_HAS_ERROR, so the AST identifies the divergent column by byte range. TSV diagnostics use the same shape with a TSV record prefix.

Two independent toggles tighten RFC 4180 conformance without switching the whole parse to strict mode. Both default to false, keeping the default lenient, and both work with or without strict:

  • rfc4180LineEndings enforces CRLF-only record terminators (RFC 4180 §2.1). A bare LF or bare CR terminator is flagged as an error on its Newline node. Newlines inside quoted fields and blank lines skipped by skipEmptyLines are unaffected.
  • whitespaceBeforeQuote rejects a field whose opening quote is preceded by whitespace (RFC 4180 §2.4). The field is flagged as an error on its Field node.
// A bare-LF record terminator is flagged when the toggle is on.
const crlf = parseCsv(encoder.encode('a,b\r\n1,2\r\n'), { rfc4180LineEndings: true });
console.log(crlf.root.hasError); // false
const lf = parseCsv(encoder.encode('a,b\n1,2'), { rfc4180LineEndings: true });
// the bare-LF Newline node carries LEX_NODE_HAS_ERROR
// A field whose opening quote is preceded by whitespace is flagged.
const ws = parseCsv(encoder.encode('a,b\n1, "two"\n'), { whitespaceBeforeQuote: true });
// the ' "two"' Field node carries LEX_NODE_HAS_ERROR

Under strict: true a toggle violation also flags the affected record and the Document root, so tree.root.hasError reads true.

Strict mode is validated against the conformance corpus under test_files/csv/:

  • RFC 4180 clauses (header row, record separation, quoting, escaping) via the conformance suite.
  • csv/invalid/ (unclosed-quote.csv, bad-quote-position.csv) carries hasError under strict: true.
  • edge-cases/ (including mismatched-columns.csv) carries hasError under strict: true where a record is over- or under-wide.

See CSV conformance for the full four-axis results.

Each record is a CsvRecord node containing CsvField children. Fields in the header row carry CSV_FLAG_HEADER; quoted fields carry CSV_FLAG_QUOTED.

const cursor = tree.cursor();
cursor.gotoFirstChild(); // CsvRecord (first row)
cursor.gotoFirstChild(); // CsvField
console.log(cursor.current.kind); // CsvField kind id
console.log(cursor.current.flags); // CSV_FLAG_HEADER | CSV_FLAG_QUOTED