Skip to content

@lanexio/parser-grammar-csv

This page documents @lanexio/parser-grammar-csv, the CSV grammar package for parsing RFC 4180 compliant CSV and TSV (tab-separated values) with configurable delimiter, quoting, and header detection.

  • Version: Stable
  • Module name: parser-grammar-csv
  • Package: @lanexio/parser-grammar-csv
  • Import path: @lanexio/parser-grammar-csv
  • Layer: 2 (Grammar)
  • Runtime: Universal
  • Module format: ESM
  • Stability: Stable
  • Primary use case: Parse CSV and TSV data into a flat AST.
  • You need to parse CSV (RFC 4180) data into a traversable AST.
  • You need to parse TSV (tab-separated values) data.
  • You need configurable delimiter, quoting, or header detection.
BoundaryDescription
InputsUint8Array (source bytes) + optional ParseCsvOptions
OutputsLexTree (flat AST rooted at CsvKind.Document)
Side effectsNone
DeterminismYes
External dependencies@lanexio/parser-core
Never-throw guaranteeYes
Security surfaceNone (parse only, no output generation)
  1. Install the package.

    Terminal window
    pnpm add @lanexio/parser-grammar-csv
  2. Import the named export.

    import { parseCsv } from '@lanexio/parser-grammar-csv';
RequirementRequiredLayerNotes
@lanexio/parser-coreYes1^1.0.0
import { parseCsv } from '@lanexio/parser-grammar-csv';
const encoder = new TextEncoder();
const bytes = encoder.encode('name,age\nAlice,30\nBob,25');
const tree = parseCsv(bytes);
console.log(tree.nodeCount);
console.log(tree.root.kind); // CsvKind.Document
import { parseTsv } from '@lanexio/parser-grammar-csv';
const encoder = new TextEncoder();
const tree = parseTsv(encoder.encode('name\tage\nAlice\t30'));
import { parseCsv } from '@lanexio/parser-grammar-csv';
const encoder = new TextEncoder();
const tree = parseCsv(encoder.encode('name|age\nAlice|30'), { delimiter: 0x7c });
ExportTypeDescription
parseCsv(bytes: Uint8Array, options?: ParseCsvOptions) => LexTreeParse CSV. Never throws.
parseTsv(source: Uint8Array, opts?: ParseTsvOptions) => LexTreeParse TSV (tab-separated values).
CsvKindconst objectNumeric kind IDs for all CSV node types.
CsvFieldconst objectNumeric field IDs for CSV value slots.
CSV_FLAG_HEADER4Per-node flag: this record is the header row. Set on the first CsvKind.Record node when header option is true.
CSV_FLAG_QUOTED2Per-node flag: field is enclosed in double quotes. Set on CsvKind.Field nodes whose value is quoted.
CSV_FIELD_NAMES_BY_IDreadonly string[]Field name lookup by numeric field ID.
CSV_KIND_NAMES_BY_IDreadonly Record<number, string>Kind-name lookup by numeric ID.
csvGrammarLanexioParserPureGrammarGrammar descriptor for CSV.
tsvGrammarLanexioParserPureGrammarGrammar descriptor for TSV.
csvRegistrationGrammarRegistrationRegistration for CSV.
tsvRegistrationGrammarRegistrationRegistration for TSV.
csvReuseOracleReuseOracleIncremental reuse oracle for CSV.
tsvReuseOracleReuseOracleIncremental reuse oracle for TSV.
csvToValue(target: LexTree | LexNode, options?: ToValueOptions) => LexValueResultProject a CSV tree or node subtree into string[][] (or Record<string, string>[] in header mode). Never throws.
tsvToValue(target: LexTree | LexNode, options?: ToValueOptions) => LexValueResultProject a TSV tree or node subtree into host values. Never throws.
csvValueExtractorLexValueExtractorThe CSV/TSV value extractor (registered under csv and tsv).

csvToValue() returns string[][] by default. When the first record carries CSV_FLAG_HEADER (the header parse option, on by default), the first record is consumed as the header and the result is Record<string, string>[]. Quoted fields strip the surrounding quotes and unescape "" to "; TSV cells are raw. Duplicate header names last-win with a duplicate_key diagnostic, rows wider than the header drop the excess fields with an excess_field diagnostic, and shorter rows omit the missing keys. Empty rows are preserved as empty arrays.

import { parseCsv, csvToValue } from '@lanexio/parser-grammar-csv';
const { value } = csvToValue(parseCsv(new TextEncoder().encode('name,age\nTom,30\n')));
console.log(value); // [ { name: 'Tom', age: '30' } ]

See the value-extraction reference for the full contract.

FieldTypeRequiredDefaultDescription
delimiternumberNo0x2C (,)Field delimiter byte. The delimiter is a byte, not a string.
quotenumberNo0x22 (")Quote character byte.
headerbooleanNotrueWhether the first record is a header row.
relaxQuotesbooleanNofalseAllow unescaped quotes inside quoted fields (non-RFC dialects).
languagestringNo"csv"Language name stamped into tree.metadata.language on every result. parseTsv passes "tsv"; the default keeps CSV trees stamped "csv".
mode"lenient" | "strict"No"lenient"Lenient recovery (default) or strict flagging (ADR 0033).
strictbooleanNoundefinedConvenience alias for mode: "strict" (ADR 0044). An explicit mode wins.
rfc4180LineEndingsbooleanNofalseEnforce CRLF-only record terminators per RFC 4180 §2.1. A bare LF or bare CR terminator is flagged on its Newline node. Independent of strict.
whitespaceBeforeQuotebooleanNofalseReject a field whose opening quote is preceded by whitespace (RFC 4180 §2.4). Flags the field on its Field node. Independent of strict.
FieldTypeRequiredDefaultDescription
headerbooleanNotrueWhether the first record is a header row.
skipEmptyLinesbooleanNofalseSkip lines that are entirely empty (zero-length or whitespace-only).
commentsnumber | readonly number[]NoundefinedByte value(s) that start a comment line. Lines starting with any of these bytes are skipped.
mode"lenient" | "strict"No"lenient"Lenient recovery (default) or strict flagging (ADR 0033). Forwarded to parseCsv.
strictbooleanNoundefinedConvenience alias for mode: "strict" (ADR 0044). An explicit mode wins. Forwarded to parseCsv.
rfc4180LineEndingsbooleanNofalseEnforce CRLF-only record terminators per RFC 4180 §2.1. A bare LF or bare CR terminator is flagged on its Newline node. Forwarded to parseCsv. Independent of strict.
whitespaceBeforeQuotebooleanNofalseReject a field whose opening quote is preceded by whitespace (RFC 4180 §2.4). TSV disables quoting, so this is a no-op for TSV and is accepted only for API symmetry. Forwarded to parseCsv.

Every tree returned by parseCsv or parseTsv carries tree.metadata.language — "csv" on parseCsv results and "tsv" on parseTsv results, including the never-throw error-tree path. The stamp matches the unified parse() dispatcher and the HTML/YAML/CSS grammar entry points.

Lanexio™ Parser keeps the default parse lenient: mode: "lenient" recovers from malformed input without throwing, producing CsvKind.Error leaves. Under mode: "strict" (or the strict: true alias), spec-level rejections are surfaced as flags, never exceptions:

  • A record that fails to parse (unclosed quote, bad quote position) carries LEX_NODE_HAS_ERROR | GRAMMAR_FLAG_NODE_ERROR on its Record node.
  • A record whose field count differs from the reference row is flagged the same way. The reference is the header row when header: true (the default), otherwise the first record.
  • On a column-count mismatch the tree’s metadata.diagnostics gains a per-row entry: CSV record <n>: expected <h> fields, found <f> (TSV trees use the TSV record prefix). The over-wide record’s surplus Field nodes also carry LEX_NODE_HAS_ERROR, identifying the divergent column by byte range.
  • When any record is rejected, the Document root carries LEX_NODE_HAS_ERROR | GRAMMAR_FLAG_NODE_ERROR and tree.root.hasError reads true.

Two independent toggles tighten RFC 4180 conformance without switching to strict mode. Both default to false, keeping the default lenient, and both work with or without strict:

OptionEffect
rfc4180LineEndings: trueEnforce CRLF-only record terminators (RFC 4180 §2.1). A bare LF or bare CR terminator is flagged on its Newline node. Newlines inside quoted fields and blank lines skipped by skipEmptyLines are unaffected.
whitespaceBeforeQuote: trueReject a field whose opening quote is preceded by whitespace (RFC 4180 §2.4). The field is flagged on its Field node.

Under strict: true a toggle violation additionally flags the affected record and the Document root, so tree.root.hasError reads true.

Strict mode is validated against the conformance corpus under test_files/csv/:

  • RFC 4180 clauses (header row, record separation, quoting, escaping, whitespace) via the conformance suite.
  • csv/invalid/ (unclosed-quote.csv, bad-quote-position.csv) carries hasError under strict: true.
  • edge-cases/ (including mismatched-columns.csv) carries hasError under strict: true where a record is over- or under-wide.
PropertyTypeDescription
rootLexNodeRoot node (CsvKind.Document or CsvKind.Error on pure invalid).
nodeCountnumberTotal nodes in the tree.
sourceUint8ArrayOriginal parsed bytes.
TypePurposeNotes
ParseCsvOptionsOptions for parseCsvSee options table above.
ParseTsvOptionsOptions for parseTsvSee options table above.
CsvKindTypeUnion of all CsvKind valuesType-safe kind reference
CsvFieldTypeUnion of all CsvField valuesType-safe field reference

parseCsv accepts any byte as delimiter. Use parseTsv for tab-delimited data (delegates to CSV parser with delimiter: 0x09, quote: -1).

  • No direct accessibility surface. The CSV parser produces flat AST data structures.
ConcernStatus
Generated output semanticsNot applicable (data format)
ARIA attributes in serialized outputNot applicable
Semantic element round-tripNot applicable
  • parseCsv and parseTsv never throw on any byte sequence.
ThreatMitigationStatus
Malformed input byte sequencePanic-free guarantee: all inputs accepted, errors produce CsvKind.Error AST nodesImplemented
PackageRelationshipLayerNotes
@lanexio/parser-coreRequires1Provides LexTree, LexNode, LexCursor
@lanexio/parserConsumes6Unified entry point
VersionDateStatusNotable changes
1.0.02026-05-29CurrentInitial stable release. Apache-2.0.
  • None.