Skip to content

Value extraction

Package: @lanexio/parser (core surface) + the grammar packs (coercion) Layer: 1 (Core) + Tier 2 grammar extractors Runtime: Universal (browser, server, edge worker).

Lanexio™ Parser parses into a flat, frozen AST. Value extraction projects that AST into the plain JavaScript value the source document represents: an object for a YAML mapping or TOML table, an array for a JSON list, a number for a TOML integer, and so on.

import { parse } from '@lanexio/parser/all';
const tree = parse('{"a": 1, "b": [true, null]}', { language: 'json' });

toValue() is the single entry point for every supported grammar. It never throws: hard failures (malformed input, depth overflow, an empty document, an unresolvable alias) produce undefined plus a structured diagnostic. A missing extractor is not a hard failure: toValue() returns the lossless source text span with an unresolved_language diagnostic, so value is defined there.

import { toValue } from '@lanexio/parser';
import { parse } from '@lanexio/parser/all';
const tree = parse('{"a": 1, "b": [true, null]}', { language: 'json' });
const { value, diagnostics } = toValue(tree);
console.log(value); // { a: 1, b: [true, null] }
console.log(diagnostics); // []

The convenience functions in each grammar pack preset the extractor, so you can skip the registry entirely:

import { yamlToValue } from '@lanexio/parser-grammar-yaml';
import { parseYaml } from '@lanexio/parser-grammar-yaml';
const { value } = yamlToValue(parseYaml(new TextEncoder().encode('name: Tom\nage: 30\n')));
// { name: 'Tom', age: 30 }

TOML and YAML are common config formats. toValue() turns a parsed document into a plain object in one call:

import { parseToml } from '@lanexio/parser-grammar-toml';
import { tomlToValue } from '@lanexio/parser-grammar-toml';
const source = `
[server]
host = "localhost"
port = 8080
[[products]]
name = "Hammer"
price = 9.99
`;
const { value, diagnostics } = tomlToValue(parseToml(new TextEncoder().encode(source)));
console.log(value);
// {
// server: { host: 'localhost', port: 8080 },
// products: [{ name: 'Hammer', price: 9.99 }]
// }

TOML table headers, dotted keys, inline tables, and arrays of tables all reconstruct into nested objects and arrays.

CSV extracts to a grid of strings by default, or to an array of row objects when the first record is a header:

import { parseCsv } from '@lanexio/parser-grammar-csv';
import { csvToValue } from '@lanexio/parser-grammar-csv';
const csv = 'name,age\nTom,30\nJay,41\n';
const { value } = csvToValue(parseCsv(new TextEncoder().encode(csv)));
console.log(value);
// [ { name: 'Tom', age: '30' }, { name: 'Jay', age: '41' } ]

With { header: false } the same input extracts to string[][]:

const { value } = csvToValue(parseCsv(new TextEncoder().encode(csv), { header: false }));
// [ ['name', 'age'], ['Tom', '30'], ['Jay', '41'] ]

TSV works the same way through tsvToValue, with raw (never quoted) cells.

Each grammar follows its own rules for turning a text token into a value:

GrammarNumbersStringsOther
JSONdecimal, JSON5 hexRFC 8259 escapes plus JSON5booleans, null
YAMLcore schema: decimal/octal/hex int, float, .inf/.nanquoted and block scalars unquotebooleans, null, timestamps stay strings
TOMLdecimal/hex/octal/binary int, float with inf/nanbasic/literal strings unescapebooleans, datetimes stay raw RFC 3339 text
CSV / TSVnone (cells are strings)quoted fields unescape "" to "n/a

YAML anchors and aliases resolve to the anchored value, with missing_anchor and circular_alias diagnostics. YAML << merge keys copy anchored mappings into the merging mapping, with explicit keys winning.

toValue() never throws. Instead you always get a result object with a diagnostics array:

const tree = parse('{"a": 1, "a": 2}', { language: 'json' }); // duplicate key
const { value, diagnostics } = toValue(tree);
console.log(diagnostics[0].code); // 'duplicate_key'
console.log(diagnostics[0].rangeStart); // byte offset of the second key

The most common codes are has_error (the subtree contains a parse error), unresolved_language (no extractor is registered; the lossless source text span is returned instead), and unsafe_integer (an integer lost precision past Number.MAX_SAFE_INTEGER). See the reference page for the full table.

Recursive projection is bounded by maxDepth (default 512). Deeply nested documents that exceed the bound produce undefined plus a max_depth diagnostic instead of overflowing the stack:

// A document that nests deeper than 64 levels:
let src = '0';
for (let i = 0; i < 65; i += 1) src = `[${src}]`;
const tree = parse(src, { language: 'json' });
const { value, diagnostics } = toValue(tree, { maxDepth: 64 });
if (diagnostics.some((d) => d.code === 'max_depth')) {
// the document nests more than 64 levels; no value was produced
}