Value extraction (toValue)
Lanexio™ Parser parses source bytes into a flat, frozen AST. toValue() projects a parsed LexTree (or any LexNode subtree) into a JSON-expressible host value: plain objects, arrays, strings, numbers, booleans, and null. It is a pure read of the frozen flat AST and it never throws.
- Version: Stable
- Module name:
parser-core(surface) + the grammar packs (coercion) - Package:
@lanexio/parser-core,@lanexio/parser-grammar-json,@lanexio/parser-grammar-yaml,@lanexio/parser-grammar-toml,@lanexio/parser-grammar-csv - Import path:
@lanexio/parser(re-exports the core surface),@lanexio/parser-core, or any grammar pack - Layer: 1 (Core) with Tier 2 grammar extractors
- Runtime: Universal (browser, server, edge worker)
- Module format: ESM
- Stability: Stable
- Primary use case: Turn a parse tree into the object/array value a consumer wants without walking the AST by hand.
When to use this module
Section titled “When to use this module”- You parse JSON, YAML, TOML, or CSV/TSV and want the equivalent JavaScript value (object, array, scalar).
- You want a deterministic, never-throwing conversion with structured diagnostics instead of exceptions.
- You need the same
toValue()entry point across every grammar that supports value extraction.
Module boundary
Section titled “Module boundary”| Boundary | Description |
|---|---|
| Inputs | A LexTree (uses its root) or any LexNode subtree |
| Outputs | LexValueResult ({ value, diagnostics }) |
| Side effects | None (registry registration is the only state, and it is opt-in) |
| Determinism | Yes (same tree + same options produce identical results) |
| External dependencies | @lanexio/parser-core only; grammar packs depend only on parser-core |
| Never-throw guarantee | Yes |
| Security surface | None (read-only projection, no code execution) |
Basic usage
Section titled “Basic usage”import { toValue } from '@lanexio/parser';import { parse } from '@lanexio/parser/all';
const tree = parse('{"a": 1, "b": [true, null]}', { language: 'json' });const { value, diagnostics } = toValue(tree);
console.log(value); // { a: 1, b: [true, null] }console.log(diagnostics); // []Importing @lanexio/parser/all registers every grammar and extractor; a bare grammar-pack import does not register. @lanexio/parser/all does not export toValue, hence the split import.
Each grammar pack also ships a convenience function that presets its extractor, so no registration is required:
import { jsonToValue, parseJson } from '@lanexio/parser-grammar-json';
const { value } = jsonToValue(parseJson(new TextEncoder().encode('{"a": 1}')));// { a: 1 }The convenience functions are jsonToValue, yamlToValue, tomlToValue, csvToValue, and tsvToValue.
The LexValue model
Section titled “The LexValue model”LexValue is the union of every value the extractors can produce:
| Kind | Example |
|---|---|
string | "hello" |
number | 42, 1.5, Infinity, NaN |
boolean | true |
null | null |
| object | { a: 1 } |
| array | [1, 2, 3] |
Dates and bigints are out of scope for v1: TOML datetimes project to their raw RFC 3339 text as strings, and integers beyond Number.MAX_SAFE_INTEGER project to number with an unsafe_integer diagnostic.
Diagnostics
Section titled “Diagnostics”Every hard failure produces undefined for value plus a structured diagnostic. Warning and fallback diagnostics (unresolved_language, duplicate_key, unsafe_integer, overflow, excess_field) accompany a defined, possibly degraded value, so value === undefined is not a reliable failure signal; inspect diagnostics.
type LexValueDiagnostic = { code: LexValueDiagnosticCode; rangeStart: number; rangeEnd: number; message: string;};| Code | Meaning |
|---|---|
has_error | The subtree contains a parse error; the value path never projects it. |
unresolved_language | No extractor is registered; the lossless text span was returned instead. |
unsafe_integer | An integer lost precision past Number.MAX_SAFE_INTEGER. |
overflow | A number overflowed to +-Infinity or NaN (for example JSON 1e400). |
duplicate_key | A duplicate object key was encountered; the last value wins. |
missing_anchor | A YAML alias references an anchor that was never defined. |
circular_alias | A YAML alias chain references itself. |
max_depth | Nesting or alias-chain depth exceeded the bound. |
excess_field | A CSV row is wider than the header; the excess fields were dropped. |
empty_document | No value is present (for example an empty YAML stream). |
Resolution order
Section titled “Resolution order”toValue() is fully deterministic:
- The target node is the tree root for a
LexTree, or the node itself for aLexNode. - Error gate (always runs). A subtree containing any parse error projects to
undefinedplus ahas_errordiagnostic at the first error range. - Extractor resolution.
options.extractorwins; otherwise the registered extractor foroptions.language(or the owning tree’smetadata.language) is looked up case-insensitively. - Fallback. With no extractor, the lossless source text span of the target node is returned with an
unresolved_languagediagnostic. This keeps every grammar without an extractor (html, markdown, mdx, sql, graphql, xml, css) round-trippable. - Otherwise the resolved extractor runs under
options.maxDepth(default 512) and its result is returned.
type ToValueOptions = { language?: string; // resolves a registered extractor when metadata.language is absent extractor?: LexValueExtractor; // explicit override, bypasses registry and fallback maxDepth?: number; // nesting/alias-chain bound, default 512};Extractor registry
Section titled “Extractor registry”ValueExtractorRegistry maps language names to extractors. Registration is case-insensitive and last registration wins, mirroring grammarRegistry.
import { registerValueExtractor, valueExtractorRegistry } from '@lanexio/parser';import { yamlValueExtractor } from '@lanexio/parser-grammar-yaml';
registerValueExtractor('yaml', yamlValueExtractor);valueExtractorRegistry.byLanguage('YAML'); // same extractor@lanexio/parser/all registers every extractor alongside the grammars: json, jsonc, json5, yaml, yml, toml, csv, tsv.
Per-grammar coercion
Section titled “Per-grammar coercion”The extractors live in the grammar packs and implement the grammar-specific rules.
| Grammar | Value shape |
|---|---|
| JSON / JSONC / JSON5 | Object, array, string, number, boolean, null. Duplicate keys last-win with duplicate_key; 1e400 overflows to Infinity with overflow. |
| YAML | Objects, arrays, scalars resolved by the YAML 1.2.2 core schema (null, bool, decimal/octal/hex int, float, .inf/.nan, else string). Anchors/aliases resolve with cycle detection; << merge keys copy anchored mappings with explicit keys winning. Multi-document streams become arrays. |
| TOML | Document object graph reconstructed by path: dotted keys, [table] headers (with implicit tables), [[array-of-tables]] appends, inline tables, arrays. Strings unescape, integers honor 0x/0o/0b, floats support inf/nan, datetimes stay raw text. |
| CSV / TSV | string[][] by default, or Record<string, string>[] when the first record carries the header flag. Quoted fields unescape "" to ". TSV cells are raw. |