Skip to content

Parsing TOML

Lanexio™ Parser supports both TOML v1.0.0 and v1.1.0 through a single package. Choose your version explicitly, or use the default (v1.1.0).

  1. Install the package.

    Terminal window
    pnpm add @lanexio/parser-grammar-toml
  2. Parse a TOML document.

    import { parseToml } from '@lanexio/parser-grammar-toml';
    const encoder = new TextEncoder();
    const tree = parseToml(encoder.encode(`
    [server]
    host = "0.0.0.0"
    port = 8080
    [database]
    url = "postgres://localhost"
    pool = 10
    `));

Use parseToml10 for TOML v1.0.0 syntax and parseToml11 for v1.1.0 constructs like dotted keys in inline tables and hex float literals. parseToml defaults to v1.1.0.

import { parseToml10, parseToml11 } from '@lanexio/parser-grammar-toml';
const v1 = parseToml10(encoder.encode('[table]\nkey = "value"'));
const v2 = parseToml11(encoder.encode('a.b.c = 42'));
OptionTypeDefaultDescription
version'1.0' | '1.1''1.1'TOML version dialect. The canonical key; wins over spec when both are given.
spec'1.0' | '1.1'undefinedDocumented alias for version (ADR 0048), matching the unified strict convention’s “spec” vocabulary. An explicit version wins over spec; spec wins over the strict-mode default.
mode'lenient' | 'strict''lenient'Parse mode. Strict mode defaults the version to '1.0' when neither version nor spec is given.
strictbooleanundefinedConvenience alias for mode: 'strict' (ADR 0044). true picks strict, false is the explicit lenient spelling; an explicit mode wins.

parseToml with no options defaults to TOML v1.1.0. Version precedence (highest to lowest): an explicit version wins over spec, and either wins over the strict-mode default, so { strict: true, spec: "1.1" } keeps the 1.1 rules and { spec: "1.0", version: "1.1" } keeps the 1.1 rules too. parseToml10 and parseToml11 are unaffected.

The default (mode: "lenient") parses with the TOML v1.1.0 grammar, so 1.1-only constructs are accepted. Strict mode defaults the grammar to the TOML 1.0.0 rule set when no explicit version or spec is given, so 1.1-only constructs are rejected by rule selection. An explicit version (or its spec alias) always wins over the strict-mode default.

import { parseToml } from '@lanexio/parser-grammar-toml';
const encoder = new TextEncoder();
// \x byte escape: a 1.1-only construct, accepted in lenient (1.1 default)
const lenient = parseToml(encoder.encode('answer = "\\x33"'));
console.log(lenient.root.hasError); // false
// Rejected in strict (1.0 default)
const strict = parseToml(encoder.encode('answer = "\\x33"'), { mode: 'strict' });
console.log(strict.root.hasError); // true
// The { strict: true } alias is the same option (ADR 0044)
const strictAlias = parseToml(encoder.encode('answer = "\\x33"'), { strict: true });
console.log(strictAlias.root.hasError); // true
// The { spec } alias selects the version the same way { version } does
const spec10 = parseToml(encoder.encode('answer = "\\x33"'), { spec: '1.0' });
console.log(spec10.root.hasError); // true

The same rule selection flags other 1.1-only constructs in strict mode: seconds-less datetimes and times, inline-table newlines and trailing commas, and the \xHH byte escape. Precedence is mode > strict > default lenient, so an explicit mode always wins over a strict boolean, and strict: false is the explicit lenient spelling. The unified parse() entry point forwards the same options through grammarOptions.mode or the top-level strict key, so parse(src, { language: "toml", grammarOptions: { mode: "strict" } }) and parse(src, { language: "toml", strict: true }) both behave identically to parseToml(bytes, { mode: "strict" }). Strict mode never throws; a rejected construct is a flag on the tree, never an exception. See Leniency and strict mode for the full policy.

Strict mode is mechanically enforced against the official toml-test reference suite (ADR 0034): every .toml file under corpus/toml-test/invalid/ (511 files across its 16 category subdirectories) must carry hasError: true under strict: true with zero false accepts. Nine of those files are TOML 1.1-only constructs that parse clean in lenient mode and are flagged by the 1.0 rule set in strict mode. See TOML conformance for the ledger.

Tables become TomlTable nodes, key-value pairs become TomlKvPair nodes, and arrays become TomlArray nodes. Values carry type flags for string, integer, float, boolean, datetime, and array-of-types.

import { TomlKind } from '@lanexio/parser-grammar-toml';
const cursor = tree.cursor();
cursor.gotoFirstChild();
console.log(cursor.current.kind === TomlKind.TomlTable); // true

TOML 1.1.0 adds the \e escape (ESC character, U+001B) and \xHH byte escape to the existing \b, \t, \n, \f, \r, \", \\, \uXXXX, and \UXXXXXXXX escapes from TOML 1.0.0.

# All valid escape sequences
string_val = "Tab:\t Newline:\n Quote:\" Backslash:\\"
hex_val = "\x48\x65\x6C\x6C\x6F" # "Hello" (TOML 1.1.0 only)
esc_val = "\e" # ESC character (TOML 1.1.0 only)
unicode = "Hello" # "Hello"
big_unicode = "\U0001F600" # Unicode beyond BMP

The \xHH escape is only valid in TOML 1.1.0 mode. In TOML 1.0.0 mode, \x is a reserved escape sequence and produces an Error node.

The \e escape is valid in both TOML 1.0.0 and 1.1.0 per the toml-test reference suite.

import { parseToml10, parseToml11 } from '@lanexio/parser-grammar-toml';
const bytes = new TextEncoder();
// \xHH accepted in 1.1.0, rejected in 1.0.0
const v11 = parseToml11(bytes.encode('val = "\\x48"'));
console.log(v11.root.hasError); // false
const v10 = parseToml10(bytes.encode('val = "\\x48"'));
console.log(v10.root.hasError); // true
// \e accepted in both versions
const ev11 = parseToml11(bytes.encode('val = "\\e"'));
console.log(ev11.root.hasError); // false
const ev10 = parseToml10(bytes.encode('val = "\\e"'));
console.log(ev10.root.hasError); // false
EscapeCode PointDescriptionTOML Version
\bU+0008Backspace1.0.0+
\tU+0009Tab1.0.0+
\nU+000ALine Feed1.0.0+
\fU+000CForm Feed1.0.0+
\rU+000DCarriage Return1.0.0+
\"U+0022Quotation Mark1.0.0+
\\U+005CBackslash1.0.0+
\eU+001BEscape (ESC)1.0.0+
\xHHvariesByte escape (2 hex digits)1.1.0+
\uXXXXU+0000-U+FFFFUnicode (4 hex digits)1.0.0+
\UXXXXXXXXU+0000-U+10FFFFUnicode (8 hex digits)1.0.0+