Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Syntax as data

The source grammar and lexical rules are ordinary x2c List values. A tool can compile them, inspect their structure, and generate another representation without first parsing the prose specification.

Three modules under etc/syntax/ own the data:

ModuleAccessorContents
grammar.xsyntax_grammar()Entry point and 179 source productions
lexical.xsyntax_lexical()Byte patterns, ordered scanning, modes, transitions, predicates, token passes, failures
extensions.xsyntax_extensions()Reviewed C11 comparison, feature groups, affected rule names, example spellings, documentation links

For example, an ordinary x2c script can consume the grammar:

#!/usr/bin/env -S x2c script
#include "etc/syntax/grammar.x"

List grammar = syntax_grammar();
printf("%s\n", grammar.repr());

Save this example at the repository root. x2c script builds and links its included local module. An ordinary x2c build invocation must list the module as a build input. The accessors construct descriptive values; the compiler’s production parser and tokenizer do not consult them yet.

Grammar model

The root has the shape (grammar (start NAME) (rules RULE...)). Each rule is (rule NAME EXPRESSION). Names are atoms and terminal spellings are Strings. The current expression algebra is deliberately small:

ExpressionMeaning
(ref NAME)Recognize another named production
(token "text")Recognize the token spelling in the applicable lexical context
(seq EXPR...)Recognize every child in order
(choice EXPR...)Recognize an alternative; this does not specify PEG priority
(optional EXPR)Recognize zero or one occurrence
(repeat 0 EXPR)Recognize zero or more occurrences
(external "contract")Delegate recognition to the named lexical or contextual contract

For example, the function-body rule contains both familiar blocks and the expression-body addition:

List rule = %(rule function-body
  (choice
    (ref block)
    (seq (token "=") (token ">") (ref expression) (token ";"))));

The EBNF is generated from this value. The grammar chapter explains binding, macro signatures, precedence, and other contextual contracts. An external leaf is an explicit boundary, not an instruction to accept arbitrary text. A parser generator would need implementations for those contracts.

Lexical model

Lexing needs more than character regular expressions. A quoted List, for example, changes the significance of punctuation and signed numbers. Braced interpolation temporarily returns to code scanning.

The generated inventory exposes every record. The lexical specification explains the operations and boundary cases. These sections carry different parts of the operation:

SectionMeaning
inputByte encoding, NUL termination, token fields, positions, and trivia
classes, rulesNamed byte recognizers and their actions
modesOrdered rule dispatch for each scanner mode
transitionsSource mode, opener/closer bytes, emitted spelling, stack push/pop
predicatesConditions on surrounding tokens and named token sets
passesFrontend preparation, scanning, layout, conditional processing, keyword retagging, parsing
keywords, operatorsExact scanner spellings and keyword normalization
failuresScanner result conventions, status, emitted errors, and stopping behavior
boundariesLiteral decoding, Symbol representation, host preprocessing, native C acceptance
sourcesImplementation owners used for the extraction

Byte-pattern primitives are byte, bytes, inclusive range, and one-of-bytes. ref names a class or rule. seq, choice, optional, and repeat MINIMUM compose patterns. ordered tries children in dispatch order; until records whether a closing delimiter is consumed; except excludes listed byte patterns or prefixes. Guards such as when-prefix and when-mode select the applicable pattern. They are structured records, not regular-expression strings.

Actions record emit, consumed width, failure behavior, and transition. A transition’s push stores a new mode; pop restores the previous stack mode. The transition table’s defaults record applies unless a row overrides it. Predicates record Boolean conditions and token-history operations. Passes record ordered transformations and the state they use. external explicitly names work owned elsewhere.

This is a descriptive operational model. The tool validates named references and renders records; it does not interpret every lexical operation. Building an executable scanner from these records requires giving each operation an implementation and comparing its token streams with the existing tokenizer.

Existing consumers

Run these commands from the repository root with a built compiler:

builds/0/x2c script tools/syntax-spec summary
builds/0/x2c script tools/syntax-spec dump grammar
builds/0/x2c script tools/syntax-spec uses expression
builds/0/x2c script tools/syntax-spec graph quoted-list
builds/0/x2c script tools/syntax-spec modes
builds/0/x2c script tools/syntax-spec write
builds/0/x2c script tools/syntax-spec check

dump also accepts lexical and extensions. It prints the runtime List representation for inspection. uses finds direct callers of a production. graph emits a Mermaid production-dependency graph. modes emits a Mermaid graph of stack-push transitions; ordered dispatch and predicates still apply, and closing delimiters return to the previous stack mode. These graphs show dependencies and transitions, not complete railroad diagrams or acceptance proofs.

write generates the EBNF, lexical inventory, and the table in the x2c syntax map. check validates references and reports stale generated outputs without changing them. These are optional commands; no recurring build or publication check was added.

The extension data identifies affected productions and the additions within them. Its grouping is reviewed metadata. It does not claim that subtracting production names computes the difference between C and x2c.

Further uses and limits

These values make several bounded next steps possible:

  • Generate railroad diagrams for productions with no external leaves.
  • Produce candidate token sequences from bounded derivations, then filter them through the real compiler and retain accepted and rejected cases.
  • Compare scanner outputs against an interpreter of the lexical records.
  • Generate highlighting tables from keyword and operator spellings.
  • Select documentation examples and diagrams by extension family.

The data is manually extracted from the implementation. Reference checks and generated-output checks prevent some forms of drift; they do not establish language equivalence. Type validity, macro expansion, binding, native C rules, and deliberately opaque recognition contracts remain outside a context-free production tree. Corpus comparison and differential tests would provide additional evidence before any generated recognizer replaces existing code.