Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Syntax specifications

x2c’s syntax is specified here as a token grammar plus lexical transition rules. This first extraction describes the implementation at revision 48715bde0d16f1a22c5a19a00e5568080a92367c on dev. It introduces no language change. The language reference supplies the meaning and constraints of the forms.

  • Source grammar: EBNF, precedence, contextual decisions, and macro extension contracts. The EBNF file is the same grammar included in that chapter, generated from etc/syntax/grammar.x.
  • Lexical specification: byte classes, ordered recognition, numeric and escape rules, nested modes, token reclassification, layout, and failure behavior.
  • Syntax as data: ordinary x2c Lists, their algebra, and a tool for generation, reference checks, dependency graphs, and lexical mode graphs.
  • x2c syntax map: the additions beyond C11, grouped for onboarding and linked to the specification data.

The grammar describes source after tokenization and indentation rewriting. It is not a claim that a context-free recognizer can decide whether an x2c program is valid. Names, types, macro signatures, and preprocessor state participate in parsing. The grammar names these dependencies explicitly. Constructed macro ASTs also enter binding without passing through source text; source productions do not restrict their provenance.

Catalog of existing material

This catalog records what existed before this extraction. Paths below are repository-relative. Function names locate the rules even when lines move.

MaterialLocationWhat it specifiesLimits
Language referencedocs/src/reference/language.mdUser-facing forms, examples, semantic constraints; a small protocol grammar under ProtocolsNo complete grammar or lexical rule inventory
Syntax guidesdocs/src/guide/{from-c,indentation,macros,match,collections,symbols,protocols}.mdExplanations and examples of individual formsExamples do not define all accepted token sequences
Declaration parsersrc/parse.xRecursive-descent rules for units, imports, types, declarations, declarators, parameters, fields, enumeratorsMixes recognition with name binding, collection, and semantic work
Expression parser and precedence ledgersrc/expressions.x, src/operator-ledger.xExpression productions, ambiguity decisions, complete binary operator precedence rowsDepends on type and macro state
Statement parsersrc/statements.xStatements, blocks, match/catch arms, directive-aware governed statementsMacro-produced statements also enter binding
Literal and lambda parserssrc/literals.x, src/lambdas.xQuoted data, interpolation, patterns, lambdas and capturesToken modes determine which spellings reach them
Macro parser and built-inssrc/macros.x, etc/builtin-macros.xCategory table, definitions, invocations, quotations, typed holes, keyword aliases; built-in class and foreachGrammar is extended by visible definitions; libraries such as lib/system-macros.x add more
Protocol parsersrc/protocol.xProtocol definitions, associated types, members, adoptionsWitness resolution and type constraints are separate
AST source patternssrc/grammar.xExecutable source-shaped macro patterns and AST accessors used by loweringDespite the filename, not a complete source grammar
Character scannerslib/scan.xAllocation-free byte recognizers, keywords, longest operator prefix, literal escapes, failure statusesA scanner in isolation does not specify tokenizer dispatch or context
Tokenizer and layoutlib/tokenizer.xOrdered scanning, mode stack, operand context, keyword reclassification, indentation rewrite, positionsContextual and stateful, not a single token regex table
Preprocessor handlingsrc/preprocess.x, src/frontend.xDirective handling, conditional groups, configured units and host-preprocessed importsHost C preprocessing has its own language and target configuration
Compiler token preparationsrc/compiler.x (Compiler.tokenize, _retag_keywords)Phase ordering and contextual in/match token typesRuns after raw tokenization and conditional scanning
Lisp readerlib/lisp.x (LispReader)Lists, reader prefixes and atomic forms over Lisp-mode tokensEvaluation and compile-time operations are not source grammar
Generated API pagesdocs/src/internals/compiler-api/, notably parse.md and grammar.mdPublic signatures and source doc commentsDerived API documentation, not an independent grammar
Architecture and ownership mapdocs/src/internals/{architecture,implementation-map}.mdPhase order and syntax ownersDescribes organization, not acceptance rules
Scanner/tokenizer testsunittest/test-scan.x, unittest/test-tokenizer.xExecutable recognition, boundary, mode and layout casesFinite observations, not a formal definition
Compiler fixturesunittest/compiler-fixtures/Token dumps, AST/translation outputs, diagnostics, successful and rejected programsExpectations cover selected cases, not all derivations
Editor grammarsetc/vsc-extension/syntaxes/{x2c,x2c-lisp}.tmLanguage.jsonTextMate regular expressions and highlighting contextsApproximate; cannot decide acceptance or reproduce binding
Editor grammar testsetc/vsc-extension/test/current-grammar.test.xExpected highlighting scopesContains syntax-shaped samples that are not accepted programs; not an acceptance corpus
Historical designsplans/archive/indentation-syntax.md, plans/archive/macro-function-style-syntax.mdDesign rationale and earlier syntax decisionsHistorical; current implementation and book take precedence
Bootstrap copiesbootstrap/lib/scan.{c,h}, bootstrap/lib/tokenizer.{c,h}Generated C implementation of lexical rulesDerived copies, not independent owners
Executable examplesexamples/manifest.txt and its programsWorking uses of the language with recorded checksDemonstrations, not exhaustive syntax rules

No standalone BNF, EBNF, PEG, yacc/lex, or tree-sitter source grammar was found in the tracked tree at the extraction revision. The parser and scanners were the closest executable specifications; the book was the principal prose specification.

Status and boundaries

The extraction has three layers:

  1. Byte recognition and token transitions in the lexical specification.
  2. Token productions in etc/syntax/grammar.x, rendered as x2c.ebnf, with explicit external recognizers.
  3. Context and semantic constraints in the grammar chapter and language reference.

This is a manually reviewed descriptive specification, not a generated parser, a proof of equivalence, or a replacement implementation. Named external recognizers cover the deliberately extensible macro surface and host-specific syntax. The lexical chapter identifies observed implementation edges separately from the supported spelling described by the book.

C translation constraints, preprocessing expressions, type checking, ownership, overload/protocol resolution, macro evaluation, AST lowering, and emitted C are outside the source grammar. The host compiler still checks the generated C. Resource limits and diagnostic wording are not productions.

Possible later uses

These are possible consumers, not work implemented by this extraction:

  • Generate positive and negative source cases, with environments for typedefs, macro signatures, and imports rather than only random terminal strings.
  • Compare token streams across scanner implementations and test each mode transition, numeric boundary, and indentation rewrite.
  • Check documentation and editor highlighting against shared terminal and precedence inventories.
  • Prototype a parser or code generator for a selected syntax category, then compare it with the existing compiler on the same corpus.
  • State proof obligations for recognition, normalization, or AST round trips, keeping lexical, syntactic, and semantic claims separate.

There is no new recurring check or generation pipeline. Future consumers must choose their own executable representation and demonstrate agreement with the compiler before relying on this extraction as an oracle.