Files
M1kep-KepPromptLang/lib/parser/grammar.py
T
Claude f4c862ec56 Add $var bindings and # comments
$NAME = arg; binds; $NAME substitutes. Single-pass: define before
use, no reassignment. Substitution is structural — refs share the
parsed action object, but actions are evaluated per occurrence (so
$r = rand(3); $r $r re-rolls each use).

Implementation: assign stores into PromptTransformer.vars and
returns None (filtered by tokenizer); ref returns the stored arg
wrapped in a transparent WeightedGroup(items, 1.0), so existing
_flatten and embedding_tensor code paths handle it without a new
container type.

NAME and WORD terminals overlap on bare identifiers; the earley
parser's dynamic lexer disambiguates by grammar context (the $
prefix forces NAME). Noted in grammar.py since this would break
under a basic/contextual lexer.

# comments run to end-of-line and are lexer-ignored.

Nine new tests cover top-level/arg-level/weighted substitution,
chaining, define-before-use error, reassignment error, comments,
and assign-only prompts.
2026-04-13 02:58:56 +00:00

39 lines
815 B
Python

grammar = r"""
?start: stmt+
?stmt: assign
| item
assign: "$" NAME "=" arg ";"
item: embedding
| WORD
| generic_function
| QUOTED_STRING
| weighted
| ref
generic_function: FUNC_NAME "(" arg ("|" arg)* ")"
weighted: "(" arg ":" SIGNED_NUMBER ")"
ref: "$" NAME
arg: item+
embedding: "embedding:" WORD
// NAME and WORD overlap on bare identifiers; the earley parser's dynamic lexer
// disambiguates by grammar context (the leading "$" forces NAME). This breaks
// under a basic/contextual lexer, so keep parser="earley" in __init__.py.
FUNC_NAME: /[A-Za-z_-]+/
NAME: /[A-Za-z_][A-Za-z0-9_]*/
WORD: /[A-Za-z0-9,_\.-]+/
QUOTED_STRING: /"([^"\\]*(\\.[^"\\]*)*)"|'([^'\\]*(\\.[^'\\]*)*)'/
SIGNED_NUMBER: /-?\d+(\.\d+)?/
COMMENT: /#[^\n]*/
%import common.WS
%ignore WS
%ignore COMMENT
"""