agent's regex docstring extractor failed on multiline examples
Fixes a docs agent whose regex-based docstring extractor broke on multiline code examples. Regexes cannot balance fences or track indentation, so examples spanning several lines get cut off or merged with the next section. Use when extracted docstrings show truncated or garbled examples. Trigger: docstring examples cut mid-block in generated docs.
TL;DR
Replace the regex extractor with a real parser. Regular expressions cannot match balanced code fences or track indentation across lines, so any example longer than one line gets truncated or bleeds into the next docstring section. Fix: extract docstrings with the language's AST (Python's ast module, tree-sitter, or the TS compiler API) and take the raw docstring text verbatim.
The error
agent's regex docstring extractor failed on multiline examplesSteps
Find a docstring that renders wrong and check the extractor's output for it. Multiline failures show the example cut at the first blank line or merged with the following parameter description. Expected: you confirm the raw docstring is fine and the extractor mangled it.
Identify the regex responsible, usually a pattern matching from the opening quotes to the first closing quotes or first blank line. Expected: you can point at the pattern that cannot span lines correctly.
Replace extraction with the language AST: in Python use
ast.get_docstring()on the parsed node; for other languages use tree-sitter or the compiler API's comment attachment. Expected: the extractor returns the complete docstring including all multiline examples, verbatim.Keep regexes only for splitting the docstring into sections (parameters, returns, examples) after extraction, and make those section splitters line-aware rather than single-line. Expected: examples stay intact; section splitting still works.
Re-run doc generation and check the previously broken docstrings. Expected: multiline examples render complete, with fences and indentation preserved.
Use this when
- docstring examples appear truncated mid-block in generated docs
- multiline examples merge with the next section's text
- the extractor is regex-based and the language has an AST available
Not for this skill when
- docstrings extract fully but render wrong (a markdown rendering issue)
- the docstring itself is malformed in source (fix the source)
- you need to parse docstring section conventions like Google vs NumPy style (a convention question)
Variant phrasings
- regex extractor cuts off multiline docstring examples
- docstring parser fails on code fences spanning lines
- extracted docstring truncated at blank line
- multiline example garbled in generated reference
Why it happens
Regexes match linear patterns; they cannot count balanced delimiters or remember indentation context across lines. A docstring with a fenced multi-line example is a nested structure, and only a parser that understands nesting extracts it correctly.
Edge cases
- Docstrings with nested fences (an example showing markdown): the AST gives you the raw text; let the markdown renderer handle the nesting, not the extractor.
- Indented examples under list items: preserve the raw indentation; dedent only the common docstring prefix.
- Non-UTF8 or mixed-indentation sources: parse with errors tolerated and log the file rather than silently dropping the docstring.
Provenance
Resolved from the public thread: https://vectle.com/posts/pst_oBLfNF0u-xhQXBuYUdKhdw