Skip to content

Doc impact

doc_impact

Doc impact analysis — doc refs, undocumented symbols, stale signatures.

DocImpactResult

Bases: TypedDict

Output shape of :func:analyze_doc_impact.

Source code in packages/axm-ast/src/axm_ast/core/doc_impact.py
Python
class DocImpactResult(TypedDict):
    """Output shape of :func:`analyze_doc_impact`."""

    doc_refs: dict[str, list[DocRefEntry]]
    undocumented: list[str]
    stale_signatures: list[StaleSignature]

DocRefEntry

Bases: TypedDict

Single documentation reference (backtick mention or heading hit).

Source code in packages/axm-ast/src/axm_ast/core/doc_impact.py
Python
class DocRefEntry(TypedDict):
    """Single documentation reference (backtick mention or heading hit)."""

    file: str
    line: int

StaleSignature

Bases: TypedDict

Stale signature record extracted from a doc code block.

actual_sig is added only after matching against the AST signatures; intermediate entries produced by :func:_match_signature_line omit it.

Source code in packages/axm-ast/src/axm_ast/core/doc_impact.py
Python
class StaleSignature(TypedDict):
    """Stale signature record extracted from a doc code block.

    ``actual_sig`` is added only after matching against the AST signatures;
    intermediate entries produced by :func:`_match_signature_line` omit it.
    """

    symbol: str
    file: str
    doc_sig: str
    line: int
    actual_sig: NotRequired[str]

analyze_doc_impact(root, symbols)

Full doc impact analysis for a set of symbols.

Combines doc refs, undocumented detection, and stale signature detection.

Caveat (canonical) — all three signals rest on a purely lexical matching, never a semantic one. A symbol counts as mentioned when its bare name appears between backticks or in a Markdown heading, and a documented signature is only compared inside a fenced code block. No meaning is read: a purely semantic change leaves this output identical byte for byte, and a bare name-drop of the symbol anywhere in the prose is enough to remove it from undocumented.

Read the result as a list of pages to read, not a proof that the documentation is correct or up to date. Never use it as a non-regression oracle: an unchanged output proves nothing about the prose still telling the truth — a human review remains the only verdict on doc correctness.

Limits — what this tool does not detect:

  1. A semantic change at unchanged name. Rewrite what a symbol means without touching its name and every signal stays identical, byte for byte: nothing here reports the prose that now lies.
  2. A bare name-drop counted as documentation. A single backticked mention, even in an unrelated sentence, is enough to drop the symbol from undocumented — presence is not coverage.
  3. undocumented is not a non-regression oracle. An empty or unchanged result proves nothing about documentation drift.

When the semantics of a symbol change while its name does not, re-read by hand every page listed in doc_refs: that manual pass is the one thing this tool cannot do for you. See docs/explanation/doc_impact_limits.md.

Parameters:

Name Type Description Default
root Path

Project root directory.

required
symbols list[str]

Symbol names to analyze.

required

Returns:

Type Description
DocImpactResult

Dict with doc_refs, undocumented, stale_signatures.

Source code in packages/axm-ast/src/axm_ast/core/doc_impact.py
Python
def analyze_doc_impact(
    root: Path,
    symbols: list[str],
) -> DocImpactResult:
    """Full doc impact analysis for a set of symbols.

    Combines doc refs, undocumented detection, and stale
    signature detection.

    Caveat (canonical) — all three signals rest on a **purely lexical**
    matching, never a semantic one. A symbol counts as mentioned when its
    bare name appears between backticks or in a Markdown heading, and a
    documented signature is only compared inside a fenced code block. No
    meaning is read: a purely semantic change leaves this output identical
    byte for byte, and a bare name-drop of the symbol anywhere in the prose
    is enough to remove it from ``undocumented``.

    Read the result as a list of pages to read, not a proof that the
    documentation is correct or up to date. Never use it as a non-regression
    oracle: an unchanged output proves nothing about the prose still telling
    the truth — a human review remains the only verdict on doc correctness.

    Limits — what this tool does not detect:

    1. A **semantic** change at unchanged name. Rewrite what a symbol means
       without touching its name and every signal stays identical, byte for
       byte: nothing here reports the prose that now lies.
    2. A bare **name-drop** counted as documentation. A single backticked
       mention, even in an unrelated sentence, is enough to drop the symbol
       from ``undocumented`` — presence is not coverage.
    3. ``undocumented`` is not a non-regression **oracle**. An empty or
       unchanged result proves nothing about documentation drift.

    When the semantics of a symbol change while its name does not, re-read by
    hand every page listed in ``doc_refs``: that manual pass is the one thing
    this tool cannot do for you. See ``docs/explanation/doc_impact_limits.md``.

    Args:
        root: Project root directory.
        symbols: Symbol names to analyze.

    Returns:
        Dict with ``doc_refs``, ``undocumented``, ``stale_signatures``.
    """
    refs = find_doc_refs(root, symbols)
    symbol_nodes = _index_symbol_nodes(root)
    return {
        "doc_refs": refs,
        "undocumented": find_undocumented(refs, symbol_nodes),
        "stale_signatures": find_stale_signatures(root, symbols),
    }

find_doc_refs(root, symbols)

Find documentation references for given symbols.

The hit is purely lexical: a reference is recorded when the bare symbol name appears between backticks or in a Markdown heading of a doc file. Nothing else is interpreted — prose outside those two forms, and the body of a fenced code block, establish no semantic link.

The returned entries are pages to read, never a non-regression oracle: a stable output does not prove the prose still describes the code.

Parameters:

Name Type Description Default
root Path

Project root directory.

required
symbols list[str]

Symbol names to search for in docs.

required

Returns:

Type Description
dict[str, list[DocRefEntry]]

Dict mapping symbol name to list of references

dict[str, list[DocRefEntry]]

(each with file and line keys).

Source code in packages/axm-ast/src/axm_ast/core/doc_impact.py
Python
def find_doc_refs(
    root: Path,
    symbols: list[str],
) -> dict[str, list[DocRefEntry]]:
    """Find documentation references for given symbols.

    The hit is **purely lexical**: a reference is recorded when the bare
    symbol name appears between backticks or in a Markdown heading of a doc
    file. Nothing else is interpreted — prose outside those two forms, and
    the body of a fenced code block, establish no semantic link.

    The returned entries are pages to read, never a non-regression oracle:
    a stable output does not prove the prose still describes the code.

    Args:
        root: Project root directory.
        symbols: Symbol names to search for in docs.

    Returns:
        Dict mapping symbol name to list of references
        (each with ``file`` and ``line`` keys).
    """
    doc_files = _collect_doc_files(root)
    refs: dict[str, list[DocRefEntry]] = {s: [] for s in symbols}
    for sym in symbols:
        for doc_file in doc_files:
            hits = _search_symbol_in_file(doc_file, sym, root)
            refs[sym].extend(hits)
    return refs

find_stale_signatures(root, symbols=None)

Detect stale code signatures in documentation.

Compares def / class signatures in doc code blocks against actual AST signatures.

The scope is strictly a fenced code block: a signature written in plain prose, in an indented block or in an inline span is never extracted. The comparison itself is a lexical string equality, so a reformatted but semantically equivalent signature still reads as stale, and a stale signature outside a fenced code block is invisible here.

Parameters:

Name Type Description Default
root Path

Project root directory.

required
symbols list[str] | None

Symbol names to check. If None, check all symbols.

None

Returns:

Type Description
list[StaleSignature]

List of dicts with symbol, file, doc_sig,

list[StaleSignature]

actual_sig, and line keys.

Source code in packages/axm-ast/src/axm_ast/core/doc_impact.py
Python
def find_stale_signatures(
    root: Path,
    symbols: list[str] | None = None,
) -> list[StaleSignature]:
    """Detect stale code signatures in documentation.

    Compares ``def`` / ``class`` signatures in doc code blocks
    against actual AST signatures.

    The scope is strictly a fenced code block: a signature written in plain
    prose, in an indented block or in an inline span is never extracted. The
    comparison itself is a lexical string equality, so a reformatted but
    semantically equivalent signature still reads as stale, and a stale
    signature outside a fenced code block is invisible here.

    Args:
        root: Project root directory.
        symbols: Symbol names to check. If ``None``, check all symbols.

    Returns:
        List of dicts with ``symbol``, ``file``, ``doc_sig``,
        ``actual_sig``, and ``line`` keys.
    """
    ast_sigs = _extract_ast_signatures(root)
    doc_files = _collect_doc_files(root)
    if symbols is None:
        sym_set = {qk.rsplit(".", 1)[-1] for qk in ast_sigs}
    else:
        sym_set = set(symbols)
    # Build reverse index: bare name → list of (qualified_key, sig)
    bare_index: dict[str, list[str]] = {}
    for qkey in ast_sigs:
        bare = qkey.rsplit(".", 1)[-1]
        bare_index.setdefault(bare, []).append(qkey)
    stale: list[StaleSignature] = []
    for doc_file in doc_files:
        doc_sigs = _extract_doc_signatures(doc_file, sym_set, root)
        for entry in doc_sigs:
            sym_name = entry["symbol"]
            qkeys = bare_index.get(sym_name, [])
            if not qkeys:
                continue
            doc_sig = entry["doc_sig"].strip()
            # Conservative: report stale only if NO qualified match agrees
            if all(ast_sigs[qk].strip() != doc_sig for qk in qkeys):
                entry["actual_sig"] = ast_sigs[qkeys[0]]
                stale.append(entry)
    return stale

find_undocumented(doc_refs, symbol_nodes)

Return public, docstring-less symbols absent from the prose docs.

A symbol is reported only when it has no prose documentation reference and the shared :func:is_documentation_required policy considers it a real gap — i.e. it is public surface (name not _-prefixed) and carries no docstring. Private/dunder symbols and symbols already documented by a docstring are never reported; this can only ever shrink the prose-missing set, never grow it (the output schema is unchanged).

A symbol absent from symbol_nodes (unresolvable in the analyzed source) keeps the legacy prose-only verdict, so a genuinely missing symbol is never silently dropped.

The prose signal it consumes is purely lexical: find_doc_refs only matches the bare name between backticks or in a Markdown heading, and never reads the meaning of a fenced code block. A single name-drop of the symbol therefore suffices to drop it from this list, without any real prose being written. See :func:analyze_doc_impact for the canonical caveat.

Parameters:

Name Type Description Default
doc_refs dict[str, list[DocRefEntry]]

Output of find_doc_refs.

required
symbol_nodes dict[str, DocSymbolNode]

Bare-name → parsed node index (see :func:_index_symbol_nodes) supplying the docstring/privacy signal.

required

Returns:

Type Description
list[str]

List of symbol names that are documentation-required gaps.

Source code in packages/axm-ast/src/axm_ast/core/doc_impact.py
Python
def find_undocumented(
    doc_refs: dict[str, list[DocRefEntry]],
    symbol_nodes: dict[str, DocSymbolNode],
) -> list[str]:
    """Return public, docstring-less symbols absent from the prose docs.

    A symbol is reported only when it has **no** prose documentation reference
    *and* the shared :func:`is_documentation_required` policy considers it a
    real gap — i.e. it is public surface (name not ``_``-prefixed) and carries
    no docstring. Private/dunder symbols and symbols already documented by a
    docstring are never reported; this can only ever *shrink* the prose-missing
    set, never grow it (the output schema is unchanged).

    A symbol absent from ``symbol_nodes`` (unresolvable in the analyzed source)
    keeps the legacy prose-only verdict, so a genuinely missing symbol is never
    silently dropped.

    The prose signal it consumes is **purely lexical**: ``find_doc_refs`` only
    matches the bare name between backticks or in a Markdown heading, and never
    reads the meaning of a fenced code block. A single name-drop of the symbol
    therefore suffices to drop it from this list, without any real prose being
    written. See :func:`analyze_doc_impact` for the canonical caveat.

    Args:
        doc_refs: Output of ``find_doc_refs``.
        symbol_nodes: Bare-name → parsed node index (see
            :func:`_index_symbol_nodes`) supplying the docstring/privacy signal.

    Returns:
        List of symbol names that are documentation-required gaps.
    """
    undocumented: list[str] = []
    for sym, refs in doc_refs.items():
        if refs:
            continue
        node = symbol_nodes.get(sym)
        if node is not None and not is_documentation_required(node):
            continue
        undocumented.append(sym)
    return undocumented