Skip to content

Architecture and persistence

Responsibilities

Module Responsibility
home Home creation/permissions and repository-path guard.
profile Active profile selection and transport overlay.
store TOML namespace reads, migration and file replacement.
resolver Validated keys, precedence, model loading and execution policies.
paths Typed runtime values, defaults, configured path guards and the profile's service ports.
doctor Provenance report without returning values.
isolation Resolver-backed profile locations and lexical containment.
tools AXMTool diagnostic boundaries.
cli Cyclopts request/response commands calling the central functions.

The SDK registry is axm.tools. No axm.commands registry or YAML hook integration is required. The legacy YAML engine is decommissioned; the remaining protocols_dir accessor is compatibility surface.

One TOML file per profile

The store uses ~/.axm/config.toml in production and ~/.axm/profiles/<name>/config.toml for a named profile. A dotted namespace maps to nested TOML tables:

TOML
[research.demo]
dataset = "sample"
timeout = 30

Each namespace's own scalar/array keys are separate from child namespace tables. write and replace_section preserve child tables. The generic resolver validates names before passing them to the store.

Lazy migration

The former <store-directory>/<namespace>.toml format remains readable. If the current section has any own keys, a read returns that section as a whole; it does not fill missing keys from the legacy file on each read. Only when the section is empty/missing does the read fall back to legacy.

On a write or delete to that namespace, legacy keys are merged first and current section keys win. The updated file is replaced, then the legacy file is removed. A delete of an absent key can still perform this migration. Only the selected profile's legacy files are considered. Unrelated namespaces are not all migrated at once; policy canonicalization has its own compatibility contract.

Atomicity, durability and concurrency

A normal write loads the whole mapping, serializes it to a same-directory temporary file and calls os.replace. The resulting file is chmod 0600 on POSIX. The swap helper removes its staged temporary path even if replace/chmod fails. The base home is tightened to 0700; profile directories use the process's normal mkdir permissions under that home.

This prevents readers observing a partly replaced file, but it is not a transaction across writers. There is no lock or version check: two overlapping read-modify-write operations can lose each other's updates, even when changing different namespaces. Serialize all writers to the same profile file in the owning application. There is no explicit file/directory fsync, so atomic replacement is not a promise of persistence through power loss.

The commit and legacy unlink are separate operations. A failure after replacement can leave changed contents despite an exception; chmod or legacy cleanup can fail after the new file is visible. Temporary-file creation/write failures before the swap helper are not covered by its cleanup guarantee.

Current limits that affect data

  • A missing, malformed or unreadable TOML file generally degrades to an empty mapping. A later write may replace that file with only the new known data. There is no automatic backup, recovery or corruption report.
  • Deleting an own key from a parent namespace can erase its child tables. Unlike write/replace_section, delete does not reattach children. This also affects set_(namespace, key, None) and the CLI delete command. Avoid parent-key deletion in stores with descendants until the product defect is corrected; the dedicated policy deletion uses replace_section.
  • The process environment selects the profile at each call; global environment changes are not safe per-thread profile contexts.
  • AXM_HOME and the typed accessor fallback do not share the store's home semantics. The profile guide details these boundaries rather than promising blanket isolation.

Who owns a listening point

A profile already owned every on-disk root a consumer needs, but not a port. So each service re-implemented the same refusal: outside production, demand an environment variable or a CLI option, else raise. That pushes onto every caller a decision the profile can take, and duplicates one fallback rule across packages that never talk to each other.

paths.service_port inverts that ownership. The production numbers are adopted -- written here as literals rather than imported, because the services carrying them today live in other repositories; duplicating the constant once is the price of moving the decision, and they drop their copy afterwards. Outside production the value is derived from the profile name with hashlib, never the builtin hash(), whose per-process PYTHONHASHSEED salt would rebind a service to a different port at every restart. The available band is cut into fixed 16-slot blocks, and each service's offset is its explicitly declared slot. Declaring another service therefore cannot move an existing service's port.

Two boundaries keep this honest. The derived number is only the default argument handed to get_int, so env > file > default is untouched and production resolves byte for byte as before -- the same additive guarantee the path accessors make. And this module computes a number: it opens no socket, reserves nothing and cannot prove the port is free. Distinctness between two services of one profile holds by construction; between two profile names it rests on a digest, so it is likely rather than guaranteed.

Ownership also implies the question can be asked about a profile, not only from inside it. The state roots already took a profile= keyword, so the port that did not was the asymmetry: deciding before launch whether two installations would collide meant exporting AXM_PROFILE -- changing the profile of the whole querying process to learn one number, and defeating a read-only report on the way. service_port therefore resolves through profile_root_for(requested) exactly as the path accessors do, a pure computation that neither consults nor mutates the active selection. Keeping the two families on one spelling -- keyword-only, named profile, None meaning the active profile -- is deliberate: a divergence between a state location and a listening point would be the next defect to explain.

Why a historical variable still wins

Moving the decision here also renamed the variable: network.mcp_port derives AXM_NETWORK_MCP_PORT, while installations in the field export AXM_MCP_PORT. Reading only the derived name would silently ignore a choice the operator made explicitly -- the trap already paid once for the warden socket, between its historical and its derived name. A service id may therefore carry historical variable names, and resolve consults them as a layer of its own.

That layer's position is the whole design. Above the derived name, an alias would shadow the canonical variable. Below the file layer -- which is what bolting the lookup on around resolve produces, since one call collapses the environment and file layers -- a stale configured value would outrank a variable the operator just posed. So the alias sits strictly between the two, under every profile rather than in production only.

The split is deliberate: the alias list is data at the call site (paths' service registry), the precedence is mechanism in the resolution layer. The parameter is keyword-only and defaults to an empty tuple, so no other accessor and no existing caller changes behaviour.

Security and secrets boundary

This is plaintext non-sensitive configuration. Keep passwords, API keys and tokens in axm-vault. Restrictive permissions do not turn TOML into a secret store. The provenance doctor does not return values but does parse file contents.

The store resolves its home, refuses git-repository ancestry and enforces that its resolved file paths stay below that home. This is containment under the base home, not a symlink-proof boundary between sibling profiles; avoid treating profile directory symlinks as isolation. On a write, axm_home() can still create/chmod the directory before the store rejects it; a read creates nothing. resolve_safe is a path check, not a lock against filesystem changes between check and use.

Why reading a location never creates the home

Computing a location and persisting state are two different rights, and the package used to conflate them: the warden-log fallback and the store's read path both went through axm_home(), so merely reading configuration materialised ~/.axm. A read-only consumer -- a report describing what a profile would use, a doctor, a dry run -- could not describe a location without creating it.

The split therefore lives in the layer that owns it. axm_home_path() computes, axm_home() creates and tightens to 0700; every read path resolves through the first, and the store's write path calls the second before staging its temp file. That call is not redundant with the ordinary mkdir next to it: dropping it would create a fresh home with the ambient umask instead of 0700, silently weakening the permissions of the file that holds the whole configuration.

The named profile root follows the same rule. profile_root_for(profile), and the profile= keyword the path accessors forward to get_path, answer for any profile without touching AXM_PROFILE and without creating a directory, so a report can describe several profiles in one pass. The convention stays rooted at Path.home() / ".axm": AXM_HOME keeps selecting only the compatibility config.toml, never this layout.

Why the isolation report asks the resolver

A report consulted before overwriting the state of a running service has to answer with the locations that service actually uses. The diagnostic used to recompose its own tree under <home>/profiles/<name>, which made it wrong in two directions at once: it ignored configured overrides, and it tested containment against a root it had just built, so the verdict was true by construction. For the default profile -- where the resolver deliberately places nothing under a profile directory -- it announced a profile tree and declared the state isolated with an empty escape list.

isolation therefore owns no path convention of its own. Its six entries are the six state accessors called with the requested profile, so one resolution path is shared by the runtime and by the report, and AXM_HOME keeps its narrow meaning here too. Because the report is read-only it resolves through axm_home_path() and never through axm_home(): describing a location still creates nothing. isolated and escapes stay derived from the reported paths through is_isolated rather than computed in parallel, so correcting the paths mechanically corrects the verdict -- there is no second truth to maintain.

The visible consequence is the point of the change. The default profile is now reported as not isolated, with its six locations named as escapes, because it owns no root. A report still calling it isolated would be the defect intact.