Skip to main content

Architecture

One sentence: episodes are evidence, nodes are things, facts are beliefs, and the context pack is the product. Everything in mecha-graph is machinery that turns raw records of your life into a token-bounded, provenance-carrying slice an agent can trust.

SOURCES EVIDENCE WIRING BELIEFS PRODUCT
(bee, cal, ┌─────────────┐ ┌──────────────┐ ┌─────────────────┐ ┌─────────┐
slack, ──▶│ episodes │──▶│ mentions │──▶│ fact candidates │──▶│ context │
github, │ (append- │ │ episode↔node │ │ → review → │ │ pack │
sessions, │ only raw │ │ (M:N) │ │ facts (bi- │ │ (ranked,│
reflect, │ records) │ │ │ │ temporal) │ │ bounded)│
mbox, …) └─────────────┘ └──────┬───────┘ └─────────────────┘ └─────────┘
│
┌──────┴───────┐
│ nodes │ people · projects · orgs ·
│ (entities) │ topics · places · tasks · …
└──────────────┘

The four layers​

1. Episodes — evidence (append-only)​

An episode is the raw record that something happened, before any interpretation: one Bee conversation, one calendar event, one Slack channel-day, one GitHub repo-day, one agent session, one Reflect note.

  • Keyed by (source, source_id) — re-ingest is idempotent; a content_hash detects edits (e.g. a re-exported Reflect note updates in place, never duplicates).
  • occurred_at is when it happened in the world (not when we saw it — that's ingested_at; the distinction matters below).
  • sensitivity tiers gate retrieval: private episodes (all Bee transcripts) are excluded from default search; callers opt in.
  • Episodes are never edited or deleted. The two exceptions are deliberate: mecha-graph redact (true delete, §10) and nothing else. When a fact turns out wrong you supersede the fact — the episode that spawned it stays, because it's the answer to "where did we learn this?"
  • For file-based sources the full original is archived to episode_raw inside the encrypted DB, verified, and only then is the plaintext file deleted (capture_delete). Stream sources never touch disk at all.

2. Nodes — things (the entity layer)​

People, projects, orgs, topics, places, goals/areas/tasks, events, documents. A closed type set — extractors return null rather than invent types, because open-ended types cause junk-node explosion.

Identity is layered, strongest first:

  • node_identifier — deterministic keys (email, phone, slack_uid, url, cwd/path). Two sources asserting the same identifier ARE the same entity; this merges with no model and no review.
  • node_alias — names and nicknames, indexed. Aliases that map to more than one node ("two Victors") never auto-link; ambiguity is surfaced to the caller, and a user's answer is written back as a permanent alias — resolution learns.
  • canonical_name — the display name, scanned alongside aliases.

Event/document nodes are retrieval targets, never query anchors: a meeting literally titled "Nadia" must not shadow the person.

3. Mentions — the wiring (episode ↔ node, M:N)​

mention rows connect episodes to the nodes they involve. This table is load-bearing far beyond what it looks like:

  • Entity timelines ("show me everything with June") are mention scans.
  • Filter-first retrieval (§8.1): when a query names an entity, the candidate set collapses to episodes mentioning that node before any ranking runs. An episode whose text says "flowmail" but has no mention edge is invisible to a flowmail-anchored query — text match is not membership. (This is why linking quality matters more than ranking.)
  • Co-occurrence statistics (NPMI) that propose edges read this table.

Mentions carry their extractor (attendee, alias, backlink, temporal_join, reflect, manual) and a confidence — provenance all the way down.

4. Facts — beliefs (interpreted claims, bi-temporal)​

A fact is a claim connecting nodes: Nadia –works_at→ Bayview Institute, with a natural-language statement ("Nadia works at Bayview Institute.") because sentences embed and retrieve well while triples traverse well — store both, embed the sentence, walk the triple.

  • Two timelines per fact: valid time (when it was true in the world: valid_from/valid_to) and system time (when we believed it: ingested_at/invalidated_at). You can ask "what was true in March" and "what did I believe in March" separately. Supersede, never delete.
  • Predicates are a controlled vocabulary (with an alias table), or works_on/working_on/is_working_on become three relations.
  • observation_count accumulates corroboration: re-asserting an existing live fact bumps the count instead of duplicating. Re-observation is evidence.
  • Every fact points at the episode it came from.
  • Beliefs have a polarity. A negative fact ("Nadia does not work at NYU") is rejection memory: it records that something was asked and answered, so nothing re-proposes it. Traversal (fact_current, the edges view, linkers, GTD, stats) is positive-only — a negative edge would be a bug — while every display path serves both polarities, because the surfaces where you'd re-ask are exactly where a denial has to be readable.

Three ways a belief stops being current, and the difference is not cosmetic — it is what the two timelines are for:

MeansSets
supersedereplaced by a better valueboth timestamps
decaytrue then, false now (the world moved)valid time only
never-truewe were wrong to believe it at allsystem time + a zero-length valid window

Only decay leaves invalidated_at NULL, so facts_as_of keeps answering correctly for the period the belief actually held, and no producing class is blamed for having been right at the time. A never-true retraction collapses valid time to a point so no as-of date serves it. All three remove a fact from retrieval; kg_timeline still shows the whole history.

How episodes become facts: the trust ladder​

Nothing skips the ladder. Each rung is cheaper-and-more-precise first (§7), and the rung determines whether the result writes directly or must be staged for review:

RungMechanismWrites
Deterministic keysemail / slack_uid / cwd / attendee listsdirect (mentions, identity)
User-authoredTUI capture, Reflect typed notes, b-bind aliasesdirect, high confidence
Alias/name scanknown names found in episode textdirect mentions (unambiguous only)
Temporal joinBee recording overlaps calendar meetingdirect mentions, confidence < 1
Statistical (NPMI)frequency-corrected co-occurrencedirect related_to, capped confidence
Embedding kNNmean-centered node centroids, similar contextsstaged as candidates
Structural (Adamic-Adar)shared graph neighborhoodsstaged as candidates
LLM extractiongemma reads episodes, proposes factsstaged as candidates

The staging queue (fact_candidate) is the membrane between "a model said so" and "the graph believes it": extraction proposes, promotion disposes. mecha-graph precheck auto-triages that queue — duplicates of known facts are rejected with an observation bump, in-queue repeats collapse, conversational recaps ("X discussed Y") are dropped as bloat since the episode already records them, contradictions on single-valued predicates are flagged and always held for a human, and (opt-in) clean novel facts on durable predicates auto-accept. With --triage (opt-in, set from the owner's own verdict history), two more lanes run: a restatement of a live fact — same subject, predicate and object, no longer, not negated, cosine 0.90–0.97 — is folded into it as an observation, and a claim about a subject the graph knows nothing about and that fewer than three claims name is rejected as a one-off (minting and the lane count one pool, so the third mention mints the node instead). What reaches the TUI review screen is meant to be only what genuinely needs a decision.

Retrieval — filter first, rank second​

Embedding "when did I last meet June?" yields a vector about the question, not about June. So the router (§8):

  1. Detects entities in the query deterministically (alias + identifier scan — sub-millisecond, not an LLM), plus #tag filters and time expressions.
  2. Classifies intent: LOOKUP ("when did I last…") is answered from the person_interaction rollup with no embeddings at all; AGGREGATE ("who do I interact with most") reads rollups/counts; RECALL does hybrid search.
  3. Filters first: entity and tag filters collapse the candidate set via mentions before ranking.
  4. Ranks with BM25 (porter-stemmed FTS5) + vector similarity fused by RRF, facts and episodes competing on one scale.
  5. Packs: the result is a context pack — ranked items with kind, source, timestamps, and provenance ids, truncated to a token budget. If entity resolution was ambiguous, the pack says so and the consumer is expected to ask, not guess.

Only live beliefs are served: the search indexes cover every row ever written, so retrieval filters retracted facts explicitly rather than trusting the index to forget them.

The envelope carries two further self-descriptions, both omitted when they have nothing to say:

  • flags (≤2) — problems mecha-graph detected in what it is about to return: a contradiction on a single-valued predicate, a denial contesting a served belief, a fact past its predicate's half-life. mecha-graph detects with provenance; the consumer judges whether to act. Same division as ambiguity, generalized.
  • scope — whether this pack could see facts, evidence, or both. A verifier has to know what an answer could have drawn on; facts_only/evidence_only are how two readers can be given deliberately blind halves of the same question.

Time, privacy, and ops​

  • Future episodes are not interactions: a scheduled meeting is not "last met". Rollups exclude occurred_at > now.
  • At rest: the DB is SQLCipher-encrypted (~/.mecha-graph/db.key, 0600); plaintext source files are deleted after verified capture; analytical snapshots come from mecha-graph decrypt and are transaction-pinned.
  • Nightly (03:30 cron): source sync → bee-facts two-way sync → linker cascade → GPU-gated embed + LLM extract → precheck → scope summaries → MEMORY.md boot context → health alerts. Everything is cursored and idempotent; a missed night just catches up.
  • Eval: ~/.mecha-graph/eval/gold.jsonl (outside the repo — it is mined from real episodes; --gold or MECHA_GRAPH_GOLD override) is a regression guard run after any router/linker change. Recall@10 = 1.00 is the floor, not a score.

Boundaries — what lives here, what lives in the agent​

mecha-graph and mecha (~/Github/mecha) are deliberately separate repositories with no compile-time dependency in either direction. The entire interface is the MCP tool namespace (kg_search, kg_upsert, …, which carry their own kg_ prefix, so a consumer registers them unprefixed); mecha's own eval suite tests against a fixture graph server, not against mecha-graph itself. Settled 2026-08-12; the reasoning is worth keeping because it will be re-litigated.

Two opt-in runtime dependencies: the TUI's closures and extraction's holds (the first ruled 2026-09-25, mecha's APPRAISAL-WIRING-DESIGN.md row 1c, option A3; the second added 2026-09-28, below). A task moved to done or dropped, or reopened, is a verdict mecha records on its closure record and appraises; the TUI's status keys used to write it straight here, where mecha never saw it. With [board] close_through = "mecha" in ~/.mecha-graph/config.toml, the TUI hands exactly those moves to mecha tasks set <task> --status <s> --surface graph-tui (--only-open on a close) and writes nothing itself; mecha's own graph server makes the write through kg_task_update, and the TUI reads the row back to confirm it landed here. What the opt-in changes, and what it does not:

  • Without it, nothing is different. No code looks for mecha; standalone installs keep the direct write. Moves between open statuses stay direct even when opted in — they are not verdicts.
  • With it, the route never degrades. A missing program, an unreadable config, or a TUI on any database but the default one (a fork, another --db) refuses the move with nothing written. The child runs without MECHA_GRAPH_DB, so the server it starts opens the default database — the one the TUI must be on. done ↔ dropped is refused too: it crosses no line, so mecha would record nothing, yet it changes a recorded closure's verdict — reopen, then close.
  • Still no compile-time dependency either way, and mecha-graph-core still knows nothing about any agent: it parses a command string; the argv and the graph-tui surface live in the mecha-graph binary (src/closure.rs). The CLI and the MCP server are unchanged — the MCP server is the write path mecha uses.
  • It is a config key, not an environment variable, on the [llm] model_path precedent: a permission granted by being explicitly configured. The environment variables here name locations and secrets; an opt-in carried in the environment would be lost by the next shell that did not export it, and silently reverting to the direct write is the failure the opt-in exists to prevent.

Extraction's holds (2026-09-28, the owner's choice after a nightly extraction undid a model switch). With [llm] holds_dir = "~/.mecha/holds", extract holds mecha's router one episode at a time, in the file protocol mecha's hold.rs defines (src/holds.rs, pinned by the_files_are_the_ones_mecha_reads). A switch then waits for the episode in flight. The same rules as the TUI's opt-in:

  • Without it, nothing is different. Nothing is held. Every request still follows the router's loaded model (ChatClient::follow_settled), which is agent-agnostic and needs no opt-in.
  • With it, the guard never degrades. A missing directory or an unreadable config stops extraction with an error.
  • Still no compile-time dependency, and mecha-graph-core knows only a gate (extract_pending_gated). The protocol lives in the binary.
  • A config key, not an environment variable, for the reason above. The first draft was an environment variable, and a hand-run extract went unheld (found on review of #24).

Why separate​

Durability asymmetry — the decisive argument. mecha-graph holds an encrypted, migration-versioned store of a life; it is irreplaceable if corrupted. mecha holds agent behaviour, which is replaceable and should be replaced as models and harnesses change. You do not fold a durable asset into a disposable tool. Expect mecha-graph to outlive whatever harness is currently in front of it.

The boundary carries the security model. mecha's taint interlock treats a mecha-graph read as arming both taint legs, and mecha's eval has cases asserting exactly that (web-then-memory: taint private+untrusted, blocked_sends: 0). That analysis only works because reading mecha-graph is an external act. Merge the two and "reading my own memory" versus "reading mecha-graph" blurs precisely where it currently needs to be sharp.

§2 is only enforceable across a crate boundary. "mecha-graph-core knows nothing about any agent" has repeatedly produced better designs by pushing orchestration out — the ask_ada route, pack flags that describe rather than decide, mecha-graph not writing into mecha's mailbox. In one workspace that constraint erodes by convenience.

Multiple consumers. Claude Code and Hermes over MCP, and the mecha-graph CLI; FlowMail is a future consumer on macOS. Even at one real consumer the MCP surface costs nothing already being paid.

The two invariants that keep the split clean​

Both are current practice. Violating either is what would actually create redundancy between the repos:

  1. mecha-graph's own interface stays non-conversational — commands, tables, keystrokes. The moment mecha-graph grows a chat surface there are two harnesses.
  2. mecha never stores facts — it produces episodes through kg_upsert and reads context packs. The moment mecha caches graph state there are two graphs and a sync problem.

Two interfaces because there are two modes​

The direct interface (mecha-graph CLI + TUI) is load-bearing, not a convenience:

  • it is the unmediated correction channel — if the only way to fix the graph is to ask an agent, there is no ground-truth path;
  • it is the audit surface for autonomy — auto-accepting classes of fact is only safe if their output can be inspected without a model in between;
  • bulk work is keyboard work — cluster review with a marked set beats conversing about a thousand candidates.
ModeSurfaceFor
conversational, in-context, reactivemechapoint-of-use flags, questions, corrections in flow
direct, bulk, deliberatemecha-graph CLI/TUIcluster review, schema authoring, health, forks

Most apparent duplication between the repos is nominal — the same word for different jobs. mecha-graph's review queue holds world facts, mecha's holds behaviour rules. mecha-graph's sensitivity is static classification on a row; mecha's taint is dynamic flow control on a conversation. mecha-graph's eval measures retrieval quality; mecha's grades agent traces. Only two overlaps are real: the class/outcome ledger (same state machine, different substrate — share the written mechanism, implement twice) and scheduling (mecha's cron.rs/trigger.rs is the better one; scripts/nightly.sh stays as the standalone fallback).

mecha-graph is the agent's declarative memory​

ACT-R — already borrowed for base-level activation (§11.5) — splits declarative memory (chunks) from procedural memory (production rules). That split answers "should mecha-graph be mecha's memory system?":

MemoryContentHome
semantic (declarative)facts about people, projects, orgsmecha-graph
episodic (declarative)what happened, including the agent's own sessionsmecha-graph (sources/sessions.rs)
proceduralreflexion rules — "check the graph before searching the web"mecha ~/.mecha/learning/
workingthe live conversation, compactionmecha runtime (state, not memory)
operationaloutbox, triggers, messages, livenessmecha files (infrastructure)

So mecha-graph already is mecha's declarative memory: the session-end distiller writes episodes through kg_upsert, sources/sessions.rs ingests sessions, and §8.3's boot digest supplies opening context.

The litmus for anything new is the one that already routes corrections: "would the user ask an assistant about this later?" → mecha-graph. "Should the agent behave differently next time?" → the harness.

Procedural memory stays out of mecha-graph even though bi-temporality, provenance and supersede would all be useful for rules — it is not world knowledge, it is model-specific, and a harness swap should not inherit the previous harness's habits. Revisit only if rules reach the hundreds.

The boundary that genuinely blurs: agent episodic memory with retrieval ("what did I try last time this build failed?"). Session content is often world knowledge and belongs here; the operational residue ("tried X, it failed") is procedural and does not.

One graph, not many​

MECHA_GRAPH_DB/--db makes a second graph free today, and it is still usually the wrong move: a second graph splits the entity space. The same person in two graphs means schema drift, identity mismatch, no cross-graph query, and a reconciliation problem needing machinery this project deliberately declined to build.

Partition inside one graph instead — sensitivity tiers, annotation tags, and the scope_id parent chain (§4.5) are the axes that already exist. Note scope_id is hierarchical containment, not tenancy.

The one case where a separate graph is right is a shared or team graph with genuinely different ownership and access. That is real federation, and it is the condition under which the parked peer-to-peer reconciliation work (Semantic Gossiping, cycle consistency — covered in the internal research notes) becomes relevant again.

A worked example​

A Reflect note titled Iris Calder with Type: #person, Email: iris.calder@example.com, Company: Westfield:

  1. mecha-graph ingest reflect streams it from the export zip → episode (reflect.note, keyed by the note's stable id), raw markdown archived, zip deleted after verification.
  2. mecha-graph reflect-process sees the type tag → resolves the email identifier → attaches to the existing Iris node rather than creating a duplicate; Company: Westfield becomes a fact (works_at, extractor reflect, pointing at this episode); the episode gets a mention of Iris.
  3. mecha-graph link scans every episode for known names → more mentions (its candidate-staging tiers — kNN, structural, rules — run only with --propose, off in the nightly by default); NPMI notices who co-occurs with Iris unusually often → related_to facts; kNN and Adamic-Adar stage speculative candidates for review.
  4. A query for "iris westfield" detects the entity, collapses to episodes mentioning her, ranks, and returns a pack whose items each carry the uid you'd need to trace any claim back to this note.