Architecture
One sentence: episodes are evidence, nodes are things, facts are beliefs, and the context pack is the product. Everything in mecha-graph is machinery that turns raw records of your life into a token-bounded, provenance-carrying slice an agent can trust.
SOURCES EVIDENCE WIRING BELIEFS PRODUCT
(bee, cal, ┌─────────────┐ ┌──────────────┐ ┌─────────────────┐ ┌─────────┐
slack, ──▶│ episodes │──▶│ mentions │──▶│ fact candidates │──▶│ context │
github, │ (append- │ │ episode↔node │ │ → review → │ │ pack │
sessions, │ only raw │ │ (M:N) │ │ facts (bi- │ │ (ranked,│
reflect, │ records) │ │ │ │ temporal) │ │ bounded)│
mbox, …) └─────────────┘ └──────┬───────┘ └─────────────────┘ └─────────┘
│
┌──────┴───────┐
│ nodes │ people · projects · orgs ·
│ (entities) │ topics · places · tasks · …
└──────────────┘
The four layers
1. Episodes — evidence (append-only)
An episode is the raw record that something happened, before any interpretation: one Bee conversation, one calendar event, one Slack channel-day, one GitHub repo-day, one agent session, one Reflect note.
- Keyed by
(source, source_id)— re-ingest is idempotent; acontent_hashdetects edits (e.g. a re-exported Reflect note updates in place, never duplicates). occurred_atis when it happened in the world (not when we saw it — that'singested_at; the distinction matters below).sensitivitytiers gate retrieval:privateepisodes (all Bee transcripts) are excluded from default search; callers opt in.- Episodes are never edited or deleted. The two exceptions are
deliberate:
mecha-graph redact(true delete, §10) and nothing else. When a fact turns out wrong you supersede the fact — the episode that spawned it stays, because it's the answer to "where did we learn this?" - For file-based sources the full original is archived to
episode_rawinside the encrypted DB, verified, and only then is the plaintext file deleted (capture_delete). Stream sources never touch disk at all.
2. Nodes — things (the entity layer)
People, projects, orgs, topics, places, goals/areas/tasks, events, documents. A closed type set — extractors return null rather than invent types, because open-ended types cause junk-node explosion.
Identity is layered, strongest first:
node_identifier— deterministic keys (email, phone, slack_uid, url, cwd/path). Two sources asserting the same identifier ARE the same entity; this merges with no model and no review.node_alias— names and nicknames, indexed. Aliases that map to more than one node ("two Victors") never auto-link; ambiguity is surfaced to the caller, and a user's answer is written back as a permanent alias — resolution learns.canonical_name— the display name, scanned alongside aliases.
Event/document nodes are retrieval targets, never query anchors: a meeting literally titled "Nadia" must not shadow the person.
3. Mentions — the wiring (episode ↔ node, M:N)
mention rows connect episodes to the nodes they involve. This table is
load-bearing far beyond what it looks like:
- Entity timelines ("show me everything with June") are mention scans.
- Filter-first retrieval (§8.1): when a query names an entity, the candidate set collapses to episodes mentioning that node before any ranking runs. An episode whose text says "flowmail" but has no mention edge is invisible to a flowmail-anchored query — text match is not membership. (This is why linking quality matters more than ranking.)
- Co-occurrence statistics (NPMI) that propose edges read this table.
Mentions carry their extractor (attendee, alias, backlink, temporal_join,
reflect, manual) and a confidence — provenance all the way down.
4. Facts — beliefs (interpreted claims, bi-temporal)
A fact is a claim connecting nodes: Nadia –works_at→ Bayview Institute, with
a natural-language statement ("Nadia works at Bayview Institute.") because
sentences embed and retrieve well while triples traverse well — store
both, embed the sentence, walk the triple.
- Two timelines per fact: valid time (when it was true in the world:
valid_from/valid_to) and system time (when we believed it:ingested_at/invalidated_at). You can ask "what was true in March" and "what did I believe in March" separately. Supersede, never delete. - Predicates are a controlled vocabulary (with an alias table), or
works_on/working_on/is_working_onbecome three relations. observation_countaccumulates corroboration: re-asserting an existing live fact bumps the count instead of duplicating. Re-observation is evidence.- Every fact points at the episode it came from.
- Beliefs have a polarity. A negative fact ("Nadia does not work
at NYU") is rejection memory: it records that something was asked and
answered, so nothing re-proposes it. Traversal (
fact_current, theedgesview, linkers, GTD, stats) is positive-only — a negative edge would be a bug — while every display path serves both polarities, because the surfaces where you'd re-ask are exactly where a denial has to be readable.
Three ways a belief stops being current, and the difference is not cosmetic — it is what the two timelines are for:
| Means | Sets | |
|---|---|---|
| supersede | replaced by a better value | both timestamps |
| decay | true then, false now (the world moved) | valid time only |
| never-true | we were wrong to believe it at all | system time + a zero-length valid window |
Only decay leaves invalidated_at NULL, so facts_as_of keeps
answering correctly for the period the belief actually held, and no
producing class is blamed for having been right at the time. A
never-true retraction collapses valid time to a point so no as-of date
serves it. All three remove a fact from retrieval; kg_timeline still
shows the whole history.
How episodes become facts: the trust ladder
Nothing skips the ladder. Each rung is cheaper-and-more-precise first (§7), and the rung determines whether the result writes directly or must be staged for review:
| Rung | Mechanism | Writes |
|---|---|---|
| Deterministic keys | email / slack_uid / cwd / attendee lists | direct (mentions, identity) |
| User-authored | TUI capture, Reflect typed notes, b-bind aliases | direct, high confidence |
| Alias/name scan | known names found in episode text | direct mentions (unambiguous only) |
| Temporal join | Bee recording overlaps calendar meeting | direct mentions, confidence < 1 |
| Statistical (NPMI) | frequency-corrected co-occurrence | direct related_to, capped confidence |
| Embedding kNN | mean-centered node centroids, similar contexts | staged as candidates |
| Structural (Adamic-Adar) | shared graph neighborhoods | staged as candidates |
| LLM extraction | gemma reads episodes, proposes facts | staged as candidates |
The staging queue (fact_candidate) is the membrane between "a model
said so" and "the graph believes it": extraction proposes, promotion
disposes. mecha-graph precheck auto-triages that queue — duplicates of known
facts are rejected with an observation bump, in-queue repeats collapse,
conversational recaps ("X discussed Y") are dropped as bloat since the
episode already records them, contradictions on single-valued predicates
are flagged and always held for a human, and (opt-in) clean novel facts
on durable predicates auto-accept. With --triage (opt-in, set from the
owner's own verdict history), two more lanes run: a restatement of a live
fact — same subject, predicate and object, no longer, not negated, cosine
0.90–0.97 — is folded into it as an observation, and a claim about a
subject the graph knows nothing about and that fewer than three claims
name is rejected as a one-off (minting and the lane count one pool, so the
third mention mints the node instead). What reaches the TUI review screen
is meant to be only what genuinely needs a decision.
Retrieval — filter first, rank second
Embedding "when did I last meet June?" yields a vector about the question, not about June. So the router (§8):
- Detects entities in the query deterministically (alias +
identifier scan — sub-millisecond, not an LLM), plus
#tagfilters and time expressions. - Classifies intent: LOOKUP ("when did I last…") is answered from
the
person_interactionrollup with no embeddings at all; AGGREGATE ("who do I interact with most") reads rollups/counts; RECALL does hybrid search. - Filters first: entity and tag filters collapse the candidate set via mentions before ranking.
- Ranks with BM25 (porter-stemmed FTS5) + vector similarity fused by RRF, facts and episodes competing on one scale.
- Packs: the result is a context pack — ranked items with kind, source, timestamps, and provenance ids, truncated to a token budget. If entity resolution was ambiguous, the pack says so and the consumer is expected to ask, not guess.
Only live beliefs are served: the search indexes cover every row ever written, so retrieval filters retracted facts explicitly rather than trusting the index to forget them.
The envelope carries two further self-descriptions, both omitted when they have nothing to say:
flags(≤2) — problems mecha-graph detected in what it is about to return: a contradiction on a single-valued predicate, a denial contesting a served belief, a fact past its predicate's half-life. mecha-graph detects with provenance; the consumer judges whether to act. Same division as ambiguity, generalized.scope— whether this pack could see facts, evidence, or both. A verifier has to know what an answer could have drawn on;facts_only/evidence_onlyare how two readers can be given deliberately blind halves of the same question.
Time, privacy, and ops
- Future episodes are not interactions: a scheduled meeting is not
"last met". Rollups exclude
occurred_at > now. - At rest: the DB is SQLCipher-encrypted (
~/.mecha-graph/db.key, 0600); plaintext source files are deleted after verified capture; analytical snapshots come frommecha-graph decryptand are transaction-pinned. - Nightly (03:30 cron): source sync → bee-facts two-way sync →
linker cascade → GPU-gated embed + LLM extract → precheck → scope
summaries →
MEMORY.mdboot context → health alerts. Everything is cursored and idempotent; a missed night just catches up. - Eval:
~/.mecha-graph/eval/gold.jsonl(outside the repo — it is mined from real episodes;--goldorMECHA_GRAPH_GOLDoverride) is a regression guard run after any router/linker change. Recall@10 = 1.00 is the floor, not a score.
Boundaries — what lives here, what lives in the agent
mecha-graph and mecha (~/Github/mecha) are deliberately separate repositories
with no compile-time dependency in either direction. The entire
interface is the MCP tool namespace (kg_search, kg_upsert, …, which carry
their own kg_ prefix, so a consumer registers them unprefixed); mecha's own eval suite tests against a fixture graph server, not
against mecha-graph itself. Settled 2026-08-12; the reasoning is worth keeping
because it will be re-litigated.
Two opt-in runtime dependencies: the TUI's closures and extraction's
holds (the first ruled 2026-09-25, mecha's APPRAISAL-WIRING-DESIGN.md
row 1c, option A3; the second added 2026-09-28, below). A task
moved to done or dropped, or reopened, is a verdict mecha records on its
closure record and appraises; the TUI's status keys used to write it
straight here, where mecha never saw it. With [board] close_through = "mecha" in ~/.mecha-graph/config.toml, the TUI hands exactly those moves
to mecha tasks set <task> --status <s> --surface graph-tui (--only-open
on a close) and writes nothing itself; mecha's own graph server makes the
write through kg_task_update, and the TUI reads the row back to confirm it
landed here. What the opt-in changes, and what it does not:
- Without it, nothing is different. No code looks for mecha; standalone installs keep the direct write. Moves between open statuses stay direct even when opted in — they are not verdicts.
- With it, the route never degrades. A missing program, an unreadable
config, or a TUI on any database but the default one (a fork, another
--db) refuses the move with nothing written. The child runs withoutMECHA_GRAPH_DB, so the server it starts opens the default database — the one the TUI must be on.done↔droppedis refused too: it crosses no line, so mecha would record nothing, yet it changes a recorded closure's verdict — reopen, then close. - Still no compile-time dependency either way, and
mecha-graph-corestill knows nothing about any agent: it parses a command string; the argv and thegraph-tuisurface live in themecha-graphbinary (src/closure.rs). The CLI and the MCP server are unchanged — the MCP server is the write path mecha uses. - It is a config key, not an environment variable, on the
[llm] model_pathprecedent: a permission granted by being explicitly configured. The environment variables here name locations and secrets; an opt-in carried in the environment would be lost by the next shell that did not export it, and silently reverting to the direct write is the failure the opt-in exists to prevent.
Extraction's holds (2026-09-28, the owner's choice after a nightly
extraction undid a model switch). With [llm] holds_dir = "~/.mecha/holds", extract holds mecha's router one episode at a time,
in the file protocol mecha's hold.rs defines (src/holds.rs, pinned by
the_files_are_the_ones_mecha_reads). A switch then waits for the
episode in flight. The same rules as the TUI's opt-in:
- Without it, nothing is different. Nothing is held. Every request
still follows the router's loaded model (
ChatClient::follow_settled), which is agent-agnostic and needs no opt-in. - With it, the guard never degrades. A missing directory or an unreadable config stops extraction with an error.
- Still no compile-time dependency, and
mecha-graph-coreknows only a gate (extract_pending_gated). The protocol lives in the binary. - A config key, not an environment variable, for the reason above. The
first draft was an environment variable, and a hand-run
extractwent unheld (found on review of #24).
Why separate
Durability asymmetry — the decisive argument. mecha-graph holds an encrypted, migration-versioned store of a life; it is irreplaceable if corrupted. mecha holds agent behaviour, which is replaceable and should be replaced as models and harnesses change. You do not fold a durable asset into a disposable tool. Expect mecha-graph to outlive whatever harness is currently in front of it.
The boundary carries the security model. mecha's taint interlock
treats a mecha-graph read as arming both taint legs, and mecha's eval has cases
asserting exactly that (web-then-memory: taint private+untrusted,
blocked_sends: 0). That analysis only works because reading mecha-graph is an
external act. Merge the two and "reading my own memory" versus
"reading mecha-graph" blurs precisely where it currently needs to be sharp.
§2 is only enforceable across a crate boundary. "mecha-graph-core knows
nothing about any agent" has repeatedly produced better designs by
pushing orchestration out — the ask_ada route, pack flags that
describe rather than decide, mecha-graph not writing into mecha's mailbox. In
one workspace that constraint erodes by convenience.
Multiple consumers. Claude Code and Hermes over MCP, and the mecha-graph
CLI; FlowMail is a future consumer on macOS. Even at
one real consumer the MCP surface costs nothing already being paid.
The two invariants that keep the split clean
Both are current practice. Violating either is what would actually create redundancy between the repos:
- mecha-graph's own interface stays non-conversational — commands, tables, keystrokes. The moment mecha-graph grows a chat surface there are two harnesses.
- mecha never stores facts — it produces episodes through
kg_upsertand reads context packs. The moment mecha caches graph state there are two graphs and a sync problem.
Two interfaces because there are two modes
The direct interface (mecha-graph CLI + TUI) is load-bearing, not a
convenience:
- it is the unmediated correction channel — if the only way to fix the graph is to ask an agent, there is no ground-truth path;
- it is the audit surface for autonomy — auto-accepting classes of fact is only safe if their output can be inspected without a model in between;
- bulk work is keyboard work — cluster review with a marked set beats conversing about a thousand candidates.
| Mode | Surface | For |
|---|---|---|
| conversational, in-context, reactive | mecha | point-of-use flags, questions, corrections in flow |
| direct, bulk, deliberate | mecha-graph CLI/TUI | cluster review, schema authoring, health, forks |
Most apparent duplication between the repos is nominal — the same word
for different jobs. mecha-graph's review queue holds world facts, mecha's
holds behaviour rules. mecha-graph's sensitivity is static classification
on a row; mecha's taint is dynamic flow control on a conversation.
mecha-graph's eval measures retrieval quality; mecha's grades agent traces.
Only two overlaps are real: the class/outcome ledger (same state
machine, different substrate — share the written mechanism, implement
twice) and scheduling (mecha's cron.rs/trigger.rs is the better
one; scripts/nightly.sh stays as the standalone fallback).
mecha-graph is the agent's declarative memory
ACT-R — already borrowed for base-level activation (§11.5) — splits declarative memory (chunks) from procedural memory (production rules). That split answers "should mecha-graph be mecha's memory system?":
| Memory | Content | Home |
|---|---|---|
| semantic (declarative) | facts about people, projects, orgs | mecha-graph |
| episodic (declarative) | what happened, including the agent's own sessions | mecha-graph (sources/sessions.rs) |
| procedural | reflexion rules — "check the graph before searching the web" | mecha ~/.mecha/learning/ |
| working | the live conversation, compaction | mecha runtime (state, not memory) |
| operational | outbox, triggers, messages, liveness | mecha files (infrastructure) |
So mecha-graph already is mecha's declarative memory: the session-end
distiller writes episodes through kg_upsert, sources/sessions.rs
ingests sessions, and §8.3's boot digest supplies opening context.
The litmus for anything new is the one that already routes corrections: "would the user ask an assistant about this later?" → mecha-graph. "Should the agent behave differently next time?" → the harness.
Procedural memory stays out of mecha-graph even though bi-temporality, provenance and supersede would all be useful for rules — it is not world knowledge, it is model-specific, and a harness swap should not inherit the previous harness's habits. Revisit only if rules reach the hundreds.
The boundary that genuinely blurs: agent episodic memory with retrieval ("what did I try last time this build failed?"). Session content is often world knowledge and belongs here; the operational residue ("tried X, it failed") is procedural and does not.
One graph, not many
MECHA_GRAPH_DB/--db makes a second graph free today, and it is still
usually the wrong move: a second graph splits the entity space.
The same person in two graphs means schema drift, identity mismatch,
no cross-graph query, and a reconciliation problem needing machinery
this project deliberately declined to build.
Partition inside one graph instead — sensitivity tiers, annotation
tags, and the scope_id parent chain (§4.5) are the axes that already
exist. Note scope_id is hierarchical containment, not tenancy.
The one case where a separate graph is right is a shared or team graph with genuinely different ownership and access. That is real federation, and it is the condition under which the parked peer-to-peer reconciliation work (Semantic Gossiping, cycle consistency — covered in the internal research notes) becomes relevant again.
A worked example
A Reflect note titled Iris Calder with Type: #person,
Email: iris.calder@example.com, Company: Westfield:
mecha-graph ingest reflectstreams it from the export zip → episode (reflect.note, keyed by the note's stable id), raw markdown archived, zip deleted after verification.mecha-graph reflect-processsees the type tag → resolves the email identifier → attaches to the existing Iris node rather than creating a duplicate;Company: Westfieldbecomes a fact (works_at, extractorreflect, pointing at this episode); the episode gets a mention of Iris.mecha-graph linkscans every episode for known names → more mentions (its candidate-staging tiers — kNN, structural, rules — run only with--propose, off in the nightly by default); NPMI notices who co-occurs with Iris unusually often →related_tofacts; kNN and Adamic-Adar stage speculative candidates for review.- A query for "iris westfield" detects the entity, collapses to episodes mentioning her, ranks, and returns a pack whose items each carry the uid you'd need to trace any claim back to this note.