Skip to main content

Changelog

All notable changes to this project will be documented in this file.

The format is based on Keep a Changelog, and this project adheres to Semantic Versioning.

Unreleased​

Added​

  • mecha-graph redact --source <S> --source-id <ID>, and redaction that reaches every table. Redact by provenance as well as by uid — how mecha purges a deleted chat (--source agent:mecha --source-id <session id>); no match is success with redacted: 0, and --json reports what went. A redaction now also takes the vectors of the facts and candidates it deletes, their co-occurrence alarms, the telemetry naming them (retrieval_touch, event_log — a correction's payload carries its text), the TUI's undo snapshots of the item (by uid and by source id), the person_interaction pointer to it, and the generated summaries that could quote it; its sightings of other facts are deleted and those facts re-derived, where the old SET NULL turned them into anonymous support. It runs as one savepoint, sets secure_delete, and merges the FTS indexes — an FTS5 delete leaves the tokens in the old segment. --vacuum checkpoints the WAL and rewrites the file. Nodes left with no mention and no fact are listed, not deleted.

  • Extraction follows a router's loaded model, and holds it one episode at a time. Every request names the model the router has loaded now (a request mid-swap waits up to 15 s for it to settle). Each episode is recorded under the model that extracted it. With [llm] holds_dir set (to ~/.mecha/holds, beside mecha), each episode takes a hold a model switch waits on. Before this, a night's extraction resolved one model at startup, named it for hours, and on 2026-09-28 undid the owner's switch to another by retrying a request the switch had cut off.

  • mecha-graph extract --charged, and why each episode was charged (extract_state.failure, migration V026). An episode charged as its own failure — marked attempted so one bad input cannot wedge every night — now records the reason, and --charged lists them (uid, date, reason), running no model; extract --episode <id> re-runs one. The nightly's ALERTS line counts tonight's charges and points here. Every write since records whether it had a reason (extract_state.reason_recorded, V027), so marks that did not — written before V026, or copied from a store older than V027 — are counted as unknown, never listed.

  • The TUI can close tasks through mecha, when you opt in. With [board] close_through = "mecha" (or a path to it) in ~/.mecha-graph/config.toml, closing a task (d, x) or reopening one runs mecha tasks set … --surface graph-tui, so mecha records and appraises the move. The route never falls back: a missing program, an unreadable config, a TUI on any database but the default, or done ↔ dropped (a verdict change mecha has no record for) refuses with nothing changed, and the row is read back after mecha answers, however it answered. Without the key nothing is different — the direct write, as before.

Fixed​

  • The embedding probe waits out a cold start. Embedder::available() gave up on /health after 1.5 s; the embedding server now runs on demand behind a systemd socket (mecha's scripts/llama/), and a cold start takes ~4 s, so a sleeping server read as absent and semantic search, kg_search and embed fell back to keyword-only without a word. The probe now waits up to 20 s (embed::AVAILABLE_TIMEOUT); only a refused connection or a 404 is a fast "no" — a 503 (loading), a reset, a timeout or a 5xx is something there, polled to the same deadline, as llm.rs's health() keeps them apart. Search no longer re-probes per arm: the caller's gate is the probe, so one query waits at most once, and an embedder that dies mid-run is an error rather than a silently empty vector arm. In the TUI that error is a status line over the keyword results — a sleeping server never closes the session. The TUI no longer probes at launch (it would add the cold start to every launch and wake the model for nothing); it probes on each Ctrl-E, and says "semantic search unavailable" rather than labelling keyword results "semantic". Permanent errors — a malformed URL, an unknown scheme — are a fast "no"; an unresolvable host is polled, since a resolver blip is transient. A query that will embed nothing — an AGGREGATE over facts, or a tag-only query — no longer probes (kg_search, mecha-graph query, and the gold eval, so the guard scores the path production runs), so it never starts the model; a LOOKUP still probes, because it can fall through to recall. Embedder::health_within keeps "nothing here" apart from "there and failing", so mecha-graph embed no longer advises starting a second server over one that holds the port, and the TUI's group view and semantic search both probe on a 2 s budget and show the reason when the server is not ready. kg_search, served one request at a time, probes on an 8 s budget and believes a "not ready" for 30 s, so a server mid-load costs the budget at most once per window instead of on every call. Still quiet: kg_search and mecha-graph query return a keyword-only pack with no flag when the probe says no — the pack's flags channel does not carry it yet.

  • LLM calls against a llama-server router (mecha's :8080 from 2026-09-27): served_model read the router's placeholder /props alias (llama-server) as the served model and sent it on every request, which the router refuses — the 2026-09-27 nightly lost 100 extractions (marked attempted, so not retried) and 30 summaries. On a router the served model is now the one resident in /models, and a configured fallback the router does not list makes connect refuse before the first request, so a batch is never burned.

  • extract no longer marks an episode attempted for the server's failure. The poison-episode mark exists so one bad episode cannot wedge every night, and it was applied to every failure — so a refusing or absent server aged out a whole batch. Who a failure belongs to is now settled by asking: after a failed episode, ChatClient::canary sends the server the same request — same model, system prompt and schema — with an empty input. If that fails too, the run stops with an error and marks nothing, every episode staying pending; if it answers, the failure was the episode's and it is marked. No answer at all (a timeout, or a 5xx or dropped connection that outlasts ~50 s of retries — a router answers 503 while it loads a model) gets one more try first, because a server that recovered looks the same as a failing episode. A server still loading at connect is waited for rather than refused or spawned over. summarize asks the same question and stops with an error instead of waiting out every node against a hung server. extract --episode is settled the same way. connect also refuses when a server that passed its health check does not answer /props (a probe now waits 10 s, not 1.5).

  • A multiword denylist term split across a line break is caught. grep reads one line at a time, and prose here is hard-wrapped at ~75 columns in docs, comments and commit messages, so a two-word term with its words on either side of a wrap passed both gates as clean. Each line is now also joined to the next — indentation and a leading comment marker stripped, whitespace collapsed — and multiword terms are looked for across the join, in the tree, in what each commit adds, and in commit messages — for any term with a space in it, of either kind. A paragraph break, a hunk or file boundary in a commit's additions, and path and ref-name lists are not joined. A file the check cannot read refuses rather than being skipped. (A term split across two breaks is still out of reach.)

  • The denylist gates can no longer pass a term they did not check. Both the pre-push hook and CI read a roster line with no trailing newline, strip CRLF, trim stray spaces, and refuse a line with a kind but no term — each of those used to shrink the check silently while it printed clean. Terms are fixed strings for both kinds, so a term holding [ is no longer a broken regex whose error read as "no match" (verified: the old hook pushed a leak past such a term). Any grep error refuses. The hook now checks every commit being pushed that the push target lacks — its own tracking refs, never a second remote's — and those commits' messages, not the checked-out tree, so pushing another branch, a leak fixed in a later commit, a name in a commit or annotated-tag message, a file, directory, branch or tag named after someone, or a short name inside snake_case (grep's -w counts _ as a word character; a w term is now bounded by any non-alphanumeric) is caught. Whole trees are checked at the pushed tips and each commit for what it adds, so a change that removes an inherited leak can pass. Run by hand (no refs from git, whatever stdin is) it checks the whole working tree from the top, untracked files included. CI keeps grep's exit status under Actions' bash -e (a missing term no longer ends the step), checks every commit a pull request or push brings in — what each adds, its message and its added paths, with full history, plus the branch name — not only the tip, and refuses (rather than narrowing to the tip) when it cannot tell what a push brought in; it prints counts and commit ids, never a term or a path (a path can carry the name); it excludes nothing but LICENSE (not .githooks, not itself), writes the roster to a private temp file removed on every exit, and its failure line says the naming tool is owner-only.

Changed​

  • nightly.sh resolves its precheck toggles before it sources nightly.env, not after. The live source could reach the helper's "clean" environment in that half only (a HOME= line), and could overwrite the resolved values; the helper also fixes HOME and its PATH when it is loaded. Its comment now says what is true: the fixed PATH is the union of what the two nightlies prepend, not what both do.

  • The two nightlies take their precheck toggles from one resolver. nightly-env.sh already gave both halves one verdict on nightly.env; now nightly_env_toggle gives them one value too. nightly.sh had read the toggles from its live source, so a line conditional on the caller's environment ([ "$USER" = lab ] && PRECHECK_TRIAGE=0) could switch triage off at 03:30 and leave it on at 08:00. Both halves now read the file in the same clean environment — including one fixed PATH that still leads with the ~/.local/bin and ~/.cargo/bin both nightlies prepend, and the syntax check runs there too, so it and the source probe are one interpreter — falling back to the process value, then to 1.

  • Both nightlies run precheck --triage. scripts/nightly.sh and scripts/nightly-mecha.sh pass the flag by default; PRECHECK_TRIAGE=0 in ~/.mecha-graph/nightly.env turns it off, beside PRECHECK_AUTO_ACCEPT. The 08:00 half never read that file, so it now takes both precheck toggles from it (and no longer hardcodes --auto-accept) — an off-switch that reached one nightly would have left the other making permanent rejects. Both halves judge that file through one helper, scripts/nightly-env.sh, so they cannot reach opposite answers: a file that is unreadable (a dangling symlink included), does not parse, or aborts when sourced — probed in one clean environment, so the two callers' own variables cannot split the verdict — is skipped and fails closed — both precheck toggles off, with an ALERT in each log — while a valid file whose last line merely returns false is fine. The 08:00 precheck also gains the blindness alarm the 03:30 one has, and both now also alarm when the embedding server was down before precheck started: with --triage on, "folded 0" beside real one-off rejects is blindness, not a clean queue. Each alarm reads only its own run's precheck output, so a same-day re-run after the fix is judged on its own. The recurrence pool is read once per sweep and shared by minting and the triage tally — it now holds every one-off reject permanently, and was being scanned twice — and the fold lane's stored reject prefix is pinned by a test, as the one-off lane's already was, since fact::class_prior matches it on disk.

Added​

  • precheck --triage: two review-queue lanes set from the owner's own verdicts, off by default. A calibration replayed 4,362 human-decided llm candidates against the graph as it stood when each was proposed. A claim about a subject the graph knows nothing about, named by only one or two candidates, was accepted 10% of the time (521 items); a resolved subject, 76%. The one-off-subject lane rejects those (llm only; never the owner however spelled; never a NEVER_AUTO social claim). Three mentions is the minting bar, and minting and the lane count ONE pool — pending claims plus the lane's own earlier rejects — so a name arriving once a night is rejected twice and minted on the third night, rather than spared by one count and starved by the other. A lane reject counts only as a distinct claim, so one sentence re-extracted three times is still one claim and mints nothing; a name minting could never mint (too short, more than five words) is left for a human. The fold lane handles the calibration's surprise: at same-subject cosine ≥ 0.90 the owner accepted more, not less (81%), and each accept minted a second copy of a fact already held — a verdict of "true" is not a verdict of "new". A restatement (0.90–0.97, same predicate, the same normalized text object — a same-node restatement is tier 1's exact dup already — no longer than the fact, not a social-standing predicate, and carrying no negation marker — "no longer uses the kiln" embeds beside "uses the kiln daily" and is the retraction, not a confirmation) is now recorded as an observation of the existing fact instead, by the corroboration rule (the counter moves once per new non-agent episode, so a re-extraction of the same episode does not inflate it; sensitivity rises to the evidence's MAX; confidence is recomputed), decided after the contradiction tier so a fold never swallows a conflict. --list prints every decision tab-separated for a spot-check. Neither lane's rejects count toward fact::class_prior: they are selections, not verdicts on the class, and ~1,900 a sweep would pull every future llm fact's prior toward the lane's own rate. The first dry run on the live queue caught the object guard's gap before anything was written: free-text objects compared as "no object" on both sides, which would have folded "uses tool X" into "uses tool Y". On the live queue: 1,922 one-off rejects, 11 folds, of 4,439.

  • A task row carries its parent project's node id beside the name. kg_task_list, kg_entity's task block and the two task echoes gain project_id, absent exactly when project is — the create echo now carries the whole row under task through the same renderer as the update echo, beside the top-level keys callers already read. And kg_task_create accepts that id back as project: a pointer the server hands out is one it accepts, so a consumer filing a task under the project it just read is not refused for citing what it was told to cite. With the id path, the name path tightened to match: a name matching several nodes is refused with their ids rather than resolved by access count (validate_about_target's rule — the guess now minted a citable pointer), and a task, person, agent, place, event, event_series, document or artifact is never a parent, whichever way it was named. A parent can now be corrected — kg_task_update takes project ("" clears), resolved exactly as capture resolves it — because a pointer another repo cites must have a correction path that keeps the task's id; and mecha-graph repair-parents surveys the rows written before the guard (a task filed under a person, say), detaching them only with --apply — stats and the nightly's alert line count the slips alone, since the plausible ones are kept, and task-project <task> <its parent's id> says a plausible filing was meant (the vouch names that parent, lapses if the parent changes or is retyped or its row is deleted — a vouch reads as standing only where the survey honours it, and a declined vouch takes a stale mark with it — and is never inferred from an MCP update that echoes the row it read) and takes it off the survey — and mecha-graph task-project re-files one from the direct interface (the survey's JSON echoes both of its flags, applied and include_plausible); it also finds a parent whose node is gone (reported as missing), and a task on the board cannot be retyped into anything, nor merged into a container — only the type would move, or the row would land on a project, and the task would be a legal parent the survey could not see — except that a finished task with nothing under it converts: retype removes its task row with the type, keeping the id, the facts and the associations a drop-and-recreate would lose, and keeps the whole row it removes on the node (properties.converted_task, every column, plus a detachment record for the parent it was filed under). A parent name is resolved over every container that matches it — exact name or alias, then substring, with no limit — so a second container cannot hide outside a window and task titles cannot fill one. project_id comes from the same join as project, so a parent whose node row is gone is handed out by neither. A parent argument that is a node id resolves to that node before any name, so the id the board hands out always comes back to the same node. Everywhere the rule is applied, the row is the fact: a node that is on the board is a task whatever its type says, and is never a parent — to the resolver, the id writer, the merge, or the survey. mecha-graph task-project <task> with no parent prints the current one and every detachment the store recorded; the survey's text output says which filings are plausible; --apply runs as one transaction, recording a detachment only where it landed, and keeps the plausible filings unless --include-plausible; and the survey also lists every task detached earlier — by an apply or a merge; a converted node has no task row and is read by task-project — and not re-filed since, so a merge's silent detach stays reviewable as a set; the JSON report echoes the flag it was computed under and says per row whether the apply it previews would detach it; task-project <task> "" on a task already under nothing marks its record reviewed ("no project is right"), so the list is not a standing pile, and a finished task is not on it; a deliberate re-file or clear records the parent it leaves, reviewed; a task node keeps its twenty newest records. On kg_task_update, a list or number where a string belongs refuses the call rather than skipping the field and answering updated, and on kg_task_create a non-string due or context refuses the create. A filing under a single event is plausible like one under a series. task-project <task> reads a converted node's record too. On kg_task_update the parent is resolved before anything is written, and so are the dates and waiting_on, so a refused one changes nothing — not a status that already landed. A name shared with a node that could never be a parent is not ambiguous: the container of that name is the parent. A refused captured_from refuses kg_task_create before the insert, where it used to leave a task the error said was never created. The rule binds every writer of a parent, not only the resolver: the id writer re-checks what it is handed, retype refuses to turn a parent into a non-container while tasks sit under it, merge detaches rather than re-points onto one, and a detached task keeps where it was on its own node (appended to properties.detached_parents) so the record is in the store rather than a terminal. The survey marks a filing under a place or a recurring event plausible — legal under the old rule — apart from a slip under a person, the agent or another task. NEVER_A_PARENT and CONTAINER_TYPES partition the closed type set, held by a test. The name is prose (spaces, and two nodes can share one); a consumer recording which project a task served cites project:<node id> and refuses whitespace in an id, so the name alone could never be cited. mecha's goal record is that consumer.

0.1.5 - 2026-09-06​

Added​

  • A task's association with an entity now outlives the task. about (task → person/project/topic) carries what a task concerns; waiting_on keeps its existing job of naming who holds the ball. One predicate could not be both, and trying made each surface wrong in a different direction: close the fact on completion and the finished work disappears from the person it was for; leave it live and their card claims they owe something they handed back months ago. facts_for_node is bidirectional, so it was their card carrying it, not the task's.

    Two read surfaces, because they answer different questions: kg_task_list and mecha-graph tasks gain entity (the precise query — unions about, waiting_on and assigned_to, plus tasks whose parent project is that node), and kg_entity gains a tasks block split into open and closed, capped at 15 a side with the total and a truncated flag. An unknown entity name is an error on all of them: "no tasks for her" and "there is nobody here by that name" are opposite findings.

    No migration — about and assigned_to were already in the seeded vocabulary with inverses. Nothing had ever written them for tasks.

  • mecha-graph scan-tasks proposes associations by scanning task titles for entities the graph already knows, landing them at tier shadow: "this title contains a word that is also a name" is an inference, and inference is served rather than asserted. Strong matches only — a bare first name has no corroboration to draw on in a one-line title — and task nodes are excluded as targets, or every task files itself under every task sharing a word. Rejection memory keys on whether the pair was ever asserted in any state, so a refuted association is not re-minted nightly. Dry by default (--apply writes, --limit bounds a pass), because a command whose output a human is meant to judge should not have written everything before it prints the count.

  • mecha-graph repair-dates finds date columns holding text that is not a date, and clears them with --apply. Reports by default.

Changed (beyond the task board)​

  • Accepting a candidate now honours subject_node/object_node. resolve_candidate_parts re-resolved both endpoints from the display strings and never read the ids, so producers that derived a pair from nodes — linkers and rules have been setting these fields all along — had that thrown away at accept time, and two same-named entities collapsed onto whichever the lookup returned first. An explicit id now wins, falling back to the name when the id no longer resolves (a merge deletes the losing row, so the name is still reachable). This changes what accepting an already-queued kNN or rule candidate resolves to.

Changed​

  • Closing a task now closes its waiting_on, in valid time — the obligation ended, it was not wrong. This changes what an existing call does to existing rows: before, the claim stayed live forever and every later read of that person carried it. The task stays findable under them, because the entity filter reads fact history rather than only what is live. Reopening deliberately does not resurrect the claim; who owes a reopened task is a new question, and guessing the old answer silently re-obligates someone.

Fixed​

  • kg_upsert now refuses a valid_from that is not a date. It wrote the string verbatim, and at confidence >= 0.9 auto-accepts, so prose reached fact.valid_from with no human in between — the same defect as the commitment path below, on the higher-volume route. Shipping repair-dates without closing this would have made the repair a treadmill: idempotent in its own test, dirty again by morning.

  • A due date is no longer stored as a valid time. accept_commitment passed the commitment's when to three columns: task_detail.due_at, where it belongs, and the valid_from of both facts it asserts, where it does not. when says when the work is owed; valid_from says when the belief became true in the world, and "X is waiting on Nadia" became true when the commitment was made — the episode's occurred_at, which the same file already uses correctly for ordinary extracted facts.

    Not cosmetic: facts_as_of filters valid_from <= as_of, so a commitment accepted in September and due in December was a belief no as-of query would answer until December — invisible for exactly the months somebody might have wanted reminding of it. A past deadline back-dated the belief to before anyone held it. repair-dates now finds and corrects rows already written that way, rewriting valid_from to the episode's time rather than nulling it, and reports before it writes.

  • A model's when is parsed before it is stored. accept_commitment wrote the extractor's raw string into three date columns — task_detail.due_at and the valid_from of both facts it asserts. A model answering the literal string "null" put that in all three, where it sorts as a date (lexically after any real one), so the task never read as overdue and it answered the wrong side of every bi-temporal --as-of query. Nothing on any surface renders valid_from, which is how it stayed invisible. Unparseable now degrades to None rather than failing the accept; the candidate payload keeps the raw value either way.

0.1.4 - 2026-08-31​

Large for a patch, and numbered one anyway to stay in step with this project's 0.1.x cadence: three schema migrations (V021 fact_tier, V022 vec_rejected, V023 candidate_embedding) and the review model's phase 2.

Added​

  • The review queue's vectors persist between runs (V023 candidate_embedding). Grouping the pending queue by similarity re-embedded every pending statement on every call — ~7,000 of them, ~40s measured — and the vectors went out with the process, so the same statements were embedded again for the next threshold the stepper visited, for the class listing after the global one, and for the TUI after the phone. A pending statement's text does not change while it waits, so its vector is immutable and re-deriving it is waste.

    A plain table rather than a vec0 virtual one, unlike its three siblings: those exist to be searched and pay an index for it, while nothing searches this — the grouping fetches vectors for a known set of ids and clusters them in process. text_hash covers the model, the embed task's instruction and the exact text, so a model swap, an instruction change or an edited statement all invalidate by construction rather than by anyone remembering an invalidation rule. Stored as little-endian f32 rather than the JSON the vec0 tables are fed: ~20MB across this queue instead of ~60MB. Falls through to a plain embed on any storage trouble — the cache is derivable and this is a read path, so a store that cannot be read must give a slow grouping, never a failed one.

    Measured on a copy of a live store: cross-class grouping 41.0s → 4.3s, a threshold step 36.9s → 4.0s, and cold and warm output byte-identical.

  • Review-on-use (phase 2). Extraction output goes live unreviewed as a shadow fact — retrievable, rank-discounted, labeled — instead of queueing for review at birth, and earns a human verdict when it is about to matter. The loop closes: retrieval feeds the ladder and gates the extractor. Rejection memory survives a paraphrase (V022), so the same wrong claim re-extracted in different words no longer re-claims the owner's attention.

  • The entity arc. relink-aliases judges the mentions already on file, the owner can file a merge proposal so every merge leaves a record, and kg_entity surfaces the identifiers rather than only the aliases. kg_upsert kind=alias learns the other direction.

  • kg_notes — the notebook view, handing back the key that can write to a note.

  • Task provenance: a task remembers what asked for it and the conversation that worked it, @owner so a harness need not know your name, and the board can say who holds a task rather than only that it waits.

Fixed​

  • Every group verdict re-grouped the class twice, on the UI thread. reload_review re-groups on its own when group_view is set, and the accept and reject keys are only reachable while it is — so the explicit reload_groups preceding it ran the class's grouping a second time after every verdict. Both passes embedded the class before V023 and both read the cache after it; either way the terminal was frozen for two where one was needed.

  • A test fixture wore a real name, in a repo whose export gate exists for exactly that. The fixtures move to the fictional cast — tracked source is one export away from public.

  • The queue's depth is the count before the page cut, so a listing that shows less than the queue says how much less.

  • Precheck could go blind without saying so, and commitments were exempt from it.

0.1.3 - 2026-08-22​

Fixed​

  • A class's displayed accept rate counted this pipeline's own rejections as the owner's. precheck::review_clusters summed every status='rejected' row, including the dedup and ephemeral rejects precheck writes itself — in the one view a person reads immediately before verdicting a whole class. Measured on a live store: llm/has displayed 18% against a true 67% over 48 human verdicts, llm/has_role 7% against 53%, llm/attended 39% against 81%, and three classes displayed a 0% accept rate on which no human had ever voted at all. ladder::human_record had carried the correct filter (reject_reason NOT LIKE 'precheck:%') since it was written; the cluster view never did. Machine rejects are now reported beside the rate as machine_rejected and never inside it — a class that mostly repeats itself is a different problem from one that is mostly wrong.

  • review --json no longer prints prose ahead of the array. An empty result emitted no pending candidates and then [], so the whole of stdout failed to parse and a caller asking for an empty set got a JSON error instead of [].

Added​

  • review --proposers — the queue rolled up by proposing mechanism, with each one's human accept rate and Wilson lower bound, and p in the TUI for the same view. A proposer spreads across many predicates (the extractor alone holds ~90), so its own hit rate is invisible in a list of 733 (proposer, predicate) rows — and mechanisms are what get switched on, tuned, and switched off. An unjudged mechanism shows a dash, never 0%: "never reviewed" and "always rejected" are opposite findings.

  • review --sample N [--seed S] — a uniform random draw from what --proposer / --predicate left. The queue is ordered, every order it could have is correlated with something, and judging the first N then reading the result as a class's accept rate measures the ordering. The seed is printed when not supplied, because a sample nobody can redraw is a sample nobody can check. Partial Fisher–Yates over a four-line splitmix64, with a uniformity test over 4,000 draws that fails on truncate(k).

  • review --proposer / --predicate filter the item view, matching on precheck::cluster_key so a drill-down can never show a different set than the cluster row it came from.

0.1.2 - 2026-08-22​

Fixed​

  • The binary stopped introducing itself as pkg. #[command(name = "pkg")] survived the 0.1.0 rename, so --version printed pkg 0.1.1 and --help read Usage: pkg — on a crate whose front door is cargo install mecha-graph. Three of the five sites were worse than cosmetic because they told the reader to run something: a usage hint after a DB move, the review nudge in render.rs (→ pkg review), and mecha-graph-mcp's open-failure prefix. All five renamed; PKG_* env vars and the ~/pkg data dir are untouched, being a migration rather than a rename.

0.1.1 - 2026-08-21​

Changed​

  • One engine, one model, chosen by measurement. ollama is gone; embeddings are served by llama-server on :8081, beside the chat model, and the embedder is harrier-oss-v1-0.6b. Four candidates were scored on 80 semantic queries against a 5,000-episode pool: harrier took MRR 0.3595 and recall@10 0.562 against the nomic incumbent's 0.2708 / 0.4375. Re-embedding the live store takes about nine minutes and 175 MB.

    MTEB did not predict this. Qwen3-Embedding-0.6B scores 70.70 on MTEB English v2 against nomic's ~62 and tied it here; harrier beat Qwen3 by 26% at identical size, which also kills the bias flagged before the run — the queries were generated by a Qwen model, and the Qwen embedder still lost. Qwen3-4B's 0.028 MRR edge over harrier is inside the standard error at n=80 with identical recall@10, so it is not separable and not worth 6.7× the size.

  • SEMANTIC_DUP_THRESHOLD 0.93 → 0.97, flag threshold unchanged at 0.83. A similarity threshold encodes one model's cosine scale and nothing about these said so: nomic's range is compressed, putting unrelated text at 0.56 where harrier puts it at 0.28, so carrying 0.93 across would have silently stopped matching anything. Recalibrated by matching the operating point — sweeping dedupe-facts over the same corpus against a pre-migration backup and the re-embedded store — which preserves the rate, not the accuracy. precheck.rs carries the calibration table at the constants it explains.

  • The README meets the stranger it now has. Install from crates.io as the front door, the synthetic eval as the try-it-with-no-data path, wiring blocks for mecha (unprefixed, marked untrusted — the interlock reasoning stated), Claude Code, and any MCP client, and the privacy story gathered into one section: encrypted at rest, stream-first ingestion, the sensitivity ladder, true delete, local by construction.

  • The self-improvement plan is a design document. ./self-improvement now states the three doctrine decisions (autonomy by exception, graph-first with session-splitting, the mechanical error contract), the five gossip roles, the mechanism catalog with build waves, and every settled design point with its rationale — in an impersonal voice.

Added​

  • embed_meta (migration 16) records which model produced the live vectors. Nothing about a vector reveals what made it, and a 768-dim nomic vector is indistinguishable from a truncated 768-dim Qwen one.
  • docs/EMBEDDING-RESEARCH.md — the candidates, the two measurement attempts that measured nothing and why, the results, and the failure analysis. Read it before changing an embedder or a threshold.

Fixed​

  • --pooling last for decoder-only embedders. mean does not error; it silently produces plausible, worse vectors. The serving flags and the numbers behind them live in ~/.local/bin/mecha-embed-server.

0.1.0 - 2026-08-16​

The first public release: a clean-room extraction of a private research repository. History starts here on purpose — the development record was a journal of the data the project exists to hold, and no filter makes a journal safe.

Added​

  • Three crates: mecha-graph-core (ingest, enrich, resolve and link, retrieve — knows nothing about any agent), mecha-graph (the CLI), and mecha-graph-mcp (a stdio MCP server any client can sit on), published to crates.io.
  • Eleven MCP tools: kg_search, kg_entity, kg_timeline, kg_related, kg_upsert, the verification family (kg_verify, kg_pending, kg_verdict), and a small task family. The tools carry their own kg_ namespace, so harnesses that support unprefixed registration can drop the server prefix.
  • A synthetic eval world: eval/synthetic/run.sh builds a throwaway graph from a fictional corpus (a twelve-message mailbox, a small calendar) and grades 24 retrieval queries — no personal data, no live store, works on a fresh clone.
  • Encrypted-at-rest store at ~/.mecha-graph/graph.db (SQLCipher; the raw key beside it, mode 0600), MECHA_GRAPH_* environment overrides, stream-first ingestion, capture_delete retention for file sources, a four-level sensitivity ladder with private-by-default retrieval exclusion, and true delete (redact purges an episode and everything derived from it).