Skip to main content

Mail and calendar

mecha-mail is a third crate: a library plus four thin MCP binaries.

The library holds Gmail and Google Calendar v3, Outlook mail and calendar over Microsoft Graph, both OAuth flows, and the token lifecycle. It is what a GUI would depend on directly.

BinaryServes
mecha-googleone Google account, its own credential store
mecha-outlookone Microsoft account, its own credential store
mecha-mailevery account in ~/.mecha/mail/ behind one provider-neutral surface
mecha-docsGoogle Docs — see Documents

mecha-mail is the one deployments should wire. Behind it, no mecha-core or mecha-cli code knows that Google or Microsoft exists — and neither does the model.

[[mcp]]
name = "mail"
command = "~/.cargo/bin/mecha-mail"
# No zone here: mecha hands every server [agent] timezone as MECHA_TZ.

[mcp.capabilities]
untrusted_input = true

[outbox]
tools = [
"mail__mail_send",
"mail__mail_reply",
"mail__calendar_create_event",
"mail__calendar_update_event",
"mail__calendar_delete_event",
]

The model names an account, never a provider​

accounts.toml in ~/.mecha/mail/ maps short names to providers:

default = "campus"
# default_mail = "campus" # optional: where new mail goes from
# default_calendar = "personal" # optional: where new events go

[[account]]
name = "personal"
provider = "google"
# grant_lifetime_days = 7 # optional: see below

[[account]]
name = "campus"
provider = "outlook"

default_mail and default_calendar override default for one surface each, because where your mail goes out from and which calendar your life is on are separate decisions. Either one left out falls back to default.

grant_lifetime_days declares how long this account's sign-in lasts when it is known to expire on a schedule — a Google OAuth client left in Testing status issues refresh tokens that die after 7 days. mecha doctor uses it to warn you before the grant expires instead of after. mecha never guesses it; leave it out and nothing warns.

Set $MECHA_MAIL_DIR to keep the registry somewhere other than ~/.mecha/mail/.

mecha-mail auth campus --provider outlook --tenant <tenant-id>
mecha-mail auth personal --provider google
mecha-mail import personal --provider google # copy a mecha-google / mecha-outlook login in
mecha-mail accounts # names, providers, addresses, the default
mecha-mail default campus # set the standing default
mecha-mail default personal --calendar # …or one surface's (--mail / --calendar)
mecha-mail serve # the MCP server (also the no-subcommand default)

auth takes the OAuth client configuration where it needs it: --client-id (the Google Desktop app id or the Entra application id, also read from GMAIL_CLIENT_ID / OUTLOOK_CLIENT_ID), --client-secret for Google's desktop pseudo-secret, --tenant for the Entra directory, and --port for the Google loopback redirect. A second mailbox on the same app registration needs none of them: with no --client-id, the client configuration is taken from this account's own stored login, or failing that from a configured sibling of the same provider. Adding your second Gmail account is one command with two arguments.

Credentials live at ~/.mecha/mail/<name>/oauth.json, one store per account. Names are lowercase letters, digits, - and _; duplicates and a default (general or per-surface) that names no configured account are refused at load.

The account names are baked into every tool schema as an enum at startup. accounts.toml is read once, and every tool's account property is emitted as {"type": "string", "enum": ["campus", "personal"]}, plus a sentence naming the default where one exists. The model picks from real names instead of guessing at them, and the schemas are built once rather than per request.

A configured account whose credentials will not load fails startup with the command that fixes it, rather than being quietly skipped:

account `campus`: run `mecha-mail auth campus --provider outlook`

Resolution: the rule that shapes the surface​

Thirteen tools — mail_search, mail_recent, mail_get_thread, mail_send, mail_reply, mail_triage, calendar_list, calendar_list_events, calendar_freebusy, calendar_create_event, calendar_hold, calendar_update_event, calendar_delete_event — and three resolution modes.

calendar_hold is the one write that is not a send. It blocks time on your own primary calendar, invites nobody, and marks the event private, so someone you share the calendar with at a reader level sees "busy" and not the title. Private is not secret: a person who can make changes to your calendar, or an Outlook delegate you allowed to see private items, sees the details. It has no attendees or calendar_id field, and the server ignores both if a model sends them anyway. Because it reaches nobody it does not need the outbox: leave it out of [outbox] tools and allow it by rule, and a reminder or a focus block lands without a draft to release. To invite anyone, the model uses calendar_create_event, which still stages.

[[rule]]
tool = "mail__calendar_hold"
decision = "allow"
# An `allow` must carry an example; the call has no command to match, so
# any plain word proves the rule loads.
match = ["hold"]

calendar_freebusy is the scheduling one: busy intervals merged across every account, with no event details in them. "When am I free on Thursday?" is answered from it; calendar_list_events is for when the events themselves matter.

Both take a window as RFC 3339 timestamps or as now, today, tomorrow, yesterday, +3d, -1d — and the model is told to prefer those. The server resolves them against its own clock in your [agent] timezone, so a model that was told the wrong date still gets today's calendar; time_min: today with time_max: today is the whole of today. Every answer states the window it covered and the clock it was resolved against, so a wrong premise is contradicted by the result rather than confirmed by it. With no [agent] timezone the relative terms are refused by name instead of resolved against the machine's clock, which on a server is usually UTC and would be the wrong day every evening. this week is not accepted: whether a week starts on Sunday or Monday is a convention the server cannot know.

Reads fan out​

No account on a search, a recents listing, a calendar list, or a calendar window means every mailbox, queried concurrently and merged. Mail rows are sorted newest first; calendar events are sorted by start time. Every row carries the account it came from:

[
{"account": "campus", "thread_id": "AAQk...", "message_id": "AAMk...",
"from": "Priya Nair <priya@example.edu>", "subject": "Retreat agenda",
"date": "2026-08-04T09:12:00Z", "snippet": "…", "unread": true,
"has_attachments": false}
]

That tagging is what makes the next rule workable: the model always already has the account by the time it needs to name one.

Item operations: threads are found, events are named​

Thread and event ids are account-scoped, and the two are handled differently.

A thread is looked up. mail_get_thread, mail_reply and mail_triage with no account ask every account for the thread and act in the one that holds it — a reply goes out from the mailbox the thread arrived in, so naming it only repeats what the id already says. A thread found in no account, or in more than one, is refused with every account's answer.

An event must be named. calendar_update_event, calendar_delete_event, and calendar_list_events with a calendar_id other than primary require account when more than one is configured, and say where to find it:

several accounts are configured (campus, personal) and this id is
account-scoped — pass `account` (every search and list row carries it)

A single-account install never needs to name anything: with one account, every mode resolves to it.

Creates use the default, or ask​

mail_send falls back to default_mail, and calendar_create_event and calendar_hold to default_calendar; either falls back to default. With several accounts and no default, the error says to ask the user:

several accounts are configured (campus, personal) and no default is set —
ask the user which account to use, then pass it as `account`.
(They can set a standing default with `mecha-mail default <name>`, or one for
this surface alone with `mecha-mail default <name> --mail`.)

(--calendar in place of --mail for a calendar create.)

The wording is deliberate and there is a test pinning it. "Ask the user" rather than "use your best judgment": the second phrasing was measured to make models invent an answer instead of stopping. The same finding shaped ask_user's decline wording elsewhere in mecha.

A failed account never sinks a fan-out​

Failures are collected separately from successes. If at least one account answered, the results are returned with the failures appended as a note:

note — some accounts could not be read:
account `personal`: request timed out

The call reports an error only when every account failed. One expired refresh token does not cost you the other mailbox.

Commands with no model in them​

mecha-mail also serves the scheduling pipeline directly, as data, on a timer:

mecha-mail freebusy --days 60 --json # merged busy intervals across every account
mecha-mail bookings --dry-run # what drained bookings would become events
mecha-mail polls --dry-run # what meeting polls are owed

freebusy deliberately inverts the rule above: it fails when any account is unreadable. The MCP surface answers a person who can see the note about which mailbox was skipped; this one feeds a public booking page. A mailbox that could not be read is not a mailbox with free time, and a slot list built from a partial answer offers strangers hours the user does not have. The one exception is a login that has been revoked: no retry will ever fix it, and you have already been told (mecha doctor, exit code 77), so its calendar is skipped with a loud warning rather than halting your booking page for days. If every login is revoked, it fails. --from/--to name an explicit window instead of --days, and --account narrows to one.

bookings is the inbound sibling: it turns drained booking records into calendar events, deterministically, with no model anywhere. It is idempotent against ~/.mecha/mail/bookings.jsonl — a record already ledgered is skipped, so re-running after a partial failure picks up exactly where it stopped — and each event is re-verified against live free/busy before it is created, because the slot was sold from a cache and home holds the fresher truth. A collision is parked loudly for a human rather than double-booked. --account names the calendar that receives the events, defaulting to the default account; an absent request store is "nothing drained yet" rather than an error, because this runs on a timer that must not cry wolf.

polls does the mail-and-calendar half of a meeting poll: it mails each person their own link, sends the one nudge the sweep queued, and creates the event for a clean winner from your account with everyone invited. It decides nothing — that is factory-publish polls sweep, on the same timer — and it is idempotent against ~/.mecha/mail/polls.jsonl. --account names the account to send and book from when a poll names none.

The plain inbox, and writing a letter yourself​

The triage queue answers what needs me. It is deliberately not an inbox — it holds what the classifier decided you should look at, in the order it thinks you should look. Sometimes the question is just what arrived:

mecha mail recent # newest first, every account, merged
mecha mail recent --account work # or one

That reads through the same mail_recent tool the model uses, so there is one definition of "recent" and no second query to keep correct. On the web surface it is the Inbox tab beside the queue.

Writing a new letter goes the one way mail leaves this system:

mecha mail compose --to someone@example.edu \
--subject "Thursday" --body "Does 2pm still work?"

It stages into the outbox; it does not send. No model runs and no MCP server starts — the draft is your own words, staged verbatim under whatever tool name [outbox] tools routes sends to, so mecha outbox send releases it exactly as it releases one the model wrote. If mail_send is not routed there, composing refuses rather than staging a draft that no release path knows how to execute.

The point is not ceremony. It is that one queue holds everything outbound regardless of who wrote it, so there is a single place to look before anything leaves.

Capability labeling: reads are untrusted sources, not send sinks​

This is the part worth not re-litigating.

Reads carry readOnlyHint and deliberately not openWorldHint. A search query travels only to googleapis.com or graph.microsoft.com — hosts that already custody the mailbox. There is no payload channel to a third party. That is precisely the difference from http_fetch, whose query string can reach any host in the world, and it is why the read tools are not trifecta sinks.

But mail bodies are other people's words. Reading mail must arm the interlock, so config forces untrusted_input = true on the server — the same treatment the knowledge graph gets. That override only ever widens: config can distrust a server further than its own annotations, never less.

Sends and calendar writes that can reach someone do reach third parties — recipients, invitees — so they carry openWorldHint, and calendar_update_event / calendar_delete_event add destructiveHint. Those names go in [outbox] tools, so they stage rather than deliver. See the outbox.

A fourth quadrant is the private write: calendar_hold. It reaches nobody. It goes on your own primary calendar, invites no one, and is marked private. So it says openWorldHint: false outright, stays out of [outbox] tools, and sits with the approver, where an allow rule lets it run unasked. The same quadrant holds the documents surface's create verbs. The guard is a test that inspects every such tool's schema against an allowlist of content fields, because nothing else reviews a call that doesn't stage.

And there is one more quadrant, which is neither. mail_triage — archive, mark read or unread, report spam, trash — mutates your own mailbox and reaches nobody. It carries destructiveHint alone:

  • Not openWorldHint, so it must never appear in [outbox] tools. Staging it would make triage circular — you would review a queue in order to fill another queue.
  • Not readOnlyHint, or an unattended run under permission_mode = "read-only" could empty your inbox at seven in the morning.

The shared surface test takes a third slice that asserts exactly that pair of negatives, so neither mistake can ship quietly. Documents land docs_trash in the same quadrant, from the other direction.

A shared test runs against each provider's tool list and asserts all of it: every read is readOnlyHint and is not openWorldHint ("reaches only the provider that already custodies this data — not a send sink"), every write is openWorldHint and is not readOnlyHint, and every tool has an object schema and a description worth reading. A new provider cannot ship a mislabelled surface. Unification did not weaken this: the same annotations ride on the unified tools, and one send name in the outbox list now covers every account it could send from.

Triage: the queue over the mailbox​

mail_triage is the verb. mecha mail is the surface you actually use, and it exists because an inbox is not a thing you read once — it is a queue you work.

mecha mail classify --account campus # read recent mail, decide what each thread is
mecha mail list # what needs you, newest first
mecha mail list --aged # day two: what you meant to answer
mecha mail show <thread_id> # read one, in full
mecha mail reply <thread_id> # draft an answer — stages, never sends
mecha mail task <thread_id> # track it on the graph's board
mecha mail correct <thread_id> --bucket respond # the classifier got it wrong
mecha mail dismiss <thread_id> # drop it from the queue without acting

Threads are named by an eight-character handle — the last eight characters of the id, and any unique suffix is accepted wherever a thread id is. A suffix rather than a prefix because Outlook conversation ids share a 57-character common prefix: in a real 68-thread store every prefix handle collapsed to the same eight characters and identified nothing. Ambiguity is an error rather than a guess, because acting on the wrong thread is silent and, for mail_triage, irreversible.

classify writes one typed verdict per thread to ~/.mecha/mail-triage/: a bucket (respond / notify / ignore), an urgency, a proposed action, tags from a closed vocabulary, a deadline if the thread implies one, and the kind of standard request it is if it recognises one. On a fifty-thread sample of real academic mail, twenty-eight were archivable and twenty-two needed attention.

A deadline must come with the words it was taken from. The classifier copies the phrase from the message that sets the date, and the date is kept only if that phrase is actually in the subject or body it was shown — and is long enough to be a date and short enough not to smuggle anything after it. Otherwise the date is dropped, the verdict otherwise stands, and the reason is kept. mecha mail list and mecha mail show show a dropped date with why, and mecha mail task says the classifier's date was dropped rather than that it found none — two different situations, because a dropped date is one you may want to type in with --due. A wrong date makes a task due when nothing is; a missing one costs you one look at the thread.

The prefilter: half the mailbox never reaches a model​

prefilter disposes of a thread from its envelope alone, ahead of the classifier: a List-Unsubscribe header, or a sender address or display name that reads as a system. Measured on a year of real mail with exactly the shipped marker list, the two rules match a little under half of all threads, and five of the threads they caught had ever received a reply — about one in a thousand.

List-Unsubscribe alone is not enough, and that is the finding underneath the rule: it catches marketing, which is obliged to offer an unsubscribe, and misses every institutional and transactional sender, which is not. If you are tuning this, a sender-address rule is worth more than a better prompt.

Three properties keep it safe rather than merely cheap, each with a test named on it:

  • It only ever produces ignore. A deterministic rule may say "nothing here" and may never say "this needs a reply" — the cases it would have to get right to do that are exactly the ones that need judgement.
  • It reads the envelope and never the body, so it is not a second place a stranger's prose gets interpreted outside the classifier's quarantine. Markers written into a subject line do not fire it.
  • The sender list is portable rather than maximal. An exploratory pass scored five points higher by matching one institution's own systems, and a shipped default tuned to one mailbox quietly underperforms in every other.

A pre-filtered thread still gets a verdict and a one-line summary, because it still appears in mecha mail list and a list you cannot recognise a thread in is not a list.

The store is an index, not a copy of your mailbox. It holds ids, envelope metadata and the verdict. Bodies are fetched on demand and never written there, so the retention question stays with your provider and there is no second place for mail to leak from.

The classifier never talks to a run that has tools​

This is the whole design, and it is the front door's shape applied one directory over:

The privileged run sees the extraction, never the prose.

Reading mail arms untrusted_input. A loop that reads fifty threads into one conversation therefore arms the interlock for all fifty, and every draft it stages comes out tainted — correct, and useless, because a warning that fires on everything has stopped being a warning.

So the prose goes to a classifier issued no tools, no history, no system prompt and no shared cache prefix. It is a fresh one-shot call per thread, and only its typed output travels. What a run with tools is given is the verdict and the sender's address; what stays behind is the subject, the sender's chosen display name, the classifier's reasoning, and its one-line summary. That last one is the tempting one to pass — it is short, and it is exactly what a summary line wants — but it is model prose derived from prose a stranger wrote, and paraphrasing an injection does not remove it. A run that genuinely needs to know what a thread says calls mail_get_thread and takes the taint honestly.

mecha mail show prints the prose, deliberately. A person reading their own mail in a terminal is the safe context: you cannot be prompt-injected into mailing your own calendar somewhere. mecha mail list --json serves the typed view instead, because a script has no human's excuse.

Snippet first, body only where it matters​

A preview settles the newsletters. The full message is read only when the verdict is respond or names a request kind — the cases where the answer changes what happens next. Roughly a quarter of threads escalate.

It is deliberately not triggered by how short the snippet looks. A provider caps its preview at a couple of hundred characters, so nearly every real email appears truncated, and escalating on that would escalate everything.

Working the queue: what you can do to a thread​

Every verb is a separate command rather than mecha mail act --action <x>, on the same reasoning that made mail_triage's actions a closed enum: a free-form label argument would put spam inside a verb that reads as harmless.

CommandWhat it doesWhere it lands
replya model reads the thread and composes an answerstaged in the outbox
forward --to <addrs>passes it on with a covering linestaged in the outbox
scheduleturns it into a calendar eventyour calendar, directly — a private hold that invites nobody; staged in the outbox only when your note names people to invite
archiveout of the inbox — reversible, nobody notifiedthe mailbox
spamtrains the provider's filterthe mailbox
tasktracks it on the knowledge graph's boardthe graph
needs-info --missing <what>parks it until somebody answersthe store
dismissdrops it from the queue without actingthe store

Everything here that reaches a third party stages rather than sends: a reply, a forward, and a schedule whose note names people to invite. A schedule with no invitees reaches nobody. It makes a private hold on your own calendar (calendar_hold), which needs the allow rule above to run without asking. Without that rule the hold is refused, and the run is told to fall back to a staged calendar_create_event with no invitees. That is today's behaviour: an ordinary, visible event you release from the outbox, not a private hold. The fallback is an instruction to the model, not a mechanism. A run that doesn't follow it adds nothing and leaves the thread alone.

reply, forward and schedule are the actions here that need an agent rather than a tool call — a model has to read the thread and write prose — and the run that does so reads the thread, which arms both interlock legs. So the draft arrives in /outbox flagged tainted, which is correct: it was written after reading a stranger's words. Drafting from the classifier's one-line summary instead would produce clean drafts written from a paraphrase, which is worse exactly where it matters.

archive and spam reach nobody outside your own mailbox, so they are not staged — staging them would make triage circular, reviewing a queue in order to fill another queue. spam is separated from archive because it is the one triage action with an effect outside your mailbox: it trains the provider's filter.

task carries the deadline the classifier already found, which is the whole point — a task somebody has to re-read the mail to schedule is a task they will schedule later or never. Its --project must already exist on the graph and is passed through untouched, never invented from a subject line: a project node conjured out of an email is a board nobody can query.

needs-info is not dismiss. Dismissing says I am not doing this; parking says I have asked and cannot proceed yet, and the thread stays your problem.

Day two: the threads you meant to answer​

mecha mail list --aged # respond threads old enough to have been answered
mecha mail list --aged --surface # …and record that they were surfaced

A thread still unanswered after a day is overwhelmingly unlikely ever to be answered, and by then the person has stopped looking at mail — so a queue that only works as a pull surface is exactly why those threads die. --aged is the list the morning trigger reads.

Two decisions in it:

  • It keys on the bucket, never on silence. Most unanswered mail correctly needed no reply, so only respond threads age into this list.
  • --surface is separate from reading the list on purpose. A list command that mutates as a side effect of being run cannot be used to look, and looking is most of what anyone does with a queue. The briefing passes it; a person checking what day two would say does not.

The default age is 30 hours rather than 24 — a working day, so an email that arrived in the evening is not nagged about at breakfast.

Corrections: telling the classifier it was wrong​

mecha mail correct <thread> --bucket respond --urgency today
mecha mail correct <thread> --deadline none # `none` clears a field

Field-level on purpose. A misread bucket, a missed deadline and a wrong request kind are different errors with different fixes, and a correction that only says "this was wrong" teaches a learner noise.

The verdict is fixed immediately, so the list you read is right straight away — and the before/after pair is kept on the record, because the mistake is what a learner has to see. A learner shown only the right answer cannot tell what to stop doing. A correction that agrees with the classifier records nothing.

Corrections become rules​

mecha mail reflect # corrections → triage-domain reflections

One tool-less, history-less model call per unmined correction — the same shape as the classifier it is reasoning about — and it is idempotent, each correction keyed into its own ledger so a nightly pass never re-argues one.

Most corrections produce nothing, deliberately. The frame asks for a rule about a kind of mail and says outright that declining is the common case: a wrong rule rides in every future classification, and a missing one costs a single verdict.

Those reflections feed the triage learning domain, whose rules ride only in the classifier's own pass and never in a run that has tools. That distinction is what makes learning from mail possible at all — see the learning page for the provenance argument, which is the subtlest thing in either subsystem.

Bulk reading is an operator verb, never a tool​

mecha-mail corpus --since 2026-07-01 --account campus

corpus downloads a span of mail for analysis into ~/.mecha/mail-corpus/<account>.jsonl, walking all folders including Sent so a reply can be joined back to the thread that prompted it. It is what score and eval read.

It is absent from the MCP surface on purpose. The model has no business reading a year of mail, and a corpus verb on the tool surface is one prompt away from being asked to.

It also stores mail unclassified, which is the subtler half. Running a corpus through the classifier projects the current tags onto it and confirms them by construction, so a taxonomy derived that way measures the labels rather than the mail. That is how the vocabulary was wrong for a month: the largest single category of mail arriving was missing from the list entirely, because the most routine thing that arrives is the thing that does not come to mind.

The analysis that produced those decisions is gitignored. One mailbox's figures are its owner's, so what the measurement decided is written down and what it counted is not.

Measuring it: score and eval​

Classification accuracy stops being a feeling. Two instruments, answering different questions:

mecha mail score # the live store, against what actually happened
mecha mail eval --account campus --out graded.jsonl # the classifier, against a known corpus

score grades the live triage store. Behaviour (did a reply actually go out) and testimony (what you corrected) are reported apart, because a reply is one-sided evidence and a correction is not. Threads younger than 48 hours are excluded: most replies that ever happen land on the first day, so a same-day thread has no outcome yet, and counting it would punish every rule equally for how recently the mail arrived. Reply evidence comes from the corpus rather than from mail_get_thread, because that tool renders prose for a model to read and a measurement keyed on a display format breaks silently the day the format changes.

eval grades the classifier against a corpus whose outcome is already known, with no human grading anything. The ground truth is one-sided and the output says so: a thread you answered proves the thread mattered, so burying it is a countable error — a thread you never answered proves nothing, because most unanswered mail correctly needed no answer and some was settled in a meeting. So it reports a false-ignore rate on the answered stratum and a volume on the other, and never a single blended accuracy. Both strata are sampled to the same size, because answered threads are rare and a uniform sample of 200 would hold a handful of the only threads carrying ground truth.

--out keeps every graded verdict. A measurement that discards its evidence has to be re-run to be re-read: the first run of this eval reported a merged figure and threw away the 120 judgements behind it, so splitting respond from notify afterwards cost another hour of inference rather than a grep. Grading the artifact is this project's rule for models; it applies to its own instruments too.

eval writes nothing to the triage store — grading year-old mail is not triaging it, and a scorecard that mutated the queue it measures would be unrepeatable.

/mail — the queue as a modal​

The TUI works the queue without leaving it, on the /outbox pattern: the store is read for display, and every mutation is a mecha mail … child process. Nothing there reimplements a verb.

That is not tidiness. The store and the CLI are the product and every front-end is one reader — the nightly, the morning briefing and the modal all act through the same commands — so anything the modal can do a script can do, and a modal-only action would be a feature no trigger could ever use. Slow work (a reply builds a whole tool surface and can take minutes) spawns detached and is watched by polling the store, never the child.

A reply's result lands in /outbox, not here. There is exactly one approval surface and this is not it: /mail decides whether something needs an answer, /outbox decides whether this answer goes.

Unlike the front door, this modal shows prose — the same reasoning as mecha mail show. What must not see the prose is a privileged run, and none happens here.

Tags are mecha's own​

Not a Gmail label, not a Graph category. Those are different objects, and a tag that means something subtly different per account fails at the one job a tag has. Keeping them internal costs no OAuth scope and works identically on both providers. The cost, stated plainly: tags are invisible in Gmail, Outlook and on your phone. Mail triaged by mecha looks untouched in every other client.

Recognising a request is not routing it​

If a thread is really a standard request arriving as an email — a recommendation letter, someone asking to join the lab — the classifier names the kind. Whether it can then be handed to the front door is a separate question, answered by whether a form for that kind actually exists. A kind with no form keeps its name, because that is evidence about what your mail actually contains, and loses only a promotion there would be nothing behind.

Running it on a schedule​

Two timers, one sweep. scripts/mecha-mail-classify.{service,timer} sweeps at 05:30 in the machine's local time (its OnCalendar= names no zone, so a host on UTC runs it at 05:30 UTC) as the after-hours catch-up; scripts/mecha-mail-classify-day.{service,timer} sweeps every 20 minutes through the working day (07:30–21:50), so a thread is sorted within about half an hour of arriving. The daytime timer names its zone (America/New_York) so the window follows daylight saving; set the zone on its two OnCalendar= lines to your own before installing. A zone in a calendar spec needs a recent systemd (this was written on 255); check yours accepts the spec with your zone in it — systemd-analyze calendar '*-*-* 08..21:10/20:00 Your/Zone' — before installing. A fresh enable --now of a timer with no valid OnCalendar= fails loudly, but editing an installed one and running daemon-reload only logs the parse error, and the timer then silently stops firing:

cp scripts/mecha-mail-classify.{service,timer} ~/.config/systemd/user/
install -D -m 755 scripts/model-idle.sh ~/.local/bin/mecha-model-idle
cp scripts/mecha-mail-classify-day.{service,timer} ~/.config/systemd/user/
systemctl --user daemon-reload
systemctl --user enable --now mecha-mail-classify.timer mecha-mail-classify-day.timer

The daytime sweep stands down rather than competing: before each run, model-idle.sh skips it when any slot on the local model server is busy (the owner is chatting, or an agent is working) or the GPU is above 30%. That check is for a local model: it asks the llama-server named by MECHA_SLOTS_URL in the service (http://127.0.0.1:8080/slots as shipped — set it to your [providers.local] base_url plus /slots). With a hosted provider there is no local server to ask; remove the service's ExecCondition= line instead — left in, it reads a refused connection as a restarting server, so the sweep skips silently for about three hours and then fails every tick. A quiet tick costs one mailbox listing and no model call, since the sweep only classifies threads it has not seen.

A timer rather than a trigger, because a trigger's action is a prompt on purpose and this is a deterministic command. The unit names its workspace explicitly: a user unit without one runs in $HOME, which contains ~/.mecha, and a workspace the mecha home sits under is refused. Failures need no special handling — mecha doctor already watches every mecha-* unit.

A failed classification is retried; a dismissed one is not. The sweep asks whether a thread needs classifying rather than whether the store has heard of it — absent or failed means classify, classified is done, and dismissed was a person's decision. The distinction is not academic: the store holds failures as well as verdicts, so a sweep skipping everything it had a record for would skip a failed thread forever, on the strength of a record saying the classification never happened. Found by the outage that produced it — a model server down overnight left 17 threads recorded failed, among them a manuscript review invitation, and every later sweep would have reported "0 to classify" like any quiet morning. A transient infrastructure failure must not become a permanent editorial one.

A thread that keeps failing for its own reason is retried on a backoff — an hour, doubling, at most a day — rather than every sweep, and an outage never counts against it. Because a sweep reads only the newest messages, a failed thread that has scrolled out of that window is retried from the store when it comes due, so a busy inbox cannot strand it. mecha mail list shows when each failure is next tried.

Signing in​

Microsoft: device code​

mecha-mail auth campus --provider outlook --tenant <tenant-id>
To sign in, open https://microsoft.com/devicelogin on any device
and enter this code:

F7KQ2XM9B

Device code needs no redirect URI, so it works with an app registration your organisation has already approved, and no forwarded port, so it works over SSH. The scopes are exactly four:

https://graph.microsoft.com/Mail.ReadWrite
https://graph.microsoft.com/Mail.Send
https://graph.microsoft.com/Calendars.ReadWrite
offline_access

Mail.ReadWrite is there because triage — archive, mark read, report spam, trash — changes messages in place; without it triage would work on Gmail and silently do nothing on Outlook. It often needs an administrator. Microsoft classes it high-impact, and a managed tenant's recommended consent policy blocks end users from granting it: instead of a consent screen you see "Need admin approval", and auth fails there rather than at first use. Ask your tenant administrator to grant the app registration Mail.ReadWrite once, then run auth again. mecha doctor reports an account whose grant does not cover the triage verbs. User.Read is not requested; the account's address is read from Sent Items instead.

When sign-in fails, the Entra error is translated into a sentence saying what to change — admin consent, an app not registered in the tenant, the wrong organisation — with the raw description kept beside it.

Google: browser loopback​

mecha-mail auth personal --provider google opens a browser sign-in that returns to 127.0.0.1:8924 (--port to change it) and waits two minutes. Four scopes: gmail.modify, gmail.send, calendar, calendar.events — stopping short of https://mail.google.com/, so permanent deletion is never granted.

When a sign-in expires​

Tokens refresh on their own; you should not see them. What you will see is a refresh token the provider has revoked or expired — a changed password, a tenant policy, or a Google client in Testing status after 7 days. That is permanent, so nothing retries it: mecha-mail exits with code 77, mecha doctor and mecha-mail accounts name the dead account, and the fix is to run mecha-mail auth <name> --provider <provider> again. Declaring grant_lifetime_days for an account that expires on a schedule gets you the warning before it happens.

Two unification wrinkles​

mail_reply takes a thread_id and answers the newest message in it (or message_id when one is named). Graph does that natively; Gmail cannot, so the library synthesizes the addressing: answer the sender, or the recipients when you are replying to your own message; keep everyone on reply-all; drop the user's own address, which is known from the credential store; and add Re: only if the subject does not already carry it.

Merged calendars sort on the raw provider stamps, before any zone rendering, because rendered strings only sort within one zone. Rendering happens afterwards in MECHA_TZ, which is your [agent] timezone as mecha hands it over (falling back to TZ, then to leaving the stamp alone; a query window never falls back) — and all-day events skip zone conversion entirely and keep their bare date, or a Monday retreat gets announced as Sunday at 8pm. See Timezones.