Switch from a v0 setup¶
ameesh v0 was a set of local scripts: agent-mail with a file mailbox under
~/.local/state/agent-mail, and a loop script that starts and resumes each
agent's harness session. This guide moves a running team from v0 to ameesh v1
without ever breaking an agent in progress. It is adapted from the
step-by-step plan in the repository (docs/BASCULE.md, in French).
Rules of the switch¶
- One change at a time, and every step has its rollback, symmetric, written below.
- v0 stays intact until the last step: the v0
agent-mail, the v0 loop and their state directories are neither modified nor deleted. - Never two loops on the same session. The v0 loop does not know v1 leases. A v0 loop is stopped, and its whole process group is dead, before a v1 runner starts for that agent (and the other way round on rollback).
- Authority is proven: no decision is recognised from its text.
- When in doubt:
ameesh export-v0 --agents <name>puts the mail back in the file mailbox and v0 takes over again.
Acts reserved to the owner are marked [O]: creating repositories and system users, enrolling their passkey, any spending, any network exposure.
Prerequisites¶
On the bench, the end-to-end test and the demo pass, on both drivers:
scripts/test.sh tests.test_bout_en_bout
scripts/demo-v1.sh
Rollback: nothing has changed.
Step 0: backups and freeze (no change)¶
Archive the v0 state (mailbox, v0 configuration, the loop's per-agent state), record the checksums of the v0 scripts, and list every agent with its harness and session id: this is the map of the switch. No agent should be in the middle of a long turn.
Rollback: nothing has changed; the backup serves every later step.
Step 1: database and package¶
A durable Postgres (a container with automatic restart and a daily
pg_dump copied off the machine, or a system package). The bench container is
not a production database.
python3 -m venv ~/.local/share/ameesh/venv
~/.local/share/ameesh/venv/bin/pip install -e <path to your ameesh clone>
# link ameesh, agent-runner and ameesh-approve into your PATH;
# the v0 agent-mail stays in place until step 8
install -m 700 -d ~/.config/ameesh
cat > ~/.config/ameesh/config.json <<'JSON'
{ "dsn": "postgresql://ameesh@127.0.0.1:5432/ameesh", "host": "<this host>",
"project": "<project>", "humans": ["<your id>"] }
JSON
chmod 600 ~/.config/ameesh/config.json # the password goes in ~/.pgpass (0600)
ameesh migrate
ameesh doctor --notify-test
ameesh import-v0 --dry-run # what would be imported, without touching anything
Do a restore drill: pg_dump, restore into a throwaway database, run
ameesh doctor on it, drop it.
Rollback: remove the links from your PATH (not the v0 agent-mail);
stop the database. Nothing depends on it yet.
Step 2: bootstrap canon [O] for the repository¶
Create the project's canon on the model of
Write a canon: a Member per human (authenticators: []
for now), a Host for this machine with its policy (including
policy.work_dirs or work_roots: each agent's current working directory),
and an Agent (with responsible, without approve) and a Placement
(hosts: [<this host>]) for every agent of the step 0 map, including
orchestrators. A human's interactive session that should keep receiving mail
is an agent in externe mode (see step 8). There is no generator yet: the cards are written by hand (or
drafted by an agent) and the owner reviews the pull request.
git clone <canon repository> ~/canon # then AMEESH_CANON=~/canon, or "canon" in config.json
ameesh canon check --fetch # 0 errors; read at the merged commit of origin/main
ameesh canon sync --fetch --bootstrap-ref main # first time only
ameesh list # canon agents: host, responsible, no lease
# every agent still runs on v0: keep them out of reach of any v1 runner until their own switch
for a in <every agent of the map>; do agent-runner stop "$a"; done
No runner runs yet, so nothing starts. canon sync writes only declarative
columns and never lifts a stop set by agent-runner stop.
Rollback: remove AMEESH_CANON (or canon from the configuration). The
registry keeps its columns and stops, with no effect while no runner runs.
Step 3: ameesh-approve [O]¶
-
A dedicated system user for ameesh-approve, distinct from the agents' user; its state directory
0700, its service token0600:sudo useradd --system --create-home ameesh-approve sudo -u ameesh-approve ameesh-approve gen-tokenThe agents' side reads a
0600copy of this token: it lets ameesh request an approval, never sign one. SetAMEESH_APPROVE_URL=http://127.0.0.1:8765(HTTPS, or HTTP on loopback only) andAMEESH_APPROVE_TOKEN_FILE=<copy>in the agents' environment. -
Expose it to the phone over HTTPS through a reverse proxy; the service itself listens on loopback only (
--bind 127.0.0.1 --port 8765). The RP ID is the public host name, the originhttps://<host>(AMEESH_APPROVE_RP_ID,AMEESH_APPROVE_ORIGINS, identical on the ameesh side, which verifies receipts; check withameesh approve-check). The service starts in the strict profile, and can also terminate TLS itself behind a passthrough gateway: see Host ameesh-approve for a team. Going online is an owner's act. Changing the host name later requires enrolling passkeys again. -
Enrol the owner's passkey, through the canon:
sudo -u ameesh-approve ameesh-approve enroll-link --approver human:<your id> # → single-use link, to open ON THE PHONE; the passkey becomes a PROPOSAL, nothing is activeCheck the fingerprint out of band, add the proposed entry to
authenticatorsof yourMembercard by pull request, merge it, thenameesh canon sync --fetchon each host.ameesh authenticator listshows the active passkey with the canon commit it comes from. -
Real test: a
shell-noopaction of classirreversiblestays blocked without a receipt ([receipt_required]);ameesh action request <id> --approver human:<your id>writes the link into the thread; once signed,ameesh action fetch-receipt <id>fetches, verifies and attaches the receipt; execution passes; replaying the receipt fails ([replay]).
Rollback: stop the service and remove the exposure. Irreversible actions
stay blocked, which is the safe behaviour. To remove a passkey: a pull request
that removes it from the canon, then canon sync.
Step 4: one pilot agent¶
In this order, without shortcuts:
- wait for the end of its current turn;
- stop its v0 loop and wait until its whole process group is dead (no
process may still hold the session; escalate
TERMthenKILLon the process group if a harness survives); -
take over mail and session on that stable state, and lift the stop set in step 2:
ameesh import-v0 --agents <agent> agent-runner register <agent> <harness> --session "<session id>" \ --prompt "Resuming under ameesh v1: continue your current lot." ameesh canon sync # restores the canon columns; canon state "ok" required ameesh show <agent> # session = <session id>, status queued, no lease -
start the v1 runner for this agent only, then a round trip:
agent-runner --agents <agent> --poll 5 ameesh mail send <agent> "switch round trip" --from <your id> ameesh fil show <project> # the message is in the thread, in clear ameesh list # lease, unread, budget
Hooks: the hooks' command line does not change (agent-mail hook
<harness>, with byte-identical JSON). A harness started by the runner receives
AGENT_MAIL_NAME, AMEESH_RUNNER_ID and AMEESH_LEASE_EPOCH; the hook
delivers only while that lease is alive.
Rollback, symmetric: agent-runner stop <agent>, stop its runner
(SIGTERM releases the lease and kills the harness group), check that the
lease is released (ameesh show <agent>) and nothing holds the session,
ameesh export-v0 --agents <agent>, then restart the v0 loop on the same
session. Messages that arrived in the database meanwhile are written back to
the v0 mailbox; none is lost.
Step 5: the other agents, in waves of one or two¶
Same procedure as step 4, agent by agent. Once two or more agents are switched,
a single runner per host serves them (agent-runner without --agents); agents
still on v0 stay stopped in the registry and are not claimed. After each wave:
ameesh list # no switched agent without a lease; no v0 agent with a lease
ameesh canon check # canon valid, state "ok" recorded for the host
ameesh fil list # threads are moving
ameesh doctor --notify-test
Rollback (per agent): as in step 4.
Step 6: orchestrators become turn-based agents¶
- The monitoring that used to feed the interactive session now drops its
findings into the orchestrator's mailbox as events:
ameesh mail send <orchestrator> "<readable finding>" --kind event [--urgent] --from monitoring. Events are coalesced (at most one wake-up perAMEESH_EVENT_COALESCEseconds, 120 by default) unless--urgent. - End the interactive session (end of turn, harness exit, group dead),
import-v0,agent-runner register … --session … --prompt …,ameesh canon sync: the runner claims it and wakes it by turns. - Set its session policy to
taille(rotation on size and turn duration, never on a change of lot), since it receives messages from every lot:ameesh set <orchestrator> session_policy=taille. - To talk to it directly:
ameesh attach <orchestrator>(refused during a turn, unless--wait). On exit, the lease goes back to the runner.
Rollback, symmetric: stop it, stop the event feed, export-v0, restart the
interactive session on the same session id.
Step 7: supervision (systemd)¶
The runner as a user service; the canon is synced before any claim (an unreadable or invalid canon does not prevent start-up: it closes claims of canon agents, fail closed):
# ~/.config/systemd/user/agent-runner.service
[Unit]
Description=ameesh runner (one per host)
After=network-online.target
[Service]
EnvironmentFile=%h/.config/ameesh/env # AMEESH_HOST, AMEESH_CANON, AMEESH_APPROVE_*
ExecStartPre=-%h/.local/share/ameesh/venv/bin/ameesh canon sync --fetch
ExecStart=%h/.local/share/ameesh/venv/bin/agent-runner --poll 5
Restart=always
RestartSec=5
[Install]
WantedBy=default.target
ameesh-approve [O] as a system service under its own user, never the agents':
# /etc/systemd/system/ameesh-approve.service
[Unit]
Description=ameesh-approve (human approvals with passkeys)
After=network-online.target
[Service]
User=ameesh-approve
EnvironmentFile=/etc/ameesh-approve/env # AMEESH_DSN, AMEESH_APPROVE_RP_ID, _ORIGINS, _PUBLIC_URL
ExecStart=/opt/ameesh/venv/bin/ameesh-approve serve --bind 127.0.0.1 --port 8765
Restart=on-failure
NoNewPrivileges=yes
ProtectHome=read-only
[Install]
WantedBy=multi-user.target
After each merged canon pull request: ameesh canon sync --fetch on each host
(by hand or with a timer). Watch: ameesh alerts --follow --json (long turn,
idle with mail, dead runner, session too big, stale lot, host pressure,
stopped with mail, orphan lot, expired delegation), ameesh decisions
(pending approvals, unknown outcomes), the canon state per host. To have those
alerts pushed to the responsible human, run ameesh notify as a user
service next to the runner (example unit deploy/systemd/ameesh-notify.service;
see Operate agents).
Rollback: disable the runner service (leases released, harness groups killed), then step 4's rollback agent by agent; disable ameesh-approve.
Step 8: end of life of v0 (after one or two weeks)¶
Point agent-mail to v1, keep the v0 scripts and file state read-only for
another month, replace v0's restart mechanism by the systemd services, archive
the step 0 backup with the switch date, and delete the v0 scripts only after a
tested restore.
Rollback: reinstall the scripts from the backup, point agent-mail back to
v0, ameesh export-v0 for the mail, then the v0 loops (step 4 rollback).
Moving the hooks to v1 ahead of step 8¶
The v0 agent-mail takes an identity from the working directory and reads only
the file mailbox, so a human's interactive session can pick up another agent's
identity and never sees its v1 mail. Since v1.4.0 (decision 0030), the hooks can
move to v1 before the rest of step 8; the v1 hook takes an identity only from
AGENT_MAIL_NAME (with the runner's lease) or from an explicit session
binding, and delivers nothing otherwise. An owner's act [O]:
- Prerequisites:
ameesh migrateapplied; no agent still run by a v0 loop (after the switch,agent-mail sendandinboxtalk to the database); pending v0 mail of external sessions imported (ameesh import-v0 --agents <names>). Back upagent-mailand the harness configurations. - If a local wrapper sets
AGENT_MAIL_NAMEfor some sessions, remove it first from the harness configurations (hooks callagent-mail hook <harness>directly), whileagent-mailis still v0. - Bind each interactive session to keep reachable:
ameesh mail bind <agent> --session <id> --harness <h> --pid <PID>(orameesh mail bind --import <file>from a wrapper's bindings file), thenameesh mail bindings. Agents born from such sessions areexterne. - Last, point
agent-mailto the v1 entry point of the venv.
Check: in a bound session, agent-mail whoami shows the agent with source
session; in an unbound session it answers identité non liée ("identity not
bound", exit 1) and
nothing is delivered. Rollback, symmetric: restore the v0 agent-mail and
the harness configurations from the backup, ameesh export-v0 --agents
<names> for mail that arrived meanwhile, and ameesh mail unbind the
bindings if you give up.
Upgrading an existing installation¶
To v1.3.x: working directories from the host¶
Since v1.3.0, an agent's working directory comes from its host's Host card
(policy.work_roots, work_root), no longer from the cwd of Placement
cards; v1.3.1 adds policy.work_dirs, the {agent} template and a
transitional fallback to the old placement cwd (warning
admission-cwd-inherited, removed in v1.5.0). A v1.3.0 binary ignores
work_dirs and copies {agent} literally, so the order matters:
- Write the new settings on a canon branch without merging it, and check
it with the new binary, outside the host's installation:
ameesh canon check --canon <checkout> --ref origin/<branch>(0 errors, noadmission-cwd-inheritednorhost-work-dir-missing). - Stop the old runner (and with it its periodic
canon sync). - Install the new version;
ameesh canon check --fetchon the current canon shows where the fallback is still used. - Merge the canon pull request,
ameesh canon check --fetch, thenameesh canon sync --fetch;ameesh list --jsonshows a distinct, existingcwdfor each agent. - Start the runner. An agent left
blocked("directory missing") resumes by itself once itscwdis right. - Later, remove
cwdfrom thePlacementcards.
Rolling back to v1.3.0 needs a canon it reads the same way: one directory per
team (work_roots without {agent}, no work_dirs) before any v1.3.0 sync.
To v1.4.0: migrations 0030 to 0036¶
v1.4.0 brings the migrations 0030 to 0036 (agent mode and stop reason,
delegations, several canons, the session's account, session bindings,
authenticators per canon, and review fixes). 0032 and 0035 change unique
keys that older code writes to: after the migration, an older runner's
canon sync and passkey sync fail. Therefore every host and every service
that shares the database (runners, ameesh-approve, ameesh notify, canon
sync timers, v1 agent-mail hooks) is stopped, upgraded together, and
restarted. Migrations are not undone: the backup taken before
ameesh migrate is the only rollback.
systemctl --user stop agent-runner ameesh-notify # on EACH host
sudo systemctl stop ameesh-approve # [O]
ameesh list # no live lease left
pg_dump --format=custom --file ~/backups/ameesh-before-1.4.0-$(date +%F-%H%M).dump "$AMEESH_DSN"
# install v1.4.0 in the venv of each host (and of ameesh-approve [O]), then ONCE:
ameesh migrate # applies 0030 … 0036
# [O] re-run the SQL roles (idempotent): ameesh-approve now reads actions.canon and
# authenticators.canon (it answers 503 otherwise); the supervisor reads the new columns
PGOPTIONS='-c search_path=<schema>' psql "<admin DSN>" -v ON_ERROR_STOP=1 -f deploy/sql/role-approve.sql
PGOPTIONS='-c search_path=<schema>' psql "<admin DSN>" -v ON_ERROR_STOP=1 -f deploy/sql/role-superviseur.sql
# 0030 gives mode=execute to every existing row: reclassify humans' sessions by hand
ameesh set <agent> mode=externe # an external agent needs a responsible human
systemctl --user start agent-runner ameesh-notify # on each host
sudo systemctl start ameesh-approve # [O]
Then check: no human session left in execute mode; ameesh canon check and
ameesh list (canon agents idle, canon ok on each host); ameesh alerts
without unexpected dead_runner or orphan_lot (an old stop without a reason
raises stopped_with_mail if it has mail: restart it or stop it again with
agent-runner stop); ameesh-approve answers (no 503);
ameesh notify --once --dry-run.
Visible changes: send all reaches only the sender's team; work add
--assignee and work assign refuse an unknown name and an external agent
without --externe; the strictest max_agents of a host's Host cards caps
the runner; ameesh mail bind refuses an execute agent without --force.
Rollback: stop every service again, restore the backup
(pg_restore --clean --if-exists -d "$AMEESH_DSN" <file>), reinstall the old
code on each host, restart. Never the old code on the migrated database.
Risks and countermeasures¶
| Risk | Countermeasure |
|---|---|
| two loops on one session | v0 stopped and group dead before any v1 start (and the reverse); exclusive v1 lease; ameesh list after each step |
| the database goes down during the switch | the CLI falls back to the v0 files with a warning; hooks stay silent |
| a message imported twice | import-v0 moves files to imported/ and does not rewrite a message already in the thread |
| invalid or unreadable canon | no new claim of canon agents (fail closed); running leases and turns intact; ameesh canon check says why |
| an agent misplaced or without a responsible human | not claimable; reason in status_text |
| an irreversible action without a human | the gate requires a receipt bound to the digest, single use, with expiry; unknown outcome → reconcile, never an automatic retry |
| ameesh-approve exposed too early | service on loopback only; exposure is an owner's act |
| forged passkey enrolment | enrolment is only a proposal; out-of-band check and reviewed canon pull request before activation |
| an unmerged pull request that declares itself canonical | the trusted branch comes from the host, the last applied commit or the bootstrap, never from the commit read |
two concurrent canon sync (two hosts, an old canon) |
registry lock taken before the check, monotonicity: an old commit never re-activates a removed passkey |
| a replayed receipt | nonce consumed at launch, in the database |
| the machine's disk dies | daily pg_dump off the machine, restore drill of step 1 |
| upgrading a host whose working directories are not declared (v1.3.x) | transitional fallback to the placement cwd; stop the old runner, install and validate with the new binary before publishing work_dirs / {agent}; canon check warns admission-cwd-inherited / host-work-dir-missing |
| an older host left running on a database migrated to v1.4.0 | all hosts and services stopped and upgraded together; backup before ameesh migrate; never old code on a migrated database |
| a human's session takes another agent's identity from its folder (v0 hooks) | hooks moved to v1: identity from AGENT_MAIL_NAME or an explicit session binding, ancestor PID checked; without a binding nothing is delivered |