Skip to main content
本页暂无完整中文版。以下先提供中文导读,随后是英文原文。
中文导读: 本页说明数据库、上传文件和其他数据卷的备份与恢复。升级前请先备份并核验结果;密钥与配置文件要单独保存,具体顺序见中文的升级部署。 Health data is irreplaceable — a reading you lose is a blood draw you cannot repeat. This page is the answer to “what do I copy, and how do I get it back”. Every command here was run against a live stack while writing it, including the restore: the dump below was restored into a scratch database and came back with 29 tables in the theta_ai schema and the vector extension in place.

What actually holds your data

compose.yaml declares four volumes. They are not equally precious, and treating them alike is how people back up a pip cache and miss their uploads. There is no mirobody_charts volume. The ChartService MCP tools that rendered PNGs through a Node toolchain were removed; the agent now writes a fenced vis-chart data block that the frontend renders, so nothing writes chart files at all.

Two things that make copy-pasted commands fail silently

Volume names are project-prefixed. The volume declared as mirobody_upload is <project>_mirobody_upload on the daemon — for a checkout in mirobody/, that is mirobody_mirobody_upload. docker volume ls shows the real names. This matters because getting it wrong does not error:
shell/backup.sh checks docker volume inspect before reading anything, so a name that does not exist is reported rather than archived as 45 empty bytes. Container names come from service names. There is no container_name: in compose.yaml and the database service is pg, so the container is mirobody-pg-1, not mirobody-db-1. Prefer docker compose exec pg …, which does not care what the container is called.

Backing up

Nightly, via cron:
What it does, and why:
  • pg_dump -Fc, not a tar of /var/lib/postgresql/data. Tarring a running cluster copies files that are moving under the tar; the result restores, or does not, depending on timing. pg_dump is consistent by construction and restores into any pgvector image rather than only a byte-identical one.
  • The dump is verified before the run is called a success. pg_restore --list parses the archive’s table of contents, which fails on a truncated or empty file. A backup nobody has read is a hope, not a backup.
  • The dump is written and verified inside the container, then copied out. A custom-format archive is not seekable through a pipe, so verification cannot read one from stdin, and pg_dump > file that dies half-way leaves a plausible-looking truncated file.
  • Uploads are archived only if the local volume exists (see the table).
  • Retention is pattern-scoped — it prunes mirobody-db-*.dump, mirobody-uploads-*.tar.gz and mirobody-redis-*.tar.gz older than RETENTION_DAYS, and leaves anything else in the directory alone.
The script is backup-only on purpose. Restore stays a human decision: the one time you need it, you want to be reading the steps.

Restoring

Stop the app first so nothing writes while the schema is being replaced. Keep pg running — it is what does the restoring.

Database

Into a fresh database (the safe path — the old one stays until you are satisfied):
Then point PG_DBNAME at holistic_db_restored in your config and start the app. To restore over the existing database instead, add --clean (it drops each object before recreating it) and be aware that there is no undo:
A note on the vector columns: they restore as ordinary data, and the pinned pgvector/pgvector image already has the extension, so no re-embedding is needed after a restore. (Re-embedding is only for changing UTILS_EMBEDDING_MODEL or the model of the MODELS entry carrying that embedding: family.)

Uploaded files

Check the real volume name with docker volume ls first. The archive is made with -C /data ., so it extracts as the volume’s contents, not as a nested data/ directory.

Then

Before upgrading

  1. Run shell/backup.sh and keep the output off this machine. This is the whole checklist, because it is also the rollback: there are no down migrations.
  2. Read CHANGELOG.md for the version you are moving to.
  3. git pull && ./deploy.sh — the script takes no arguments and already does compose down, then up -d --remove-orphans, then tails the logs.
  4. Watch that log tail: the schema DDL runs on the first start (see below), and a file that fails is logged with its sql_filename.

How the schema actually changes

server/bootstrap.py::create_schema replays every file in mirobody/schema/*.sql, in filename order, on each start. Every file is written to be safely re-runnable — there is no ledger of what has been applied, and no schema-version table by design. Verified by running the whole set three times against a clean database: zero errors. A failing file is rolled back on its own and logged; the rest still run, so one bad increment cannot leave half a schema with no explanation. Two consequences worth knowing:
  • Upgrades are additive. A new release adds tables or columns on first start; you do not run a migration command.
  • There is no automated downgrade. Rolling back a release that added a column means restoring the dump you took in step 1. That is why step 1 is step 1.