Backup and restore

Create verified database backups and synchronize persistent OpenDataBio files without duplicating large storage on the application server.

An OpenDataBio backup has two parts: a logical database dump and the persistent files referenced by that database. The supplied tools deliberately do not compress or duplicate the potentially large media storage on the application server. They create small, verified database snapshots and optionally let an administrator’s workstation, NAS, or backup server pull the current files with incremental rsync.

Keeping only database dumps on the application server is a minimum safeguard, not disaster recovery. If that server or disk is lost, its original storage is lost too. Whenever possible, regularly pull or push a second copy to another machine.

What must be protected

Back up together:

  1. the MySQL/MariaDB database;
  2. storage/app/public/media;
  3. storage/app/public/datasets, including published dataset versions;
  4. the production environment and secret files in a separate secure store;
  5. especially APP_KEY, which is required to decrypt values stored by the application, such as personal Pl@ntNet keys.

Redis, caches, sessions, logs, vendor, node_modules, and regenerable downloads are not part of the durable backup.

Included tools

The repository provides:

scripts/backup/
├── backup-opendatabio.sh
├── backup.env.example
├── pull-opendatabio-backup.sh
├── pull.env.example
├── verify-opendatabio-backup.sh
└── systemd/

backup-opendatabio.sh runs on the application server. It creates a compressed database dump, verifies it, writes checksums and a manifest, marks the snapshot COMPLETE, and atomically updates LATEST. It never copies or deletes storage.

pull-opendatabio-backup.sh runs on another machine. It downloads the latest complete database snapshot and synchronizes media and datasets directly from their original directories.

Configure the application server

Create a root-owned configuration:

sudo install -d -m 700 /etc/opendatabio /var/backups/opendatabio
sudo cp scripts/backup/backup.env.example /etc/opendatabio/backup.env
sudo chmod 600 /etc/opendatabio/backup.env
sudo editor /etc/opendatabio/backup.env

Always set the real checkout and backup paths. The backup destination must be an absolute path other than /.

Docker production

Use:

ODB_DEPLOYMENT=docker
ODB_INSTALL_ROOT=/opt/opendatabio
ODB_ENV_FILE=/opt/opendatabio/.env.production
ODB_COMPOSE_FILE=/opt/opendatabio/docker-compose.prod.yml
ODB_COMPOSE_PROJECT=odb-prod
ODB_BACKUP_ROOT=/var/backups/opendatabio
ODB_BACKUP_GROUP=
ODB_STORAGE_PATH=/srv/opendatabio/storage

Docker production uses a named odbstorage volume by default. A remote rsync client cannot safely address files inside that private Docker volume. For a new production installation that will use pull backups, prepare a host directory:

sudo install -d -m 755 /srv/opendatabio/storage

Then set in .env.production before the first start:

ODB_STORAGE_SOURCE=/srv/opendatabio/storage

The Compose services mount this directory at the normal application storage location, and the container entrypoint applies the image’s configured www-data ownership. Do not hard-code a host UID/GID because image build settings may change it. Existing installations using the named volume must copy its contents to the bind-mounted directory during a planned maintenance window before changing ODB_STORAGE_SOURCE; do not merely change the variable or the files will appear missing.

Native Apache/nginx installation

Use:

ODB_DEPLOYMENT=native
ODB_INSTALL_ROOT=/var/www/opendatabio
ODB_BACKUP_ROOT=/var/backups/opendatabio
ODB_DB_NAME=opendatabio
ODB_DB_DEFAULTS_FILE=/root/.opendatabio-backup.cnf
ODB_BACKUP_GROUP=
ODB_STORAGE_PATH=/var/www/opendatabio/storage/app/public

Store database credentials outside the script:

[client]
host=localhost
port=3306
user=opendatabio_backup
password=REPLACE_WITH_A_PRIVATE_PASSWORD
sudo chmod 600 /root/.opendatabio-backup.cnf

The database account needs enough read access to dump all application tables and triggers. Credentials are not passed as host command-line arguments.

Run and verify a backup

Test manually before scheduling:

sudo /opt/opendatabio/scripts/backup/backup-opendatabio.sh \
  /etc/opendatabio/backup.env

The result is:

/var/backups/opendatabio/
├── LATEST
└── snapshots/20260813T021700Z/
    ├── database.sql.gz
    ├── SHA256SUMS
    ├── manifest.env
    └── COMPLETE

Verify it independently:

sudo scripts/backup/verify-opendatabio-backup.sh \
  /var/backups/opendatabio/snapshots/20260813T021700Z

Only directories with COMPLETE are eligible for automatic retention. The default retains daily database snapshots for seven days. Set ODB_RETENTION_DAYS as appropriate.

Schedule on the server

Copy and review the supplied examples:

sudo cp scripts/backup/systemd/opendatabio-backup.service.example \
  /etc/systemd/system/opendatabio-backup.service
sudo cp scripts/backup/systemd/opendatabio-backup.timer.example \
  /etc/systemd/system/opendatabio-backup.timer
sudo systemctl daemon-reload
sudo systemctl enable --now opendatabio-backup.timer
sudo systemctl list-timers opendatabio-backup.timer

Inspect executions with:

sudo systemctl status opendatabio-backup.service
sudo journalctl -u opendatabio-backup.service

Persistent=true runs a missed backup after the server returns.

cron alternative

Edit root’s crontab with sudo crontab -e:

17 2 * * * /opt/opendatabio/scripts/backup/backup-opendatabio.sh /etc/opendatabio/backup.env >> /var/log/opendatabio-backup.log 2>&1

List scheduled jobs with:

sudo crontab -l
sudo ls -la /etc/cron.d /etc/cron.daily
sudo systemctl list-timers --all

The script uses flock, so overlapping executions fail safely.

Pull to an administrator-controlled machine

Install openssh-client and rsync on the receiving workstation, NAS, or server. Configure an SSH user that can read only the backup and persistent storage paths. Prefer a dedicated SSH key and verify the server host key. On the application server, create the configured reader group and add that SSH user to it:

sudo groupadd --system opendatabio-backup
sudo usermod -aG opendatabio-backup backup-reader

Then set ODB_BACKUP_GROUP=opendatabio-backup in backup.env and run a new backup so the completed snapshot receives the group permissions.

backup-opendatabio.sh grants this group read access only to completed snapshots and LATEST. Separately grant the reader traverse/read access to the configured media and datasets paths using an appropriate group or filesystem ACL; do not make private configuration files world-readable.

Copy the client configuration:

sudo install -d -m 700 /etc/opendatabio /srv/backups/opendatabio
sudo cp scripts/backup/pull.env.example /etc/opendatabio/pull-backup.env
sudo chmod 600 /etc/opendatabio/pull-backup.env
sudo editor /etc/opendatabio/pull-backup.env

Important values:

ODB_REMOTE_HOST=backup-reader@example.org
ODB_REMOTE_BACKUP_ROOT=/var/backups/opendatabio
ODB_REMOTE_STORAGE_ROOT=/srv/opendatabio/storage
ODB_LOCAL_ROOT=/srv/backups/opendatabio
ODB_SSH_IDENTITY=/home/admin/.ssh/opendatabio-backup
ODB_STORAGE_POLICY=archive

Test:

scripts/backup/pull-opendatabio-backup.sh \
  /etc/opendatabio/pull-backup.env

The client reads LATEST from the server instead of guessing a folder from its local date. It reuses one SSH connection, validates the downloaded checksum and gzip stream, and then runs incremental rsync directly against the original storage. It does not use compression because images and many dataset artifacts are already compressed.

Storage policies:

  • archive (default) never deletes a local file that disappeared remotely;
  • mirror makes the current copy match the server, but moves deleted or replaced files into history/SNAPSHOT/ instead of discarding them.

Do not add plain --delete without a snapshot or backup directory. A mistaken deletion on the application server would otherwise propagate to the backup.

The supplied opendatabio-backup-pull.service.example and opendatabio-backup-pull.timer.example can be installed and edited on the receiving machine. Their Persistent=true setting is useful for a workstation that is not always on. Cron is acceptable when the receiving machine runs continuously.

Restore

Test restoration periodically on an isolated installation. At minimum:

  1. verify COMPLETE, SHA256SUMS, and gzip -t;
  2. use the OpenDataBio version recorded in manifest.env;
  3. place the target application in maintenance and stop workers;
  4. create a safety backup of the target;
  5. recreate the target database and import database.sql.gz;
  6. restore media and datasets to storage/app/public;
  7. restore the matching APP_KEY and required secrets;
  8. correct ownership and permissions;
  9. run php artisan migrate --force, php artisan optimize, and php artisan locales:audit;
  10. inspect logs and verify login, record counts, media, dataset versions, and a small queued job before reopening access.

For MariaDB-to-MySQL migration, recent mariadb-dump may prepend a sandbox directive unsupported by Oracle MySQL. Remove only that exact directive; never blindly remove the first line of an SQL dump.

Operational checks

Monitor at least:

  • age and size of LATEST;
  • exit status and logs of the server backup;
  • state/last-success on the pull destination;
  • free space on both machines;
  • periodic checksum verification;
  • a real test restoration at least quarterly.

Backups have only been proven when a restoration succeeds.