Backup and restore
6 minute read
An OpenDataBio backup has two parts: a logical database dump and the persistent
files referenced by that database. The supplied tools deliberately do not
compress or duplicate the potentially large media storage on the application
server. They create small, verified database snapshots and optionally let an
administrator’s workstation, NAS, or backup server pull the current files with
incremental rsync.
Keeping only database dumps on the application server is a minimum safeguard, not disaster recovery. If that server or disk is lost, its original storage is lost too. Whenever possible, regularly pull or push a second copy to another machine.
What must be protected
Back up together:
- the MySQL/MariaDB database;
storage/app/public/media;storage/app/public/datasets, including published dataset versions;- the production environment and secret files in a separate secure store;
- especially
APP_KEY, which is required to decrypt values stored by the application, such as personal Pl@ntNet keys.
Redis, caches, sessions, logs, vendor, node_modules, and regenerable
downloads are not part of the durable backup.
Included tools
The repository provides:
scripts/backup/
├── backup-opendatabio.sh
├── backup.env.example
├── pull-opendatabio-backup.sh
├── pull.env.example
├── verify-opendatabio-backup.sh
└── systemd/
backup-opendatabio.sh runs on the application server. It creates a compressed
database dump, verifies it, writes checksums and a manifest, marks the snapshot
COMPLETE, and atomically updates LATEST. It never copies or deletes storage.
pull-opendatabio-backup.sh runs on another machine. It downloads the latest
complete database snapshot and synchronizes media and datasets directly
from their original directories.
Configure the application server
Create a root-owned configuration:
sudo install -d -m 700 /etc/opendatabio /var/backups/opendatabio
sudo cp scripts/backup/backup.env.example /etc/opendatabio/backup.env
sudo chmod 600 /etc/opendatabio/backup.env
sudo editor /etc/opendatabio/backup.env
Always set the real checkout and backup paths. The backup destination must be an
absolute path other than /.
Docker production
Use:
ODB_DEPLOYMENT=docker
ODB_INSTALL_ROOT=/opt/opendatabio
ODB_ENV_FILE=/opt/opendatabio/.env.production
ODB_COMPOSE_FILE=/opt/opendatabio/docker-compose.prod.yml
ODB_COMPOSE_PROJECT=odb-prod
ODB_BACKUP_ROOT=/var/backups/opendatabio
ODB_BACKUP_GROUP=
ODB_STORAGE_PATH=/srv/opendatabio/storage
Docker production uses a named odbstorage volume by default. A remote rsync
client cannot safely address files inside that private Docker volume. For a new
production installation that will use pull backups, prepare a host directory:
sudo install -d -m 755 /srv/opendatabio/storage
Then set in .env.production before the first start:
ODB_STORAGE_SOURCE=/srv/opendatabio/storage
The Compose services mount this directory at the normal application storage
location, and the container entrypoint applies the image’s configured
www-data ownership. Do not hard-code a host UID/GID because image build
settings may change it. Existing installations using the named volume must copy its contents
to the bind-mounted directory during a planned maintenance window before
changing ODB_STORAGE_SOURCE; do not merely change the variable or the files
will appear missing.
Native Apache/nginx installation
Use:
ODB_DEPLOYMENT=native
ODB_INSTALL_ROOT=/var/www/opendatabio
ODB_BACKUP_ROOT=/var/backups/opendatabio
ODB_DB_NAME=opendatabio
ODB_DB_DEFAULTS_FILE=/root/.opendatabio-backup.cnf
ODB_BACKUP_GROUP=
ODB_STORAGE_PATH=/var/www/opendatabio/storage/app/public
Store database credentials outside the script:
[client]
host=localhost
port=3306
user=opendatabio_backup
password=REPLACE_WITH_A_PRIVATE_PASSWORD
sudo chmod 600 /root/.opendatabio-backup.cnf
The database account needs enough read access to dump all application tables and triggers. Credentials are not passed as host command-line arguments.
Run and verify a backup
Test manually before scheduling:
sudo /opt/opendatabio/scripts/backup/backup-opendatabio.sh \
/etc/opendatabio/backup.env
The result is:
/var/backups/opendatabio/
├── LATEST
└── snapshots/20260813T021700Z/
├── database.sql.gz
├── SHA256SUMS
├── manifest.env
└── COMPLETE
Verify it independently:
sudo scripts/backup/verify-opendatabio-backup.sh \
/var/backups/opendatabio/snapshots/20260813T021700Z
Only directories with COMPLETE are eligible for automatic retention. The
default retains daily database snapshots for seven days. Set
ODB_RETENTION_DAYS as appropriate.
Schedule on the server
systemd timer (recommended)
Copy and review the supplied examples:
sudo cp scripts/backup/systemd/opendatabio-backup.service.example \
/etc/systemd/system/opendatabio-backup.service
sudo cp scripts/backup/systemd/opendatabio-backup.timer.example \
/etc/systemd/system/opendatabio-backup.timer
sudo systemctl daemon-reload
sudo systemctl enable --now opendatabio-backup.timer
sudo systemctl list-timers opendatabio-backup.timer
Inspect executions with:
sudo systemctl status opendatabio-backup.service
sudo journalctl -u opendatabio-backup.service
Persistent=true runs a missed backup after the server returns.
cron alternative
Edit root’s crontab with sudo crontab -e:
17 2 * * * /opt/opendatabio/scripts/backup/backup-opendatabio.sh /etc/opendatabio/backup.env >> /var/log/opendatabio-backup.log 2>&1
List scheduled jobs with:
sudo crontab -l
sudo ls -la /etc/cron.d /etc/cron.daily
sudo systemctl list-timers --all
The script uses flock, so overlapping executions fail safely.
Pull to an administrator-controlled machine
Install openssh-client and rsync on the receiving workstation, NAS, or
server. Configure an SSH user that can read only the backup and persistent
storage paths. Prefer a dedicated SSH key and verify the server host key.
On the application server, create the configured reader group and add that SSH
user to it:
sudo groupadd --system opendatabio-backup
sudo usermod -aG opendatabio-backup backup-reader
Then set ODB_BACKUP_GROUP=opendatabio-backup in backup.env and run a new
backup so the completed snapshot receives the group permissions.
backup-opendatabio.sh grants this group read access only to completed snapshots
and LATEST. Separately grant the reader traverse/read access to the configured
media and datasets paths using an appropriate group or filesystem ACL; do not
make private configuration files world-readable.
Copy the client configuration:
sudo install -d -m 700 /etc/opendatabio /srv/backups/opendatabio
sudo cp scripts/backup/pull.env.example /etc/opendatabio/pull-backup.env
sudo chmod 600 /etc/opendatabio/pull-backup.env
sudo editor /etc/opendatabio/pull-backup.env
Important values:
ODB_REMOTE_HOST=backup-reader@example.org
ODB_REMOTE_BACKUP_ROOT=/var/backups/opendatabio
ODB_REMOTE_STORAGE_ROOT=/srv/opendatabio/storage
ODB_LOCAL_ROOT=/srv/backups/opendatabio
ODB_SSH_IDENTITY=/home/admin/.ssh/opendatabio-backup
ODB_STORAGE_POLICY=archive
Test:
scripts/backup/pull-opendatabio-backup.sh \
/etc/opendatabio/pull-backup.env
The client reads LATEST from the server instead of guessing a folder from its
local date. It reuses one SSH connection, validates the downloaded checksum and
gzip stream, and then runs incremental rsync directly against the original
storage. It does not use compression because images and many dataset artifacts
are already compressed.
Storage policies:
archive(default) never deletes a local file that disappeared remotely;mirrormakes the current copy match the server, but moves deleted or replaced files intohistory/SNAPSHOT/instead of discarding them.
Do not add plain --delete without a snapshot or backup directory. A mistaken
deletion on the application server would otherwise propagate to the backup.
The supplied opendatabio-backup-pull.service.example and
opendatabio-backup-pull.timer.example can be installed and edited on the
receiving machine. Their Persistent=true setting is useful for a workstation
that is not always on. Cron is acceptable when the receiving machine runs
continuously.
Restore
Test restoration periodically on an isolated installation. At minimum:
- verify
COMPLETE,SHA256SUMS, andgzip -t; - use the OpenDataBio version recorded in
manifest.env; - place the target application in maintenance and stop workers;
- create a safety backup of the target;
- recreate the target database and import
database.sql.gz; - restore
mediaanddatasetstostorage/app/public; - restore the matching
APP_KEYand required secrets; - correct ownership and permissions;
- run
php artisan migrate --force,php artisan optimize, andphp artisan locales:audit; - inspect logs and verify login, record counts, media, dataset versions, and a small queued job before reopening access.
For MariaDB-to-MySQL migration, recent mariadb-dump may prepend a sandbox
directive unsupported by Oracle MySQL. Remove only that exact directive; never
blindly remove the first line of an SQL dump.
Operational checks
Monitor at least:
- age and size of
LATEST; - exit status and logs of the server backup;
state/last-successon the pull destination;- free space on both machines;
- periodic checksum verification;
- a real test restoration at least quarterly.
Backups have only been proven when a restoration succeeds.