Every file restic could copy out of Fractal's data directory was safe on its own. A backup made of those files could still restore the wrong canvas, and load without a single error while doing it.
I found this while planning the backups, before an agent had written any of the scripts (how I work with agents). Nothing was lost, and the bug only shows up if you read the server code in the order restic reads the disk.
The server, written in Go, keeps each canvas in memory and persists it to that directory. I wanted hourly, encrypted, off-host copies of it, using the same tools as the laptop that runs our TV: restic (a backup tool), a Backblaze B2 bucket, and a systemd timer.
Each file on its own#
A canvas on disk is mostly two files.
canvas.snapshot is the full canvas state, rewritten every 30 seconds while the canvas is active. The server writes it to a temp file, fsyncs it, moves the current snapshot to canvas.snapshot.bak, and renames the temp file into place. Nobody ever modifies a snapshot file after it's written. If restic has one open, it reads a complete snapshot even if a newer one replaces it mid-read.
operations.log is newline-delimited JSON, one record per operation, appended in place. When compaction (below) rewrites it, it uses the same temp file, fsync, and rename. The server applies an operation only after the fsync for its batch returns. A reader that catches the log mid-append sees a prefix with a possibly torn last line. On load, the server cuts that line off. It belonged to an operation that hadn't been applied or confirmed to anyone.
Whatever restic opens, the loader accepts.
What a restore needs#
Each snapshot records OpSeq, the sequence number of the last operation it includes. Loading a canvas means reading the snapshot, then replaying every log record with a higher sequence number.
The log can't grow forever, so after each snapshot write the server compacts it. Compaction keeps only the records newer than the previous snapshot, the one that just became .bak. That margin exists so .bak can still be replayed if the primary snapshot is corrupt.
An example with made-up sequence numbers. Say three snapshot writes land at OpSeq 100, 160, and 230:
| Snapshot write | canvas.snapshot | .bak | operations.log keeps |
|---|---|---|---|
| 1st | 100 | previous | newer than previous |
| 2nd | 160 | 100 | 101 onward |
| 3rd | 230 | 160 | 161 onward |
After the second write, a copy of the 100 snapshot still restores with the log. After the third, records 101 to 160 exist only inside the newer snapshots.
The 100 snapshot needs 101 to 160. Paired with the log from after the third write, it can't get them. A snapshot and a log are only a valid pair if they were read close enough together.
The order restic reads in#
restic doesn't freeze the filesystem. Its archiver lists each directory, sorts the names, and walks them in order. It descends into a subdirectory right away and finishes walking it before moving to the next name. Regular files are opened and handed to a pool of saver workers, and that handoff blocks when the workers are busy uploading.
The default canvas lives at the root of the data directory. Every other canvas lives under canvases/<id>/. Sorted, the root looks like this (trimmed):
canvas.json
canvas.snapshot <- opened here
canvas.snapshot.bak
canvases/ <- every other canvas: opened, queued for upload
operations.log <- opened only after the canvases/ walk finishes
session.key
With the example numbers from above, a live backup of the default canvas could go like this:
The walk through canvases/ opens each file and hands it to the savers, and it slows down whenever they're busy uploading. The first run uploads everything. Later runs skip files that haven't changed since the last backup, so what stretches the walk is the amount of changed data under canvases/.
Two snapshot writes during the walk are enough. Depending on where the 30-second timer was when restic opened the snapshot, that takes anywhere from about 30 to about 60 seconds of an active default canvas. Right after a clear it's shorter, because a clear forces an immediate snapshot.
Why it would look fine#
That backup restores without a complaint. The snapshot decodes, and so does the log. Replay applies the records newer than OpSeq, but the oldest of them are missing, so strokes drawn in the gap are just absent. A later delete or undo that names one of those strokes finds nothing to remove. The canvas counts it as missing, replay ignores the count, and nothing is logged.
A clear would be worse, if someone drew after it. The clear is a log record, and the server writes an extra snapshot right after it. If anyone draws afterwards, the next periodic write follows within 30 seconds, and its compaction drops the clear record. (On an untouched canvas that write is skipped, so nothing compacts past the clear.) A backup that paired the snapshot from before the clear with that log would restore the pre-clear canvas with the post-clear strokes on top.
No file is corrupt. Every check that asks "does this load?" says yes.
The fix: back up a copy, not the live directory#
The backup script first copies each canvas directory into a local staging directory and checks that the copy is consistent. restic never sees the live directory.
stage_dir() {
local src=$1 dst=$2 attempt before
for attempt in 1 2 3; do
before=$(snapshot_id "$src") # inode + mtime of canvas.snapshot
if copy_dir_once "$src" "$dst" && [[ $before == "$(snapshot_id "$src")" ]]; then
return 0
fi
echo "fractal-backup: $src changed during copy (attempt $attempt), retrying" >&2
done
die "could not stage a consistent copy of $src"
}
copy_dir_once copies every regular file except *.tmp and the log, then copies operations.log last. The check is a bracket. If canvas.snapshot has the same inode and mtime after the copy as before, no snapshot write finished during the copy. Without a new snapshot there's no new compaction, so the copied log still holds every record the copied snapshot needs. If the identity changed, the directory is recopied, and three failures fail the run. (Account data lives in SQLite and is staged separately with SQLite's .backup.)
Periodic writes are 30 seconds apart, with an extra one after each clear, and a local copy of one directory is a much smaller window than a network upload of the whole tree, so retries should be rare.
restic then backs up the staging directory, which nothing else writes to. A timer runs this hourly, and restic forget --prune keeps 24 hourly, 7 daily, 4 weekly, and 6 monthly snapshots. The fix needed no change to the server.
Staging also covers two smaller problems from the same reading. restic quietly skips a file that disappears between listing a directory and opening it, and a snapshot write has exactly that window between its two renames. In the staging copy, that moment either fails the copy or changes the snapshot's identity, so the directory is copied again. Half-written *.tmp files are skipped by name, so they never reach the backup.
What I didn't do#
Controlling restic's read order isn't possible. It sorts names itself, and when it opens a file depends on how fast uploads drain.
Stopping the server during the backup would make the read trivially consistent. It would also disconnect everyone drawing, every hour.
An LVM snapshot of the volume would give a true point-in-time view. It also adds mount and cleanup steps that can fail on their own. I hadn't checked that the volume group had free space for one. And the staging copy gives the same guarantee for the files that matter.
One risk is accepted. If an append fails its fsync, the server truncates the log back to where it was. A copy taken in that instant can keep a complete record the server then rolls back. That takes a disk error, and a crash at the same moment would leave the same record behind.
What the restore check can't see#
A second timer runs a restore check daily, and it passes. The check runs restic check, fails if the newest backup is more than two hours old, and restores it to a temp directory. It then starts a throwaway server on localhost, sends one WebSocket handshake per canvas, and watches the server's log for up to five seconds for that canvas's "ready" line.
That catches a broken repository, a stale backup, or a snapshot that won't decode. It can't catch the bug in this post. A snapshot paired with an over-compacted log loads fine, and "ready" is all the check asks for.
The restored files do carry the evidence. Sequence numbers come from a counter that goes up by one per operation, so an over-compacted pair leaves a gap: the first operation record newer than the snapshot has a sequence number above OpSeq + 1. In the example above, the 100 snapshot paired with the trimmed log would show its first record at 161.
That check isn't built. The snapshot's OpSeq sits inside its gob-encoded header and the loader never logs it, so reading it needs a small server change, and the backup plan ruled out app changes. A gap isn't always a bug either: the rolled-back append from "What I didn't do" has already spent its sequence numbers. For now, the bracket in the staging copy is the only thing keeping the pair consistent, and nothing checks its work after the fact.