Operations
The runbook. Most of this is doable from the web UI at /$/ui; the curl equivalents are given because they're what you'll want in a script or over SSH.
Throughout, $ is escaped in URLs (localhost/\$/volumes) so your shell doesn't eat it.
Deployment
STRUBS runs as root — it mounts filesystems and reads raw block devices.
# /etc/systemd/system/strubs.service
[Unit]
Description=strubs
Wants=network-online.target
After=network-online.target
StartLimitIntervalSec=0
[Service]
ExecStart=/usr/bin/node service.js
WorkingDirectory=/opt/strubs/dist
Restart=always
RestartSec=10
User=root
[Install]
WantedBy=multi-user.targetTuning goes in a drop-in, so it survives upgrades:
# /etc/systemd/system/strubs.service.d/tuning.conf
[Service]
Environment=STRUBS_DRAIN_CONCURRENCY=8Deploying a change:
cd /opt/strubs
npx tsc # -> dist/
cd ui && npx vite build # -> ui/dist/ (only if the UI changed)
systemctl restart strubsRestarting is safe, but not free. STRUBS unmounts its volumes on shutdown, so a restart will fail to unmount cleanly if anything else holds a file open on them (an
rsync, a script). In-flight relocations are aborted (IOABORT) — safe, because a slice is only removed from its source after the copy is verified and the reference flipped — but they'll need to be redone. Long-running jobs (scrub, drain, rebalance) checkpoint their progress and resume automatically on boot.
What STRUBS needs from the host
Worth knowing before you try to sandbox it, because it explains why it needs root.
| Binaries | lsblk, parted, mkfs.ext4, partprobe, mount, umount, smartctl, journalctl, udevadm |
| Syscalls | mount/umount — it mounts each volume itself |
| Devices | Raw reads on /dev/sd* (identity, SMART, LED flashing), and /dev/fuse |
| Hotplug | udev events, with a periodic scan as backstop |
| Logs | The systemd journal, for the kernel/smartd watcher |
Running in Docker
You can. Be clear about what it does and doesn't buy you.
STRUBS needs the host's block devices, the mount syscall, raw device access, and /dev/fuse. Once you've handed a container all of that, it is no longer a security boundary — it's a packaging format. That's a perfectly good reason to do it (a pinned Node version, and the native modules here are the fiddly kind: fuse-native, @ronomon/reed-solomon, diskusage), just not the reason people usually reach for a container.
What it takes:
docker run -d --name strubs \
--privileged \ # or: --cap-add SYS_ADMIN --cap-add SYS_RAWIO --cap-add MKNOD
-v /dev:/dev \ # NOT --device: STRUBS must SEE new disks appear
--device /dev/fuse \
-v /run/udev:/run/udev:ro \ # udevadm
-v /var/log/journal:/var/log/journal:ro \ # journalctl (kernel + smartd watcher)
--mount type=bind,source=/run/strubs,target=/run/strubs,bind-propagation=rshared \
-v /var/lib/strubs:/var/lib/strubs \ # instance identity
-p 80:80 \
strubsTwo of those are easy to get wrong:
-v /dev:/dev, not--device=/dev/sdb.--devicemaps a fixed node. STRUBS is built around disks appearing and vanishing; a static mapping means hotplug never works and a replaced drive never shows up.bind-propagation=rsharedon/run/strubs. Every volume STRUBS mounts lands under/run/strubs/mounts/, and the FUSE mount is at/run/strubs/data. Without shared propagation those mounts exist only inside the container's namespace — the host, and anything else on it, sees an empty directory.
The image needs smartmontools, parted, e2fsprogs, util-linux, and fuse.
If you don't grant everything, two features degrade rather than crash — both have an off switch:
| Missing | What stops working | Set |
|---|---|---|
| udev access | Hotplug detection | STRUBS_DISABLE_UDEV=true (falls back to periodic scanning) |
| Journal access | Kernel/smartd errors triggering a targeted verify | STRUBS_SYSLOG_WATCH_INTERVAL_MS=0 |
The recommendation: run STRUBS on the host, under systemd. It's a single-host appliance that already fits that shape, and every sharp edge above simply doesn't exist. Containerise MongoDB separately if you like — that one is a clean win, and it's the piece you most want to be able to move, upgrade, and back up independently.
Watching it
curl -s localhost/\$/status | jq # capacity, volumes by state
curl -s localhost/\$/volumes | jq # every drive: flags, SMART, fill
curl -s localhost/\$/faults | jq # outstanding slice faults
curl -s localhost/\$/verify-volumes | jq # scrub progress
curl -s localhost/\$/rebalance | jq # rebalance progress
journalctl -u strubs -fNotifications go to the log always, and to Slack if STRUBS_SLACK_WEBHOOK_URL is set (see Configuration). POST /$/notify/test sends a real one, which is the only way to know your webhook works.
Adding a drive
Provisioning wipes the disk. Partitions and formats it. Get the path right.
Find the device, sanity-check its SMART, then provision it:
curl -s localhost/\$/blockDevices?sort=sysfsPath | jq '.[] | {name, size, model, serial, volumeId}'
smartctl -d sat -H -A /dev/sdxcurl -X POST localhost/\$/volumes -H 'Content-Type: application/json' \
-d "{\"blockPath\":\"/dev/sdx\",\"wipe\":$(date +%s%3N)}"wipe is a timestamp, and it must be within 10 seconds of now — that freshness window is the only thing standing between a replayed request and a formatted disk.
The new volume is mounted, started, and writable immediately, and new writes start landing on it (it has the most free space). Existing data doesn't move by itself — rebalance if you want it levelled.
Encryption (LUKS)
STRUBS can encrypt every platter with LUKS2 — off by default, turned on one disk at a time. The full story (the keyfile and recovery passphrase, turning it on, converting a drive, rotating and auditing the passphrase, and recovering an encrypted fleet) has its own page: Encryption.
The one operational rule to carry away: back up /var/lib/strubs/luks.key off the machine the day you encrypt the first disk, and keep the recovery passphrase somewhere that is not this box. Those two secrets, kept apart, are the entire recovery kit.
Replacing or removing a drive
The sequence is: drain → confirm empty → identify → pull.
1. Drain it
curl -X POST localhost/\$/volumes/13/drainThis marks the volume read-only + draining and relocates every slice it holds onto healthy volumes, rewriting object references as it goes. It's move-then-flip: a slice is written and verified on its new home, the reference is flipped, and only then is the source copy deleted. Cancelling or crashing mid-drain loses nothing.
It works on a dead disk too — with the drive offline it reconstructs each slice from parity instead of copying it. That's the whole point: you don't need the failing drive to be readable to get its data off it.
Watch it:
curl -s localhost/\$/volumes | jq '.[] | select(.id==13) | {isDraining, bytesFree, bytesTotal}'
journalctl -u strubs -f | grep -i drainSlices that genuinely can't be rebuilt (below quorum) are left in place and reported, and they block removal until you accept the loss. That's deliberate.
2. Confirm it's actually empty
Do not skip this. Drain reports "complete" when its scan finishes, and a scan that skipped objects (because of an error mid-run) can still report complete.
curl -s localhost/\$/volumes/13 | jq '{isDraining, bytesFree, bytesTotal}'The authoritative check is that nothing references it any more:
db.content.countDocuments({ $or: [{ dataVolumes: 13 }, { parityVolumes: 13 }] }) // must be 0DELETE /$/volumes/13 will refuse while live slices remain, and tell you how many — so a refusal here is the system doing its job, not an obstacle to route around.
3. Find the physical drive
In a 30-bay chassis, the difference between sdf and sdg is somebody's afternoon. Flash its LED:
curl -X POST localhost/\$/volumes/13/identify # re-POST every ~1s to keep it going
curl -X DELETE localhost/\$/volumes/13/identify # stopThe UI does the heartbeat for you — open the volume menu and pick Identify. Reads stop by themselves ~3 seconds after the last ping, so a closed tab can't leave a drive spinning.
4. Remove it
curl -X DELETE localhost/\$/volumes/13 # soft-delete; refused if slices remainThen physically pull it. Keep the drive on a shelf until a full verify confirms the relocated copies are good — it costs nothing and it's the difference between an inconvenience and an incident.
Rebalancing
After adding or removing drives, fill is uneven. Rebalance levels it.
curl -X POST localhost/\$/rebalance -H 'Content-Type: application/json' -d '{}'
curl -s localhost/\$/rebalance | jq
curl -X DELETE localhost/\$/rebalance # cancel any time; safeIt computes a capacity-weighted target fill across the pool, sheds from volumes above it, and lands on volumes below it. Some things worth knowing:
- It shows up in the UI, with live progress, rate, and ETA.
bytesToMoveis recomputed from actual volume fills, so it's honest across restarts. - It works biggest-objects-first. Byte mass concentrates in a small minority of large objects while every move costs the same fixed latency, so this reaches the target in far fewer moves.
- It never copies parity. Parity is always recomputed from the data — a byte-copy would faithfully preserve parity that's silently wrong. This is why it's slower than a plain file copy, and it isn't negotiable. See Data integrity.
- It parks the scrub while it runs, and releases it afterwards. A scrub of a volume whose slices are being relocated is verifying a moving target.
Tuning it
Concurrency is a live setting — it applies to a running job at the next batch, no restart:
curl -X PUT localhost/\$/rebalance -H 'Content-Type: application/json' -d '{"concurrency":8}'Relocation is latency-bound (open, read, write, commit, ref flip, delete), not bandwidth-bound, so concurrency is the lever that matters. Raise it in steps and watch bytesPerSec in the status. Back off if the drives start thrashing or kernel I/O errors appear — seek-bound disks in external enclosures have a real ceiling.
Verification
# Full scrub of everything (reads every byte; takes a long time)
curl -X POST localhost/\$/verify-volumes -H 'Content-Type: application/json' -d '{"mode":"full"}'
# Light pass — existence + headers only. Hours, not weeks. Low stress.
curl -X POST localhost/\$/verify-volumes -H 'Content-Type: application/json' -d '{"mode":"light"}'
# Just one drive
curl -X POST localhost/\$/verify-volumes -H 'Content-Type: application/json' \
-d '{"volumeIds":[13],"mode":"full"}'
# One object
curl -X POST localhost/\$/verify-file/6a4e9b8f3b1e7a0049000001 \
-H 'Content-Type: application/json' -d '{"mode":"full"}'
curl -s localhost/\$/verify-volumes | jq
curl -X DELETE localhost/\$/verify-volumes
# History: every run ever started, most recent first — including why it started
curl -s localhost/\$/verify-runs | jqA scrub runs automatically every 90 days by default. Use full. A light verify will happily tell you every slice is present and correctly labelled while your parity is worthless — only a full pass recomputes parity and compares it. See Data integrity.
If a rebalance is running, a verify request is queued, not rejected: the response carries deferred: true and it starts when the rebalance finishes. The UI shows it as Waiting for rebalance.
/$/verify-volumes only ever shows the run in flight right now. /$/verify-runs is the durable history: one record per run (scheduled, manual, or syslog-watcher-triggered), with the trigger that started it — for a fault-driven run, the device, the signal kind (pending/ioerror), and the kernel/smartd detail line that armed it — plus status (running/completed/stopped) and its error totals once finished. This is what to check when a targeted verify is running and the question is "why."
When something is wrong
First: consider freezing
curl -X PUT localhost/\$/maintenance-freeze -H 'Content-Type: application/json' -d '{"frozen":true}'The freeze stops all background maintenance — scrub, repair, drain, rebalance — while reads and writes carry on normally. It's persisted and survives restarts.
Reach for it when you don't yet understand what's happening. Automatic repair is usually what you want, but repair reconstructs slices, and reconstructing from sources you haven't validated is exactly how a recoverable problem becomes a permanent one. Freeze, diagnose, then unfreeze. Unfreezing resumes work in order: drains, then scrub and repair, then rebalance.
A drive is throwing errors
journalctl -k | grep -iE 'i/o error|medium error|offline'
smartctl -d sat -H -A /dev/sdx
curl -s localhost/\$/volumes/13 | jq .smartInfoDistinguish two very different things:
- Media errors — reallocated or pending sectors, uncorrectable reads. The platter is going. Drain it.
- Link errors — the device dropping off the bus and re-enumerating (
device offline error, USB CRC counts,Buffer I/O errorfollowed by a re-attach). SMART will be spotless because the disk is fine — it's the cable, the bridge, or the enclosure. Reseat before you replace.
STRUBS itself treats a kernel error as a hint: the log watcher triggers a targeted verify of that drive rather than concluding anything. If the fault count crosses the threshold, the health monitor degrades the volume to read-only + unhealthy — reads keep working, writes stop. It never evicts a drive on its own.
Objects are flagged
curl -s localhost/\$/faults | jq '.faults[] | {objectId, sliceIndex, volumeId, code, repairStatus, repairBlockedReason}'db.content.countDocuments({ isFile: true, sliceErrors: { $exists: true } })
db.content.find({ isFile: true, sliceErrors: { $exists: true } }).limit(5)Blocked reasons, and what they mean:
| Reason | Meaning |
|---|---|
insufficient-slices | Below quorum. Repair is refused — with too few good sources it would produce plausible-looking wrong bytes and overwrite what's left. |
unrecoverable | Marked recoveryComment by an operator: accepted loss. |
target-unwritable | Nowhere healthy to put the rebuilt slice. Add capacity or clear a volume's flags. |
reconstruction-mismatch | Serious. The rebuild didn't reproduce the object's MD5, so it was refused rather than committed. The surviving slices are foreign or corrupt. Nothing was overwritten — this is the safety gate doing exactly its job. Investigate before touching it. |
If a fault looks stale — the object is fine and the slices are all present — re-verify that single object (POST /$/verify-file/{id}). A clean result clears it.
A drive vanished
The device reconciler handles disks appearing, disappearing, and coming back under a different kernel name. If a volume is showing isMissing, the disk STRUBS expects is genuinely not there.
Identity lives on the disk, not in the bay, so you can move drives between ports and enclosures freely. To identify an unmounted disk out-of-band:
debugfs -R 'dump /strubs/.identity /tmp/id' /dev/sdX 2>/dev/null && xxd /tmp/id
# byte 37 (0x25) = volume idOn USB and SAS enclosures,
lsblkandsmartctloften report the enclosure bridge's serial rather than the drive's — two bays can look like the same device. Don't identify drives by serial on those setups. The.identityfile is the ground truth.
The tools/ directory
Standalone scripts, run with node tools/<name>.js. They talk to Mongo and the disks directly, and several require a compiled dist/. Most default to a dry run; read the script before running it.
They exist because a diagnosis you can run out-of-process, against a frozen system, is worth a lot more than one you have to redeploy the service to get.
full-verify.js, light-verify.js | Out-of-process verification passes, gated on whole-object MD5. |
parity-verify.js | Read-only parity audit — recompute and compare. |
classify-slice-errors.js | Bucket flagged objects: readable-now / recoverable / lost. |
recover-residual.js | Reconstruct and relocate slices from dead volumes, MD5-gated. |
restamp-headers.js | Repair mis-stamped slice headers in place, data verified first. |
verify-dead.js, archive-dead.js | Prove unrecoverability, then archive records and reclaim the slices. |
resync-stats.js | Rebuild the cached storage statistics. |
The parity-*.js scripts are forensic tools from a specific incident. They're kept because the analysis they encode is hard to reproduce, not because you'll need them routinely.
Backups
Say it once more: back up MongoDB.
The slices are useless without the content records that say which volumes hold them. Every disk can be perfectly healthy and the array still lost. mongodump on a schedule, off the box.
And STRUBS survives disks, not buildings. It is not an off-site backup.