ReedβSolomon, not RAID
Every object is split into 4 data + 2 parity slices on 6 independent drives. Any 2 can die and the object is still readable β and STRUBS rebuilds the missing slices onto healthy drives by itself.
A single-host, fault-tolerant object store. Every object is erasure-coded across many ordinary disks β so no controller, filesystem, or single component can ever take all your data at once.
object "photo.jpg"
β
βββββββββββββββββββ΄ββββββββββββββββββ
β ReedβSolomon encode (4 + 2) β
βββββββββββββββββββ¬ββββββββββββββββββ
β
βββββββββ¬ββββββββ¬ββββββββ¬βββ΄βββββ¬βββββββββ¬βββββββββ
βΌ βΌ βΌ βΌ βΌ βΌ
data0 data1 data2 data3 parity0 parity1
β β β β β β
disk 4 disk 17 disk 23 disk 30 disk 9 disk 41 β 6 independent ext4 filesystemsSTRUBS was written after a client lost data to a failed RAID controller β not to failed disks. Making that failure mode impossible is the whole point.
Hardware RAID, Linux md, btrfs, and ZFS all build one big thing out of your disks. That thing is fast and convenient, and it is a shared fate: a controller or backplane in front of it can take every disk behind it offline at once, and "but the disks are fine" is cold comfort when the data is unreachable. STRUBS refuses to have a single component whose failure loses everything.
It is aimed at a specific shape of problem: one machine, a pile of mismatched drives that grows a disk or two at a time, tens to hundreds of terabytes, a small budget, and no tolerance for a total-loss event. In production since 2017, currently holding 130+ TB across ~30 disks of assorted sizes.
dd and a ReedβSolomon library.