Skip to content

STRUBSStriping & Redundancy Using Basic Disks

A single-host, fault-tolerant object store. Every object is erasure-coded across many ordinary disks β€” so no controller, filesystem, or single component can ever take all your data at once.

In one picture ​

                        object "photo.jpg"
                               β”‚
             β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
             β”‚    Reed–Solomon encode (4 + 2)    β”‚
             β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                               β”‚
    β”Œβ”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”΄β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”
    β–Ό       β–Ό       β–Ό       β–Ό        β–Ό        β–Ό
  data0   data1   data2   data3  parity0  parity1
    β”‚       β”‚       β”‚       β”‚        β”‚        β”‚
  disk 4  disk 17 disk 23 disk 30  disk 9   disk 41   ← 6 independent ext4 filesystems

Why it exists ​

STRUBS was written after a client lost data to a failed RAID controller β€” not to failed disks. Making that failure mode impossible is the whole point.

Hardware RAID, Linux md, btrfs, and ZFS all build one big thing out of your disks. That thing is fast and convenient, and it is a shared fate: a controller or backplane in front of it can take every disk behind it offline at once, and "but the disks are fine" is cold comfort when the data is unreachable. STRUBS refuses to have a single component whose failure loses everything.

It is aimed at a specific shape of problem: one machine, a pile of mismatched drives that grows a disk or two at a time, tens to hundreds of terabytes, a small budget, and no tolerance for a total-loss event. In production since 2017, currently holding 130+ TB across ~30 disks of assorted sizes.

Where to go next ​

  • Architecture β€” the volume model, how a write is placed and a read is served, and the background jobs.
  • Data integrity β€” the checksum layers, quorum, verification, and how STRUBS refuses to make a bad situation worse.
  • On-disk format β€” the exact byte layout of a slice, and how to read your data with nothing but dd and a Reed–Solomon library.
  • Running STRUBS β€” deployment, adding and removing drives, rebalancing, and disaster recovery.
  • Encryption β€” the full LUKS story: keyfile, recovery passphrase, the identity model, and the guarantees.
  • HTTP API β€” the object API and the management API.

AGPL-3.0-only. In production since 2017.