# Drive Operational Runbook & Disaster Recovery Guide

## 1. System Overview & Architecture

Drive is a production-grade, zero-knowledge encrypted object storage system designed for high durability, strict confidentiality, and verifiable data integrity.

### Architecture Components
- **CLI & Core Binary (`drive`)**: Unified management and server binary providing HTTP/Unix socket daemon, background maintenance workers, and administrative subcommands.
- **SQLite Index (`drive.db`)**: Authoritative database tracking logical nodes, folders, physical objects, replica states, upload sessions, and GC queues. Uses WAL mode and optimistic revision locking (`expected_revision`).
- **Cryptographic Engine (`drive-crypto`)**: Zero-knowledge encryption using HKDF-SHA256 derived keys:
  - `data-v1` (16 bytes): AES-128-CTR object payload encryption with monotonic big-endian counter.
  - `index-v1` (32 bytes): AES-256-GCM authenticated snapshot envelope encryption.
  - `XXH3-128`: Fast checksum calculation for corruption and bit-rot detection.
- **Storage Backends (`drive-storage`)**: Parallel fanout across $N$ provider buckets (local filesystem or S3-compatible: Wasabi, Backblaze B2, MinIO, Cloudflare R2). Writes succeed once $W$ of $N$ quorum is reached.
- **Dedicated Snapshot Destination**: Completely isolated storage destination storing encrypted SQLite database snapshots with monotonically increasing generation numbers ($1001, 1002, \dots$).
- **Reverse Proxy (Angie / NGINX)**: Fronts Drive over a Unix domain socket (`/run/drive/drive.sock`), terminating TLS, HTTP/2, and HTTP/3 (QUIC), and forwarding client headers (`X-Forwarded-Proto`, `X-Forwarded-For`).

---

## 2. Installation & Deployment

### 2.1 Native Systemd Deployment (Recommended for Bare-Metal & VMs)

1. **Install Binary**:
   ```bash
   install -m 755 target/release/drive /usr/local/bin/drive
   ```

2. **Create System User and Directories**:
   ```bash
   useradd -r -s /usr/sbin/nologin -d /var/lib/drive drive
   mkdir -p /var/lib/drive/staging /var/lib/drive/buckets /var/lib/drive/snapshots /run/drive /etc/drive
   chown -R drive:drive /var/lib/drive /run/drive /etc/drive
   chmod 750 /var/lib/drive /run/drive
   ```

3. **Deploy Configuration**:
   ```bash
   cp config/drive.example.toml /etc/drive/drive.toml
   chmod 600 /etc/drive/drive.toml
   chown drive:drive /etc/drive/drive.toml
   ```

4. **Install Systemd Unit**:
   ```bash
   cp deploy/drive.service /etc/systemd/system/drive.service
   systemctl daemon-reload
   systemctl enable --now drive.service
   ```

5. **Verify Status & Logs**:
   ```bash
   systemctl status drive.service
   journalctl -u drive.service -f
   ```

### 2.2 Alpine Linux (OpenRC) Deployment

Alpine Linux uses the OpenRC init system. Use the provided OpenRC service script and configuration.

1. **Run Automated Installer**:
   ```bash
   cd deploy/alpine
   chmod +x setup-alpine.sh drive.initd
   ./setup-alpine.sh /usr/local/bin/drive
   ```

2. **Or Manual Installation**:
   ```bash
   # Create user and group
   addgroup -S drive
   adduser -S -D -H -h /var/lib/drive -s /sbin/nologin -G drive -g "Drive Storage Service" drive

   # Create directories
   mkdir -p /var/lib/drive /run/drive /var/log/drive /etc/drive
   chown -R drive:drive /var/lib/drive /run/drive /var/log/drive
   chmod 0750 /var/lib/drive /var/log/drive
   chmod 0755 /run/drive

   # Install binary and configs
   install -m 0755 target/release/drive /usr/local/bin/drive
   install -m 0755 deploy/alpine/drive.initd /etc/init.d/drive
   install -m 0600 deploy/alpine/drive.confd /etc/conf.d/drive
   cp config/drive.example.toml /etc/drive/drive.toml
   chown root:drive /etc/drive/drive.toml
   chmod 0640 /etc/drive/drive.toml
   ```

3. **Manage Service with OpenRC**:
   ```bash
   # Add to default runlevel to start on boot
   rc-update add drive default

   # Start, stop, restart, or check status
   rc-service drive start
   rc-service drive status
   rc-service drive restart
   rc-service drive stop

   # Tail service logs
   tail -f /var/log/drive/drive.log
   tail -f /var/log/drive/drive.err
   ```

### 2.3 Docker Container Deployment

1. **Build Container Image**:
   ```bash
   docker build -t drive:latest .
   ```

2. **Run Container**:
   ```bash
   docker run -d \
     --name drive \
     --restart always \
     -p 8080:8080 \
     -v /srv/drive/data:/var/lib/drive \
     -v /srv/drive/config/drive.toml:/etc/drive/drive.toml:ro \
     -e DRIVE_MASTER_SECRET="your-64-character-hex-secret" \
     drive:latest
   ```

### 2.3 Angie Reverse Proxy Integration

When fronting Drive with Angie / NGINX, bind Drive to the Unix domain socket `/run/drive/drive.sock`.

1. Copy `config/angie.conf` to `/etc/angie/http.d/drive.conf` or `/etc/angie/angie.conf`.
2. Ensure the `angie` user has group permissions to `/run/drive/drive.sock`:
   ```bash
   usermod -aG drive angie
   ```
3. Test and reload Angie:
   ```bash
   angie -t && systemctl reload angie
   ```

---

## 3. Configuration Reference (`drive.toml`)

| Section | Setting | Type | Description |
| :--- | :--- | :--- | :--- |
| `[server]` | `bind_tcp` | `string` | Optional TCP listen address (e.g. `"127.0.0.1:8080"`). |
| `[server]` | `bind_socket` | `string` | Optional Unix Domain Socket path (e.g. `"/run/drive/drive.sock"`). |
| `[server]` | `log_format` | `string` | `"pretty"` for human readable console, `"json"` for structured logs. |
| `[server]` | `log_level` | `string` | `"trace"`, `"debug"`, `"info"`, `"warn"`, or `"error"`. |
| `[database]` | `path` | `string` | Filesystem path to SQLite index file (`"data/drive.db"`). |
| `[crypto]` | `master_secret`| `string` | 64-char hex key. Can be supplied via `DRIVE_MASTER_SECRET` env var. |
| `[storage]` | `write_quorum` | `integer`| Number of replicas required before an upload commit succeeds. |
| `[storage]` | `gc_grace_period_secs` | `integer`| Grace period (in seconds) before deleted objects are purged (default 86400). |
| `[storage]` | `upload_timeout_secs` | `integer`| Per-bucket replica timeout before declaring bucket slow/pending. |
| `[storage]` | `staging_dir` | `string` | Directory holding uncommitted chunks and parts before encryption fanout. |
| `[[buckets]]` | `id`, `kind` | `string` | Unique bucket identifier and backend type (`"local"` or `"s3"`). |
| `[snapshots]`| `enabled` | `bool` | Whether snapshot operations are activated. |
| `[snapshots.backend]` | `id`, `kind` | `string` | Dedicated backend configuration isolated from file buckets. |
| `[workers]` | `gc_interval_secs` | `integer`| Seconds between background garbage collection passes. |
| `[workers]` | `scrub_interval_secs` | `integer`| Seconds between replica bit-rot scrub passes. |
| `[workers]` | `reconciliation_interval_secs`| `integer`| Seconds between orphan & stale upload reconciliation sweeps. |

---

## 4. Key Management & Security Best Practices

### Master Secret
- The Master Secret is a 256-bit random key (64 hex characters).
- Generate a new cryptographically secure secret:
  ```bash
  openssl rand -hex 32
  ```
- **Secrecy Guarantee**: Provider buckets NEVER receive the Master Secret or decrypted keys. Plaintext files never leave the host unencrypted.
- **Environment Variable**: To avoid storing secrets on disk, set the environment variable:
  ```bash
  export DRIVE_MASTER_SECRET="<64-hex-chars>"
  ```
  Or place it in `/etc/default/drive` with `chmod 600`.

---

## 5. Routine Operations & Maintenance

### 5.1 User & Token Administration

- **Add an Admin or Member User**:
  ```bash
  # Interactive password prompt:
  drive user add admin_user --role admin --config /etc/drive/drive.toml

  # Scripted with explicit password:
  drive user add alice --password "SuperSecretPassphrase123!" --role user
  ```

- **Generate an API Token**:
  ```bash
  # Token expires in 90 days:
  drive token create alice "backup-automation" --expires-in-days 90

  # Never-expiring token:
  drive token create admin_user "cluster-sync"
  ```
  Save the printed `drv_...` bearer token. It cannot be retrieved later.

### 5.2 Storage Maintenance CLI Commands

- **Run Garbage Collection**:
  ```bash
  # Respects configured 24h grace period:
  drive gc --config /etc/drive/drive.toml

  # Force immediate purge (bypassing grace periods):
  drive gc --force
  ```

- **Run Bit-rot Scrubber**:
  ```bash
  # Scrub all replicas across all buckets:
  drive scrub

  # Scrub a specific provider bucket:
  drive scrub --bucket wasabi-us-east
  ```

- **Run Storage Reconciliation**:
  ```bash
  # Reconcile all provider buckets:
  drive reconcile

  # Reconcile specific bucket:
  drive reconcile --bucket backblaze-b2-west
  ```

---

## 6. Backup, Snapshots & Disaster Recovery

### 6.1 Creating Point-in-Time Snapshots

Drive uses SQLite's online backup API to take a non-blocking, point-in-time consistent snapshot while the server is active. The snapshot image is encrypted with AES-256-GCM using `index-v1` and written to the dedicated snapshot storage destination.

```bash
drive snapshot create --config /etc/drive/drive.toml
```

Example Output:
```text
Snapshot created successfully:
  Generation:     1005
  Size:           4582912 bytes
  Object ID:      000000000000000000000000000003ed
  Schema Version: 1
```

### 6.2 Disaster Recovery: Full Database Restoration

When the host SQLite database file is corrupted, lost, or the entire host server is destroyed:

1. **Stop Drive Server**:
   ```bash
   systemctl stop drive
   ```

2. **Restore Database from Snapshot Storage**:
   ```bash
   drive snapshot restore --target-db /var/lib/drive/drive.db --config /etc/drive/drive.toml
   ```

   - Drive automatically downloads candidates from snapshot storage sorted by generation descending.
   - It verifies magic header `DRVSNAP1`, decrypts and authenticates AES-256-GCM tag.
   - Runs `PRAGMA integrity_check` on the restored database.
   - **Automatic Corruption Fallback**: If the latest snapshot generation was damaged or corrupted during upload, Drive logs a warning and automatically falls back to the previous valid generation (e.g. falling back from 1005 to 1004).
   - **Post-Restore Reconciliation**: Drive immediately runs reconciliation across all configured provider buckets, detecting missing replicas, healing inconsistencies, and quarantining orphans.

3. **Start Server**:
   ```bash
   systemctl start drive
   ```

### 6.3 Restoring a Specific Historical Generation
```bash
drive snapshot restore --generation 1002 --target-db /var/lib/drive/drive.db
```

---

## 7. Troubleshooting & Incident Runbooks

### Incident 1: One Storage Bucket Goes Offline
- **Symptom**: Logs show network errors or connection refused connecting to bucket `wasabi-us-east`.
- **System Behavior**: Because $W=2$ of $N=3$, writes continue uninterrupted on healthy buckets `local-primary` and `backblaze-b2-west`. Replicas for `wasabi-us-east` are recorded as `pending`.
- **Action**:
  1. Verify provider connectivity.
  2. If the outage is planned, set the bucket status to `offline`:
     The storage engine automatically excludes offline buckets from quorum and reads.
  3. When the bucket comes back online, run:
     ```bash
     drive reconcile --bucket wasabi-us-east
     ```
     The repair worker backfills all pending objects onto the bucket.

### Incident 2: Scrubber Detects Bit-rot / Checksum Mismatch
- **Symptom**: Scrubber log outputs `corrupted` replica detected for object ID on bucket `local-primary`.
- **System Behavior**:
  1. The replica state in SQLite is transitioned from `ok` to `bad`.
  2. The SQLite authoritative checksum is NEVER altered.
  3. Read operations immediately fail over to healthy replicas on other buckets.
  4. The repair worker automatically copies the valid ciphertext from a healthy replica, overwrites the damaged object on `local-primary`, verifies the XXH3-128 checksum, and restores replica state to `ok`.
- **Verification**: Run `drive scrub --bucket local-primary`. Verified OK count should increase; corrupted count should be 0.

### Incident 3: Complete Host Failure & Reconstruction
1. Provision new server with same architecture.
2. Install Drive binary and Angie reverse proxy.
3. Configure `drive.toml` with the same `master_secret` and remote bucket credentials.
4. Execute `drive snapshot restore`. The entire index database, user accounts, and replica states are reconstituted and reconciled against the remote buckets.
5. Launch `drive server`. All data is immediately accessible.
