# Spec v5 — Revised Security, Idempotency, Failure Semantics, and Build Order

## 0. Threat model

The storage providers are treated as **honest-but-curious infrastructure providers**.

The providers:

* store ciphertext supplied by the server;
* may observe object IDs, object sizes, request timing, request frequency, and bucket-level metadata;
* may experience outages, transient failures, incomplete uploads, or accidental corruption;
* are not assumed to know the master encryption key;
* are not expected to deliberately modify ciphertext;
* must not be able to determine plaintext file contents from stored objects.

The Rust server is trusted with plaintext while files are being uploaded, downloaded, indexed, or previewed.

The storage providers therefore do **not** receive:

* plaintext file contents;
* plaintext filenames through provider object keys;
* the SQLite database in plaintext;
* the master encryption key;
* plaintext index snapshots.

The provider can still infer some metadata from ciphertext and request behavior:

* ciphertext size;
* number of objects;
* object creation/deletion timing;
* approximate access patterns;
* which objects are replicated together.

The system does **not** attempt to hide traffic analysis or storage size from the provider.

### Separate snapshot storage

SQLite snapshots are **not stored in the ordinary file buckets**.

They are encrypted locally and uploaded to a separate snapshot-storage location using separate credentials.

Ideally this snapshot location is:

* a different bucket;
* a different credential set;
* preferably a different account;
* and optionally a different provider.

The normal file buckets must not be required for index recovery.

---

# 1. Encryption model

## 1.1 Object encryption

Keep AES-128-CTR for the file data.

The reason for retaining CTR is that the system requires efficient random-access reads and range requests.

The storage provider receives only:

```text
random object ID
+
ciphertext
```

The ciphertext length equals the plaintext length.

There is no plaintext metadata in the object body.

## 1.2 CTR uniqueness invariant

The most important crypto invariant is:

> The same encryption key must never be used with the same CTR keystream for two different plaintexts.

Every object therefore receives a fresh random 128-bit object ID.

The object ID defines the initial counter:

```text
counter = object_id + block_number
```

The counter is interpreted as a big-endian 128-bit integer.

The implementation MUST detect counter exhaustion.

For an object of length `L` bytes:

```text
required_blocks = ceil(L / 16)
```

The object is rejected if:

```text
object_id + required_blocks - 1
```

would overflow the 128-bit counter space.

In practice this is unreachable for realistic files, but the condition must still be explicitly handled rather than relying on integer wraparound.

## 1.3 Key derivation

The master secret is never used directly as the AES key.

Use HKDF-SHA256:

```text
master secret
    │
    ├── "data-v1"  → AES data key
    │
    └── "index-v1" → snapshot encryption key
```

The derivation labels are versioned.

Future algorithms receive new labels rather than silently changing the meaning of existing keys.

## 1.4 What CTR does and does not protect

AES-CTR provides confidentiality.

It does **not** authenticate ciphertext.

This is acceptable under the system threat model because storage providers are not considered malicious editors of objects.

XXH3-128 is therefore retained as an accidental-corruption detector.

It is not described anywhere in the system as cryptographic authentication.

If the threat model is later expanded to include malicious storage providers, the object format must be upgraded to an authenticated-encryption format.

That would be a separate format version rather than an in-place change.

---

# 2. Object integrity

Every object has:

```text
object_id
size
XXH3-128(ciphertext)
```

The checksum is stored in SQLite.

The provider does not receive a plaintext checksum separately unless needed for repair operations.

The checksum is checked by:

* upload verification;
* background scrubbing;
* repair verification;
* reconciliation.

Normal downloads do not need to checksum the entire object.

For a ranged download, only the requested ciphertext range needs to be downloaded from the provider.

---

# 3. Object immutability

Objects are immutable.

A logical file replacement never modifies an existing object.

Instead:

```text
old node
   │
   └── old object

new upload
   │
   └── new object
```

The metadata transaction changes which object the node references.

This eliminates overwrite races at the storage layer.

---

# 4. Logical copies

A copy operation should NOT duplicate the physical object.

Instead:

```text
file A ─────┐
            ├── object 123
file B ─────┘
```

Multiple nodes may reference the same immutable object.

The object is eligible for GC only when no logical node references it.

This prevents copying a 500 GB file from causing another 500 GB of S3 traffic.

The object metadata therefore needs reference accounting or an equivalent queryable invariant.

A simple implementation is:

```text
objects(
    object_id PRIMARY KEY,
    size,
    checksum,
    mime,
    ...
)

nodes(
    ...
    object_id REFERENCES objects(object_id)
)
```

GC only considers an object when:

```text
SELECT COUNT(*)
FROM nodes
WHERE object_id = ?
```

is zero.

---

# 5. Database source of truth

SQLite is the authoritative source of logical filesystem state.

S3 objects are storage artifacts.

The system must never infer the logical directory tree by enumerating S3.

Therefore:

```text
SQLite says object exists
        +
replica says object exists
```

means the object is live.

An unexpected provider object is treated as an orphan and is eligible for reconciliation/GC.

---

# 6. Replication state machine

Replica state is explicitly modeled as:

```text
pending
uploading
ok
missing
bad
repairing
deleting
```

Valid examples:

```text
pending → uploading
uploading → ok
uploading → pending

ok → missing
ok → bad

missing → repairing
bad → repairing

repairing → ok

ok → deleting
missing → deleting
bad → deleting
```

No operation may directly invent a final state without performing the required verification.

For example:

```text
upload completed
```

does not automatically mean:

```text
ok
```

The server must verify that the expected object exists with the expected size and checksum.

---

# 7. Upload transaction state

Uploads are represented by an explicit durable state machine.

```text
created
  ↓
uploading
  ↓
quorum_reached
  ↓
metadata_committing
  ↓
committed
```

Failure states:

```text
aborted
failed
```

An interrupted upload may remain in:

```text
created
uploading
quorum_reached
```

and be recovered after restart.

The server must never assume that an incomplete database transaction implies that provider-side multipart uploads do not exist.

Provider-side leftovers are normal and are eventually cleaned.

---

# 8. Upload commit ordering

The safe ordering is:

```text
1. Allocate object ID.
2. Create durable upload record.
3. Encrypt and upload to buckets.
4. Verify quorum.
5. Commit SQLite metadata transaction.
6. Mark upload committed.
7. Queue remaining replicas for repair.
8. Cleanup temporary multipart/orphan objects asynchronously.
```

If the process crashes:

### Crash before quorum

The upload remains incomplete.

It can be resumed or abandoned.

### Crash after quorum but before SQLite commit

The object is physically present but not logically referenced.

It becomes an orphan.

The reconciliation/GC system eventually removes it.

### Crash during SQLite commit

SQLite's transaction guarantees determine the outcome.

The recovery code examines the upload state and reconciles provider state.

### Crash after SQLite commit

The object is logically live.

Replica repair continues after restart.

---

# 9. Idempotency

All operations that can be retried after a network failure must be idempotent.

## 9.1 API idempotency keys

Mutating API requests accept:

```http
Idempotency-Key: <random value>
```

The server stores:

```text
idempotency_key
request_hash
response_status
response_body
created_at
expires_at
```

The same key with the same request returns the original result.

The same key with a different request body is rejected.

This applies to operations such as:

```text
create upload
complete upload
create folder
copy
move
rename
delete
```

## 9.2 Upload part idempotency

A part is identified by:

```text
upload_id
part_number
```

Uploading the same part again replaces/revalidates that part rather than creating another logical part.

If the same part number is submitted with different data, the server either:

* replaces it before completion, or
* rejects it if the upload has already advanced past that state.

The behavior must be deterministic.

## 9.3 Complete-upload idempotency

Calling:

```text
POST /uploads/:id/complete
```

twice must not create two objects or two nodes.

If the upload is already committed, the second request returns the existing result.

## 9.4 Delete idempotency

Deleting an already deleted node returns success or an explicit already-deleted result.

It must never resurrect or duplicate GC work.

GC itself is also idempotent.

Deleting an object that has already disappeared from a provider is considered successful.

## 9.5 Repair idempotency

Repairing:

```text
object X → bucket B
```

multiple times must produce the same final state.

A repair worker may safely crash and retry.

---

# 10. Bucket failures

A bucket may be:

```text
active
offline
draining
```

`offline` is an operational condition.

`draining` is an administrative state.

An offline bucket:

* does not participate in quorum;
* does not serve normal reads;
* retains its existing objects;
* receives repair work when available.

A draining bucket:

* receives no new writes;
* serves no new reads;
* is removed from quorum;
* is eventually emptied after replacement replicas are verified.

---

# 11. Slow bucket handling

A slow bucket may be excluded from an individual upload if quorum is still reachable.

However, the server must distinguish:

```text
slow
```

from:

```text
failed
```

The replica remains `pending` rather than being considered corrupt.

The server must never delete a slow replica merely because it missed the upload deadline.

---

# 12. Read failover

Normal reads use only:

```text
replica.state = ok
```

If a provider connection fails:

```text
provider A
    ↓ failure
provider B
```

the server may resume using another healthy replica.

For a range request:

```text
requested plaintext range
        ↓
corresponding ciphertext range
        ↓
provider GET
        ↓
AES-CTR seek
        ↓
plaintext
```

No preceding plaintext bytes need to be downloaded.

---

# 13. Repair

Repair always copies ciphertext.

The repair worker must not decrypt/re-encrypt an object.

```text
healthy bucket
      │
      │ ciphertext
      ▼
repair worker
      │
      │ ciphertext
      ▼
target bucket
```

After copying:

```text
HEAD target
verify size
verify checksum
state = ok
```

This preserves the exact same object on every replica.

---

# 14. Scrubbing

The scrubber verifies:

```text
provider ciphertext
        ↓
XXH3-128
        ↓
SQLite checksum
```

On mismatch:

```text
ok → bad → repairing → ok
```

The scrubber must never silently overwrite a good checksum with a newly observed checksum.

A checksum mismatch is an error condition.

---

# 15. Reconciliation

Add a separate reconciliation worker.

It compares logical metadata against physical provider state.

It detects:

```text
SQLite says object exists
provider object missing

provider object exists
SQLite has no reference

wrong size

wrong checksum

stuck multipart upload

stale delete

stale replica state
```

Reconciliation does not immediately delete unexpected objects.

It should first place them into a quarantine/orphan state.

This protects against bugs and interrupted transactions.

After a configurable grace period, confirmed orphans become GC candidates.

---

# 16. Garbage collection

GC is deliberately asynchronous.

A deleted logical node does not immediately imply:

```text
DELETE object
```

Instead:

```text
node deleted
    ↓
object has zero references
    ↓
GC candidate
    ↓
grace period
    ↓
delete from buckets
    ↓
verify deletion
```

The grace period protects against recovery from a recently committed metadata state and makes operational debugging safer.

GC records should contain:

```text
object_id
bucket
attempt_count
last_attempt_at
next_attempt_at
created_at
```

Deletion must be idempotent.

---

# 17. SQLite snapshots

Snapshots are encrypted using the separate index key.

They are uploaded only to the dedicated snapshot-storage destination.

Each snapshot contains:

```text
snapshot_generation
schema_version
created_at
database contents
```

The snapshot has an authenticated encryption envelope.

Snapshot generations are monotonic:

```text
1001
1002
1003
1004
```

Recovery always selects the newest valid authenticated snapshot.

An incomplete snapshot upload is ignored.

A corrupted snapshot is ignored.

A valid older snapshot remains usable if the newest snapshot is invalid.

Snapshots are never considered part of the normal file-replica quorum.

---

# 18. Snapshot isolation

The snapshot destination must use credentials that are different from the ordinary file buckets.

The ordinary bucket credentials must not have permission to delete or overwrite snapshot data.

Where supported, enable:

* object versioning;
* retention/version protection;
* restricted delete permissions.

This protects the index-recovery path from ordinary bucket failures.

---

# 19. Bucket add

Adding a bucket:

```text
1. Register bucket.
2. Verify credentials.
3. Verify bucket identity.
4. Mark bucket active-for-backfill.
5. Create pending replica rows.
6. Backfill objects.
7. Verify every copied object.
8. Mark individual replicas ok.
9. Make bucket eligible for reads.
```

The new bucket must not participate in normal reads for an object until that object's replica is verified.

It also must not suddenly become part of write quorum before it has been initialized appropriately.

---

# 20. Bucket removal

Bucket removal has two phases.

### Drain

```text
draining
```

No new writes.

Existing replicas remain available for administrative inspection.

### Decommission

Only after replacement replicas are confirmed:

```text
delete objects
verify deletion
remove replica records
remove bucket
```

A bucket must never be removed from the configuration merely because the operator stopped seeing it.

The state transition must be durable.

---

# 21. Node concurrency

Add a metadata revision:

```text
revision INTEGER NOT NULL
```

Mutations can specify the expected revision.

Example:

```text
rename node 123
expected_revision = 8
```

If the node is already revision 9:

```text
409 Conflict
```

This prevents two browser tabs from silently overwriting each other's metadata changes.

---

# 22. SSE semantics

SSE events receive monotonically increasing IDs.

```text
event_id
sequence
event_type
payload
```

Clients reconnect with:

```http
Last-Event-ID: <id>
```

The server either:

* replays events since that ID, or
* sends a resynchronization event if the event is too old.

SSE is therefore treated as a synchronization aid rather than the authoritative data source.

The dashboard can always reload state from the API.

---

# 23. Revised build order

The build order should be changed substantially.

## Phase 1 — Cryptographic/object primitive

Implement only:

```text
object ID generation
AES-CTR
HKDF
range seeking
XXH3
object metadata
```

Tests:

```text
empty object
1 byte
15 bytes
16 bytes
17 bytes
large object
random ranges
range at every offset modulo 16
counter carry
maximum supported object size
```

Do not build S3 or the UI yet.

---

## Phase 2 — Local durable storage engine

Implement:

```text
objects
nodes
references
uploads
GC
SQLite transactions
```

Build a local filesystem/mock-provider backend.

Prove:

```text
upload
resume
replace
copy
delete
GC
crash recovery
```

before introducing network storage.

---

## Phase 3 — Single S3 bucket

Implement:

```text
multipart upload
ranged GET
HEAD
delete
retry
provider errors
```

Now prove that the entire object pipeline works against one real S3-compatible provider.

Acceptance test:

```text
kill process
at every upload state
restart
reconcile
```

---

## Phase 4 — Idempotent API

Implement:

```text
idempotency keys
upload IDs
part IDs
request hashes
revision numbers
```

Every mutating endpoint gets retry tests.

Specifically test:

```text
request succeeds
response is lost
client retries
```

The second request must not duplicate the operation.

---

## Phase 5 — Replication

Add:

```text
N buckets
quorum
replica states
parallel fan-out
slow bucket handling
repair
scrubbing
read failover
```

Start with:

```text
N=2
N=3
```

before testing arbitrary N.

---

## Phase 6 — Crash/reconciliation engine

Add:

```text
reconciliation
orphan detection
multipart cleanup
GC retry
repair retry
bucket outage recovery
```

Run fault injection continuously.

Examples:

```text
kill during upload
kill after quorum
kill before metadata commit
kill after metadata commit
kill during repair
kill during GC
kill during bucket removal
```

---

## Phase 7 — Snapshot/recovery system

Only after the storage engine is stable:

```text
SQLite online snapshots
encryption
separate snapshot bucket
snapshot generations
restore
reconciliation after restore
```

Test:

```text
destroy SQLite
restore newest snapshot
reconcile all buckets
```

Also test:

```text
newest snapshot corrupted
```

and confirm the previous valid generation is selected.

---

## Phase 8 — Public HTTP server

Implement:

```text
directory listing
GET
HEAD
Range
conditional requests
ETag
MIME handling
download disposition
```

Test large files and random ranges against the storage engine.

---

## Phase 9 — Authentication and admin API

Implement:

```text
Argon2id
sessions
CSRF
API tokens
rate limiting
trusted proxy handling
audit/activity
```

Then add:

```text
folders
move
rename
copy
delete
search
usage
health
events
```

---

## Phase 10 — Upload protocols

Add:

```text
dashboard resumable API
plain streaming PUT
tus
```

All three must use the same underlying upload state machine.

There must be exactly one storage implementation, not three subtly different ones.

---

## Phase 11 — Angie integration

Only after the application works correctly without the proxy:

```text
Unix socket
forwarded headers
timeouts
streaming
compression
TLS
HTTP/2
HTTP/3
```

Then test:

```text
large upload
large download
Range
disconnect/reconnect
slow clients
slow S3 providers
```

---

## Phase 12 — Dashboard

Finally integrate the Proton-derived UI.

The frontend should consume the already-tested API rather than becoming part of the storage-development loop.

Implement:

```text
browse
upload
download
preview
move
copy
rename
delete
search
Photos
health
activity
settings
```

SSE comes last because the application must remain correct without it.

---

## Phase 13 — Production hardening

Final stage:

```text
systemd
container
Angie configuration
metrics
structured logs
backup/restore runbook
CLI
upgrade/migration system
```

Then run the full failure matrix.

---

# 24. Required failure matrix

Before declaring the system production-ready, test at least:

| Failure                     | Expected behavior                                          |
| --------------------------- | ---------------------------------------------------------- |
| One bucket offline          | Quorum continues if possible                               |
| Multiple buckets offline    | Writes stop only when quorum is unavailable                |
| Bucket becomes slow         | Upload continues if quorum remains                         |
| Bucket dies during upload   | Other replicas complete; failed replica becomes pending    |
| Bucket dies during read     | Read fails over                                            |
| Process dies during upload  | Upload is resumable/recoverable                            |
| Process dies after quorum   | Orphan is reconciled or committed                          |
| Process dies during GC      | GC retries safely                                          |
| Process dies during repair  | Repair resumes                                             |
| Ciphertext bit flips        | Scrubber detects checksum mismatch                         |
| Provider loses object       | Replica becomes missing and is repaired                    |
| SQLite is lost              | Valid encrypted snapshot restores it                       |
| Latest snapshot is corrupt  | Previous valid snapshot is selected                        |
| Snapshot bucket unavailable | Existing local SQLite continues operating                  |
| Copy retried                | No duplicate physical object                               |
| Delete retried              | No failure/duplicate GC                                    |
| Complete-upload retried     | No duplicate object/node                                   |
| Rename races                | Revision conflict rather than silent overwrite             |
| SSE disconnects             | Client reconnects and resynchronizes                       |
| Server restart              | All durable work resumes                                   |
| Bucket added                | Backfill occurs without disrupting existing reads/writes   |
| Bucket drained              | No new writes; replicas remain recoverable                 |
| Old object replaced         | New object created; old object GC'd only when unreferenced |

---

# 25. Core invariants

The implementation should treat these as non-negotiable.

### Encryption

```text
No plaintext file contents are stored in provider buckets.
```

### Key secrecy

```text
No provider receives the master key.
```

### CTR safety

```text
No two objects reuse the same key + initial counter.
```

### Immutability

```text
An object ID refers to exactly one immutable ciphertext.
```

### Logical consistency

```text
SQLite is authoritative for which objects are live.
```

### Replication

```text
A committed object has reached the configured write quorum.
```

### Repair

```text
Repair copies ciphertext and does not change object identity.
```

### GC

```text
An object cannot be physically deleted while it has a live logical reference.
```

### Idempotency

```text
Every retryable operation can safely be executed more than once.
```

### Recovery

```text
A crash at any durable state transition eventually converges to a consistent state.
```

### Snapshots

```text
Index recovery does not depend on the ordinary file buckets.
```

### Provider privacy

```text
Provider-side object names contain no user-controlled filenames or paths.
```

### Metadata limitation

```text
The system protects file contents, not storage-side metadata such as
object size, timing, count, or access patterns.
```
