Officina · contract audit

Ten invariants have a file behind them. The two-hour promise has a habit.

Every rule in docs/CONTRACT.md was traced to the code, config or test that would fail if the rule were broken. Ten of the eleven have a candidate control that can be pointed at. Whether those checks actually run on every change was not verified here. Rule 3 — work resumes within two hours from any surviving device, the promise the whole fleet exists to keep — is checked by nothing. It holds because people remember it, which is another way of saying it is not enforced.

Invariants
11
As written, one page, unchanged since 17 Jul.
Backed by an artifact
10
A file, a unit, a hook or a test that fails on breach.
Backed by convention
1
Rule 3. Nothing fails if it stops being true.
Rules with a cited artifact
10/11
Counted as one per rule. Rule 3 cites none.

The contract, as written, with marks

11 rules · 1 unbacked

The document below is the real one, in its own order. Each rule carries its mark in place: an underlined clause is the load-bearing one, and the note beside it (below it, on a phone) names what enforces that clause.

The Officina Architecture Contract

docs/CONTRACT.md · one page · eleven rules that code and convenience bend to

One page. These are the invariants. Code, tooling, and convenience all bend to them — never the reverse. Change this file only with a deliberate, documented decision.

1.One canonical filesystem

All development state — repos, worktrees, uncommitted changes, running dev servers, agent sessions — lives on one server (dev-primary). Devices are windows, never homes. There is no file syncing between devices: not Syncthing, not Dropbox, not rsync habits. A device that dies loses nothing but itself.

Offline escape clones on a laptop are permitted, treated as disposable, and reconciled only through git push/pull.

2.Git is durability, not sync

Work moves through branch → commit → push. Nothing important sleeps uncommitted on the server. GitHub is the durable record — but GitHub is a service, not a guarantee: critical repos are additionally mirrored (bare) to the watcher node and included in filesystem backups.

3.“Nothing can crash” means two-hour recovery

Not high availability. Any single component — the primary server, the phone, a laptop, the ADE, a provider account — can die, and work resumes within two hours from any surviving device. This is drilled, not assumed: rebuild drills and restore tests are scheduled work, and a backup that hasn’t been restored from is not a backup.

4.Production is a different planet

The production server is not a dev machine, not a playground, and not managed by Officina beyond read-only observation. Real production deploys go through CI (GitHub Actions), triggerable from anywhere — never from an interactive shell as the normal path. Officina tooling may trigger CI; it may not be the deploy.

5.No public ingress on dev infrastructure

dev-primary and the watcher accept no unsolicited traffic from the public internet. All access rides the tailnet. Provider consoles are the break-glass path. Any service that must be reachable gets there via Tailscale Serve or an authenticated tunnel — never a public port. Pairing URLs and tunnel tokens are treated as root credentials.

6.Heavy compute is rented, not hosted

Burst workloads run on on-demand compute with a reproducible image, an input/output contract, and a budget guardrail. dev-primary is sized for the steady state, and steady-state sizing is revisited with evidence, not fear.

7.Machines prepare, the human approves

Updates are semi-automatic: security patches apply unattended; everything else is prepared by machines and applied by a human — pinned, staged, badged, PR’d. The primary interface never auto-updates. Every pinned component keeps its previous binary for rollback.

8.The custom layer stays narrow

Officina’s own code does only what boring tools can’t: project-aware git state, whitelisted actions, workflow glue. Vitals, dashboards, uptime, logs-at-rest, admin — boring, battle-tested tools. The agent exposes no arbitrary shell endpoint, binds to localhost, and audit-logs every action. If a feature could be described as “a small Datadog”, it is out of scope.

9.Secrets are inventoried, scoped, and revocable

Every credential has a listed owner, scope, and revocation path (see threat model). Machine identities are narrow, never a copied personal key. Secrets move over encrypted channels only — never through chat logs, never into git unencrypted. Losing any one device must be an inconvenience, not a breach.

10.Everything is rebuildable from two repos

The public engine (officina) plus the private fleet config (officina-fleet) plus off-site backups must be sufficient to rebuild the entire fleet. If a setup step exists only in shell history or in someone’s memory, it doesn’t exist.

11.Agents get knowledge, not a wider door

AI agents are first-class users of this fleet. Agents running on a node already have a shell; what they need is knowledge — a skill and per-repo AGENTS.md that teach the contract. Agents off the box get exactly one way in: the officina-agent action registry, exposed as MCP alongside the Deck’s HTTP.

The rule that ties it together: the action registry is the single control plane; Deck, skills and MCP are all thin clients over it. No agent may do anything the human Deck cannot.

Why the one matters more than the ten

Rule 3 is not a detail of the contract. It is the thing the other ten are arranged to deliver: one filesystem, git as durability, mirrors, two repos, off-site backups — all of it is machinery for getting back to work within two hours. Ten enforced rules produce a fleet that should recover in two hours. Nobody has shown that it does since January, and nothing on the fleet would notice if it stopped being true.

  1. Close it — 1

    Make the drill a scheduled unit, not an intention: officina-fleet/systemd/officina-drill.timer, monthly, restoring the latest off-site snapshot into a throwaway node and recording wall-clock time to first usable shell.

  2. Close it — 2

    Write the number down where a test can read it. A drill log at runbooks/drills/ with a duration field, and tests/fleet/test_recovery.py::test_last_drill_under_two_hours_and_under_35_days going red on either failure.

  3. Close it — 3

    Drill from the phone at least once a year. “From any surviving device” is the half of rule 3 that a CI runner can never exercise, and it is the half most likely to be false.

What “enforced” meant here, and what this audit could not see

A rule counted as enforced only if breaking it makes something fail on its own: a play errors, a test goes red, a request is refused, a hook rejects a commit. Documentation that describes the rule, and code that happens to comply with it, both counted as not enforced — that is the distinction rule 3 falls on.

Not covered: whether the tests named here actually run (CI history was not read), whether the off-site backups restore, and the escape-clone clause in rule 1, which is deliberately unenforceable by design. Two repos were traced at the commits in the masthead; docs/THREAT_MODEL.md was read for context but its own rules were not audited.