fugl.dev
Merge PR #8: Unbrick the compose project, narrow what the site builder can reach (pr/site-deploy-least-privilege)
merged by Russ T. Fugalopened by Russ T. Fugal5 files+156 −81453fd367 merged here in total
description
Post-merge follow-up to PR #5. Four of the review findings, plus one that landing PR #5 caused on the live host. All five are in compose.yaml, the hook, and the runbook; nothing here touches deploy.sh.
P0 — the live stack is currently unmanageable
~/srv/stack/softserve/compose.yaml is a symlink into this repo's working tree, so the site-deploy service definition went live on the host the moment main was checked out — before a single install step ran. Its build context used ${SITE_DEPLOY_SRC:?...}, and Compose interpolates the whole file before selecting a service. A required-variable marker on a service nobody has installed yet is therefore an error for every command on every service.
Confirmed against the live project:
$ cd ~/srv/stack/softserve && docker compose config --services
error while interpolating services.site-deploy.build.context:
required variable SITE_DEPLOY_SRC is missing a value
The containers keep running — softserve, meta-sidecar, cloudflared and ssh-deprecated are all still up — but ps, logs, restart, and stop all fail. Also reproduced on a minimal two-service file, to confirm it is interpolation order and not something specific to this project.
Both :? markers on this service become :- with a self-describing sentinel path. A missing value now degrades to "up site-deploy fails, and the error names the variable" instead of "the whole project is bricked". Verified: the project parses with SITE_DEPLOY_SRC unset.
Immediate unblock without this PR: add SITE_DEPLOY_SRC to ~/srv/stack/softserve/.env. That is install step 1 here anyway.
The same footgun already exists on cloudflared (CF_TUNNEL_CREDS_FILE:?). It is set on the live host so it is latent, and cloudflared genuinely cannot start without it — deliberately left alone rather than widening this change. Worth a card.
P1 — env_file: .env gave third-party build code the stack's credentials
site-deploy is the service that runs bun install over the site's entire dependency tree, on a network where softserve and meta-sidecar answer to their own names. env_file: .env put every variable in that file into its environment — SS_SIDECAR_SECRET among them, which is the bearer token for the sidecar API one DNS name away.
site-builder, forty lines above in the same file, carries an explicit comment refusing exactly this and interpolating only the secret it needs. This restores that precedent. CLOUDFLARE_API_TOKEN is named explicitly; nothing else from .env enters the container. Verified by resolving the project and confirming SS_SIDECAR_SECRET is absent from the result.
P2 — the mount exposed private keys and every private repo
The service bound all of ~/srv/softserve read-only. That directory holds ssh/soft_serve_host_ed25519 and ssh/soft_serve_client_ed25519 (both private keys), soft_serve.db, meta.db, and all nine bare repos including the private ones — all readable by the same build code. read_only: true is not a defence when reading is the exposure.
Replaced with two narrow read-only binds: the site's bare repo and the trigger directory, both keyed off SITE_REPO / SITE_TRIGGER_FILE. PR #5's file:// guarantee (the deployer cannot publish an unpushed commit) is untouched — only the blast radius moves.
P5 — the trigger override moved the writer but not the watcher
The hook wrote to ${SITE_TRIGGER_FILE}; the supervisor watched a hardcoded path. Setting the documented override silently killed the fast path while every log line still looked healthy. One variable now drives both halves — verified by resolving with an override set and watching both move together.
P6 / P7 — runbook and stale references
SITE_DEPLOY_SRCbecomes install step 1, since until it is set no compose command in the runbook runs at all. It was step 4, after two steps that invoke compose.- Step 3 now states it is a prerequisite rather than a sanity check: the narrowed binds need
deploy-triggers/to exist first, or Docker invents it as an empty root-owned directory. - The stuck-lock section no longer points at
~/srv/stack/softserve-site-deploy-state/_data/. Named volumes do not live in the project directory, and under Docker Desktop for Mac they are not on the host filesystem at all. .pathunit andTriggerLimitBurstreferences removed from the hook header,touch_site_trigger, and hook.test.ts. Those have not existed since 0ff5e67 moved the scheduler into the container, and the hook is what an operator reads inside it.
Deferred
The lock's pid-file write window (a reclaim can fire between mkdir and the pid write) and the double ensure_checkout in main() are both in deploy.sh, which had uncommitted work in flight in the primary checkout when this branch was cut. Neither is affected by anything here.
Verification
make check green: shellcheck 8 files, sqlcheck, gocheck, tscheck 135 tests across 15 suites. Compose behaviour verified by resolving the project under several env combinations rather than by reading the YAML. Nothing here is in the code.fugl.dev request path, and site-deploy has still never been started on the host.
Refs: kanban #26
discussion
Russ T. Fugalcommented
Rebased onto c502f01 (PR #6). No conflicts — PR #6 is entirely in deploy.sh and supervisor.ts, this branch is entirely in compose.yaml, the hook, the runbook, and hook.test.ts. make check green after the rebase: 137 tests across 15 suites.
Russ T. Fugalcommented
Independent agent review of this branch found five issues. All five are folded in commit
dffc70b. Two were real regressions introduced by the narrowing in the first commit, and both are worth recording because they share a cause.The narrowing turned latent config drift into hard failure. The old wholesale
/soft-servemount was hiding two variables that had quietly stopped agreeing with each other:-
SITE_REPOdid not driveSITE_REPO_URL. The bind source already followed${SITE_REPO}; the URL was a hardcodedruss/codeliteral. SettingSITE_REPOalone gave you a container that mounted one repo and cloned from another. Under the wholesale mount every repo was mounted, so the hardcoded URL still resolved and nothing looked wrong — which is exactly why it survived. Now derived:file:///soft-serve/repos/${SITE_REPO:-russ/code}.git. -
SITE_TRIGGER_FILEmoved the env vars but not the mount. My previous commit claimed one variable drove writer and watcher together. True of the two env vars, false of the bind, which stayed hardcoded atdeploy-triggers/. Any override outside that directory left the supervisor watching an unmounted path —trigger dir absentforever, six-hour tick silently the only route, every other log line healthy. That is the same failure mode the commit set out to eliminate, reintroduced one layer down. Fixed withSITE_TRIGGER_DIRas a single knob driving bind source, bind target, and the defaultSITE_TRIGGER_FILEon both containers. -
create_host_path: falsewas missing on both new binds.ssh-deprecatedhas carried it for the same reason with a comment calling it load-bearing, thirty lines below where I added the narrow binds. Compose invents an empty directory at a missing bind source, and here that lands inside softserve's data root: a typo inSITE_REPOfabricatesrepos/<typo>.git/in the tree softserve serves;up site-deploybefore the site repo exists fabricates the real path as an empty directory. A service this branch presents as read-only least-privilege should not be able to corrupt the server's repo list by starting. Now it refuses to start instead. -
Stale cross-reference — the compose comment pointed at install step 2 for creating
deploy-triggers/, which the first commit renumbered to step 3. -
README overclaimed the missing-bind consequence — it described one outcome for both mounts. A missing repo bind means nothing to clone; a missing trigger bind costs only the fast path. Now distinguished, since the runbook uses this to justify step 3 being load-bearing.
Findings 1 and 3 are the same mistake twice: a precedent already in this file that I did not apply.
env_fileon site-builder in the original review,create_host_pathon ssh-deprecated here.Verified by resolving the project under each override rather than by reading YAML —
SITE_REPOmoves URL and bind together,SITE_TRIGGER_DIRmoves both env vars and both bind sides,create_host_path: falsesurvives into the resolved config.make checkgreen: 137 tests across 15 suites.The reviewer also confirmed several things I had asserted but not proven:
git clone/fetchoverfile://goes throughgit-upload-packand writes nothing on the source side, so the read-only bind is genuinely sufficient; the resolved environment contains noSS_SIDECAR_SECRET; and the:-sentinel does unbrickdocker compose configwithSITE_DEPLOY_SRCunset. It could not empirically testfs.watchpropagation across the bind on Docker Desktop VirtioFS without starting the container, so that remains reasoning rather than measurement — install step 6 is still the place it gets proven.-
5 files changed
This view needs a browser with declarative shadow DOM: Chrome 111, Safari 16.4, or Firefox 123. Read the source instead.
infra/site-deploy/README.md
29 unmodified lines303132333433343536371 unmodified line394041424344454647484950515253545556575859426061626311 unmodified lines75767760787980818283848562868764888990918 unmodified lines1001011027910310410510610 unmodified lines1171181199697989910010110212012110412212312412541 unmodified lines1671681691521701711721732 unmodified lines17617717816117918018118247 unmodified lines23023123221521623323423523623729 unmodified linesThe live host is a Mac Mini running Docker Desktop. The convergence layer runs entirely inside a container in the existing `~/srv/stack/softserve` compose project, so the host contribution is:
- Docker Desktop already running (the softserve/website/cloudflared containers already run here — this adds one more).- A Cloudflare API token with `Workers Scripts:Edit` scope for the `code-fugl-dev` script (see step 3 below). Stored in `.env`, never in the tree.- One extra env var in `.env`, `SITE_DEPLOY_SRC`, pointing at the absolute path of this directory in the operator's checkout. See "Sibling paths, and why not just `../site-deploy`" below.- A Cloudflare API token with `Workers Scripts:Edit` scope for the `code-fugl-dev` script (see step 4 below). Stored in `.env`, never in the tree.- One extra env var in `.env`, `SITE_DEPLOY_SRC`, pointing at the absolute path of this directory in the operator's checkout — see step 1, which must be done before any other compose command, and "Sibling paths, and why not just `../site-deploy`" below for why the variable exists at all.
There is no launchd, no systemd, and no host cron; `restart: unless-stopped` on the compose service handles boot recovery, and the 6h interval lives inside the container's supervisor process.
1 unmodified line
`[user]` marks a step that needs a human hand or an explicit go-ahead. Everything else is a `docker compose` invocation that any operator with shell access to the host can perform.
**Step 1 is not optional and not reorderable.** `~/srv/stack/softserve/compose.yaml` is a symlink into this repo's working tree, so the `site-deploy` service definition goes live on the host the moment it is checked out — before anything is installed. Compose interpolates the whole file before it selects a service, which means an unset variable in `site-deploy` is an error for `docker compose ps`, `logs`, and `restart softserve` too. `SITE_DEPLOY_SRC` therefore has to exist in `.env` before any compose command in this runbook will run. (The service definition uses `:-` defaults precisely so a missing value degrades to "`up site-deploy` fails" rather than "the whole project is unmanageable" — but the sentinel path is not a working build context, so step 1 is still required before step 5.)
### 1. `[user]` Point `SITE_DEPLOY_SRC` at this directory
The `site-deploy` compose service bind-mounts the scripts from `${SITE_DEPLOY_SRC}` into `/opt/site-deploy` and uses the same path as its build context. Set it to the absolute path of this directory in the operator's checkout:
```shprintf 'SITE_DEPLOY_SRC=%s\n' "$HOME/code/personal/fugl.dev/infra/site-deploy" >> ~/srv/stack/softserve/.env```
(Substitute the actual checkout path if it lives elsewhere.) Confirm the project parses again before continuing — this is also the recovery step if the stack is already locked out:
```shcd ~/srv/stack/softservedocker compose config --services# → softserve, meta-sidecar, backup, website, site-builder, site-deploy, ssh-deprecated, cloudflared```
### 1. Recreate the softserve container to pick up the updated hook### 2. Recreate the softserve container to pick up the updated hook
The `infra/softserve/hooks/post-receive` change and the compose `SITE_TRIGGER_FILE` env var ship with this PR. Once merged and pushed to `origin`, the hook file inside the running container is stale.
11 unmodified lines
Nothing else moves at this step; a push after this point writes an aggregate-only ISO timestamp to `~/srv/softserve/deploy-triggers/pending` on the host, but the supervisor container is not up yet, so nothing reacts.
### 2. Verify the trigger file mount### 3. Create and verify the trigger directory
The trigger path inside the softserve container maps to `~/srv/softserve/deploy-triggers/pending` on the host through the existing `~/srv/softserve` bind mount, and *also* into the site-deploy container as `/soft-serve/deploy-triggers/pending`.
site-deploy does **not** mount `~/srv/softserve` wholesale. It gets two narrow read-only binds — `repos/${SITE_REPO}.git` and `${SITE_TRIGGER_DIR}/` — and nothing else. Both follow the same variables the rest of the stack uses, so re-pointing either is one line in `.env` and the mount, the URL, and the hook's write path stay in agreement. Overriding `SITE_TRIGGER_FILE` directly still works for renaming the leaf file, but it must stay inside `SITE_TRIGGER_DIR` — the narrow bind means a path outside it is not mounted on the reading side at all. That is deliberate and it is a security boundary, not tidiness: `~/srv/softserve` also holds `ssh/soft_serve_host_ed25519` (soft-serve's private host key), `soft_serve.db`, `meta.db`, and every bare repo on the server including private ones, while this container is the one that runs `bun install` over the site's full dependency tree. `read_only: true` does not help against that — reading is the exposure. Keep the binds narrow when editing the service.
Because they are binds and not volumes, **both source paths must exist on the host before `docker compose up`.** Both carry `create_host_path: false`, so a missing source makes the container refuse to start rather than letting Compose invent an empty directory inside softserve's data root — a fabricated `repos/<typo>.git/` would show up in the tree softserve actually serves. `repos/russ/code.git` already exists; `deploy-triggers/` does not until something creates it, which is what the verification below does — so this step is a prerequisite for step 5, not just a sanity check.
The trigger path inside the softserve container maps to `~/srv/softserve/deploy-triggers/pending` on the host through the existing `~/srv/softserve` bind mount, and *also* into the site-deploy container as `/soft-serve/deploy-triggers/pending` through a second read-only bind of the same host path. The same read-only `~/srv/softserve` bind into site-deploy also exposes softserve's bare repos at `/soft-serve/repos/<org>/<name>.git` — the site-deploy pass reads the site's tip directly from `/soft-serve/repos/russ/code.git` over a `file:///` URL, not by fetching over HTTP from the softserve container.The two failures look different, which is useful when diagnosing: a missing repo bind means the deploy has nothing to clone, while a missing trigger bind leaves the supervisor logging `trigger dir absent; watching for it` — the startup pass and the 6h tick still run, so only the fast path is lost.
Two structural properties fall out of that choice, both deliberate:The pass reads the site's tip directly from `/soft-serve/repos/russ/code.git` over a `file:///` URL, not by fetching over HTTP from the softserve container. Two structural properties fall out of that choice, both deliberate:
- **The deployer cannot publish an unpushed commit.** The bare repo contains exactly what softserve has accepted; there is no working tree side-channel. This closes the failure mode of an operator running `bun run site:deploy` from a laptop while `main` is ahead of `origin/main` — the exact scenario that shipped a colophon fix that was not on the remote (kanban #26).- **A softserve outage does not stop the deployer.** With HTTP fetches, a `docker compose stop softserve` in the middle of a six-hour tick would fail the pass; with the file:// read, the deploy still resolves to whatever softserve's on-disk bare last had. Anonymous clone requirements are moot too — the bare files on disk are readable regardless of softserve's ACL configuration. `depends_on: [softserve]` in the compose file governs startup ordering only, not runtime.8 unmodified lines
If the file on the host does not contain that line, the bind mount is not what this design assumes — stop and investigate before continuing.
### 3. `[user]` Create the Cloudflare API token### 4. `[user]` Create the Cloudflare API token
Nothing in this repository issues API tokens; that has to be done in the Cloudflare dashboard with the account that owns the `code-fugl-dev` Worker.
10 unmodified lines
Edit the placeholder with the real value. Do not paste terminal history to a shared channel afterwards.
### 4. `[user]` Point `SITE_DEPLOY_SRC` at this directory
The `site-deploy` compose service bind-mounts the script from `${SITE_DEPLOY_SRC}` into `/opt/site-deploy` inside the container. Set it to the absolute path of this directory in the operator's checkout:
```shprintf 'SITE_DEPLOY_SRC=%s\n' "$HOME/code/personal/fugl.dev/infra/site-deploy" >> ~/srv/stack/softserve/.env```Note that the site-deploy service reads this one variable by name (`CLOUDFLARE_API_TOKEN=${CLOUDFLARE_API_TOKEN:-}`) rather than loading `.env` wholesale. That is why the rest of what lives in `.env` — the sidecar bearer token in particular — does not end up in the environment of a container that executes the site's third-party install scripts. If you add a secret this service genuinely needs, add it by name; do not reintroduce `env_file:`.
(Substitute the actual checkout path if it lives elsewhere.) Compose fails at parse time with a clear message if this is unset — no silent misconfiguration.If the variable is unset the project still parses and every other service stays manageable; the failure surfaces as a wrangler auth error in `docker compose logs site-deploy` on the first pass.
### 5. Start the site-deploy service
41 unmodified lines
If a push does not appear to have triggered a deploy:
1. Confirm the file appeared: `cat ~/srv/softserve/deploy-triggers/pending | tail -3`. Missing → the hook did not run in the softserve container (unlikely) or the container is stale from step 1.1. Confirm the file appeared: `cat ~/srv/softserve/deploy-triggers/pending | tail -3`. Missing → the hook did not run in the softserve container (unlikely) or the container is stale from step 2.2. Confirm the site-deploy container is running: `docker compose ps site-deploy`. Not running → `docker compose up -d site-deploy` and revisit the logs.3. Confirm the watcher is active: `docker compose logs --tail=200 site-deploy | grep 'watching trigger file'`. Missing → the trigger directory did not exist at container start; the supervisor retries attaching the watcher every 60s (`watchRetryMs` default), so a push landing in the first-boot window before that retry fires is picked up only by the 6h fallback tick. `[fugl-site-deploy] trigger dir absent; watching for it (6h tick still active)` is the log line that says the retry loop is active. Once you see `watching trigger file`, subsequent pushes take the fast path.4. If all three pass and no deploy fired, wait for the next 6h tick — the fingerprint gate will decide whether there is anything to do. A silent no-op is the correct outcome for a push that changed nothing observable in the public corpus.2 unmodified lines
### Stuck lock — no passes are running and none can start
`deploy.sh` serialises passes through a portable `mkdir` lock at `/state/deploy.lock.d/` (`~/srv/stack/softserve-site-deploy-state/_data/deploy.lock.d/` on the host, via the `site-deploy-state` named volume). The `pid` file inside contains three whitespace-separated fields:`deploy.sh` serialises passes through a portable `mkdir` lock at `/state/deploy.lock.d/`, inside the `site-deploy-state` named volume. Reach it with `docker compose exec site-deploy`, as the commands below do — a named volume is not a path under `~/srv/stack/`, and on Docker Desktop for Mac it is not on the host filesystem at all (it lives inside the Linux VM). `docker volume inspect softserve_site-deploy-state` will report a `Mountpoint`, but that path is only meaningful inside the VM. The `pid` file inside contains three whitespace-separated fields:
```<pid> <boot-cookie> <start-time>47 unmodified lines
Consolidated from the install section, so the operator can scan before starting:
1. **Create the Cloudflare API token** (step 3) — a browser step, done once per rotation. `CLOUDFLARE_API_TOKEN` in `~/srv/stack/softserve/.env`.2. **Set `SITE_DEPLOY_SRC`** in the same `.env` — one line, absolute path to this directory.1. **Set `SITE_DEPLOY_SRC`** in `~/srv/stack/softserve/.env` (step 1) — one line, absolute path to this directory. Do this first: until it is set, `docker compose` cannot manage *any* service in the project, including the ones already running.2. **Create the Cloudflare API token** (step 4) — a browser step, done once per rotation. `CLOUDFLARE_API_TOKEN` in the same `.env`.3. **Rollback decisions** — any `docker compose stop`, `rm`, or token revocation is the operator's call.
Everything else is `docker compose up -d`.infra/site-deploy/test/hook.test.ts
163 unmodified lines164165166167168169167168169170171172163 unmodified lines
test("touches the trigger exactly once for a multi-ref push", () => { // A push carrying five branches at once (including the site branch) must // produce one trigger, not five. The file-append semantics would fire // the .path unit's TriggerLimitBurst repeatedly and delay the actual // deploy. // produce one trigger, not five. The supervisor coalesces anyway, but // deduping at the source keeps the trigger file readable as a push log // and keeps the watcher from firing five times for one push. const trigger = join(tmp, "deploy-triggers", "pending"); const result = runHook( {infra/softserve/compose.yaml
20 unmodified lines212223242526272829302425262728293031323334353637217 unmodified lines25525625725825926026126226326426526626726826925627027127227312 unmodified lines2862872882892902912922932942752952962972782982993003013023033043053063073083093103113123133142842852862873153163173183193203213223233243253263273283293303313323333343353362 unmodified lines33934034129634234334429930034534634734834935035135235335435535635735835936036136236336436536636736836937037137237337437537637730230337837938038138238338438520 unmodified lines - SS_BUILDER_SECRET=${SS_BUILDER_SECRET-} - SITE_REPO=${SITE_REPO:-russ/code} - SITE_BRANCH=${SITE_BRANCH:-main} # Where the hook drops a trigger file for the host-side deploy service # (fugl-site-deploy.path — see infra/site-deploy/). This path is # in-container; the same inode is at ~/srv/softserve/deploy-triggers/ # on the host through the /soft-serve bind mount below. Overridable so a # future host layout that moves the trigger elsewhere does not need a # hook edit. - SITE_TRIGGER_FILE=${SITE_TRIGGER_FILE:-/soft-serve/deploy-triggers/pending} # Where the hook drops a trigger file for the site-deploy service below, # whose supervisor fs.watches it. This path is in-container; the same # inode is at ~/srv/softserve/${SITE_TRIGGER_DIR}/ on the host through # the /soft-serve bind mount below, and site-deploy narrow-binds that # same host directory. # # Both halves derive from SITE_TRIGGER_DIR, so moving the trigger is one # variable and the writer, the watcher, and the mount stay in agreement. # site-deploy's mount is narrow, so a path outside that directory is not # merely unconventional — it is unmounted on the reading side. - SITE_TRIGGER_FILE=${SITE_TRIGGER_FILE:-/soft-serve/${SITE_TRIGGER_DIR:-deploy-triggers}/pending} # Public URLs shown in TUI clone hints. Env vars override config.yaml, # so the hostname lives HERE (or in .env) — not in the container's # config.yaml. Without these Soft Serve advertises ssh://localhost:23231.217 unmodified lines # # The supervisor/deploy scripts stay bind-mounted at runtime, so a script # edit is `docker compose restart site-deploy` and does not need `--build`. # # SITE_DEPLOY_SRC uses `:-` with a self-describing sentinel path, NOT `:?`. # Compose interpolates the ENTIRE file before it selects services, so a # `:?` here fails every command for every service — `docker compose ps`, # `logs`, `restart softserve`, all of it — on any host where the variable # is unset. This compose.yaml is symlinked live into ~/srv/stack/softserve/, # so that is not hypothetical: it locked the operator out of the running # stack the moment it merged. With the sentinel, the rest of the project # stays manageable and only `up site-deploy` fails, printing the sentinel # path (which says what to do) as the missing build context. site-deploy: build: context: ${SITE_DEPLOY_SRC:?set SITE_DEPLOY_SRC in .env to the absolute path of infra/site-deploy} context: ${SITE_DEPLOY_SRC:-/SITE_DEPLOY_SRC-unset-see-infra-site-deploy-README} dockerfile: Dockerfile restart: unless-stopped # bun runs as PID 1 directly; the Dockerfile has already installed the12 unmodified lines # The `file://` scheme, not the bare path, disables git's hardlink-based # local-clone shortcut so softserve's next repack cannot mutate our # checkout's object DB. # Derived from SITE_REPO so a repo override moves the URL and the bind # mount below together. Hardcoding the slug here while the bind source # followed ${SITE_REPO} meant an operator who set SITE_REPO alone got a # container that mounted one repo and cloned from another — which the # old wholesale /soft-serve mount silently absorbed, and the narrow # binds turn into `does not appear to be a git repository`. - SITE_REPO_URL=${SITE_REPO_URL:-file:///soft-serve/repos/russ/code.git} - SITE_REPO_URL=${SITE_REPO_URL:-file:///soft-serve/repos/${SITE_REPO:-russ/code}.git} - SITE_BRANCH=${SITE_BRANCH:-main} - STATE_DIR=/state - TRIGGER_FILE=/soft-serve/deploy-triggers/pending # SITE_TRIGGER_DIR is the single knob: it drives the bind source below, # the bind target, and the default trigger path on BOTH containers (the # softserve service's SITE_TRIGGER_FILE is derived from it too). The # hook is the writer, this supervisor is the watcher, and a narrow bind # means the directory has to be mounted, not merely named — setting a # path the mount does not cover leaves the supervisor logging # `trigger dir absent` forever while every other line looks healthy. # # SITE_TRIGGER_FILE may still be overridden directly to rename the leaf, # but it MUST stay inside SITE_TRIGGER_DIR. Point it outside and the # watcher is watching an unmounted path. - TRIGGER_FILE=${SITE_TRIGGER_FILE:-/soft-serve/${SITE_TRIGGER_DIR:-deploy-triggers}/pending} - DEPLOY_SCRIPT=/opt/site-deploy/deploy.sh - LOG_TAG=fugl-site-deploy # Six hours in ms. Overridable so a smoke test can shorten it without # a compose edit; production stays at the default. - INTERVAL_MS=${SITE_DEPLOY_INTERVAL_MS:-21600000} env_file: # CLOUDFLARE_API_TOKEN, scoped to the code-fugl-dev Worker. Never commit # this file. See ../site-deploy/README.md for the token creation step. - .env # Scoped to the code-fugl-dev Worker. Interpolated from .env, which is # never committed. See ../site-deploy/README.md for the creation step. # # `:-` and not `:?` for the same file-wide reason as SITE_DEPLOY_SRC # above: a required-variable marker anywhere in this file locks the # operator out of every other service. Unset is therefore allowed at # parse time and surfaces where it belongs — the pass runs and wrangler # exits non-zero with an auth error in `docker compose logs # site-deploy`. Loud, local to the service, and not a stack-wide brick. - CLOUDFLARE_API_TOKEN=${CLOUDFLARE_API_TOKEN:-} # Deliberately no `env_file:` — for exactly the reason spelled out on # site-builder above, which applies here word for word. This service runs # third-party build code: every install script and every plugin in the # site's dependency tree, on a network where softserve and meta-sidecar # are both reachable by name. `env_file: .env` handed all of that code # the whole stack's credentials — SS_SIDECAR_SECRET among them, which is # the bearer token for the sidecar API one DNS name away. `environment:` # above interpolates the one secret this service needs from that same # file without loading the rest. volumes: # The script (bind mount, not build context). SITE_DEPLOY_SRC is set in # .env to the absolute path of infra/site-deploy — required because2 unmodified lines # documented in kanban #2. An env-provided absolute path avoids the trap # without forcing a new sibling symlink in ~/srv/stack. - type: bind source: ${SITE_DEPLOY_SRC:?set SITE_DEPLOY_SRC in .env to the absolute path of infra/site-deploy} source: ${SITE_DEPLOY_SRC:-/SITE_DEPLOY_SRC-unset-see-infra-site-deploy-README} target: /opt/site-deploy read_only: true # The softserve mount — read-only from this side. softserve is the sole # writer of the trigger file; we only watch it. # Two narrow binds, NOT the whole ~/srv/softserve tree. # # That directory holds soft-serve's private host and client keys # (ssh/soft_serve_host_ed25519 — the same file ssh-deprecated mounts # below), soft-serve.db, meta.db, and every bare repo on the server # including the private ones. Mounting all of it put that within reach # of `bun install`, which executes arbitrary install scripts from the # site's dependency tree. `read_only: true` is no defence when reading # IS the exposure — it stops the container corrupting soft-serve, not # a postinstall script exfiltrating a host key. # # The pass needs exactly two things, so it gets exactly two things: the # bare repo whose tip it builds, and the trigger file it watches. Both # follow SITE_REPO/SITE_TRIGGER_FILE so there is one place to re-point. # # create_host_path:false on both, for the same load-bearing reason it is # set on ssh-deprecated below: Compose otherwise invents an empty # DIRECTORY at a missing source. Here that is worse than a crash-loop, # because the invented directory lands inside softserve's own data root. # A typo in SITE_REPO would fabricate `repos/<typo>.git/` in the tree # softserve serves, and running `up site-deploy` before the site repo # has ever been pushed would fabricate the real repo's path as an empty # directory — a container this PR presents as read-only least-privilege # quietly corrupting the server's repo list. Refusing to start is the # correct failure, and install step 3 in ../site-deploy/README.md is # what creates deploy-triggers/ beforehand. - type: bind source: ~/srv/softserve/repos/${SITE_REPO:-russ/code}.git target: /soft-serve/repos/${SITE_REPO:-russ/code}.git read_only: true bind: create_host_path: false - type: bind source: ~/srv/softserve target: /soft-serve source: ~/srv/softserve/${SITE_TRIGGER_DIR:-deploy-triggers} target: /soft-serve/${SITE_TRIGGER_DIR:-deploy-triggers} read_only: true bind: create_host_path: false # Persistent checkout + node_modules + fingerprint state. A wipe costs # one slow rebuild plus a first-run clone; nothing else depends on it, # which is why it is a named volume and not under ~/srv.infra/softserve/hooks/post-receive
8 unmodified lines9101112131412131415161718192021221920212223242526272820 unmodified lines4950514950525354555617 unmodified lines74757674777879803 unmodified lines8485868487888990918792939495968 unmodified lines# refs/meta/config → meta-sidecar /sync, so the SQLite cache picks# up the repo's metadata without waiting for its# poll loop.# refs/heads/$SITE_BRANCH → touch the shared trigger file so the host-side# on $SITE_REPO fugl-site-deploy.path unit wakes the deploy# service. Every other repo and every other# refs/heads/$SITE_BRANCH → touch the shared trigger file, which the# on $SITE_REPO site-deploy container's supervisor is# fs.watching. Every other repo and every other# ref is a no-op — meta refs never trigger a# site deploy, and neither does a push to any# repository other than $SITE_REPO.## The site-builder container is retired. The deploy runs on the host now# (bun + wrangler + a fetch of the site source), so the hook cannot invoke# it directly from in-container; instead it signals across the softserve# bind mount and systemd on the host takes it from there.# The site-builder container is retired. The deploy now runs in the# site-deploy container (bun + wrangler over a read-only bind of this# repo's bare directory), which has no HTTP endpoint for us to call;# instead we signal across the shared bind mount and its supervisor# process takes it from there. Nothing on the host schedules anything —# there is no systemd unit, no launchd job, and no cron. See# ../../site-deploy/README.md.## Deployment:# Copied to ~/srv/softserve/hooks/post-receive inside the Soft Serve container.20 unmodified lines# A post-receive hook must NEVER abort a push, and must never HANG one either:# every notify below is failure-tolerant and time-limited, because the fallback# for each is a poll loop (or, for the site, the six-hour timer) in the# service being called.# for each is a poll loop (or, for the site, the supervisor's six-hour# reconciliation tick) in the service being called.readonly NOTIFY_TIMEOUT=5readonly site_repo="${SITE_REPO:-russ/code}"17 unmodified lines set -e}# touch_site_trigger: signal the host-side deploy service that the site's# touch_site_trigger: signal the site-deploy supervisor that the site's# source has moved.## Only reached when the push we are handling is (a) on the configured site3 unmodified lines## The file is appended, not overwritten, so a rapid double-push shows up as# two entries in the same file rather than being lost to a rename race. The# path unit's TriggerLimitBurst is what caps this.# supervisor's in-process coalescing is what caps the resulting deploys: a# trigger arriving mid-pass sets a pending flag and produces exactly one# follow-up pass, however many writes landed.## The write is failure-tolerant: an atomic append that could not land does# not fail the git push, and the six-hour timer catches whatever we lose.# not fail the git push, and the supervisor's six-hour reconciliation tick# catches whatever we lose.# We deliberately do not log any repo name here — we're already inside a# filter that only lets the SITE_REPO past, and the sidecar branch above is# the only place a private slug can currently reach stderr.prs.snapshot.json
123456789101112131415161718192021222324252627282923{ "open": [ { "number": 7, "branch": "pr/site-deploy-least-privilege", "base": "main", "title": "Unbrick the compose project, narrow what the site builder can reach", "body": "Post-merge follow-up to PR #5. Four of the review findings, plus one that landing PR #5 caused on the live host. All five are in compose.yaml, the hook, and the runbook; nothing here touches deploy.sh.\n\n## P0 -- the live stack is currently unmanageable\n\n`~/srv/stack/softserve/compose.yaml` is a symlink into this repo's working tree, so the site-deploy service definition went live on the host the moment main was checked out -- before a single install step ran. Its build context used `${SITE_DEPLOY_SRC:?...}`, and Compose interpolates the whole file before selecting a service. A required-variable marker on a service nobody has installed yet is therefore an error for every command on every service.\n\nConfirmed against the live project:\n\n $ cd ~/srv/stack/softserve && docker compose config --services\n error while interpolating services.site-deploy.build.context:\n required variable SITE_DEPLOY_SRC is missing a value\n\nThe containers keep running -- softserve, meta-sidecar, cloudflared and ssh-deprecated are all still up -- but `ps`, `logs`, `restart`, and `stop` all fail. Also reproduced on a minimal two-service file to confirm it is interpolation order and not something specific to this project.\n\nBoth `:?` markers on the service become `:-` with a self-describing sentinel path. A missing value now degrades to \"`up site-deploy` fails, and the error names the variable\" instead of \"the whole project is bricked\". Verified: the project parses with `SITE_DEPLOY_SRC` unset.\n\n**Immediate unblock without this PR:** add `SITE_DEPLOY_SRC` to `~/srv/stack/softserve/.env`. That is install step 1 here anyway.\n\nThe same footgun already exists on cloudflared (`CF_TUNNEL_CREDS_FILE:?`). It is set on the live host, so it is latent, and cloudflared genuinely cannot start without it -- deliberately left alone rather than widening this change. Worth a card.\n\n## P1 -- `env_file: .env` gave third-party build code the stack's credentials\n\nsite-deploy is the service that runs \\`bun install\\` over the site's entire dependency tree, on a network where softserve and meta-sidecar answer to their own names. \\`env_file: .env\\` put every variable in that file into its environment -- \\`SS_SIDECAR_SECRET\\` among them, which is the bearer token for the sidecar API one DNS name away.\n\nsite-builder, forty lines above in the same file, carries an explicit comment refusing exactly this and interpolating only the secret it needs. This restores that precedent. \\`CLOUDFLARE_API_TOKEN\\` is named explicitly; nothing else from \\`.env\\` enters the container. Verified by resolving the project and confirming \\`SS_SIDECAR_SECRET\\` is absent from the result.\n\n## P2 -- the mount exposed private keys and every private repo\n\nThe service bound all of \\`~/srv/softserve\\` read-only. That directory holds \\`ssh/soft_serve_host_ed25519\\` and \\`ssh/soft_serve_client_ed25519\\` (both private keys), \\`soft_serve.db\\`, \\`meta.db\\`, and all 9 bare repos including private ones -- all readable by the same build code. \\`read_only: true\\` is not a defence when reading is the exposure.\n\nReplaced with two narrow read-only binds: the site's bare repo and the trigger directory, both keyed off \\`SITE_REPO\\`/\\`SITE_TRIGGER_FILE\\`. PR #5's file:// guarantee (the deployer cannot publish an unpushed commit) is untouched -- only the blast radius moves.\n\n## P5 -- the trigger override moved the writer but not the watcher\n\nThe hook wrote to \\`\\${SITE_TRIGGER_FILE}\\`; the supervisor watched a hardcoded path. Setting the documented override silently killed the fast path while every log line still looked healthy. One variable now drives both halves -- verified by resolving with an override set and seeing both move together.\n\n## P6/P7 -- runbook and stale references\n\n- \\`SITE_DEPLOY_SRC\\` becomes install step 1, since until it is set no compose command in the runbook runs at all. It was step 4, after two steps that invoke compose.\n- Step 3 now states it is a prerequisite rather than a sanity check: the narrowed binds need \\`deploy-triggers/\\` to exist first, or Docker invents it as an empty root-owned directory.\n- The stuck-lock section no longer points at \\`~/srv/stack/softserve-site-deploy-state/_data/\\`. Named volumes do not live in the project directory, and under Docker Desktop for Mac they are not on the host filesystem at all.\n- \\`.path\\` unit and \\`TriggerLimitBurst\\` references removed from the hook header, \\`touch_site_trigger\\`, and hook.test.ts. Those have not existed since 0ff5e67; the hook is what an operator reads inside the container.\n\n## Deferred\n\nThe lock's pid-file write window (a reclaim can fire between \\`mkdir\\` and the pid write) and the double \\`ensure_checkout\\` in \\`main()\\` are both in deploy.sh, which has uncommitted work in flight in the primary checkout. Neither is affected by anything here.\n\n## Verification\n\n\\`make check\\` green: shellcheck 8 files, sqlcheck, gocheck, tscheck 135 tests across 15 suites. Compose behaviour verified by resolving the project under several env combinations rather than by reading the YAML. Nothing here is in the code.fugl.dev request path, and site-deploy has still never been started on the host.\n\nRefs: kanban #26", "status": "open", "openedBy": "Russ T. Fugal", "createdAt": "2026-08-06T08:19:13.954Z", "headSha": "0c0d192d5df198026acbdd6e03eea29652f0b6f0", "mergedAt": null, "mergeSha": null, "events": [ { "at": "2026-08-06T08:19:13.954Z", "actor": "Russ T. Fugal", "type": "opened" }, { "at": "2026-08-06T08:20:27.340Z", "actor": "Russ T. Fugal", "type": "comment", "body": "Closing in favour of a re-open from the same branch and commit: the body I opened this with went through a shell heredoc that left 50 literal backslash-backticks in it, and PR bodies here are world-readable and get embedded into the merge commit. No content or code change — same branch, same commit, readable body." } ] } ] "open": []}This view needs a browser with declarative shadow DOM: Chrome 111, Safari 16.4, or Firefox 123. Read the source instead.
clone
$ git clone https://git.fugl.dev/fugl.devanonymous, no account$ git clone ssh://git.fugl.dev/fugl.devneeds the bastion ProxyCommand