Files
granthi-sync/README.md
T
claude 49ab2a7408 feat: sign in from any computer, no private network needed
granthi-link is now public at https://granthi-link.shre.ai (cloudflared,
origin still tailnet-only), so the client defaults there instead of a tailnet
IP. Internal machines pass --server or GRANTHI_LINK_SERVER.

Two rollout traps recorded in the README: a third-level hostname
(link.granthi.shre.ai) fails TLS because Cloudflare Universal SSL covers
shre.ai and *.shre.ai only; and exposure REQUIRES trust_forwarded_for with
the tunnel as the sole trusted proxy, or every request looks like the tunnel
and one abuser spends everyone's rate budget.

178 tests.
2026-08-23 12:33:19 -04:00

580 lines
33 KiB
Markdown

# granthi-sync v1.2
The signup → download → link-folders → cloud product spine for the Granthi
forge, tested against the BETA forge (granthi-beta.shre.ai). Python 3 stdlib +
git CLI only — same portability heritage as the estate's `gitea_sync.py` mesh.
```
┌──────────────┐ device flow ┌─────────────────┐
│ granthi-sync │ ───────────────▶ │ shre-id Zitadel │
│ (client, │ ◀─────────────── │ id.shre.ai │
│ Mac/laptop) │ access token └─────────────────┘
│ │
│ │ POST /v1/link {zitadel_access_token, device_name}
│ │ ───────────────▶ ┌───────────────────────────────┐
│ │ ◀─────────────── │ granthi-link :3042 │
│ │ {login, token} │ (granthi VPS, tailnet-only) │
│ │ │ · userinfo validation │
│ │ POST /v1/repos │ · ensure Gitea user (admin) │
│ │ ───────────────▶ │ · mint scoped user token │
│ │ └──────────────┬────────────────┘
│ │ git push/fetch (user token │ admin API
│ │ via credential helper) ▼
│ │ ───────────────▶ ┌───────────────────────────────┐
└──────────────┘ │ BETA forge :3041 │
│ granthi-beta.shre.ai │
└───────────────────────────────┘
```
## Quickstart (invited user)
You need a shre-id account — an operator creates it; there is no open signup
(see "Invite-only story"). You do **not** need to be on any private network:
the provisioning service answers at `https://granthi-link.shre.ai`, which is
the client's default server. Internal machines can still pass
`--server http://100.111.127.127:3042` to reach it over the tailnet.
```sh
git clone https://granthi.shre.ai/nirpa/granthi-sync.git
cd granthi-sync
# 1. Link this device. Prints a URL + code; approve it in a browser within
# 5 minutes. Creates your forge account and stores a scoped token in
# ~/.granthi-sync/config.json (0600). You never see a forge password.
./bin/granthi-sync link
# 2. See what is already yours on the forge.
./bin/granthi-sync list
# 3. Either pull an existing repo down...
./bin/granthi-sync get <repo> # or <owner>/<repo>, --into DIR
# 3b. ...or push a local folder up. It becomes a private repo.
./bin/granthi-sync add ~/work/notes
# 3c. ...or pull down everything this account is allowed to see.
./bin/granthi-sync get --all --into ~/granthi
# 4. Keep everything synced. Autocommits, ff-pulls, pushes; skips anything
# that has diverged rather than merging or forcing.
./bin/granthi-sync watch # --once for a single pass
./bin/granthi-sync status # what is linked, mode, last sync
# 5. Go back to how a folder looked at some point in time.
./bin/granthi-sync snapshots ~/work/notes
./bin/granthi-sync restore ~/work/notes --at 20260823T142530Z
```
Run `watch` as a background daemon on macOS with
`client/launchd/ai.granthi.sync.plist` (edit the script path, then
`launchctl bootstrap gui/$UID <plist>`).
**If the device code expires** (5 minutes, unapproved), nothing is created —
no account, no token, no partial state. Just run `link` again.
**Merging is a forge action, not a client one.** `watch` deliberately refuses
to merge; when a folder shows `DIVERGED` in `status`, resolve it in git or on
the forge web UI. The client will never force or auto-merge your work.
## Two modes, because two very different folders ask for this
A folder people sync is either *their documents* or *their git project*, and
the correct behaviour is opposite in each case. Each linked folder therefore
carries a `mode`.
| | `mirror` | `snapshot` |
|---|---|---|
| chosen for | a plain folder `add` turned into a repo | a folder that was already a git repo, and anything `get` clones |
| commits on your behalf | yes, `sync: <ISO ts>` | **never** |
| where work lands | the branch | `refs/granthi-backup/<device>/<ts>` |
| a restore point is | every commit | every snapshot |
`snapshot` mode is what "the work may not be committed, but it is still
backed up" means in git terms. Each pass loads a scratch index from HEAD,
stages the working tree into *that* index, writes a tree, and commits it with
`commit-tree`. HEAD, your index, your stash and every file on disk are
untouched — you can be mid-rebase with a dirty tree and the backup still
records exactly what is on the disk right now. The user's history stays the
user's.
Why a custom ref namespace: verified on the beta forge (Gitea 1.27.2) that
`refs/granthi-backup/...` is accepted, is readable through `ls-remote`, and
does **not** appear in the branch list. Under `refs/heads` a machine taking a
backup every 30 seconds would bury the branches a person actually made.
Snapshots are parented on HEAD and deliberately **not** chained to the
previous snapshot: chaining would keep every old snapshot reachable from the
newest, so pruning a ref would free nothing and retention would be
decorative.
**Retention** (or 30-second backups become a disk leak nobody can navigate):
everything is kept for 24 h, then thinned to hourly for 7 days, then daily.
Pruning runs at most hourly, per device, and only over that device's own
refs. A ref whose timestamp this version cannot parse is **kept** — deleting
the unrecognised is how a backup system loses the one thing someone needed.
**Restore never writes over the working tree.** `restore` materialises a
restore point into a *new* directory and refuses a non-empty destination.
Someone restoring a backup is already having a bad day; overwriting the files
they still have would make the recovery tool the second disaster.
## The credential helper must be the ONLY helper (found by live QA)
`credential.helper` is a list that accumulates across system, global and repo
config, and git asks every helper in it. A stock mac already has two —
`osxkeychain` from Xcode's gitconfig, and `store` from many people's
`~/.gitconfig` — and they lose in both directions:
* **reading:** a stale entry for the forge host answers before our helper, so
pushes fail `remote: Failed to authenticate user` long after the token was
rotated, and nothing in this tool's config explains why. This is exactly
how the first live-QA run failed;
* **writing:** git calls `approve` on every helper after a successful auth,
so `store` copies the forge token into `~/.git-credentials` **in
plaintext**. Keeping the token in a 0600 file and out of remote URLs buys
nothing if git then hands it to a plaintext store.
So `install_credential_helper` (and the `git clone` in `get`) sets an **empty**
`credential.helper` first, which resets the inherited list, then adds ours.
Exactly one helper serves this repo.
Corollary worth remembering: a token embedded in a remote URL gets saved by
`store` on first use. During QA a verification clone with a URL-embedded
token re-created the very entry that had just been cleaned out. That is the
whole reason this client passes tokens through a helper and never a URL.
## Device identity
`link` mints a uuid on first run and persists it in `~/.granthi-sync/config.json`
as `device_id`, and sends it to `/v1/link`. Hostnames are neither stable
(people rename laptops) nor unique (every new mac is "Mac mini"), so a
hostname cannot key a backup ref or a device registry — two machines would
overwrite each other's snapshots. The service-side device registry is the
next phase; the client leads so the id already exists when it lands.
## What "add the computer to the network" means — and does not
The onboarding shape is: download → login → **the device is federated to the
account** → the device can reach its repos.
The middle step is a *device registration*, not a network membership. Those
sound like one step and must not be built as one: this estate's tailnet is a
single flat private network carrying the granthi VPS, aros-vps, the Shadow
box and the Mac. Putting a customer's laptop on it to let them sync a folder
would hand that laptop L3 reach to every piece of infrastructure we run.
So:
* **Our own machines** may join the tailnet — that is an operator action with
an operator's judgement behind it.
* **Customer devices never do.** Their transport is public HTTPS to
granthi-link and the forge through cloudflared. **Done 2026-08-23:**
`https://granthi-link.shre.ai` fronts `:3042` on the `pulse-granthi-edge`
tunnel, so a new computer signs itself in with one command and never
touches the private network.
Two things that bit during that rollout, worth not rediscovering:
`link.granthi.shre.ai` fails TLS — Cloudflare's Universal SSL covers
`shre.ai` and `*.shre.ai`, **not** a third-level `*.granthi.shre.ai`, so
the hostname has to be second-level. And exposure REQUIRES
`trust_forwarded_for: true` with `trusted_proxies: ["100.107.37.98/32"]`
in the same change: behind the tunnel every request otherwise looks like
the tunnel itself, and one abuser would spend everybody's rate budget.
## Components
### `server/granthi_link.py` — provisioning service (granthi VPS)
* `GET /health`
* `POST /v1/link {zitadel_access_token, device_name}` → validates the token
against `https://id.shre.ai/oidc/v1/userinfo`, applies the **identity
binding rules** (below), mints a token scoped
`write:repository,write:user`, returns
`{gitea_base, login, token, token_name}`. POST bodies are capped at
**64 KB** (413 beyond; missing `Content-Length` → 411, invalid → 400).
* `POST /v1/repos {token, name, private}` → creates the user repo with the
USER token; clone/html URLs are rebased onto `public_gitea_base` because the
container `ROOT_URL` (https://granthi-beta.shre.ai) does not resolve for
tailnet-only clients.
Deployment: `/opt/granthi-link/{granthi_link.py,config.json,state.json}` +
systemd unit `granthi-link.service`; binds `127.0.0.1:3042` **and**
`100.111.127.127:3042` (tailnet). **Not publicly exposed** — see promotion
window. The service **refuses to start** (exit 2) unless `config.json` is
mode 0600/0400 and owned by the user it runs as — the config carries the
forge admin password, so permissive perms fail closed, not open.
#### Rate limiting
Sliding-window, per `(route, client)`, **on by default**`/v1/link` creates
accounts and mints tokens, so unlimited has to be a deliberate config act,
never an omission. Defaults: `/v1/link` 5/hour, `/v1/repos` 60/hour. `GET
/health` is never limited. Over the limit → **429** with a `Retry-After`
header, decided *before* the body is read so an abusive caller costs nothing.
```json
"rate_limit": {
"enabled": true,
"trust_forwarded_for": false,
"trusted_proxies": [],
"rules": {"/v1/link": [5, 3600], "/v1/repos": [60, 3600]}
}
```
* Anything malformed **refuses startup** rather than silently meaning
unlimited: a bad rule, a non-boolean `enabled`/`trust_forwarded_for` (JSON
`null`, `0`, or the *string* `"false"` — every non-empty string is truthy),
or a `rate_limit` that is not an object. `[0, N]` disables an endpoint
outright.
* Denied requests are *not* recorded, so a client that keeps hammering cannot
push its own window forward and lock itself out permanently.
* State is an in-process dict behind a lock — granthi-link is one
`ThreadingHTTPServer`, so that is the entire store. **If this ever runs
multi-process or multi-host, the limiter must move with it.** The clock is
read *inside* the lock; taken outside, racing threads append out of order
and both `Retry-After` and window reclamation silently go wrong.
* The key store is capped (`MAX_RATE_KEYS`) **per route**, not globally. At
capacity it reclaims expired windows, and if every window is still live it
**refuses the new key** — fail closed. Evicting a live window would let an
attacker who can mint many distinct keys clear their *own* limit on demand.
The per-route budget matters just as much: with one shared table, a flood
of cheap `/v1/repos` keys would exhaust it and lock brand-new `/v1/link`
clients out, turning the fail-closed guard into a cross-route DoS.
* Rule values must be real integers. `bool` subclasses `int` in Python, so
`[5, true]` would otherwise pass as a **1-second** window — an hourly limit
quietly becoming ~5/sec.
* `trust_forwarded_for` is **off** by default, and turning it on **requires a
non-empty `trusted_proxies`** — the header is honored only when the socket
peer is in that list. Without it, anyone reaching the origin directly (it
also listens on the tailnet) could pick and rotate their own rate-limit key
just by sending a header. Turn it on when exposing behind cloudflared,
where every request otherwise arrives from the tunnel and one abuser would
starve everyone. A caller can *prepend* anything to `X-Forwarded-For`; a
trusted proxy *appends* the peer it actually saw, so the service reads the
**last** entry, never the first, and requires it to parse as a real IP.
`trusted_proxies` is validated at startup: it must be a list (a bare string
would be iterated character by character), every entry a valid network, and
wildcards (`0.0.0.0/0`, `::/0`) are refused outright — they would restore
exactly the "trust anyone" hole the setting exists to close.
#### Identity binding (`state.json`)
`/v1/link` originally bound purely by `preferred_username` / email
local-part — any Zitadel identity whose derived login collided with an
existing account got a token for that account (account takeover). The
service now persists a map of Zitadel `sub` → Gitea login in
`/opt/granthi-link/state.json` (0600, atomic tmp+rename writes) and applies:
1. **Mapped sub** → always the mapped login, regardless of what the current
userinfo claims. If the mapped login was deleted from the forge it is
re-created only when the service created it originally; adopted accounts
are refused (409).
2. **Unmapped sub, login free** → create the user, record the mapping
(`created_by_service: true`). A concurrent-create 409 from Gitea is
handled idempotently: the user is re-fetched and accepted only if its
primary email is exactly the one this request would have set.
3. **Unmapped sub, login taken** → bind ONLY when the Gitea user's primary
email equals the Zitadel userinfo `email` **and** `email_verified` is
true (recorded with `created_by_service: false`); anything else →
`409 login exists and is not linked to this identity`.
A Gitea token is **never minted before the binding rule passes**, and a
corrupt/unreadable `state.json` fails closed (500) instead of falling back
to an empty map.
**Migration (pre-1.1.0 accounts):** identities whose userinfo carries no
verified `email` (e.g. Zitadel *machine* users) cannot self-adopt an
existing forge login under rule (c) — for them the first link after the
upgrade would 409 forever. Any forge account the v1 service created before
this change must be seeded into `state.json` once, as
`{"<sub>": {"login": "<login>", "created_by_service": true, ...}}`, written
0600 atomically. As of the beta rollout the only such account is the E2E
machine user `granthi-sync-e2e` (seeded); the forge's human admin `nirpa`
is never provisioned through `/v1/link`, so nothing else needed seeding.
Empirically verified mechanics on Gitea **1.27.1**, re-probed and still
holding on **1.27.2** (beta forge, 2026-08-22 — all three results below
matched; the probe minted a token on the `granthi-sync-e2e` machine user and
deleted it again, `DELETE …/tokens/{id}` returning 204 under basic auth):
* Token minting: `POST /api/v1/users/{login}/tokens` returns **401 for
token-authenticated sudo** (both `Sudo:` header and `?sudo=`); it only works
with **admin basic auth + `Sudo: <login>` header** (201). The config
therefore carries `admin_login`/`admin_password` (root-only, 0600) in
addition to `admin_token` (minted via
`gitea admin user generate-access-token`, used for all other admin calls).
* User creation: `source_id` is omitted → **local user** with a random
30-char password, `must_change_password=false`, `visibility=private`.
Rationale: `/api/v1/admin/identity-auth-sources` 404s on 1.27.1; the
shre-id OAuth2 source is ID 1 (via `gitea admin auth list`), but users
attached to an OAuth2 source cannot basic-auth and admin-created users get
no `external_login_user` row anyway — first OIDC web login links by email
regardless of this choice.
* `test_mode` (config flag, **never in production**): allows `/v1/link` to
accept `test_userinfo` in the body instead of a Zitadel round-trip, so E2E
can exercise the ensure-user + mint path headlessly. It is honored **only
when the service environment also sets `GRANTHI_LINK_ALLOW_TEST_MODE=1`**;
a config flag without the env gate is logged as an ERROR and ignored.
### `client/granthi_sync_client.py` (+ `bin/granthi-sync`) — client daemon
* `link [--server URL] [--token TOK]` — Zitadel **device flow** (native app
`granthi-sync-device`, client_id `386909715541590022`, project
granthi-forge `386906525790109702`; grants: device_code + refresh_token):
prints the verification URL + user code, polls the token endpoint
(`authorization_pending`/`slow_down` handled), then calls `/v1/link`.
`--token` skips the device flow with a ready Zitadel token (headless/dev).
Result stored in `~/.granthi-sync/config.json` (0600).
* `list [pattern]` — every repo the linked token can see, with the local
folder each is already synced to. `pattern` narrows the table by name
(substring, or a glob like `work-*`), case-insensitively, against both
`owner/name` and the bare name. Filtering is display-only: the set already
came from the forge under this account's token. Reads
`GET /api/v1/user/repos` on the forge
**directly** with the scoped user token — no granthi-link round-trip, so
the read path needs no service change. Pagination is followed to a short
page; if the `FORGE_MAX_PAGES` guard trips, the output says the list is
incomplete rather than letting a bounded page read as the whole set.
* `get --all [--into DIR] [--mode M]` — clone every repo this account can
see, skipping the ones already linked here. **"Only the repos they are
granted" needs no client-side permission logic**: `/api/v1/user/repos` is
evaluated by the forge against this account's own scoped token, so the
list *is* the grant. A client-side filter would be a second opinion about
someone else's authorisation. One repo failing does not abandon the rest,
and a truncated listing is reported loudly — `--all` must never quietly
mean "the first 2000". `alice/notes` and `bob/notes` both want
`<base>/notes`; the second is cloned to `<base>/bob-notes` and the clash is
logged, because reporting it as "already present" would leave the user
believing they had pulled both.
* `get <repo|owner/repo> [--into DIR] [--mode M]` — the download half of
`add`. Defaults to `snapshot` mode unless the repo carries a
`.granthi-sync.json` marker saying otherwise, so a plain synced folder
behaves the same on the second machine while someone's real project is
never autocommitted onto. Clones
with `--origin granthi` (the remote name `watch` looks for) and
`-c credential.helper=…` (the repo does not exist yet, so the helper
cannot be installed first; git also persists it into the new config), then
registers the folder in `config.json` with the same shape `add` writes —
so a cloned repo is picked up by `watch` immediately. Refuses a non-empty
destination. Branch is read with `symbolic-ref` (an empty repo has an
unborn HEAD) and falls back to `main`.
* `add <folder> [--name N] [--private|--public] [--mode M] [--force]`
`git init -b main` if needed, creates the cloud repo via `/v1/repos`, adds
remote `granthi`, initial commit + push. The token is delivered by a **git
credential helper** (the client's hidden `git-credential` subcommand
reading the 0600 config) — never embedded in the remote URL (estate rule).
Pushes the folder's **current** branch, not a hardcoded `main`: an existing
repo may sit on `master` or a feature branch, and publishing that work
under the wrong name is not a cosmetic error.
Two guards, because `add -A` takes whatever it is given: a starter
`.gitignore` is seeded when the folder has none (an existing one is never
touched — it is the user's), and a folder over 20 000 files / 512 MB is
refused unless `--force`. The seeded ignore file covers `.env`, `*.key`,
`*.pem`, `id_rsa` and friends, and it governs snapshots too — the scratch
index honours `.gitignore` exactly as a normal commit does.
The size guard measures what git *would* sync, ignore rules included
(including the machine's global excludes), because its own advice is "add
a .gitignore for what should not sync" and advice that changes nothing is
worse than none. It asks git through a **throwaway git dir outside the
folder**, so a refused `add` leaves no `.git` behind in a directory the
user never agreed to turn into a repo.
* `watch [--interval 30] [--once]` — per folder, by mode. `mirror`:
autocommit (`sync: <ISO ts>`) → fetch → ff-pull if remote strictly ahead →
push if local strictly ahead. `snapshot`: fetch → push a snapshot of the
working tree to this device's backup ref → ff-pull only when the tree is
clean (local edits are already safe on the backup ref, so it reports and
leaves the tree alone rather than failing) → push the user's own commits
when they are strictly ahead. **DIVERGED → log + record + SKIP. Never
force, never merge** — the same policy as the mesh — **but the backup
still happens**, because divergence is when work is most at risk.
Retention pruning runs at most hourly. SIGTERM-clean.
* `snapshots <folder> [--limit 20]` — restore points, newest first, **across
every device**, with the device that took each one. Read from the
**remote**, not a local cache: the feature exists for the case where this
machine is gone.
The three scopes differ deliberately. Writing is device-scoped, so two
machines never overwrite each other. Pruning is device-scoped, so machine A
never applies its clock to machine B's refs. **Reading is not scoped** — a
replacement laptop has a new id, and scoping the read to it would print
"no restore points yet" while the backups sit on the forge. That defect
was live in the first draft and is now pinned by a test that restores a
dead machine's work from a fresh clone.
* `restore <folder> --at <ts|sha> [--into DIR]` — materialise one restore
point into a new directory; refuses a non-empty destination. Accepts what
`snapshots` printed in either mode, including a mirror-mode `%cI`
timestamp. Two commits inside the same second share that timestamp, so an
ambiguous `--at` is **refused with the candidate ids** rather than
resolved by guessing.
* `status` — table of linked folders, mode, last sync, divergence flags.
* Run as a daemon on macOS with `client/launchd/ai.granthi.sync.plist`
(edit the script path, then `launchctl bootstrap gui/$UID <plist>`).
## Tests
* `python3 -m unittest discover -s tests` — 172 tests. The v1.2 additions
cover: a snapshot capturing uncommitted work while HEAD, the index and the
working tree stay byte-identical; snapshots landing outside `refs/heads`;
an unchanged tree not being re-pushed; a diverged folder still being backed
up; a dirty tree blocking the ff-pull but not the backup; retention keeping
everything recent, thinning to hourly then daily, and **keeping**
unparseable timestamps; prune deleting only the thinned refs; `.gitignore`
seeding never overwriting an existing one and keeping `.env` out of
snapshots; the folder-size guard being bounded rather than walking the
disk; mode detection; `list` filtering; `get --all` skipping what is
already present, defaulting to snapshot mode, and shouting about
truncation; `restore` writing a new folder, refusing a non-empty
destination, and leaving the working tree alone; and the credential helper
being the only one the repo consults, proven by driving
`git credential fill` against a deliberately poisoned outer helper.
* Earlier suite: autocommit/ff/diverged
logic against real temp git repos (including "diverged never touches the
remote"), config 0600 handling (including umask-proof creation and a
no-chmod guard), credential-helper quoting/injection, mocked device-flow
polling, the full `/v1/link` + `/v1/repos` service flows against an
in-process stub playing Zitadel + Gitea, all identity-binding rules
(collision 409, verified-email adoption, deleted-login re-create/refuse,
concurrent-create race, corrupt-state fail-closed), the test_mode env
gate, config-permission refusal, and the 64 KB body cap. The `list`/`get`
set covers pagination-to-a-short-page, truncation being reported rather
than hidden, HTTP errors being fatal instead of a silent empty list,
non-empty-destination refusal, owner-qualified names, unborn-HEAD branch
fallback, no token in the remote URL, and — the one that matters — that a
`get` folder is actually picked up by a subsequent `sync_folder` pass.
* Live E2E against the beta forge is recorded in the delivery notes
(link → add → watch ff/push → forced divergence → DIVERGED skip verified
via API, remote sha untouched).
## Invite-only story (v1)
There is no open signup. An operator invites a user by creating them in
shre-id (Zitadel org). The user downloads the client, runs
`granthi-sync link`, signs in at id.shre.ai with the printed device code, and
the provisioning service creates their forge account + scoped token on the
fly — the forge never sees a password and the user never sees the forge admin.
Every linked folder becomes a private repo under their account.
## Devices, sign-out, and the audit trail
**`granthi-sync devices`** lists every computer signed in to the account —
name, id, when it linked, and whether it is still active. The registry hangs
off the identity that owns it in `state.json`, so "which computers can reach
my files" has one answer that cannot drift from the identity map.
**`granthi-sync logout [--device ID]`** signs a computer out. Revocation
happens **at the forge**: granthi-link deletes that device's Gitea token
(admin basic auth + `Sudo`, the only mechanism Gitea 1.27 accepts — verified
204, after which the token returns 401 immediately). It is not a flag a
client could ignore, which is the only kind of sign-out worth having for a
laptop somebody lost. Signing out the current computer also deletes the local
token; linked folders are left on disk untouched.
Two properties worth keeping:
* The device id is part of the **token name**, because revocation deletes by
name. Without it, one user linking two machines called "macbook" in the
same second would collide, and signing one out would kill the other.
* A failed forge deletion is **not** recorded as revoked. A registry that
says "revoked" while the token still works is worse than an honest error.
**`granthi-sync activity`** shows the security events for the account:
`device.link`, `device.revoke`, `device.revoke.denied`, with timestamp,
device, and client IP. The log is append-only JSONL (0600, rotated at 64 MB),
kept **separate** from `state.json` on purpose — state is rewritten
atomically on every change, and an audit trail the audited thing can rewrite
is not an audit trail. A failed audit write is logged loudly and never breaks
the request it was auditing.
**What the audit log cannot see, and this matters.** Every event in it is one
granthi-link handled. **Git pushes and pulls do not pass through this
service** — they go straight to the forge — so:
| you want to know | where it actually lives |
|---|---|
| who linked/revoked a computer, from what IP | `granthi-sync activity` (this log) |
| who pushed what, and when | Gitea: the repo's activity feed and commit history |
| who *pulled* or cloned | **nowhere by default** — Gitea does not record fetches unless its router access log is enabled |
| last time a device used its token | Gitea `access_token.updated_unix` |
Reading the audit log and believing it lists file activity would be a real
mistake, so `/v1/audit` returns that caveat in its own response.
### Endpoints added
* `POST /v1/devices {token}` → the caller's devices.
* `POST /v1/devices/revoke {token, device_id}` → kills that device's forge
token. **Never rate-limited** — nobody should be throttled out of signing
a lost laptop out.
* `POST /v1/audit {token, limit}` → the caller's own events.
Authorisation on all three is the same: the caller proves who they are by
holding a working forge token, and **the forge decides** whose it is
(`GET /api/v1/user`). No login is ever read from the request body, so a body
claiming another account changes nothing.
## Next phase — invites and per-repo access (designed, not built)
Today `/v1/link` creates an account and every folder becomes a private repo
under it. What is missing is the multi-person case: an existing account
inviting somebody, and that person's device waking up with access to *some*
repos and not others.
Shape this should take, so the next session does not re-litigate it:
* **Where grants live: granthi-link's own store, not Zitadel orgs.** The
estate's house pattern is app-side tenancy tables with the IdP only
providing identity (see the AROS `tenants` / `tenant_members` split).
Grants therefore sit beside `state.json`, and **Gitea is the enforcement
point** — a grant is materialised as a repo collaborator or an org team
membership, so the forge itself refuses unauthorised reads. Nothing in the
client decides access, which is why `get --all` needs no permission logic.
* **Token scope does not change.** `write:repository,write:user` stays; per
repo permission is collaborator/team state, not a token property.
* **`POST /v1/invite`** (account admin → new member): creates the shre-id
user (Zitadel admin PAT, on aros-vps at
`/opt/shre-id/deploy/secrets/shre_id_zitadel_pat`), records the intended
grants, and returns an invite the person redeems by running `link`. Until
they redeem it, nothing exists on the forge.
* **`POST /v1/grants`** (account admin): add/remove repo access for a member;
applies the change to Gitea and records it. Removing a grant must also
remove the collaborator — a grant store that drifts from the forge is
worse than no store.
* Both endpoints are account-admin-only and rate-limited like `/v1/link`.
* The grant store inherits the same fragility already noted for the rate
limiter: a flat JSON file behind an in-process lock, fine for one
`ThreadingHTTPServer` and **not** fine the day this runs multi-process.
## Promotion window (beta → prod)
1. ~~**Expose :3042** behind cloudflared~~ **DONE 2026-08-23**
`https://granthi-link.shre.ai` (second-level, see above), origin stays
tailnet-only, `trust_forwarded_for` on with the tunnel as the only
trusted proxy.
2. **Swap forge base URLs** in `/opt/granthi-link/config.json`:
`gitea_base` → prod forge, `public_gitea_base`
`https://granthi.shre.ai`; the client default server URL moves to the
public endpoint.
3. The `granthi-web` OIDC app already lists the prod callback; the device
app is host-independent. Rotate the beta admin token/password out of the
config when pointing at prod (prod forge is READ-ONLY to this estate —
promotion is an operator action, not an agent action).
4. ~~Add rate limiting / abuse controls before public exposure.~~ **DONE**
see "Rate limiting" above (5 links/hour per client by default). When
exposing behind cloudflared, set `trust_forwarded_for: true` in the same
change, or every request will look like the tunnel and one abuser will
throttle everybody.
5. **Hardening checklist (must all hold before exposing):**
- [ ] `config.json` is 0600 (or 0400) and owned by the service user —
the service refuses to start otherwise; verify with
`systemctl status granthi-link` after any config edit.
- [ ] `test_mode` is absent from the production config **and**
`GRANTHI_LINK_ALLOW_TEST_MODE` is not set in the unit environment.
- [ ] `/opt/granthi-link/state.json` exists, is 0600, and is included in
VPS backups — losing it orphans sub→login bindings (existing users
would need verified-email re-adoption).
- [ ] `GET /health` returns 200 on both binds after restart.
- [ ] Spot-check the identity map: a repeat link for a known sub returns
the same login; a colliding username with a different sub gets 409.
<!-- review-service live check 20260819T131430Z -->