Runtime architecture
How the dashboard, Node-RED, Restreamer, TrailBase, and Yugabyte fit together at runtime.
R2-D2's runtime is a small number of long-lived services with one clear talker (the dashboard) on top of them. There's no event bus between services — every cross-plane interaction goes through the dashboard's Next.js API routes. That's intentional: it makes the system legible even when there are six things moving at once.
Service map
┌──────────────────────────┐
│ Operator (web browser) │
└────────────┬─────────────┘
│ HTTPS (TrailBase cookie)
▼
┌──────────────────────────────────────────────────────────────────────┐
│ r2d2-fleet (Next.js 15) │
│ │
│ Pages ◄──── React/Zustand ────► API routes (~44, OpenAPI spec) │
│ │ │ │ │ │ │ │
└──────┼──────────────────────────────────────────┼──┼──┼──┼──┼────────┘
│ /the-force/* iframes │ │ │ │ │
▼ │ │ │ │ │
┌─────────────┐ │ │ │ │ │
│ HoloChron / │ │ │ │ │ │
│ HoloNet / │ │ │ │ │ │
│ Probe etc. │ │ │ │ │ │
└─────────────┘ │ │ │ │ │
▼ │ │ │ │
┌─────────────┐ │ │ │
│ Restreamer │ │ │ │
│ (datarhei) │ │ │ │
└─────────────┘ │ │ │
▼ │ │
┌────────────┐
│ TrailBase │ (auth + Republic + Archives)
└────────────┘
│ │
▼ │
┌───────────────┐
│ YugabyteDB │ (postgres-compatible YSQL)
└───────────────┘
│
▼
┌─────────────────────────────────────┐
│ Node-RED instances (helm-released) │
│ — registered in instances/ │
│ registry.yaml │
└─────────────────────────────────────┘
│
▼
┌─────────────────┐
│ TBMQ (MQTT) │
└─────────────────┘What lives where
r2d2-fleet — the dashboard
A Next.js 15 app with React 19 client components for the interactive bits (Zustand stores for fleet state, restreaming KPIs, the Force readiness checks). Every API call from the client goes through a same-origin Next route handler — there is no direct browser → Restreamer or browser → Node-RED traffic. That keeps the auth story simple (TrailBase cookie all the way down) and means the OpenAPI spec is the contract for everything external.
Node-RED instances — the workers
Each instance is a helm release running upstream Node-RED with a fixed admin endpoint and
an LB IP from the kube-vip pool. The registry of instances lives in
instances/registry.yaml — that file is the source of truth for what exists, and the
dashboard's /api/instances route probes them on every page load to surface live status.
Flows live as JSON files in flows/<instanceId>/flows.json and get committed to git so
diffs between staging and prod are reviewable.
Restreamer (datarhei) — the video plane
A separate container running datarhei/core. The
dashboard's /api/restreaming/* routes are a typed wrapper over the upstream /api/v3/*
surface: every mutation reads the current process config first, diffs, writes back the
merged result, and appends to data/restreamer-audit.jsonl. Layouts, panels, groups, and
the magic URLs all live in TrailBase, not Restreamer — Restreamer only knows feeds.
TrailBase — auth + small relational state
TrailBase serves three roles:
- Identity provider —
/api/auth/loginproxies to it; the session cookie is the only thing the dashboard checks forAUTH_ENABLED=true. - Small relational store —
restreamer_layouts,restreamer_panels,restreamer_groups,restreamer_view_bindings, plus per-feed metadata (alias/platform/tags/notes) and snapshot thumbnails. - HTTP-callable record API — the dashboard hits TrailBase REST endpoints for those tables rather than holding a DB connection of its own.
YugabyteDB — the large relational store
YSQL (postgres-compatible) cluster. The Matrix-side Synapse install runs against it, and it's the long-term destination for relational data that doesn't belong in TrailBase. The fleet historically ran on CNPG (Cloud Native PG); that's being retired — see the Auth (Keycloak) plan for one of the migrations gated on finishing the Yugabyte cutover.
TBMQ — MQTT for telemetry
The Temple Archives plane listens to MQTT topics from the broker and persists them. The
HoloNet page is just the MQTT Explorer iframe wired up against this broker. The
broker uses a redis-cluster for HA — if a host node restarts hard, the recovery path
involves FORGET + MEET + REPLICATE against the broken peer's nodes.conf (covered in
the operations runbook the team keeps in bd).
How requests flow
A typical "deploy a flow change" round-trip:
- Operator edits the flow in the in-cluster Node-RED editor.
- Operator clicks Pull in r2d2-fleet →
POST /api/flows/pullfetches the live flows, writes them toflows/<instanceId>/flows.json, and returns a diff summary. - Operator inspects the diff, commits it via
POST /api/registry/commit. - Operator promotes from staging to prod via
POST /api/flows/promote, which copies the selected tabs into the target instance's flow tree and deploys them via the Node-RED admin API.
Everything is one TrailBase cookie, four route handlers, and a git commit. There's no queue, no broker, no callback.
Where state lives — and only lives
| What | Source of truth | Cached / projected to |
|---|---|---|
| Which Node-RED instances exist | instances/registry.yaml | Zustand store, /api/instances response |
| Each instance's live status | live HTTP probe to the instance | Zustand store (poll on focus) |
| Flow content | the running Node-RED instance | flows/<id>/flows.json (git), Restreamer? no — flows do not touch Restreamer |
| Camera feeds | Restreamer process config | dashboard cache + TrailBase enrichment |
| Camera-feed metadata (alias, tags) | TrailBase | dashboard fetches lazily |
| Layouts, panels, groups (Holoprojector) | TrailBase | dashboard fetches lazily, magic-URL bypasses auth |
| Nations / ShopKeepers / holdings | republic-config.yaml + TrailBase | dashboard |
| The Force embed URLs | env vars / holocron-config.yaml | sidebar readiness check polls |
| Operator identity | TrailBase | session cookie |
The rule is: if you can't write down which row/file the data comes from, the architecture is wrong somewhere. Add the row.
Failure modes worth knowing
- A Node-RED instance is unreachable. The instance disappears from "online" counts on the sidebar. The dashboard keeps working; you lose the ability to push flows to that instance until the probe goes green again. The sidebar shows a Vader icon next to the instance in the nav.
- Restreamer is unreachable. Restreaming KPIs in the sidebar go red and the HoloVids / Holoprojector pages show a placeholder. Other planes are unaffected.
- TrailBase is unreachable. Login fails with
auth_backend_unreachable(502). The dashboard keeps serving anonymous routes (the sidebar nav, public Magic URLs) but every authenticated route returns 401. Operators with valid session cookies stay logged in until expiry. - YugabyteDB tserver wipe. Stale Raft UUIDs in tablet configs block recovery; the
recipe is
yb-ts-cli unsafe_config_change. Yugabyte's own page in this handbook lives in the platforms section (Republic → Platforms & Buildouts), not here. - kube-vip lease churn.
.200is the control-plane VIP and is excluded from the LB pool. If you see a service stuck Pending despite the LB range having room, double-check it didn't try to claim.200.
The redis-cluster behind TBMQ is the most common recovery scenario. If the host
node carrying a redis pod restarts hard, nodes.conf can be left in a state where
the cluster can't re-form — the Temple Archives plane loses subscriptions until the
cluster is repaired. Recovery is a FORGET + MEET + REPLICATE sequence against
the broken peer; the operational runbook is maintained by the team and is the right
place to follow exact commands.
What's not here
Anything to do with the off-cluster Matrix install, the FOSS weather archive, the wargame deploy, or the rusty-ais experiment. Those are sibling repos. R2-D2 is the operator console — anything that isn't operated through this console doesn't belong on this map.