Fleet
The Node-RED swarm — registry, status probes, flow lifecycle, and the staging→prod promotion path.
The Fleet plane is the original reason R2-D2 exists. It manages a small number of Node-RED instances running upstream Node-RED in containers, each with its own admin endpoint and its own set of flows. Each instance is one helm release; the dashboard's job is to know they exist, know if they're healthy, and let an operator move flows between them without SSH'ing into anything.
The registry
instances/registry.yaml is the source of truth for what exists:
instances:
- id: nr-staging-01
name: "Staging — N-1"
purpose: "Pre-prod testbed"
profile: medium # heavy | medium | light (resource template)
namespace: nodered
release: nr-staging-01
ingress: "nr-staging-01.office.ilab.zone"
lb_ip: "192.168.3.215" # from the kube-vip pool
flows:
- "ingest"
- "telemetry"The dashboard never invents instances. Adding a Node-RED instance means:
- Writing a registry entry (the Settings → Fleet page does this through
POST /api/registry/instances). - Committing the change (
POST /api/registry/commit— stage, commit, push). - Running the helm release in-cluster (handled by
cicd/, separately from the dashboard).
Status probes
GET /api/instances walks the registry and probes each instance:
- HTTP healthcheck against the Node-RED admin endpoint.
- Latency measurement.
- Pulls diagnostics (memory, uptime, Node-RED version) where available.
- Joins NetBird peer info so we know whether the instance is even on the overlay.
The response is cached briefly. The dashboard's Zustand store re-fetches on tab focus
and on a 30s background interval. The sidebar's "Node Red" tile (I / ● / ○ /
optional !) reflects this same data — total / online / offline / degraded.
A "degraded" instance is one that's reachable but reporting elevated memory or a
diagnostics signal we don't trust. A "down" instance disappears from online and turns
the sidebar tile red.
Flow lifecycle
Flows are JSON arrays — that's how Node-RED itself stores them. R2-D2 keeps the canonical copy in git:
flows/
├── nr-staging-01/
│ └── flows.json <- the entire flow tree for that instance
├── nr-prod-01/
│ └── flows.json
└── ...The four lifecycle routes:
| Route | What it does |
|---|---|
POST /api/flows/pull | Live → git. Pulls flows from the instance and writes them to flows/<id>/flows.json. The response includes tab labels and counts so the operator can sanity-check before committing. |
POST /api/flows/deploy | Git → live. Reads flows/<id>/flows.json and pushes it into the instance via the Node-RED admin API. |
POST /api/flows/promote | Live → live (cross-instance). Copies the selected flow tabs from a source instance into a target instance, with mode: 'merge' | 'replace'. Merge keeps tabs not in the source; replace replaces the whole tree. |
PUT /api/flows/toggle | Enable/disable a flow tab on a live instance. |
The promotion path is the workhorse. Typical sequence: build on staging → pull to git → commit → promote selected tabs to prod. Replace mode is rare; merge mode is the default because it doesn't surprise anyone.
What about the editor?
The dashboard does not host a Node-RED editor. It links to the upstream editor at
/instances/<id>/editor (a Magic redirect to the instance's own admin UI). The flows
page shows a static rendering of the flow tree — tabs, node counts, MQTT topic
inventory — so operators can do diff-style reviews without round-tripping through the
editor.
Sidebar Vader-icon
When the registry has instances but none are responding (every probe failed), the
sidebar nav items fleet, instances, flows, and registry all show a Vader
icon — the entire plane is dark. Page navigation still works; the icon is informational
so the operator doesn't waste time trying to deploy into a fleet that isn't there.
Known operational quirks
- Worker-node kubelet proxy 502. If a Node-RED instance is on a worker whose
rke2-agentlost its tunnel back to the control plane, the diagnostics probe will return 502s even though the instance is healthy. Recovery is to restart therke2-agenton that node (in-cluster nsenter job — no SSH required). The fleet plane's diagnostics page will go green again once the tunnel re-establishes. - registry-pod drift. The dashboard reads
registry.yamlfrom a baked-in copy at build time and from a mounted ConfigMap. If the ConfigMap drifts from the in-pod registry the operator sees confusing duplicate IDs. The Settings → Fleet page calls this out when it detects the drift; the fix is a redeploy or a CM sync. - Diagnostics returns
memoryUsageMB: 0. Some Node-RED versions don't expose the process-memory endpoint the dashboard reads. The diagnostics page falls back to the container's cgroup memory when this happens.
Built on
| Layer | Tech | Page |
|---|---|---|
| Instance / flow tables | @tanstack/react-table via DataView | TanStack Table |
| Live status feed | zustand Fleet store | Zustand |
| Registry editor (YAML) | @monaco-editor/react | Monaco |
| Reachability probe | Netbird daemon socket | Netbird |
| Source of truth (instances) | instances/registry.yaml in repo | (config) |
| Flow lifecycle persistence | Git + Node-RED admin API | (Node-RED upstream) |
| LB IP allocation | kube-vip pool | kube-vip |
| Public ingress | nginx-ingress via Caddy | Ingress |