The bot runs in the public cloud and can't reach the 10.x/8 network
where DBS-210 bases live. This phase adds a data-center-resident relay
agent that dials outbound over WSS to the bot, and lets /phonestatus
post a follow-up message with per-base health after its main output
has already shipped.
Bot side (services/):
- dectRelayHub.js: WebSocket upgrade handler on /dect-relay/ws with
bearer-token auth (constant-time compare, header + Sec-WebSocket-
Protocol fallback for header-stripping proxies). Promise-based RPC
API with per-call timeouts, mid-flight-disconnect rejection, and
clean replacement of a stale agent socket when a newer one connects.
- dectDiscovery.js: pure filter that turns a phoneService result into
a list of reachable bases. Enforces the "must be on 10.0.0.0/8"
guardrail per requirements, dedups by IP + MAC, prefers Meraki-live
IP over Webex-cached IP.
- dectCollectorService.js: fan-out layer over the hub. collectAll()
runs one RPC per base in parallel with per-base error isolation —
one bad base never fails the batch.
Phone-status integration:
- Renderer gets a dectFollowUpBaseCount opt that emits an italic
"diagnostics loading for N base(s)..." hint inside the DECT section
of the main message.
- New exported renderDectDiagnosticsMarkdown() renders the follow-up
message: healthy/warning icon per base, uptime + firmware summary,
structured Power Loss reboot line, and per-base failure hints (e.g.
"relay accepted the request but the base did not respond in time").
- commands/phoneStatus.js discovers reachable bases synchronously
(pure), sends the main message, then fires collectAll() and posts
the follow-up as a separate message. Failures logged, never thrown
back to the user.
- Chat only: HTTP callers keep their single-message contract.
Agent side (dect-relay-agent/):
- Standalone Node process with its own package.json (only ws, axios,
dotenv). Reuses the shared integrations/cisco-dect/{client,probes,
statusXml}.js modules from the parent workspace so there's no code
duplication.
- Auto-reconnect with exponential backoff + jitter.
- Dispatches collect / reboot / force-reboot / reboot-chain /
force-reboot-chain / factory-reset / reconfigure-tree.
- DECT admin credentials live ONLY on the agent (never on the bot).
Shared bearer token gates the WSS handshake.
- README.md covers install, config, wire protocol, and safety model.
Env / infra:
- .env.example: adds DECT_RELAY_AGENT_TOKEN + optional DECT_RELAY_PATH
and DECT_COLLECT_TIMEOUT_MS. Reframes DECT_TEST_* as the local-dev
test harness rather than the production path.
- index.js: captures the http.Server from app.listen() and attaches
the relay hub when DECT_RELAY_AGENT_TOKEN is set; graceful shutdown
now closes the hub so in-flight RPCs get rejected cleanly.
- Adds "ws" to bot dependencies.
Tests (99 -> 113):
- tests/dectDiscovery.test.js: 13 cases covering the 10.x guardrail,
MAC normalization, IP source preference, dedup, and warning shape.
- tests/dectRelayHub.test.js: 14 integration cases using a real
ws pair on an ephemeral 127.0.0.1 port — auth (missing / wrong /
correct via header / correct via protocol fallback), hello frame,
RPC round-trip with correlation, agent error surfacing, concurrent
out-of-order replies, timeout, mid-flight disconnect, replacement
of a stale socket, and execAction routing.
- tests/renderers.test.js: 8 new cases for the DECT-follow-up loading
hint (plural / singular / off) and the diagnostics renderer (empty,
healthy, warning, power-loss dedup, active RTP, error hint, footer).
162 lines
5.4 KiB
JavaScript
162 lines
5.4 KiB
JavaScript
// src/services/dectCollectorService.js
|
|
//
|
|
// Fan-out layer over the DECT relay hub. Callers hand it a list of
|
|
// bases (from services/dectDiscovery.js), it dispatches one RPC per
|
|
// base in parallel and returns a normalized per-base result array.
|
|
//
|
|
// Kept intentionally thin: it doesn't render, it doesn't decide what
|
|
// to do with warnings, it doesn't touch Meraki. Whoever calls this
|
|
// (the /phonestatus follow-up, the /dectstatus command in Phase 2,
|
|
// the Jira poller in Phase 3) owns presentation.
|
|
|
|
import { getDectRelayHub, RelayErrorCodes } from './dectRelayHub.js';
|
|
import { logger } from '../utils/logger.js';
|
|
|
|
const LOG_SCOPE = 'dect:collector';
|
|
|
|
const DEFAULT_TIMEOUT_MS = Number(process.env.DECT_COLLECT_TIMEOUT_MS) || 15_000;
|
|
|
|
/**
|
|
* @typedef {object} BaseTarget
|
|
* @property {string} mac
|
|
* @property {string} ip
|
|
* @property {string} name
|
|
*/
|
|
|
|
/**
|
|
* @typedef {object} BaseCollectResult
|
|
* @property {BaseTarget} base
|
|
* @property {boolean} ok
|
|
* @property {object|null} data parsed status object (when ok)
|
|
* @property {object|null} verdict { healthy, warnings, info } (when ok)
|
|
* @property {number|null} elapsedMs
|
|
* @property {object|null} error { code, message } (when !ok)
|
|
*/
|
|
|
|
/**
|
|
* Run `collect` against every base in the list, in parallel. One base
|
|
* failing (timeout, offline, bad creds) does NOT fail the batch —
|
|
* that base's entry just has ok:false. Ordering of returned entries
|
|
* matches the input.
|
|
*
|
|
* @param {BaseTarget[]} bases
|
|
* @param {object} [opts]
|
|
* @param {number} [opts.timeoutMs] per-base RPC timeout override
|
|
* @param {object} [opts.hub] inject a hub for tests
|
|
* @returns {Promise<BaseCollectResult[]>}
|
|
*/
|
|
export async function collectAll(bases, opts = {}) {
|
|
const list = Array.isArray(bases) ? bases : [];
|
|
if (list.length === 0) return [];
|
|
const hub = opts.hub || getDectRelayHub();
|
|
const timeoutMs = opts.timeoutMs || DEFAULT_TIMEOUT_MS;
|
|
|
|
logger(LOG_SCOPE, `Fanning out collect() to ${list.length} base(s)`, 'debug');
|
|
|
|
const results = await Promise.all(list.map((base) => collectOne(hub, base, timeoutMs)));
|
|
|
|
const okCount = results.filter((r) => r.ok).length;
|
|
logger(LOG_SCOPE, `Collect finished: ${okCount}/${list.length} succeeded`, 'debug');
|
|
return results;
|
|
}
|
|
|
|
/**
|
|
* Single-base variant. Mostly here for the eventual /dectstatus
|
|
* command's individual "refresh this base" flow — collectAll uses it
|
|
* internally.
|
|
*/
|
|
export async function collectOne(hub, base, timeoutMs = DEFAULT_TIMEOUT_MS) {
|
|
if (!base?.ip) {
|
|
return {
|
|
base, ok: false, data: null, verdict: null, elapsedMs: null,
|
|
error: { code: 'NO_IP', message: 'base has no IP address' },
|
|
};
|
|
}
|
|
const started = Date.now();
|
|
try {
|
|
const { result, elapsedMs } = await hub.collect(base.ip, { timeoutMs });
|
|
return {
|
|
base,
|
|
ok: true,
|
|
data: result?.parsed || result || null,
|
|
verdict: result?.verdict || null,
|
|
elapsedMs: elapsedMs ?? (Date.now() - started),
|
|
error: null,
|
|
};
|
|
} catch (err) {
|
|
// We keep the code+message split so renderers can decide whether
|
|
// to show a hint ("relay is offline" vs "wrong password" are very
|
|
// different remediations).
|
|
const code = err?.code || 'UNKNOWN';
|
|
return {
|
|
base,
|
|
ok: false,
|
|
data: null,
|
|
verdict: null,
|
|
elapsedMs: Date.now() - started,
|
|
error: {
|
|
code,
|
|
message: err?.message || String(err),
|
|
// For NOT_CONNECTED there's no per-base fix — surface a hint.
|
|
hint: hintFor(code),
|
|
},
|
|
};
|
|
}
|
|
}
|
|
|
|
/**
|
|
* Run one of the mutating actions against a base. Same envelope shape
|
|
* as collectOne (ok / error / elapsedMs) so callers can log it
|
|
* uniformly. Actions handled here mirror the CLI script's subcommands.
|
|
*
|
|
* @param {BaseTarget} base
|
|
* @param {string} action 'reboot' | 'force-reboot' | 'reboot-chain' |
|
|
* 'force-reboot-chain' | 'factory-reset' |
|
|
* 'reconfigure-tree'
|
|
* @param {object} [opts]
|
|
* @param {number} [opts.timeoutMs]
|
|
* @param {object} [opts.hub]
|
|
*/
|
|
export async function execAction(base, action, opts = {}) {
|
|
if (!base?.ip) {
|
|
return {
|
|
base, ok: false, elapsedMs: null,
|
|
error: { code: 'NO_IP', message: 'base has no IP address' },
|
|
};
|
|
}
|
|
const hub = opts.hub || getDectRelayHub();
|
|
const timeoutMs = opts.timeoutMs || DEFAULT_TIMEOUT_MS;
|
|
const started = Date.now();
|
|
try {
|
|
const { result, elapsedMs } = await hub.execAction(base.ip, action, {}, { timeoutMs });
|
|
return {
|
|
base, ok: true, action,
|
|
elapsedMs: elapsedMs ?? (Date.now() - started),
|
|
result: result || null,
|
|
error: null,
|
|
};
|
|
} catch (err) {
|
|
const code = err?.code || 'UNKNOWN';
|
|
return {
|
|
base, ok: false, action,
|
|
elapsedMs: Date.now() - started,
|
|
error: {
|
|
code, message: err?.message || String(err),
|
|
hint: hintFor(code),
|
|
},
|
|
};
|
|
}
|
|
}
|
|
|
|
function hintFor(code) {
|
|
switch (code) {
|
|
case RelayErrorCodes.NOT_CONNECTED:
|
|
return 'DECT relay agent is not connected. Check that dect-relay-agent is running in the data center.';
|
|
case RelayErrorCodes.TIMEOUT:
|
|
return 'Relay accepted the request but the base did not respond in time. The base may be offline, rebooting, or unreachable.';
|
|
case RelayErrorCodes.DISCONNECTED:
|
|
return 'Relay agent disconnected while this command was in flight. Try again in a moment.';
|
|
default:
|
|
return null;
|
|
}
|
|
}
|