collabSupport/services/voiceDiag/README.md
jmcqueen 90a56c4640 Fix Prisma overlay tunnel discovery with site-scoped query_params.
Use topology/links and anynetlinks filters with eq/in operators so
store WAN follow-ups return real peer tunnels instead of unscoped dumps.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-07-28 09:01:40 -04:00

338 lines
21 KiB
Markdown

# /voicediag
Rules-engine style diagnostic for the store's Webex Calling
per-user configuration. The orchestrator (`voiceDiagService.js`)
walks a registry of check modules, gathers their `CheckResult`
objects, and hands them to the renderer + adaptive-card layer in
`commands/voiceDiag.js`.
## Command
```
/voicediag <storeNum> default: hides OK checks, posts fixable cards
/voicediag <storeNum> detail include OK checks + expand every details block
/voicediag <storeNum> --only dnd,callForwarding
/voicediag <storeNum> --window 24h narrow the WAN look-back (default 7d)
/voicediag list-checks enumerate every registered check + its scope
```
HTTP path: `GET /voicediag?storeNum=<n>[&detailed=true][&only=dnd,callWaiting][&window=15m]`.
`--window` accepts `Nm` / `Nh` / `Nd` shorthand (e.g. `15m`, `1h`, `6h`,
`24h`, `1d`, `7d`) or a bare integer of minutes; hard-capped at 7d.
Applies to the WAN healthscore + LQM (per-path latency/jitter/loss/MOS)
+ per-app voice-DPI fetches + alarms. Alarms are floored at 60m
regardless — sub-hour alarm queries are usually too noisy to be
actionable. Global default is `WAN_STANDARD_WINDOW_MINUTES` (env,
defaults to `10080` / 7 days). The 7d default was chosen because
per-app DPI metrics (Webex_Calling_RTP etc.) only get datapoints when
calls actually happen — sporadic stores need a wider window for
worst-window statistics to be meaningful. For live-incident triage
where you want a fresh snapshot, pass `--window 1h` or `--window 24h`.
The same default is used by the `/phonestatus` WAN follow-up (no CLI
override on that surface — set the env if you want a different value
globally).
HTTP callers get the markdown snapshot only — remediation cards are
chat-only.
## Voice path attribution (WAN tunnels / which pipe)
See [VOICE_PATH_ATTRIBUTION.md](./VOICE_PATH_ATTRIBUTION.md) for Phase 1
surfaces (runtime circuit status, overlay tunnels, DPI by `path_type`)
and Phase 2 discovery status for per-circuit / per-handset attribution.
Enable optional per-waninterface app DPI fan-out with:
```
PRISMA_VOICE_PATH_ATTRIBUTION=1
```
## Prerequisites
All checks call `/v1/people/{personId}/features/{feature}` (with a
few Webex-namespace variants noted inline in each check module).
The scopes required are:
- `spark-admin:people_read` — every read (all checks)
- `spark-admin:people_write` — remediations (disable DND, clear
forwarding, enable call waiting, disable call intercept)
These are the same scopes `/webexhost` and `/offboarduser` already
use, so no operator work is needed to enable `/voicediag` on an
existing deployment. If the token ever loses those scopes,
each check self-reports as `skipped: scope missing` — the run still
completes, and the phone-online summary keeps working.
### About the URL scheme
Early drafts of this feature documented `/v1/telephony/config/people/{id}/callSettings/*`.
That path family returns 404 "no static resource" from the Webex API
gateway — it is not a live surface. The correct admin path is
`/v1/people/{id}/features/{feature}`, per Webex's User Call Settings
docs. A couple of features use different feature-name segments than
their check id — most notably `intercept` (not `callIntercept`);
see each check module for the exact endpoint. The runner
distinguishes routing-404s ("URL moved, update the check") from
"not applicable" 404s ("this person isn't a calling user") using
the `no static resource` marker in the response body.
## Store voice standards
The store phone standard is enforced by the checks below. Every
check descriptor exports a `standards` object so the desired state
is legible from the check file without reading `run()`. Changes
here are meant to be a two-step change: update `standards`, then
teach `run()` to interpret it — the checks compare the live state
against `standards` and emit the severity in the "when non-compliant"
column.
### Per-user Webex Calling standards
| Check | Standard | Non-compliant severity | Auto-remediation |
| -------------------- | -------------------------------------------------------- | ---------------------- | ------------------------------------ |
| `dnd` | `enabled: false` | warn | `disable_dnd` |
| `callForwarding` | `{always, busy, noAnswer}.enabled: false` | **error** | `clear_call_forwarding` |
| `callWaiting` | `enabled: true` | warn | `enable_call_waiting` |
| `callIntercept` | `enabled: false` | error | `disable_call_intercept` |
| `voicemail` | `enabled: true`, all three `send*Calls.enabled: false`, `mwiEnabled: true` | error on send-to-VM, warn on MWI-off / disabled / off-org email | `stop_sending_to_voicemail` for the send-to-VM path only |
| `hoteling` | `enabled: false` | warn | `disable_hoteling` |
| `executiveAssistant` | `type: 'UNASSIGNED'` | warn | none — Control Hub cleanup |
| `outgoingPermission` | no high-impact call type BLOCKED (LOCAL, NATIONAL, TOLL_FREE, TOLL, INTERNATIONAL) | warn | none — location-scoped fix |
`callForwarding` is intentionally at `error` severity: forwarding
active on a store line silently drops customer calls, and it's the
single most common voice ticket. Voicemail's `error` path is
narrower — only the three send-* triggers upgrade to error, since
that also silently swallows calls. MWI-off, off-org email
forwarding, and voicemail-disabled all stay at warn.
### Switchport / port-hygiene standards
These reuse the phone-status snapshot from `/phonestatus` — no
extra Webex API calls — and cross-reference against the Meraki
port config that already flows through `services/enrichment/merakiEnrichment.js`.
Trunk uplinks are treated as a visibility boundary: the phone is
behind a non-Meraki switch (typically a Cisco stack in a store)
and per-port policy isn't ours to enforce, so downstream VLAN /
PoE / admin-state checks defer to the operator.
| Check | Standard | Non-compliant severity | Notes |
| ------------- | ---------------------------------------------------- | ---------------------- | ----------------------------------------------- |
| `portType` | `portType: 'access'` on every wired phone / DECT base | warn | Trunk uplinks flagged so operator checks the downstream switch |
| `portVlan` | `vlan === VOICE_STANDARD_PHONE_VLAN` (default `102`) | warn | Env-configurable — VLAN may move from data-side to a proper voice VLAN in the future |
| `portPoe` | `poeEnabled: true` | warn | Skipped for trunk-uplinked devices |
| `portEnabled` | `portEnabled: true` | **error** | Admin-disabled port → phone is dead |
No auto-remediation on any port check — Meraki port-config PUTs
are a separate scope of work; the operator handles fixes in the
Meraki Dashboard.
### Environment overrides
Both port-hygiene knobs live in `.env`:
| Var | Default | Purpose |
| ------------------------------ | ------- | ---------------------------------------------------------------------------------------------- |
| `VOICE_STANDARD_PHONE_VLAN` | `102` | Expected VLAN for a store phone. Set per-site if the fleet moves onto a proper voice VLAN. |
| `VOICE_STANDARD_ENABLED` | `true` | Global kill-switch for the port-hygiene bucket. `false` silences portType / portVlan / portPoe / portEnabled while Meraki cleanup is in progress. Feature-config checks always run. |
### WAN standards (Prisma SD-WAN)
Sources Prisma SD-WAN metrics (see `integrations/paloalto/`) for
the store's site and grades per-path LQM + healthscore + alarms
against ITU-T G.114 / RFC 3550 references. Diagnostic-only — no
auto-remediation, since WAN config PUTs are the operator's job in
the Prisma portal, not the bot's. All checks skip cleanly with
"not applicable — missing sdwanSite" for non-Prisma-managed sites,
so the WAN bucket adds no noise to stores that live entirely
behind Meraki.
Data flow: `voiceDiagService.buildContext()` calls
`collectSdwanForStore(storeNum)` in parallel with the Webex +
Meraki fetches. That composer resolves the store to a Prisma site
(`CG${pad5(storeNum)}` — e.g. store 782 → `CG00782`), pulls
elements + healthscore + LinkState + 4 LQM metrics + recent alarms
in parallel, and normalises to a stable shape the checks + the
`/phonestatus` follow-up renderer share.
| Check | Standard | Non-compliant severity | Notes |
| ---------------- | --------------------------------------------------- | ----------------------------------------------- | ------------------------------------------------------------------------ |
| `wanSite` | Prisma site resolves for the store | skipped when unresolved | Info-only anchor — surfaces site + element + link counts |
| `wanHealthscore` | >= `WAN_STANDARD_HEALTHSCORE_WARN` (default 80) | warn < 80, **error** < 60 | Prisma AIOps composite (0-100) |
| `wanLinkState` | Every WAN path up | **error** if any path down | No remediation physical / carrier work |
| `wanLatency` | <= `WAN_STANDARD_LATENCY_WARN_MS` (default 150ms) | warn > 150ms, **error** > 400ms | ITU-T G.114 one-way reference. Worst path drives severity. |
| `wanJitter` | <= `WAN_STANDARD_JITTER_WARN_MS` (default 30ms) | warn > 30ms, **error** > 50ms | RFC 3550 jitter-buffer envelope |
| `wanLoss` | <= `WAN_STANDARD_LOSS_WARN_PCT` (default 1%) | warn > 1%, **error** > 3% | G.711 PLC tolerance |
| `wanMos` | >= `WAN_STANDARD_MOS_WARN` (default 4.0) | warn < 4.0, **error** < 3.5 | ITU-T P.800 MOS scale. LOW is bad the accessor flips the comparison. |
| `wanAppRtpMos` | Worst 5-min window >= `WAN_STANDARD_APP_MOS_WARN` (default 4.0) | warn < 4.0, **error** < 3.5 | Real DPI measurement on the configured voice app's RTP frames (e.g. `Webex_Calling_RTP`, `rtp-base`). Graded against the WORST window in the series catches transient degradation link-probe averages hide. Skipped when `PRISMA_APP_ID_VOICE` unset. |
| `wanAppRtpLoss` | Worst 5-min window <= `WAN_STANDARD_APP_LOSS_WARN_PCT` (default 5%) | warn > 5%, **error** > 15% | DPI packet-loss for the same voice app. Same feature gate. |
| `wanAppRtpJitter` | Worst 5-min window <= `WAN_STANDARD_APP_JITTER_WARN_MS` (default 30ms) | warn > 30ms, **error** > 50ms | DPI jitter for the same voice app. Same feature gate. |
| `wanAlarms` | Zero critical + zero major alarms in the last hour | warn on major, **error** on critical | Minor alarms are info-only. Pass-through of Prisma severity. |
| `wanTunnels` | All overlay / VPN tunnels up | **error** if any down; warn if recent alarms | Distinguishes overlay-down vs physical-up. |
Environment overrides (all optional — defaults match the standards
above):
| Var | Default | Purpose |
| ------------------------------------- | ------- | ----------------------------------------------------------- |
| `WAN_STANDARD_LATENCY_WARN_MS` | `150` | Latency warn threshold, ms |
| `WAN_STANDARD_LATENCY_ERROR_MS` | `400` | Latency error threshold, ms |
| `WAN_STANDARD_JITTER_WARN_MS` | `30` | Jitter warn threshold, ms |
| `WAN_STANDARD_JITTER_ERROR_MS` | `50` | Jitter error threshold, ms |
| `WAN_STANDARD_LOSS_WARN_PCT` | `1` | Packet-loss warn threshold, % |
| `WAN_STANDARD_LOSS_ERROR_PCT` | `3` | Packet-loss error threshold, % |
| `WAN_STANDARD_MOS_WARN` | `4.0` | MOS warn threshold (LOW is bad) |
| `WAN_STANDARD_MOS_ERROR` | `3.5` | MOS error threshold |
| `WAN_STANDARD_HEALTHSCORE_WARN` | `80` | Site healthscore warn threshold (0-100) |
| `WAN_STANDARD_HEALTHSCORE_ERROR` | `60` | Site healthscore error threshold |
| `WAN_STANDARD_APP_MOS_WARN` | `4.0` | Per-app worst-window MOS warn threshold |
| `WAN_STANDARD_APP_MOS_ERROR` | `3.5` | Per-app worst-window MOS error threshold |
| `WAN_STANDARD_APP_LOSS_WARN_PCT` | `5` | Per-app worst-window packet-loss warn threshold, % |
| `WAN_STANDARD_APP_LOSS_ERROR_PCT` | `15` | Per-app worst-window packet-loss error threshold, % |
| `WAN_STANDARD_APP_JITTER_WARN_MS` | `30` | Per-app worst-window jitter warn threshold, ms |
| `WAN_STANDARD_APP_JITTER_ERROR_MS` | `50` | Per-app worst-window jitter error threshold, ms |
| `PRISMA_APP_ID_VOICE` | _(unset)_ | Numeric Prisma app id for the tenant's voice app. Recommended: `Webex_Calling_RTP` for Webex Calling shops (excludes non-Webex UDP noise), otherwise `rtp-base` for the generic RTP signature. When set, /phonestatus + /voicediag pull real DPI voice-quality metrics from Prisma's Application Path Details endpoint. Discover with `npm run prisma:probe -- appdefs webex` (or `rtp`) or from the Prisma UI URL. Leave unset to skip the extra API calls. |
| `PRISMA_APP_NAME_VOICE` | `voice` | Display label for the app configured above. Set to the exact Prisma UI label (e.g. `Webex_Calling_RTP`) so section headers and check messages cross-reference cleanly. |
| `PRISMA_APP_ID_RTP_BASE` | _(unset)_ | **Deprecated.** Backwards-compat fallback for the original release. Still honored if set and `PRISMA_APP_ID_VOICE` is unset (with a one-time deprecation warning in the log); display name defaults to `rtp-base` in that case. Rename to `PRISMA_APP_ID_VOICE` at your convenience. |
| `WAN_STANDARD_ENABLED` | `true` | Global kill-switch for the WAN bucket. `false` silences all 11 WAN checks while a Prisma cleanup / integration validation is in flight. |
Prisma credentials and auth-mode selection live under the
`Palo Alto Prisma SD-WAN` block in `.env.example` (SASE OAuth 2.0
recommended, legacy CloudGenix session-token supported for
backward compatibility). See that block for the full env
inventory.
### Debugging Prisma schema drift
Prisma SD-WAN's monitor API schemas vary between tenants — request
body shapes documented on pan.dev or in LIVEcommunity examples
frequently fail schema check on this tenant with cryptic 400
messages like `"$.filter.wan_interfaces: is not defined in the
schema"` or `"$.view: does not have a value in the enumeration
[summary, timeseries]"`. Iterating on those in-band via
`/phonestatus` is painful (full DECT + Meraki + WAN pipeline runs
every attempt).
Use the standalone probe script instead:
```
node scripts/prismaProbe.js discover 782
node scripts/prismaProbe.js try-shapes health <siteId>
node scripts/prismaProbe.js try-shapes lqm <siteId> <wiId1,wiId2>
node scripts/prismaProbe.js raw POST /sdwan/v3.7/api/events/query --body '{"limit":{"count":5}}'
```
`try-shapes` fires a curated list of candidate request bodies and
prints which pass schema check (with response-shape summary) and
which fail (with the exact SCHEMA_CHECK_FAIL field). When you find
a winning shape, port it into `integrations/paloalto/metrics.js`
and add a regression test in `tests/paloalto.metrics.test.js` so
the tenant's quirk doesn't get "helpfully cleaned up" by a future
refactor. Same auth stack as the bot, no side effects.
## Apply-all-N-fixes card
When two or more checks return fixable results, `/voicediag`
posts one extra adaptive card at the bottom offering to apply the
whole batch in a single click. The individual per-issue cards stay
on-screen so operators can still pick and choose; the combined
card is a shortcut for the common "everything looks right, do it
all" case. The batch executes each fix in sequence (not parallel)
so audit lines stay readable and per-person Webex API write
throttling doesn't stack; failures accumulate into a final summary
line rather than aborting the run.
Dispatcher-side, the combined card uses the same
`pendingVoiceFixes` map and the same voicediag branch in
`index.js` — the payload's `combined: true` flag + the
`confirm_voicediag_all` / `cancel_voicediag_all` action ids are
what pick the batch handler over the single-fix handler.
## Adding a new check
1. Create `services/voiceDiag/checks/<newCheck>.js` and export a
descriptor:
```js
// The standards block is *the* source of truth for the desired
// state. Keep the run() comparison in sync — regression tests
// in tests/voiceDiag.checks.test.js assert every check has one.
export const MY_NEW_STANDARDS = Object.freeze({ enabled: false });
export const myNewCheck = {
id: 'myNew',
label: 'My New Check',
requires: ['personId'], // subset of ['personId','phoneStatus','telephonyProfile']
scope: 'spark-admin:people_read',
standards: MY_NEW_STANDARDS,
async run(ctx) {
const data = await ctx.webex.request('GET', `people/${ctx.personId}/features/whatever`);
if (!!data?.enabled === MY_NEW_STANDARDS.enabled) {
return { status: 'ok', message: 'Compliant.', details: null, remediation: null };
}
return {
status: 'warn', // 'ok' | 'warn' | 'error' | 'skipped'
message: 'Human sentence for the chat row.',
details: { relevant: 'facts' },
remediation: { // optional
action: 'fix_my_thing', // globally unique remediation id
title: 'Fix My Thing',
summary: 'One-line description of what the fix will do.',
payload: { personId: ctx.personId, personLabel: ctx.personLabel, storeNum: ctx.storeNum },
},
};
},
remediations: { // optional
async fix_my_thing(bot, data, requester) {
// do the PUT, audit, reply
},
},
};
```
2. Register it in `services/voiceDiag/checks/index.js` at the
position you want the renderer to output it.
3. Add a `describe`-level block to `tests/voiceDiag.checks.test.js`
covering the enabled / disabled / 403 branches.
That's it — no changes to the runner, renderer, command handler,
`index.js` dispatcher, or `commands/registry.js`. The single
`confirm_voicediag` / `cancel_voicediag` action namespace routes
through the shared dispatcher, which looks the remediation up in the
central registry built at module load. Remediation id collisions
throw at import so you'll notice immediately.
### CheckResult shape
```ts
type Severity = 'ok' | 'warn' | 'error' | 'skipped';
type Remediation = {
action: string; // globally unique across all checks
title: string; // button label
summary?: string; // optional card body line
payload: Record<string, unknown>; // handed to remediations[action](bot, payload, requester)
};
type CheckResult = {
id: string; // set from check.id automatically
label: string; // set from check.label automatically
status: Severity;
message: string;
details?: object | null; // shown under `detail` mode
remediation?: Remediation | null;
};
```
### Error handling
Never throw from `run()`. The runner wraps every check in a
try/catch and converts:
- HTTP 401/403/404 → `status: 'skipped'` with a "missing scope"
hint pointing at `check.scope`.
- Any other error → `status: 'error'` with the API message.
So if you _can't_ reason about a specific 5xx, just let it bubble —
the runner will surface it.