Commit graph

26 commits

Author SHA1 Message Date
1eaabbcee1 Add /voicereport store digest and harden Webex CDR feed handling.
Groups cdr_feed legs by Correlation ID for call-level summaries, fixes report-column field parsing and Docker proxy routing, and adds /calltest post-call Twilio and CDR enrichment.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-07-24 13:43:11 -04:00
7aa8c37d1d Add Twilio /calltest for store and direct-dial voice path testing.
Enables outbound PSTN probes via TwiML webhooks with Webex result cards, status polling, and optional store CDR enrichment.
2026-07-23 17:59:10 -04:00
a807f22b43 Use Jira component icons in poller enrichment summary bullets.
Show Mobility, Communication Services, and AV with distinct emoji
in the hourly Webex summary instead of a generic Phone/AV label.
2026-07-21 15:53:14 -04:00
7c7d0a898c Add /dectstatus command with full base dump and reboot cards.
Pulls complete status.xml from each reachable DBS-210 via the relay,
renders the full CLI-style dump, and posts per-base adaptive cards for
reboot and factory-reset (confirm flow + audit logging). Wired into
registry, help, and attachmentAction dispatch.
2026-07-21 15:51:07 -04:00
1117be40cc Format chat footers in DISPLAY_TIMEZONE instead of UTC.
Docker hosts default to UTC, so bare toLocaleTimeString() showed
wrong "Last checked" times in avstatus and other commands. Add
formatDisplayTime() (default America/New_York, overridable via
DISPLAY_TIMEZONE) and use it across renderers and command footers.
2026-07-21 14:41:26 -04:00
9d0dbb071e Add per-app DPI voice-quality checks + widen WAN window to 7d
Adds three new SD-WAN checks (wanAppRtpMos/Loss/Jitter) that measure
REAL voice-traffic quality on actual RTP frames via Prisma DPI, not
synthetic link probes. Graded against the WORST 5-minute window so
transient degradation the 24h link-probe averages smooth away
actually surfaces.

Voice-app selection is tenant-configurable via PRISMA_APP_ID_VOICE +
PRISMA_APP_NAME_VOICE (Webex_Calling_RTP recommended for Webex
Calling shops — the Webex-specific DPI signature excludes non-Webex
UDP noise). Legacy PRISMA_APP_ID_RTP_BASE still honored with a
one-time deprecation warning.

Widens the default WAN look-back from 24h to 7 days: per-app metrics
only get datapoints when calls actually happen, so sporadic Webex
Calling stores (3-4 calls/day) need a wider window for worst-window
statistics to be meaningful. Interval picker snaps 7d to 1hour
buckets (168 pts) to keep payloads bounded while preserving
worst-hour granularity. Hard-capped at 7d — beyond that Prisma
downsamples to 1-day buckets and the signal collapses.

Also:
- Client-side concurrency limiter (PRISMA_MAX_INFLIGHT, default 3)
  to prevent 429 cascades when /voicediag fans out 10+ parallel
  metric fetches
- "View in Prisma UI" deep links in both /phonestatus WAN follow-up
  and /voicediag details, threading through a new
  integrations/paloalto/urls.js builder
- humanizeMetricUnit maps raw API unit strings ("percentage",
  "milliseconds") to display symbols ("%", "ms") to fix
  "11.83percentage" leaking to the UI
- getAppAudio envelope distinguishes not-configured / fetch-failed /
  no-traffic states so misleading "set env var" messages don't fire
  when the real problem is a 429

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-07-09 13:54:15 -04:00
21fa8f1436 Clean up SD-WAN alarm rendering + explain overlay vs physical
The wanAlarms check was pasting raw Prisma `info` JSON blobs (nested
`vpn_reasons` arrays with element/site/vpnlink ids) into the chat
message field. On a store with 20 NETWORK_ANYNETLINK_DOWN flaps this
produced a wall of unreadable stringified JSON where the actual
signal ("SD-WAN overlay tunnels are flapping") was lost.

Introduces a shared alarmSemantics module that:
  - Buckets each code into overlay / physical / device / other so
    the check + renderer stay consistent
  - Humanizes codes (NETWORK_ANYNETLINK_DOWN → "SD-WAN overlay
    tunnel down") with a Title-Cased fallback for unknown codes
  - Rolls up (code + severity) tuples so 20 identical alarms show as
    a single line with ×20 and a "just now / Nm / Nh / Nd" age

Rewrites wanAlarms.run() to use those helpers + cross-reference the
site's physical link state so operators aren't left wondering why 20
alarms fired while every metric shows green: overlay flaps get a
"physical WAN paths are all up per Link State" clarifier, and
physical alarms point back at the Link State check. The label loses
its hardcoded "(last 1h)" suffix since the alarm window is now
dynamic (defaults to 24h to match the WAN window).

The follow-up renderer used by /phonestatus imports the same helpers
so the two surfaces cannot drift.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-07-09 10:19:15 -04:00
44ab6559ac Default WAN window to 24h + surface /voicediag in /help
Bumps the shared default WAN look-back from 15m to 24h (1440m) so
both /phonestatus follow-ups and /voicediag pick up a full day of
voice-quality signal by default — better for after-the-fact ticket
triage than a live snapshot. Operators wanting real-time behavior
can set WAN_STANDARD_WINDOW_MINUTES=15 or pass --window 15m to
/voicediag.

Also fills in a long-standing help gap: /voicediag was fully
implemented but never listed in /help. Adds it under "AV & phones"
with usage, examples, and notes covering detail mode, --only,
--window, thresholds, kill switch, HTTP shape, and audit logging.
Updates /help phonestatus to call out the 24h WAN follow-up +
cross-link to /voicediag.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-07-09 09:50:30 -04:00
b802383441 Add Prisma SD-WAN voice-quality enrichment for /phonestatus + /voicediag
Introduces a full Palo Alto Prisma SD-WAN integration (dual-mode SASE
OAuth 2.0 / legacy CloudGenix auth, pagination, 429 backoff, session
priming) that surfaces per-path latency/jitter/loss/MOS, site
healthscore, link state, and alarm data for a store. Wired into the
/phonestatus WAN follow-up and eight new /voicediag WAN checks graded
against ITU-T G.114 / RFC 3550 defaults (env-overridable via
WAN_STANDARD_*).

Also adds a shape-aware detail renderer for /voicediag (per-link
tables with verdict icons instead of a stringified JSON dump) and a
--window flag (15m / 1h / 6h / 24h / 1d, env default via
WAN_STANDARD_WINDOW_MINUTES) so operators can widen the look-back
without redeploying. scripts/prismaProbe.js is bundled as a CLI for
schema iteration against a live tenant.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-07-09 09:45:29 -04:00
d12723d010 Voicediag: store voice standards + port-hygiene checks + apply-all card
Refactor every /voicediag check to declare a top-level `standards`
object so the desired state is legible without reading run() logic
and can drive a documented reference table. Upgrade callForwarding
to error severity, tighten voicemail with three send-to-VM error
paths + a `stop_sending_to_voicemail` remediation, and add a
`disable_hoteling` remediation.

Add a port-hygiene check bucket under services/voiceDiag/checks/port
(portType, portVlan, portPoe, portEnabled) that reuses the phone-
status snapshot to enforce switchport standards. Configurable via
VOICE_STANDARD_PHONE_VLAN (default 102) and VOICE_STANDARD_ENABLED
(kill-switch). Preserve Meraki `portType`/`voiceVlan`/`dataVlan`
through the enrichment chain so the checks have clean data to read.

Add an "apply all N fixes" combined card that shows up when 2+
remediations are available. New confirm_voicediag_all /
cancel_voicediag_all actions run each fix in sequence (readable
audit trail, no per-person write-throttle stacking), accumulate
individual failures into a summary rather than aborting.

Adds regression tests asserting every check exposes .standards,
plus coverage for port checks, kill-switch, and combined-card
iteration. 63 tests in the checks file, 188 total, all green.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-07-08 14:14:38 -04:00
2eb31a2ddc Add /voicediag rules-engine command with 9 per-user calling checks
Introduces a new diagnostic command that walks a registry of check
modules against a store user's Webex Calling configuration and
surfaces per-issue adaptive-card remediation for the fixable ones.

Checks (services/voiceDiag/checks/): dnd, callForwarding, callWaiting,
callIntercept, voicemail, hoteling, executiveAssistant,
outgoingPermission, phoneOnline. Remediations offered for DND,
forwarding, waiting, and intercept.

Uses the /v1/people/{id}/features/* admin surface (spark-admin:people_read
+ spark-admin:people_write scopes we already hold) — the earlier
telephony/config/people/*/callSettings/* path scheme returns 404 from
the Webex gateway and is not a live surface. Runner distinguishes
routing-404s ("URL moved") from "not applicable" 404s ("no calling
license") via the response body.

Arg parser accepts detail/detailed/--detail/--detailed and normalises
macOS smart-dashes so --detailed doesn't die when auto-correct
turns it into an em-dash.

Wires a VOICEDIAG_ACTIONS dispatcher in index.js mirroring the IGMP
branch, and registers /voicediag in commands/registry.js. 170 tests
pass (52 new: 39 check + 12 renderer + 5 arg-normalization).

Docs updated in .env.example, services/phoneService.js:467, and a new
services/voiceDiag/README.md that includes a "how to add a check"
recipe plus a note on the earlier wrong URL scheme.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-07-07 18:51:31 -04:00
1705c88ba2 Widen empty-locations CSV ignore rule
The default --report filename is empty-locations.csv (no suffix),
which the earlier empty-locations-*.csv pattern missed. Broaden
to empty-locations*.csv so the un-suffixed default is also
ignored.
2026-07-07 17:44:19 -04:00
f58628e4a8 Add scripts/findEmptyLocations.js — closed-store cleanup discovery
Finds Webex Calling locations with zero PEOPLE-owned phone numbers
(the "we lost the store users, no one told us" pattern) by diffing
two paginated pulls:
  GET /v1/telephony/config/locations       → every location
  GET /v1/telephony/config/numbers         → every provisioned number
                                             with owner + location

Verdicts per location:
  - has-users           → ≥1 PEOPLE owner (excluded from report)
  - needs-cleanup-first → 0 PEOPLE but workspaces / AA / HG / etc
                          still present; needs Control Hub attention
                          before delete
  - safe-to-delete      → 0 of everything, ghost location shell

Console prints a per-location inventory table (top 20) plus a
summary. --report writes the full set to CSV with location id,
name, address, timezone, and per-owner-type counts.

--execute deletes safe-to-delete locations via
DELETE /v1/telephony/config/locations/{id}. Guarded by a mandatory
--i-am-sure flag and run serial (concurrency=1) with the shared
callWithRetry so 429/503 gets a Retry-After-aware backoff.
--limit / --offset let the operator pilot on a subset. Every
delete produces an audit line under `webex:emptyloc:audit`.

Also added a generic fetchAllPaginated helper to
scripts/lib/webexBulk.js (Link-header cursor pagination, configurable
array key) so subsequent bulk scripts can reuse it.

.gitignore: whitelisted findEmptyLocations.js; added
empty-locations-*.csv to the report-artifact ignore list.
2026-07-07 16:21:46 -04:00
07152a467b Add scripts/removeAdvancedMessaging.js + extract shared bulk lib
New scripts/removeAdvancedMessaging.js reads a Users Export CSV
and bulk-removes the Advanced Messaging and Advanced Space Meetings
licenses from every listed user (with optional add of a Basic
Messaging license, though in most Webex orgs Basic Messaging is a
derived entitlement and no explicit add is required).

Detection is authoritative like the reclaim script: the assignee
rosters of the two Advanced licenses are fetched once up-front,
unioned by email, and the CSV is cross-referenced. PersonIds come
straight off the roster (no per-user /people lookup). Only the
remove ops the user actually still needs are emitted — the PATCH
body is trimmed per user based on which licenses they hold.

Dry-run enumerates every org license whose name matches
/message|advanced|space|basic/i so the operator can discover the
three ids without prior knowledge. --advanced-messaging-license-id,
--advanced-space-meetings-license-id, and --basic-messaging-license-id
also read WEBEX_ADV_MSG_LICENSE_ID / WEBEX_ADV_SPACE_MTG_LICENSE_ID
/ WEBEX_BASIC_MSG_LICENSE_ID from .env if set.

Also extracted the CSV parsing, format detection, pool/retry
helpers, and Webex license helpers from reclaimWebexHosts.js into
a shared scripts/lib/webexBulk.js module. reclaimWebexHosts.js now
imports from it — no behavior change (verified against both CSV
formats: 1028 candidates on the Meetings Inactive Users report,
665 on the Users Export report). Net -106 lines from the reclaim
script.

.gitignore updates:
  - whitelist scripts/lib/ and the new removeAdvancedMessaging.js
    file so they get tracked
  - exclude reclaim-*.csv and remove-*.csv (per-user report CSVs
    generated by --report contain PII and must never be committed)
2026-07-07 15:54:37 -04:00
8777e51d10 Filter users-export by User Status (Inactive or Verified), not days
For the "Users Export" CSV, the target population is any account
Webex has flagged as not currently in use — that's status=Inactive
(previously active, now idle) or status=Verified (never signed in).
"Days since Last Service Accessed" is dropped as a filter criterion
because a Verified user has never signed in and therefore has a
blank days value. --min-days is documented as ignored for this
format.

The candidate record still carries days (nullable) so the sample
line and --report CSV can show it as informational context. Added
a "status" column to the report and to the audit-friendly console
sample.

Also prints every distinct User Status seen with counts, so the
operator can spot surprise values (e.g. the one "FALSE" row in the
current export) before hitting --execute.
2026-07-07 15:37:08 -04:00
8c6de65bb0 Support Control Hub "Users Export" CSV in reclaimWebexHosts
Auto-detect CSV format from the header:
  • meetings-inactive: EMAIL / IS_HOST / DAYS_SINCE_LAST_ACTIVE
    (Analyzer → Meetings → Inactive Users)
  • users-export: "User ID/Email (Required)" /
    "Days since Last Service Accessed"
    (Users → Manage users → Export)

The users-export report has no host flag, but we don't need one —
the authoritative host-holder set comes from the live assignee
roster fetched from the Webex API. Format-B rows with blank
"Days since Last Service Accessed" (never-signed-in accounts,
often generic mailroom/store logins) are intentionally skipped so
they aren't silently reclaimed.

Also stopped upper-casing the header so we can preserve the
punctuation-rich column names Users Export uses verbatim.
2026-07-07 15:34:11 -04:00
c996d5d32e Add scripts/reclaimWebexHosts.js — bulk host license reclaim
Reads a Control Hub "Meetings Inactive Users" CSV, filters to
IS_HOST=Y AND DAYS_SINCE_LAST_ACTIVE > --min-days (default 120),
cross-references against the current holders of --host-license-id
(so no per-user /people lookup), then PATCHes /v1/licenses/users to
atomically remove the host license and either (a) add a specific
free-tier license (--free-license-id) or (b) add attendee-only
siteUrl on --site (--free-attendee).

Dry-run by default; enumerates every license on the site so the
operator can pick the free tier. Bounded concurrency with 429/503
retry, optional --offset/--limit for staged rollouts, per-user
outcome CSV via --report, and full audit trail via the existing
webex:reclaim:audit log scope.

Whitelisted in .gitignore so it stays version-controlled alongside
the other tracked operational scripts.
2026-07-07 15:02:04 -04:00
59460b849b Fix DECT relay container: install node_modules at /workspace, not under agent
The runtime container failed with ERR_MODULE_NOT_FOUND: axios when
integrations/cisco-dect/client.js tried to load. Root cause is
Node's ESM resolver: it walks UP from the IMPORTING file looking
for node_modules, never sideways into siblings.

Container filesystem before:
  /workspace/dect-relay-agent/node_modules/       <- axios lives here
  /workspace/dect-relay-agent/index.js            <- ok, finds it by walking up
  /workspace/integrations/cisco-dect/client.js    <- walks up to /, never sees axios

Node 20.20 has --experimental-detect-module ON by default, so
client.js is still treated as ESM (starts with `import`), and the
resolver correctly reports "cannot find package 'axios'" rather
than syntax-erroring on the import keyword. But it still can't find
the package — the location is wrong.

Fix: install node_modules at /workspace/ so BOTH the agent AND the
shared modules can find it by walking up.

  /workspace/node_modules/                        <- axios here now
  /workspace/package.json                         <- also here, "type":"module" for all descendants
  /workspace/dect-relay-agent/index.js            <- walks up to /workspace/node_modules ✓
  /workspace/integrations/cisco-dect/client.js    <- walks up to /workspace/node_modules ✓
  /workspace/utils/httpDigestAuth.js              <- same ✓

WORKDIR moves from /workspace/dect-relay-agent to /workspace, and
CMD changes accordingly:
  node --enable-source-maps dect-relay-agent/index.js

Rebuild + reship the bundle with `./dect-relay-agent/bundle.sh` and
`./install.sh` on the DC host — it's an idempotent upgrade.

Verified locally (agent modules import cleanly using the same
directory shape as the container).
2026-07-03 10:15:57 -04:00
4f9ebdb5fb Rework DECT relay bundle to ship a pre-built Docker image
The previous packager (scripts/packageDectRelayAgent.js) shipped a
source-only bundle and expected the DC host to build the image with
`docker compose up --build`. That fails hard in corporate DCs with
TLS-intercepted egress: Alpine's apk fetch of dl-cdn.alpinelinux.org
can't verify the intercepted certificate ("apk: TLS: server
certificate not trusted"), and npm install would fail the same way
if apk had succeeded.

New approach: build the image ONCE on the dev machine (where TLS
works), save it as a gzipped tarball, and ship a ZIP whose install
step is `docker load` + `docker compose up -d`. Zero network calls
inside the DC container, ever.

Bundling (dev-machine):
- dect-relay-agent/bundle.sh: build → docker save → gzip → zip.
  Auto-derives version from package.json, records git sha + dirty
  flag + build date into image labels. Cross-arch friendly
  (--platform=linux/amd64 by default; --platform linux/arm64 for
  ARM DCs). Output: dect-relay-agent-bundle-<YYYYMMDD-HHMMSS>.zip
  at repo root (typically 40-60MB).
- dect-relay-agent/Dockerfile: multi-stage node:20-alpine build.
  No apk add. No runtime npm install. Non-root `node` user (uid
  1000). Node handles SIGTERM natively via index.js handlers, so
  no tini/dumb-init needed. Designed to build from the REPO ROOT
  (not the agent folder) because the agent imports shared modules
  from ../integrations/cisco-dect and ../utils.
- dect-relay-agent/Dockerfile.dockerignore: per-Dockerfile ignore
  (BuildKit ≥ 23.0) with a whitelist that keeps the build context
  to ~50KB. Older Docker daemons fall through to the repo-root
  .dockerignore, which already excludes secrets — nothing sensitive
  can leak either way.
- package.json: `npm run package:relay` now invokes bundle.sh.

Runtime (DC-host):
- dect-relay-agent/docker-compose.yml: pins IMAGE_TAG from .env
  (install.sh writes it there — never falls back to :latest), reads
  the rest of the config via env_file, restart: unless-stopped,
  host networking (needed to reach 10.x/8 without userland proxy
  translation, and the agent doesn't listen on anything). Hardened:
  read_only: true rootfs with a 16MB /tmp tmpfs, cap_drop: ALL,
  no-new-privileges, log rotation at 10MB × 5 files.
- dect-relay-agent/install.sh: preflight (docker + compose present,
  daemon reachable, bundle files intact), docker load, pin loaded
  tag into .env, validate .env has the three required values not
  still set to placeholder strings, docker compose up -d, tail last
  40 log lines. Idempotent — safe to re-run on upgrades.

Cleanup:
- scripts/packageDectRelayAgent.js: deleted (superseded).
- .gitignore: drops the scripts/* + !packageDectRelayAgent.js dance
  since we no longer need to whitelist that one file; add pattern
  for the datestamped bundle zips + staging dirs at repo root.
- dect-relay-agent/README.md: replaces the deploy section with the
  new dev-machine-build → DC-host-load workflow, plus a
  troubleshooting section keyed on the exact error messages seen
  during the failed in-DC build (TLS cert not trusted, docker perm
  denied, DIGEST_401).

Verified: all 113 existing tests still pass. Docker build itself
requires a Docker daemon (dev machine) so can't be exercised in
this sandbox — the bash scripts pass `bash -n` syntax checks.
2026-07-03 10:05:05 -04:00
e8500b4324 Package dect-relay-agent as a Docker deploy bundle
Adds a one-command packager (`npm run package:relay`) that produces a
self-contained zip ready to transfer into the data center and start
with `docker compose up -d --build`. Three commands on the DC host:
unzip, edit .env, docker compose up.

Why a packager instead of `docker build` in the repo:
The agent's index.js imports the shared cisco-dect + httpDigestAuth
modules via `../integrations/...` paths, so a naive
`docker build dect-relay-agent/` would fail because those files live
outside the build context. The packager copies them into a
`workspace/` tree inside the bundle so the Dockerfile sees them as
local paths without any source rewriting.

Docker artifacts (in dect-relay-agent/):
- Dockerfile: multi-stage node:20-alpine build (~55MB final image),
  non-root `dect` user (UID/GID 1500), tini as PID 1 for clean
  SIGTERM propagation to node's graceful-shutdown path,
  `npm install --omit=dev --ignore-scripts` in the deps stage.
- docker-compose.yml: restart:unless-stopped, JSON log rotation
  (10MB × 5 files), pgrep-based health check. No `ports:` block
  because the agent is outbound-only (dials the bot).
- .dockerignore: defensive — the bundle already excludes cruft, but
  this hardens against a stray manual build.

Packager (scripts/packageDectRelayAgent.js):
- Assembles agent code + shared modules + deploy artifacts into a
  timestamped staging dir (.package-relay-tmp/, git-ignored).
- Generates a bundle README with three-command deploy instructions,
  ongoing-ops table, no-internet-DC fallback (docker save/load), and
  troubleshooting for the most common failure modes.
- Generates BUNDLE_INFO.txt with build metadata (git sha + dirty
  flag + timestamp + size) so the DC operator can trace deployed
  bundles back to source.
- Emits `dist/dect-relay-agent-bundle-<YYYYMMDD-HHMMSS>.zip` (30KB).
- Cleans staging in a finally block so failed runs don't leak.

Bundle layout (matches Dockerfile expectations):
  dect-relay-agent-bundle-<version>/
    Dockerfile, docker-compose.yml, .dockerignore
    .env.example, README.md, BUNDLE_INFO.txt
    workspace/dect-relay-agent/{package.json, index.js}
    workspace/integrations/cisco-dect/{client,probes,statusXml}.js
    workspace/utils/httpDigestAuth.js

Wiring:
- package.json: new `package:relay` and `test` npm scripts.
- .gitignore: `scripts/` changed to `scripts/*` so `!scripts/
  packageDectRelayAgent.js` can re-include just the packager
  (git forbids re-including files under a fully-excluded directory,
  hence the glob form).
- dect-relay-agent/README.md: rewrites deployment section to show
  the Docker path as the recommended production route, with the
  node-directly path kept for local dev.

Verified end-to-end: `npm run package:relay` produces a valid zip
that unpacks to the expected layout in <2s. All 113 existing tests
still pass.
2026-07-03 09:32:26 -04:00
96b26a5aca DECT relay Phase 1: WSS hub + agent + /phonestatus follow-up
The bot runs in the public cloud and can't reach the 10.x/8 network
where DBS-210 bases live. This phase adds a data-center-resident relay
agent that dials outbound over WSS to the bot, and lets /phonestatus
post a follow-up message with per-base health after its main output
has already shipped.

Bot side (services/):
- dectRelayHub.js: WebSocket upgrade handler on /dect-relay/ws with
  bearer-token auth (constant-time compare, header + Sec-WebSocket-
  Protocol fallback for header-stripping proxies). Promise-based RPC
  API with per-call timeouts, mid-flight-disconnect rejection, and
  clean replacement of a stale agent socket when a newer one connects.
- dectDiscovery.js: pure filter that turns a phoneService result into
  a list of reachable bases. Enforces the "must be on 10.0.0.0/8"
  guardrail per requirements, dedups by IP + MAC, prefers Meraki-live
  IP over Webex-cached IP.
- dectCollectorService.js: fan-out layer over the hub. collectAll()
  runs one RPC per base in parallel with per-base error isolation —
  one bad base never fails the batch.

Phone-status integration:
- Renderer gets a dectFollowUpBaseCount opt that emits an italic
  "diagnostics loading for N base(s)..." hint inside the DECT section
  of the main message.
- New exported renderDectDiagnosticsMarkdown() renders the follow-up
  message: healthy/warning icon per base, uptime + firmware summary,
  structured Power Loss reboot line, and per-base failure hints (e.g.
  "relay accepted the request but the base did not respond in time").
- commands/phoneStatus.js discovers reachable bases synchronously
  (pure), sends the main message, then fires collectAll() and posts
  the follow-up as a separate message. Failures logged, never thrown
  back to the user.
- Chat only: HTTP callers keep their single-message contract.

Agent side (dect-relay-agent/):
- Standalone Node process with its own package.json (only ws, axios,
  dotenv). Reuses the shared integrations/cisco-dect/{client,probes,
  statusXml}.js modules from the parent workspace so there's no code
  duplication.
- Auto-reconnect with exponential backoff + jitter.
- Dispatches collect / reboot / force-reboot / reboot-chain /
  force-reboot-chain / factory-reset / reconfigure-tree.
- DECT admin credentials live ONLY on the agent (never on the bot).
  Shared bearer token gates the WSS handshake.
- README.md covers install, config, wire protocol, and safety model.

Env / infra:
- .env.example: adds DECT_RELAY_AGENT_TOKEN + optional DECT_RELAY_PATH
  and DECT_COLLECT_TIMEOUT_MS. Reframes DECT_TEST_* as the local-dev
  test harness rather than the production path.
- index.js: captures the http.Server from app.listen() and attaches
  the relay hub when DECT_RELAY_AGENT_TOKEN is set; graceful shutdown
  now closes the hub so in-flight RPCs get rejected cleanly.
- Adds "ws" to bot dependencies.

Tests (99 -> 113):
- tests/dectDiscovery.test.js: 13 cases covering the 10.x guardrail,
  MAC normalization, IP source preference, dedup, and warning shape.
- tests/dectRelayHub.test.js: 14 integration cases using a real
  ws pair on an ephemeral 127.0.0.1 port — auth (missing / wrong /
  correct via header / correct via protocol fallback), hello frame,
  RPC round-trip with correlation, agent error surfacing, concurrent
  out-of-order replies, timeout, mid-flight disconnect, replacement
  of a stale socket, and execAction routing.
- tests/renderers.test.js: 8 new cases for the DECT-follow-up loading
  hint (plural / singular / off) and the diagnostics renderer (empty,
  healthy, warning, power-loss dedup, active RTP, error hint, footer).
2026-07-02 17:03:32 -04:00
17a8469592 Add DBS-210 status.xml parser + health verdict
Second half of the DECT spike: the read-side "collector" that turns
a raw /admin/status.xml body into a normalized JS object plus a
pure health verdict. This is what will feed the /phonestatus base-
station diagnostics section once we wire it in.

- integrations/cisco-dect/statusXml.js:
  - xmlToObject(): 60-line hand-rolled parser targeted at the
    DBS-210's flat XML shape. No attributes, no CDATA, no comments
    — so we avoid pulling in a generic XML lib. Throws loudly on
    malformed input.
  - parseRebootLine(): decodes the reboot-log entries the device
    keeps in Reboot_Line_1..6, extracting timestamp + sequence #
    + reason name/code + firmware version. Unrecognized shapes come
    back marked `unrecognized:true` instead of being dropped.
  - parseStatusXml(): grouped, camelCased view of the device state
    (device / firmware / time / multiCell / rebootLog / rtp /
    network / security / emergencyNumbers / features). Every field
    is null-safe.
  - summarizeBaseHealth(): pure-function verdict. Flags recent
    reboots (uptime < 10 min), power-loss events in the log,
    DECT RF conflicts, non-zero rx/tx errors. Splits into
    warnings vs info so consumers can render at the right severity.
- tests/statusXml.test.js: 23 tests covering the parser, the
  reboot-line decoder, the higher-level normalizer, and the health
  verdict — using a REDACTED inline copy of a real status.xml
  captured from a lab base. MAC/IP/RFPI/firmware-server URL are
  all obviously-fake so the fixture is safe to commit.
2026-07-02 15:35:34 -04:00
bc56b0a0fb Add Cisco DBS-210 DECT base spike (HTTP Digest client + safe probes)
Spike scaffolding for reverse-engineering the local admin UI on a
Cisco DBS-210 DECT base station. Not wired into the bot yet -- the
plan is a status.xml data-collector next, then a per-store relay
that fronts these calls over a websocket back to the bot.

- utils/httpDigestAuth.js: dependency-free HTTP Digest MD5/qop=auth
  header builder + WWW-Authenticate parser. Preserves empty realm,
  which the DBS-210 sends and which most libs silently drop.
- integrations/cisco-dect/client.js: axios wrapper with self-signed
  TLS bypass and a single-shot Digest challenge/response interceptor.
- integrations/cisco-dect/probes.js: verified-safe read paths only in
  READ_PROBE_PATHS. Every mutating path is quarantined in the
  MUTATING_ACTION_PATHS map and exposed only via explicit trigger
  helpers (reboot/force-reboot/reboot-chain/factory-reset/reconfigure-
  tree) that fetch and attach the CSRF token from /main.html. The
  legacy /admin/reboot.htm alias -- which triggered a real reboot
  during our first blind probe -- is intentionally NOT reachable.
- tests/httpDigestAuth.test.js: 6 unit tests, including the RFC 2617
  canonical example and the DBS-210 empty-realm quirk.
- .env.example: adds DECT_TEST_BASE_IP / _USER / _PASSWORD /
  _TIMEOUT_MS for the local test harness (script itself lives under
  scripts/, which stays gitignored).
- .gitignore: adds .dect-samples/ so lab captures don't leak.
2026-07-02 15:30:42 -04:00
d112e45eb6 Detect and inline-remediate Meraki IGMP snooping in /phonestatus
DECT basestations rely on multicast for handset discovery/registration.
When Meraki switches have IGMP snooping enabled without a querier —
or per-switch overrides that deviate from a DECT-safe policy — those
frames are pruned and DECT handsets silently fail to register.

- integrations/meraki/switches.js: getSwitchMulticastSettings,
  setSwitchMulticastSettings, and pure summarizeMulticast verdict fn.
- services/phoneService.js: fires the multicast fetch as soon as the
  network id is known, overlapping DECT enrichment; attaches
  data.multicast summary to the collector output. Non-fatal on error.
- services/renderers/phoneStatusRenderer.js: emits a single warning
  line inside the DECT Basestations section only when needsFix is
  true, itemizing which parts deviate (default snoop, default flood,
  N overrides).
- commands/igmpFix.js: frozen DECT_SAFE_MULTICAST_PAYLOAD constant
  ({snoop:false, flood:true, overrides:[]}), adaptive-card builder,
  and confirm/cancel handlers.
- commands/phoneStatus.js: appends the adaptive card when needsFix
  and the trigger came from chat (skipped on HTTP path).
- index.js: IGMP_FIX_ACTIONS set + dispatch branch mirroring the
  HOST_ASSIGN pattern (one-shot pending lookup, censorActionCard,
  domain call).
- utils/pendingIgmpFixes.js: 15-min TTL pending-card store.
- tests: 11 summarizer cases (all deviation permutations + malformed
  input), 4 renderer cases (each warning-line shape + silence when
  needsFix false or no DECT), and a frozen-constant regression guard
  on the PUT payload. Full suite: 47/47 passing.
2026-07-02 14:10:33 -04:00
c15a471959 Tag skipped Jira tickets with bot-skipped label
Prevents the hourly poller from re-classifying the same "not for us"
tickets every hour and burning AI tokens on them forever.

New SKIP_LABEL='bot-skipped' is applied whenever the AI classifier
decides a ticket is out-of-scope (kind='skip', no store number, or an
unroutable kind). The JQL now excludes both bot labels while
preserving the `IS EMPTY OR` union so brand-new unlabeled tickets
still match. Transient failures (AI down, collector 5xx, comment
5xx) intentionally stay unlabeled so they retry next poll.

Labeling is wrapped in a non-throwing helper — a Jira 5xx on the
label call can't abort the batch; the ticket just gets one duplicate
classification next hour, which is far cheaper than dropping the poll.
2026-07-01 17:14:17 -04:00
351f89a9a4 Initial commit: CollabFinder Webex bot
Multi-integration Webex chat/HTTP bot that unifies phone, AV, and
network status for retail store support. Consolidates data from
Webex Calling, Meraki, Workspace ONE (MDM), Atlas AMP, RED digital
signage, and OptiSigns into rich per-store status commands.

Key surfaces:
- /phonestatus, /avstatus — per-store phone & AV device reports with
  clickable Meraki deep-links and per-port detail.
- /webexhost — check/assign Webex Meetings host licenses via the
  Service App; adaptive-card confirmation flow, HTTP-API-gated.
- /offboarduser — full Webex Admin offboarding (auth revoke, device
  wipe, license removal); adaptive-card confirmation.
- /jirapoll — on-demand trigger for the hourly Jira poller.
- /bulkavstatuscsv — bulk store CSV export with concurrency limits.

Automation:
- Hourly Jira poller (node-cron) with an X.AI (Grok) ticket classifier
  that categorizes unassigned tickets as phone/av/skip, extracts store
  numbers from free-text, and enriches Jira with the same detailed
  markdown the chat commands emit (converted to Jira ADF, preserves
  bold + Meraki links). Idempotent via a `bot-enriched` Jira label.

Architecture:
- Node.js 20+, ESM, Express 5, webex-node-bot-framework.
- Layered integrations (integrations/*), services (services/*),
  commands (commands/*), utils (utils/*).
- Shared markdown renderers (services/renderers/*) feed both chat
  handlers and the Jira poller so the two surfaces stay in sync.
- Hand-rolled markdown-to-ADF converter (utils/markdownToAdf.js) —
  no new npm dependency.
- Node built-in test runner (`node --test tests/*.test.js`), 30 tests
  covering the converter, renderers, and poller ADF assembly.

Docker + docker-compose deployment. Config via .env
(see .env.example for the full option surface).
2026-07-01 16:55:03 -04:00