Current-state review
Events, metrics, logs, traces, flow, SNMP, collectors, and alert routes. We read one recent incident and mark where correlation failed and where a page had no owner.
Deliverable: findings memo, risk list, recommended next slice. Usually one to two weeks.
Correlation and noise
Symptoms grouped to a cause. Deduplication before a ticket exists. Resource attributes and device identity that survive the hop between tools.
Deliverable: a smaller firing list with reasons, and proof against a past outage.
Predictive and agentic operations
Where the data supports it: earlier warning, recommended actions, and bounded automation. Where it does not: we say so. An agent that pages the wrong team is worse than no agent.
Deliverable: a working path in your tools, guardrails, and what stays human.
Network discovery and device certification
Scan ranges that match what ops actually manages. Standard polling first. A written path for gear that is not on a vendor supported list: test lab, MIBs, trap samples, then production.
Deliverable: discovery notes, certified device list, and devices that still need a custom poll.
Telemetry and cost
OpenTelemetry, Prometheus, Grafana, Loki, Tempo, and commercial operations platforms when they are already the system of record. Cut volume nobody queries. Keep the signal an agent can use.
Deliverable: before/after volume notes and the exact drop or filter.
Controlled platform builds
Operations platforms that have to live in a locked-down account. Assumptions, firewall reachability, and exclusions in the SOW so access delays do not surprise week two.
Deliverable: hardened install notes, acceptance tests tied to a service-impacting event, next-tenant onboarding slice.