Approach

Start with the question the tools cannot answer.

We look at the running system, not the architecture slide. The work is done when both sides agree the next step, not when a deck is sent.

Five steps from gap to next agreed slice.
  1. Name the gap

    What can you not prove during an incident? Latency by customer? Whether a device in a remote POP is even in inventory? Write that down before anyone touches config.

  2. Read the live path

    Collectors, scrape jobs, event rules, SNMP credentials, scan ranges, log volume, sampling, boards people actually opened, and one recent ticket. The last outage is the spec.

  3. Write the constraints early

    Access delays. Firewall approvals. License keys that have not landed. Cardinality you cannot drop without an app change. On-call that is already underwater. Those go in the SOW so week two is not a surprise.

  4. Change the smallest set that closes the gap

    A recording rule, a resource attribute, a drop filter, a better board, an alert that pages the right person, a discovery range that matches production. We will say no to a platform swap if the current tools can do the job.

  5. Leave artifacts the team can own

    Config in Git when the repo is ready. A short runbook. An alert list with a reason for each page. If the next person cannot change it without us, the engagement failed.

Patterns we reuse

Event catalog with owners

Every page that still fires has a reason and a team. Rules without either get retired. Proof is a past outage, not a dashboard screenshot.

Device certification

Standard polling first. Gear not on a vendor supported list goes through a test lab: MIBs, OID targets, trap samples, then a written map into the event path. Production devices are not the lab.

Reachability as a written interface

Collectors only watch what they can reach. Ports, sources, and approval time go on one page so discovery scope is not a guess after kickoff.

Tenant onboarding checklist

Subnets, credentials, ticket groups, and what “accepted” means, sized by how much of the estate is in the first scan. The client owns the network facts. We own the platform slice.

Acceptance tied to a real failure

“Platform is up” is not acceptance. A service-impacting event has a name: dashboard gone, ticket bridge dead, or discovery blind on a stated share of the estate.

Two-room writeup

A short executive summary, then the technical notes. Same facts. Different length. Recommendations and next steps at the end of both.

Services · Contact