Procedure

Site-to-site VPN incident: collect diagnostics and restore service

Collect useful evidence during a site-to-site VPN incident, isolate whether the failure is connectivity, negotiation, routing, policy, or the peer, and restore service without destroying diagnostic evidence.

Objective

Collect useful evidence during a site-to-site VPN incident, isolate whether the failure is connectivity, negotiation, routing, policy, or the peer, and restore service without destroying diagnostic evidence.

Prerequisites

  • Known addressing, VLAN, routing and firewall context for the affected flow.
  • A representative test client and the exact protocol or port involved.
  • Administrative access appropriate to the system being changed or diagnosed.
  • A clearly identified scope: affected users, systems, addresses, services and the time of the observed problem.
  • A maintenance or test window when the procedure can affect production traffic or availability.
  • A copy of the current configuration or other recovery material before any irreversible action.

Step-by-step procedure

1

Establish the baseline and scope

Before changing anything, reproduce the issue or document the requested change on a representative system. Record the affected users or services, exact time, current configuration, recent changes and a known-good comparison point. This baseline is the reference used to decide whether each later step improves the situation.

Expected result
  • The scope and current state are documented well enough to reproduce or verify the procedure.
2

Collect useful evidence during a site-to-site VPN incident

Work through this operational element in a controlled sequence: Collect useful evidence during a site-to-site VPN incident. Record the current state, verify dependencies, perform the smallest necessary action, and validate its effect before continuing. If the result differs from the expected state, stop and reassess rather than stacking additional changes.

Expected result
  • Evidence for this area is explicit, reproducible and consistent with the intended design.
3

Isolate whether the failure is connectivity

Work through this operational element in a controlled sequence: isolate whether the failure is connectivity. Record the current state, verify dependencies, perform the smallest necessary action, and validate its effect before continuing. If the result differs from the expected state, stop and reassess rather than stacking additional changes.

Expected result
  • Evidence for this area is explicit, reproducible and consistent with the intended design.
4

Negotiation

Work through this operational element in a controlled sequence: negotiation. Record the current state, verify dependencies, perform the smallest necessary action, and validate its effect before continuing. If the result differs from the expected state, stop and reassess rather than stacking additional changes.

Expected result
  • Evidence for this area is explicit, reproducible and consistent with the intended design.
5

Routing

Work through this operational element in a controlled sequence: routing. Record the current state, verify dependencies, perform the smallest necessary action, and validate its effect before continuing. If the result differs from the expected state, stop and reassess rather than stacking additional changes.

Expected result
  • Evidence for this area is explicit, reproducible and consistent with the intended design.
6

Policy

Work through this operational element in a controlled sequence: policy. Record the current state, verify dependencies, perform the smallest necessary action, and validate its effect before continuing. If the result differs from the expected state, stop and reassess rather than stacking additional changes.

Expected result
  • Evidence for this area is explicit, reproducible and consistent with the intended design.
7

Or the peer

Work through this operational element in a controlled sequence: or the peer. Record the current state, verify dependencies, perform the smallest necessary action, and validate its effect before continuing. If the result differs from the expected state, stop and reassess rather than stacking additional changes.

Expected result
  • Evidence for this area is explicit, reproducible and consistent with the intended design.
8

Restore service without destroying diagnostic evidence

Work through this operational element in a controlled sequence: restore service without destroying diagnostic evidence. Record the current state, verify dependencies, perform the smallest necessary action, and validate its effect before continuing. If the result differs from the expected state, stop and reassess rather than stacking additional changes.

Expected result
  • Evidence for this area is explicit, reproducible and consistent with the intended design.
9

Validate the complete service

Repeat the original user, system or application workflow from the real source and verify the complete result, not only one command or one local check. Confirm that logs and monitoring show the expected behavior and that no temporary debug, bypass, test account, rule or maintenance setting remains enabled.

Expected result
  • The end-to-end service works or the remaining failure is isolated to a clearly identified component.

Technical commands from the original procedure

These technical blocks are preserved byte-for-byte from the historical procedure and kept in their original order. Review names, addresses, paths and parameters before use.

Technical block 1
get vpn ipsec tunnel summary
Technical block 2
diagnose vpn ike gateway list
Technical block 3
diagnose vpn tunnel list
Technical block 4
get router info routing-table details <REMOTE-HOST>
Technical block 5
diagnose sniffer packet any 'host <TEST-HOST> or host <PEER-IP>' 4 50 l
Technical block 6
diagnose vpn tunnel list name <TUNNEL-NAME>

Validation

The procedure is validated when:

  • The original symptom or change request has been tested end to end.
  • The effective configuration matches the intended design and no unexplained error remains in the relevant logs.
  • Temporary troubleshooting controls have been removed and monitoring remains normal.
  • The result, evidence and any follow-up action are documented.

Rollback

  • Restore the configuration, policy, binding, route, credential assignment or service state recorded in the baseline when the change does not meet its success criteria.
  • Remove temporary rules, test objects and diagnostic settings that were introduced only for the procedure.
  • After rollback, repeat the minimum health checks to confirm that the previous service level has been restored.

Troubleshooting / common errors

  • A successful ping does not prove that the application protocol is allowed or listening.
  • When packet and policy evidence disagree, capture the same test flow at more than one point in the path.
  • If the result changes between tests, compare source, destination, identity, time and policy context before changing additional settings.
  • If a command succeeds but the application still fails, continue at the next protocol or application layer instead of widening access.
  • If the expected evidence is missing, verify that logging, auditing and the test path actually cover the failing component.
  • If the change does not improve the measured symptom, restore the previous state and reassess the working hypothesis.

Official and vendor references preserved from the original procedure

♡ 0