{"id":19243,"plugin_id":"plugins_6a94afc8a6688191858b8a0f24f04c2c","kind":"skill","collection_source":null,"comparison_source":null,"observed_at":"2026-09-30T23:15:24.416Z","digest":"4c09686e0c63e2dc066b73d14b35176c09a896c4e64b4a7c3d60f6a817a77c99","against":null,"payload":{"description":"Discover and map a network read-only, starting from one device the user names. Crawls outward hop by hop using LLDP, CDP, MNDP, ARP, DHCP leases, routing and MAC tables across MikroTik, Cisco, Aruba, HP, Fortinet, Juniper, Ubiquiti, Linux and Windows, and writes a structured scan record. Use when the user asks to scan, survey, crawl, inventory, audit or 'see everything on' a network or a site, theirs or a customer's.","included_files":[{"relative_path":"agents/openai.yaml","size_in_bytes":185}],"name":"netwalk-scan","skill_md_contents":"---\nname: netwalk-scan\ndescription: Discover and map a network read-only, starting from one device the user names. Crawls outward hop by hop using LLDP, CDP, MNDP, ARP, DHCP leases, routing and MAC tables across MikroTik, Cisco, Aruba, HP, Fortinet, Juniper, Ubiquiti, Linux and Windows, and writes a structured scan record. Use when the user asks to scan, survey, crawl, inventory, audit or 'see everything on' a network or a site, theirs or a customer's.\n---\n\n# netwalk-scan\n\nPart of the **netwalk** read-only network survey toolkit. Toolkit lives at `{{TOOLKIT}}`.\n\nThe goal is a complete, honest picture of the target network — and honesty includes writing down\nwhat you could not reach. A survey that quietly stops at the first ring of neighbours and presents\nitself as complete is worse than one that says \"6 of 9 devices reached, here is why\".\n\n## Read-only, and it is enforced\n\nEvery command goes through `netwalk_exec.py`, which checks it against a per-vendor read-only\nallowlist before it leaves the machine. Config writes, counter clears, service restarts, reboots\nand shell metacharacters are refused by the tool. You never SSH to a surveyed device directly —\nalways go through the wrapper, so the guarantee holds and every command lands in the evidence log.\n\nIf the gate blocks something you believe is genuinely read-only, do not work around it. Either pick\na different command, or add it to the allowlist in `scripts/netwalk_policy.py`, run\n`python3 {{TOOLKIT}}/tests/test_policy.py`, and note the change. Never bypass the wrapper.\n\n## Before you touch anything\n\nAsk, and do not guess:\n\n1. **What is the target?** An entry device (IP), a subnet, a site name, \"my whole office\"? You need\n   at least one device you can log into — netwalk crawls *from* a device. It can also sweep an\n   address range, but only one the owner has explicitly authorised; see **Sweeping a range** below.\n2. **Whose network is it?** If it is a customer's, confirm the user is authorised to log into this\n   equipment today. Record what they say in `site.scope_note` — it goes in the report.\n3. **Anything off limits?** Production boxes that must not even be logged into, a maintenance\n   window, a device that falls over when you open a session. Respect it and record it under\n   `coverage.not_covered`.\n4. **How far?** Default is exhaustive: keep hopping until every reachable neighbour has been\n   visited. On a big site, say up front roughly how many hops that might be and check in.\n\nPick a site slug (`acme-hq`). Everything for the engagement lands in `~/.netwalk/sites/<slug>/` — outside the installed toolkit, so an upgrade cannot delete it.\n\n## The shape of the whole thing\n\nThe crawl and the login form are one loop, not two phases. Every hop turns up devices nobody\nmentioned, and the person who knows what they are is at the browser, not in the conversation:\n\n```\nscan a device  ──►  new neighbours discovered\n      ▲                      │\n      │                      ▼\n      │            put them ALL on the login form  (netwalk_cred.py request --round N)\n      │            each card: credential, or \"I don't know\", or \"not ours\"\n      │                      │\n      └──────  read the answers back (answers) ──┘   log into what you were given\n```\n\n**Never drop a discovered host because you do not recognise it.** It is tempting to filter an ARP\ntable down to the entries whose OUI looks like infrastructure and quietly skip the rest. Do not:\nan unrecognised MAC is not evidence that a device is uninteresting, it is evidence that *you* cannot\nidentify it — which is precisely the question the form exists to ask. This has already gone wrong\nonce in the field: a hypervisor was left off the form because its OUI was not in a lookup table,\nand it was the single most important host on that VLAN.\n\nIf a subnet has more hosts than one page can sensibly carry, put the infrastructure-looking ones on\nfirst, then **say in the same message exactly how many you left off and on what basis**, and offer\nto add them. Silent filtering and honest triage look identical in the output; only one of them is.\n\nStart `netwalk_cred.py serve` once at the beginning and leave it running for the whole survey. Each\nround, push that round's discoveries into the open page with `add` — the user keeps one tab, keeps\none URL, and fills things in at their own pace while the crawl carries on. Never stop and restart\nthe form to add a device: the URL changes and the user loses the page they had open.\nCards the user has already answered are marked as such and carry their previous answer; new ones\nare badged **NEW this round**. Pass `--round N` so the header says which pass this is.\n\nThree answers are not credentials and all three are useful:\n\n| The user picks | What it means | What you do |\n|---|---|---|\n| **I don't know what this device is** | it is on their network and they cannot identify it | record `reachable: false`, `unreachable_reason: \"the site owner could not identify this device\"`, and raise it as a finding — an unidentified device is one of the more valuable things a survey turns up |\n| **Not ours / out of scope** | someone else's equipment | never connect. `netwalk_exec.py` refuses this one in code, so you cannot do it by accident. Keep it on the diagram as a boundary device |\n| **Skip for now** | ask again later | record it, carry on, and put it back on the form next round |\n\nRead every answer back with `answers`, which prints IP, port, username, management URL, jump host,\ntenant and any question you attached — and no secret:\n\n```bash\npython3 {{TOOLKIT}}/scripts/netwalk_cred.py answers --site acme-hq\n```\n\nKeep going until a round turns up no device the user has not already ruled on. That is the\ntermination condition — not \"enough devices\", and not \"the first ring of neighbours\".\n\n## The loop\n\nFor each device, in this order:\n\n1. **Get in.** No credential yet → hand off to `netwalk-login`. Confirm with:\n   `python3 {{TOOLKIT}}/scripts/netwalk_exec.py probe --site <slug> --host <id>`\n\n2. **Identify the vendor** before running anything else. The probe output usually tells you.\n   Wrong vendor means wrong commands and a pile of syntax errors that look like access problems.\n\n3. **Run the discovery pack** for that vendor:\n\n   ```bash\n   python3 {{TOOLKIT}}/scripts/netwalk_exec.py run \\\n     --site acme-hq --host gw01 \\\n     --cmd-file {{TOOLKIT}}/scripts/packs/mikrotik.discovery.txt \\\n     --evidence ~/.netwalk/sites/acme-hq/evidence.jsonl\n   ```\n\n   Packs exist for `mikrotik`, `cisco`, `aruba`, `hp`, `fortinet`, `linux`, `windows`.\n\n   **Controller-managed estates do not get crawled device by device.** UniFi and Omada both know\n   every device they adopted, so read the controller once instead of SSHing into a hundred APs:\n\n   ```bash\n   python3 {{TOOLKIT}}/scripts/netwalk_unifi.py collect --site acme-hq --host unifi-controller --out unifi.json\n   python3 {{TOOLKIT}}/scripts/netwalk_omada.py info    --site acme-hq --host omada-controller\n   python3 {{TOOLKIT}}/scripts/netwalk_omada.py collect --site acme-hq --host omada-controller --out omada.json\n   ```\n\n   Run Omada's `info` first: it hits `/api/info`, which needs no credential, so it separates \"wrong\n   address\" from \"wrong credential\" - the two failures that look identical from the outside. Both\n   adapters take `--via user@host` when the controller only answers from inside the site. For a vendor\n   with no pack, use the `unknown` profile (`show`/`display`/`get`/`print` only) and add commands\n   one at a time with `--cmd`, checking each with `netwalk_exec.py check --vendor ... --cmd ...`.\n   Some commands in a pack will not exist on a given model — a failed command is normal, not a\n   reason to stop.\n\n   For a **config export**, always add `--out <file>`. That writes the full text to a 0600 file\n   and prints only a summary — a config is full of PSKs, community strings and password hashes, and\n   anything printed reaches you, the model API and the transcript. Secret-shaped values in printed\n   output are masked as `<redacted>` as a backstop, but `--out` is the actual control.\n\n4. **Map the output into the scan record**, `~/.netwalk/sites/<slug>/scan-<YYYY-MM-DD>.json`, against\n   `{{TOOLKIT}}/schema/netwalk-record.schema.json`. Per device, capture at minimum:\n\n   - identity: hostname, model, serial, OS + version, uptime, role\n   - every interface: name, description, admin/link state, speed, IPs, VLAN/PVID, error and drop\n     counters, **link-down count** (the best cable-fault signal there is), and the **MAC/CAM table\n     learned on that port**\n   - the device's **own** neighbour table (`/ip neighbor print detail`, `show lldp neighbors detail`\n     …). Run it *on* each device — do not just record what its parent saw about it. Skipping this\n     is how phantom topology gets into diagrams.\n   - ARP table, DHCP leases, VLANs with tagged/untagged ports, routing table\n   - for APs: every SSID with its security mode, VLAN, band/channel/width and client count\n   - for Linux/Windows: running and failed services, listening sockets\n\n5. **Hop.** For every neighbour not yet visited, resolve its vendor and credential and repeat.\n   Try the credential that worked on the previous hop first — one account per site is the common\n   case. If it fails, that one device goes to `netwalk-login`; do not stall the whole crawl.\n   Keep going until the frontier is empty. A neighbour found at hop 4 gets visited at hop 5.\n\n6. **Record dead ends honestly.** Unreachable, no CLI, credential refused, user said don't:\n   `reachable: false` plus a real `unreachable_reason`, and move on.\n\n   Before writing a device off, collect the open questions and send them back through\n   `netwalk-login` in one batch — the credential form takes `--ask`, so \"which port is the\n   controller on\", \"is there a jump host\", \"what is this device on ether16\" are all questions the\n   user answers in the browser in one pass. Asking them one at a time in the chat is the slow way,\n   and the answers are not secrets so `answers` reads them straight back.\n\n## Sweeping a range\n\nThe crawl finds the managed estate — whatever speaks LLDP, appears in an ARP table, or holds a DHCP\nlease. It misses the printer nobody remembers, the old server on a static address, the second\nfirewall someone left plugged in. A sweep is the other half of the picture, and on most sites it is\nwhere the surprises are.\n\n**It cannot run until the owner has authorised the range, by name.** That is enforced in code:\n\n```bash\npython3 {{TOOLKIT}}/scripts/netwalk_sweep.py authorize --site acme-hq \\\n  --range 10.2.30.0/24 --range 10.2.40.0/24 \\\n  --authorized-by \"Khun Somchai, IT manager, by phone 2026-08-22\" \\\n  --exclude 10.2.30.99          # a box he asked us to leave alone\n```\n\nThe name goes in `scope.json` and comes out again in the report. There is no `--force`: a range\noutside the scope is refused, a supernet of an authorised range is refused, public address space\nneeds a second explicit flag, and anything larger than a /16 is refused outright. If you find\nyourself wanting to get past the gate, the answer is to ask the owner, not to edit the file.\n\n```bash\n# which addresses answer at all - TCP connect to a few common ports, plus one ping\npython3 {{TOOLKIT}}/scripts/netwalk_sweep.py hosts --site acme-hq --range 10.2.30.0/24\n\n# what those addresses are listening on - ~68 well-known TCP ports by default\npython3 {{TOOLKIT}}/scripts/netwalk_sweep.py ports --site acme-hq --target 10.2.30.99 \\\n  --profile standard            # or --profile quick, or --ports 22,80,8000-8010\n\n# fold the results into the scan record and list what is not in devices[] yet\npython3 {{TOOLKIT}}/scripts/netwalk_sweep.py record --site acme-hq \\\n  --record ~/.netwalk/sites/acme-hq/scan-2026-08-22.json\n```\n\nThree things to know before you read the output:\n\n- **A refused connection proves a host is there**, exactly as well as an open one does. A box with\n  every port closed still appears in the results, and that is deliberate.\n- **The sweep is blind to UDP.** SNMP, DNS over UDP, IPMI, syslog and IKE do not show up at all, and\n  neither does a host whose firewall drops instead of rejecting. `record` writes that limitation\n  into `coverage.not_covered` for you. Do not let the report imply the list is exhaustive.\n- **Ports carry a `risk` note when finding them open is a finding by itself** — telnet, SMB, RDP,\n  VNC, Redis, Winbox, a database answering on a user VLAN. `ports` prints them; you still have to\n  write them into `findings[]` with the port as evidence.\n\n**Where the sweep runs from matters.** By default it runs from your machine, so it only sees what\nyour machine can route to — a management VLAN you are not on will look empty rather than absent:\n\n| Situation | What to use |\n|---|---|\n| you are on the LAN, or on a VPN into it | the default, no extra flags |\n| the range is only reachable from inside the site | `--via user@linux-host` — an `ssh -D` SOCKS tunnel through a Linux box you already have a credential for |\n| the only way in is a MikroTik | RouterOS refuses dynamic forwarding, so `--via` cannot work. Run the sweep on the router instead: `netwalk_exec.py run --cmd '/tool ip-scan address-range=10.2.30.0/24 duration=30s'` — the allowlist requires `duration=`, so it cannot run unbounded |\n\nEvery address that answers and is not already in `devices[]` goes on the credential form, including —\nespecially — the ones you cannot identify. `record` prints that list for you. The rule from the\ncrawl applies unchanged here: an unrecognised host is not evidence the host is dull.\n\n## Deriving topology\n\n`topology_edges` is what the diagram is drawn from, so build it deliberately rather than dumping\nevery neighbour sighting:\n\n- One entry per **physical link**, not per protocol. If A sees B over both CDP and LLDP, that is one\n  edge. Record the port on each end — a diagram without port labels cannot be used to trace a cable.\n- A discovery frame arriving on a *bridge* or *VLAN* interface tells you the device is somewhere in\n  that broadcast domain, not that it is directly attached. Prefer the physical port when one is named.\n- **Suspected unmanaged switch**: one physical port that has learned 2+ MAC addresses but reports 0\n  or 1 LLDP/CDP neighbours. A dumb switch floods discovery frames instead of terminating them, so\n  the two managed devices see each other directly and the thing between them is invisible. Add it as\n  a device with `role: \"unmanaged-switch\"`, `reachable: false`, and edges marked\n  `discovered_via: \"inferred\"`. Say it is inferred; never draw it as if you logged into it.\n- **WAN links**: one entry in `wan_links` per internet uplink, with ISP, interface, IP, speed and\n  whether it is primary or backup. Never merge several uplinks into a single \"INTERNET\" cloud —\n  which link is which is exactly what someone reads the diagram to find out.\n\n## Keep the artefacts fresh as you go\n\nAfter each hop or small batch, re-run the map and report rather than waiting for the crawl to end:\n\n```bash\npython3 {{TOOLKIT}}/scripts/netwalk_map.py    ~/.netwalk/sites/acme-hq/scan-2026-08-22.json -o ~/.netwalk/sites/acme-hq/map.svg\npython3 {{TOOLKIT}}/scripts/netwalk_report.py ~/.netwalk/sites/acme-hq/scan-2026-08-22.json -o ~/.netwalk/sites/acme-hq/report.html\n```\n\nA partial map beats no map, and the user can correct a wrong assumption at hop 2 instead of hop 9.\n\n## Fill in coverage before you finish\n\n`coverage.not_covered` is not optional. Write down anything a reader could reasonably assume was\nchecked and was not: subnets never entered, a Wi-Fi RF survey that did not happen, a device the user\nasked you to leave alone, a vendor whose CLI you could only partly read. This is the section that\nkeeps the report honest.\n\n## Never\n\n- Change configuration. Report the fix; the site owner applies it.\n- Sweep a range that is not in `scope.json`, or hop into a device outside the agreed scope. If the\n  gate refuses a range, get the owner's authorisation and record it — never route around it.\n- Put a credential anywhere in the scan record.\n- Present an incomplete crawl as complete.\n\nNext: `netwalk-diag` for health and root cause, `netwalk-map` for the diagram,\n`netwalk-fullreport` for the deliverable.\n"},"changes":[],"summary":"First saved snapshot. No earlier version is available for comparison.","summary_kind":"deterministic","summary_metadata":{}}