{"id":19235,"plugin_id":"plugins_6a94afc8a6688191858b8a0f24f04c2c","kind":"skill","collection_source":null,"comparison_source":null,"observed_at":"2026-09-30T23:15:24.180Z","digest":"7b3bc36e5d2b54dd295b713796fb3709dcdf50c1aef27af80a5e3c785772eaa3","against":null,"payload":{"description":"Read a network device or Linux/Windows server end to end and work out what is wrong with it. Exports config read-only, collects CPU, memory, storage, temperature, PoE, interface counters, error and flap counts, throughput, sessions, services and logs, then reasons from that evidence to concrete findings with severity and recommendations. Use when the user asks what is wrong with a device or site, wants a health check, or has a symptom (slow, dropping, rebooting, flapping) to chase.","included_files":[{"relative_path":"agents/openai.yaml","size_in_bytes":185}],"name":"netwalk-diag","skill_md_contents":"---\nname: netwalk-diag\ndescription: Read a network device or Linux/Windows server end to end and work out what is wrong with it. Exports config read-only, collects CPU, memory, storage, temperature, PoE, interface counters, error and flap counts, throughput, sessions, services and logs, then reasons from that evidence to concrete findings with severity and recommendations. Use when the user asks what is wrong with a device or site, wants a health check, or has a symptom (slow, dropping, rebooting, flapping) to chase.\n---\n\n# netwalk-diag\n\nPart of the **netwalk** read-only network survey toolkit. Toolkit lives at `{{TOOLKIT}}`.\n\n`netwalk-scan` answers *what is out there*. This skill answers *what is wrong with it*.\n\n## Read-only, and it is enforced\n\nConfig is **exported**, never imported or edited. Not one command changes device state — that\nincludes `clear counters`, `dmesg -C`, `systemctl restart` and `debug`, all of which the gate\nrefuses. If a device needs a change, that is a finding with a recommendation, and the site owner\ndecides. Run everything through `netwalk_exec.py` so the guarantee holds and the command lands in\nthe evidence log.\n\n## 1. Collect\n\nAssumes a credential exists (`netwalk-login`) and the vendor is known.\n\n```bash\nT={{TOOLKIT}}\npython3 $T/scripts/netwalk_exec.py run --site acme-hq --host gw01 \\\n  --cmd-file $T/scripts/packs/mikrotik.config.txt \\\n  --out ~/.netwalk/sites/acme-hq/configs/gw01.conf \\\n  --evidence ~/.netwalk/sites/acme-hq/evidence.jsonl\n\npython3 $T/scripts/netwalk_exec.py run --site acme-hq --host gw01 \\\n  --cmd-file $T/scripts/packs/mikrotik.health.txt \\\n  --evidence ~/.netwalk/sites/acme-hq/evidence.jsonl\n```\n\nPacks: `mikrotik`, `cisco`, `aruba`, `hp`, `fortinet`, `linux`, `windows`, each with a\n`.config.txt`, a `.health.txt` and a `.security.txt`. Record the exported config path in\n`config_export_path` — relative to the site folder, never pasted into the report.\n\nRun the security pack too, and run it with `--out` — it deliberately pulls the parts of the\nconfiguration that hold community strings and key material, so its output must land on disk\nrather than in the conversation:\n\n```bash\npython3 $T/scripts/netwalk_exec.py run --site acme-hq --host gw01 \\\n  --cmd-file $T/scripts/packs/mikrotik.security.txt \\\n  --out ~/.netwalk/sites/acme-hq/configs/gw01.security.txt \\\n  --evidence ~/.netwalk/sites/acme-hq/evidence.jsonl\n```\n\n**Always use `--out` for a config export.** With `--out` the full text is written straight to a\n0600 file and only a one-line summary is printed. Without it, the whole config comes back through\nyou — and a vendor config is packed with PSKs, SNMP communities, RADIUS secrets, password hashes and\ncontainer environment variables. Anything that reaches you is transmitted to a model API and written\ninto the session transcript, and neither can be un-sent.\n\n`netwalk_exec.py` masks secret-shaped values in whatever it prints (`<redacted>`), so an accidental\nread is survivable — but that is a safety net, not the control. The control is `--out`.\n\nRecord the export path in `config_export_path`. Never `cat` it, never paste it into the record, and\nwarn the user before they forward it to anyone.\n\n### What to pull, whatever the vendor\n\n| Area | What matters | Why |\n|---|---|---|\n| CPU | current load, and *which process* if the vendor shows it | 90% CPU on one control-plane process is a very different fault from 90% forwarding load |\n| Memory | used vs total, and free trend if available | small routers OOM quietly and reboot |\n| Storage | free space, and flash write counts on RouterOS | a full disk stops logging first and routing second |\n| Temperature / fan / PSU | current reading, thresholds | the cause behind \"it reboots in the afternoon\" |\n| Interfaces | rx/tx errors, CRC, drops, **link-down count**, duplex mismatch, SFP dBm | the single richest source of real faults |\n| Throughput | per-interface bps at sample time | tells you whether \"slow\" is saturation or something else |\n| Sessions | conntrack/session count vs max | a firewall at its session ceiling drops new flows while looking idle |\n| PoE | used vs budget | an AP that reboots under load is often a PoE budget problem |\n| Logs | errors, auth failures, link flaps, resets, DHCP exhaustion | the timeline that connects the rest |\n| Services (servers) | running / failed / restart counts, listening sockets | a service that restarts 40 times an hour is not \"running\" |\n\nFor Linux/Windows also take failed units, `journalctl -p err`, disk health and the process list.\n\n## 2. Analyse — evidence first\n\nEvery finding needs the observation that produced it. No evidence, no finding: write it as a\nquestion for the user instead.\n\nRead the numbers before reaching for a story:\n\n- **Correlate before concluding.** \"CPU is 95%\" is an observation. \"CPU is 95% because the firewall\n  is doing connection tracking for 60k sessions on a box rated for 20k\" is a finding. If you cannot\n  bridge the two, mark `confidence: \"suspected\"` and say what would settle it.\n- **Counters are cumulative.** 5,000 CRC errors over 400 days of uptime is noise; 5,000 since a\n  reboot yesterday is a bad cable. Always divide by uptime before calling something a fault.\n- **Link-down counts beat error counters** for finding a flapping port, and a flapping access port\n  with a phone or AP on it explains a lot of \"the network is slow\" tickets.\n- **Duplex mismatch** shows as late collisions and CRC on one side only.\n- **Match the symptom.** If the user reported something specific, say explicitly whether the data\n  supports it. \"You reported afternoon slowness; the WAN interface is at 94% of 100M every day from\n  13:00\" is useful. So is \"nothing in this data explains it, here is what to capture next.\"\n- **Do not invent severity.** `critical` = it is broken or actively unsafe now. `high` = it will\n  break or is exploitable. `medium` = real risk, not urgent. `low`/`info` = hygiene.\n\n## 2b. Hardening — run the catalogue, do not rely on remembering\n\nHealth is about what is broken. Hardening is about what is configured against the vendor's own\nadvice, and it is not left to memory: `netwalk_audit.py` holds the checklist as data, so the same\nchecks run on every site whether or not anyone thought of them.\n\n```bash\npython3 $T/scripts/netwalk_audit.py guide --vendor mikrotik      # the checklist itself\npython3 $T/scripts/netwalk_audit.py run --site acme-hq \\\n  --record ~/.netwalk/sites/acme-hq/scan-2026-08-22.json --dry-run\n```\n\n`run` reads the config exports **off the disk** and writes findings into `findings[]`. The config\ntext never passes through you; the excerpt attached to each finding is one line, redacted. Drop\n`--dry-run` to write.\n\nThree things about the output matter more than the findings themselves:\n\n- **`NOT CHECKED` is part of the result.** A device with no export on disk, a check whose command is\n  missing from the pack output, and every manual item are all listed by name and written into\n  `coverage.not_covered`. Repeat them in the report. A hardening section that shows six findings and\n  hides that ten checks never ran is worse than no hardening section.\n- **A `config_absent` check is `confidence: \"suspected\"`.** It fires on the *absence* of a line, and\n  absence can mean \"not configured\" or \"not in the part of the config we pulled\". Read it before\n  you put it in front of a customer.\n- **Every finding is `public_safe: false` by default.** A hardening list is a route map. Promote one\n  to the public copy only deliberately.\n\nThe catalogue does not replace judgement. Anything you spot that it has no check for still belongs\nin `findings[]` with `category: \"security\"` — and if it is a class of problem rather than a one-off,\nadd a check to `netwalk_audit.py`, run `python3 $T/tests/test_audit.py`, and it is closed for every\nfuture site instead of just this one.\n\n### The baseline\n\nChecks come from the vendor's own hardening guidance plus what actually goes wrong on sites, and\neach check names the guidance it came from. Where a vendor publishes no guidance worth citing, the\ncheck says it is common practice rather than dressing itself up as a standard.\n\n\n## 3. Write findings into the record\n\nAppend to `findings[]` in the scan record (schema: `{{TOOLKIT}}/schema/netwalk-record.schema.json`):\n\n```json\n{\n  \"id\": \"F7\", \"severity\": \"high\", \"category\": \"availability\", \"host_id\": \"sw-core\",\n  \"title\": \"Port Gi1/0/12 has flapped 214 times since the last reboot\",\n  \"detail\": \"214 link-downs over 26 days of uptime, roughly 8 a day, on the port feeding AP-FL2. Every flap drops that AP's clients.\",\n  \"evidence\": [\n    {\"source\": \"show interfaces Gi1/0/12\", \"excerpt\": \"214 interface resets\", \"observed_at\": \"2026-08-22T10:14:00+07:00\"},\n    {\"source\": \"show version | include uptime\", \"excerpt\": \"uptime is 26 days\"}\n  ],\n  \"confidence\": \"confirmed\",\n  \"recommendation\": \"Replace the patch lead and re-seat both ends, then re-check the counter after a week. If it keeps climbing, move the AP to a different port to isolate port vs cable.\",\n  \"public_safe\": true\n}\n```\n\nMake the recommendation something a technician can act on — which port, which cable, what to check\nafterwards. \"Investigate the interface\" is not a recommendation.\n\nHealth numbers go in `devices[].health`, log lines worth keeping in `devices[].log_excerpts`.\n\n## 4. Report back\n\nSummarise for the user: what is broken now, what will break, what is only untidy — and say plainly\nwhat the data does *not* explain. Then `netwalk-fullreport` turns it into the deliverable.\n\n## Never\n\n- Fix anything. Not even \"while I'm in here\" — an undocumented change during a survey is the worst\n  kind of change.\n- Clear a counter or a log. That destroys the evidence the next engineer needs.\n- Report a finding you cannot point at evidence for.\n- Copy a config export, PSK, community string or password hash into the report.\n"},"changes":[],"summary":"First saved snapshot. No earlier version is available for comparison.","summary_kind":"deterministic","summary_metadata":{}}