# http:crawl

The http:crawl collect plugin walks a site's link graph starting from a root url, following same-host links (plus any hosts in include-domains) up to a bounded depth and request count, and returns a map of every non-2xx response or transport error encountered along the way.

It is a leaf collector (no input) - analogous to http:fetch, but for auditing a site's link graph rather than fetching a single resource, e.g. detecting broken links across a site.

# Plugin fields

Field Description Required Default
url The root url to start crawling from. Yes ""
include-domains Extra hosts to follow links to, in addition to the root url's own host. No []
include-urls Extra paths/urls to fetch regardless of whether they were discovered by crawling - useful for pages not reachable by any link. Resolved relative to url. No []
max-depth How many link-following hops from the root url to traverse. The root url itself is depth 0. No 3
limit Hard cap on the total number of requests issued across the whole crawl. No 100
timeout Per-request timeout, as a Go duration string (e.g. 30s). No 30s
allow-insecure Permits http:// urls. Leave unset for production audits - see the ISM-1139 note below. No false
resolve-env Resolve ${VAR} references in url from a .env file. See the caveat below. No false
env-file Path to the .env file used when resolve-env is true. No <project-dir>/.env

# Common fields

Field Description Required Default
name The name/identifier of the plugin - this is the yaml key in the config file when defining the fact. Yes -
connection The connection to use for collecting the fact. No ""
input A previous input to use when collecting the fact. No ""
additional-inputs Additional previous inputs to use when collecting the fact. No []

# HTTPS required by default

https:// is required unless allow-insecure: true is set. ISM-1139 recommends encrypting data in transit, and defaulting to plain HTTP for an operator-supplied url risks silently auditing unencrypted traffic. Set allow-insecure only for local/CI use against a plain-http test server - never for a production audit target.

# resolve-env reads a .env file, not the OS environment

resolve-env (pkg/env/envresolver.go) reads only a .env file in the project directory (or the path given by env-file) via godotenv.Read. It never consults the shipshape process's own OS environment (os.Environ()). A variable exported in the shell that runs shipshape run will not be substituted into url - it must be written to a .env file instead. This is the single most likely source of confusion when parameterising the crawl root across environments.

# Domain allowlist

Only the root url's host, plus any hosts listed in include-domains, are followed - matched as an exact host string: subdomains and www vs apex variants of the same site are not implied, so a redirect from www.example.gov.au to example.gov.au needs the target added to include-domains explicitly if you want it crawled.

  • A link (<a href>) to a host outside the allowlist is skipped entirely: not visited, not recorded, not even as a breach.
  • A redirect to a host outside the allowlist is refused before it is ever issued, enforced on every hop of a redirect chain, not only the initial request - so the allowlist cannot be bypassed by an on-allowlist page redirecting elsewhere. A refusal on a discovered link is logged at warning level and skipped - not recorded as a breach - because sites commonly redirect off-host for legitimate reasons (SSO, www/apex normalisation, link-tracking wrappers), and treating every such redirect as a broken link would make this check impractical on real sites.
  • The one exception is the root url itself: if it redirects to a host outside the allowlist, the crawl fails outright rather than silently visiting nothing and reporting no broken links. Fix by pointing url at the post-redirect host, or adding the target to include-domains.

This keeps a compromised or malicious link (or redirect) on the audited site from causing shipshape to fetch arbitrary third-party infrastructure during an audit run.

# Return format

A map of url to status/error, containing only non-2xx responses and transport failures:

Key Value
<crawled url> The HTTP status code as a string (e.g. "404"), or "error: <message>" for a transport-level failure (DNS, connection refused, timeout).

2xx responses - including 203 and 204 - are never recorded. A site with no broken links emits an empty map. Pair the output with the not:empty analyser, which breaches whenever its input is non-empty.

# Example

# Audit a site for broken links by crawling it and flagging any non-2xx
# response or transport error encountered along the way.
#
# `http:crawl` walks the link graph starting from `url`, following only
# same-host links (plus any hosts in `include-domains`) up to `max-depth`
# and a total request `limit`. It emits a map of url -> status/error for
# every non-2xx response or failed request; 2xx responses (including 203
# and 204) are not recorded. This pairs naturally with `not:empty`, which
# breaches whenever its input is non-empty.
#
# `url` here is `${CRAWL_BASE_URL}`, resolved via `resolve-env`. Note this
# reads a `.env` file in the project directory (see
# pkg/env/envresolver.go) - it does NOT read the process's OS environment.
# A real audit would point `url` directly at the site under test, e.g.
# `https://www.example.gov.au`; this example uses an env var so the same
# config can run against a disposable test server in CI (see
# tests/e2e/crawl_test.go) as well as a real site locally.
#
# allow-insecure is set to true here only because the e2e test target
# (tests/e2e/crawl_test.go) is a plain-http httptest server. A real audit
# target should always be https (ISM-1139); omit allow-insecure entirely
# when auditing production infrastructure so an accidental http:// url
# fails loudly instead of being silently crawled unencrypted.
collect:
  site-links:
    http:crawl:
      url: ${CRAWL_BASE_URL}
      resolve-env: true
      allow-insecure: true
      max-depth: 2

  clean-links:
    http:crawl:
      url: ${CRAWL_BASE_URL}/clean
      resolve-env: true
      allow-insecure: true
      max-depth: 2

analyse:
  # Breaches: the site links to a page that returns 404.
  site-has-no-broken-links:
    not:empty:
      input: site-links
      severity: high
      description: "Site has no broken links"

  # Passes: the /clean section is a closed subgraph of 2xx pages that never
  # reaches the broken /missing link reachable from the site root.
  clean-section-has-no-broken-links:
    not:empty:
      input: clean-links
      description: "All crawled pages returned a 2xx response"