# http:crawl
The http:crawl collect plugin walks a site's link graph starting from a
root url, following same-host links (plus any hosts in include-domains)
up to a bounded depth and request count, and returns a map of every
non-2xx response or transport error encountered along the way.
It is a leaf collector (no input) - analogous to http:fetch, but for
auditing a site's link graph rather than fetching a single resource, e.g.
detecting broken links across a site.
# Plugin fields
| Field | Description | Required | Default |
|---|---|---|---|
| url | The root url to start crawling from. | Yes | "" |
| include-domains | Extra hosts to follow links to, in addition to the root url's own host. | No | [] |
| include-urls | Extra paths/urls to fetch regardless of whether they were discovered by crawling - useful for pages not reachable by any link. Resolved relative to url. | No | [] |
| max-depth | How many link-following hops from the root url to traverse. The root url itself is depth 0. | No | 3 |
| limit | Hard cap on the total number of requests issued across the whole crawl. | No | 100 |
| timeout | Per-request timeout, as a Go duration string (e.g. 30s). | No | 30s |
| allow-insecure | Permits http:// urls. Leave unset for production audits - see the ISM-1139 note below. | No | false |
| resolve-env | Resolve ${VAR} references in url from a .env file. See the caveat below. | No | false |
| env-file | Path to the .env file used when resolve-env is true. | No | <project-dir>/.env |
# Common fields
| Field | Description | Required | Default |
|---|---|---|---|
| name | The name/identifier of the plugin - this is the yaml key in the config file when defining the fact. | Yes | - |
| connection | The connection to use for collecting the fact. | No | "" |
| input | A previous input to use when collecting the fact. | No | "" |
| additional-inputs | Additional previous inputs to use when collecting the fact. | No | [] |
# HTTPS required by default
https:// is required unless allow-insecure: true is set. ISM-1139
recommends encrypting data in transit, and defaulting to plain HTTP for an
operator-supplied url risks silently auditing unencrypted traffic. Set
allow-insecure only for local/CI use against a plain-http test server -
never for a production audit target.
# resolve-env reads a .env file, not the OS environment
resolve-env (pkg/env/envresolver.go) reads only a .env file in
the project directory (or the path given by env-file) via
godotenv.Read. It never consults the shipshape process's own OS
environment (os.Environ()). A variable exported in the shell that runs
shipshape run will not be substituted into url - it must be
written to a .env file instead. This is the single most likely source of
confusion when parameterising the crawl root across environments.
# Domain allowlist
Only the root url's host, plus any hosts listed in include-domains, are
followed - matched as an exact host string: subdomains and www vs
apex variants of the same site are not implied, so a redirect from
www.example.gov.au to example.gov.au needs the target added to
include-domains explicitly if you want it crawled.
- A link (
<a href>) to a host outside the allowlist is skipped entirely: not visited, not recorded, not even as a breach. - A redirect to a host outside the allowlist is refused before it is
ever issued, enforced on every hop of a redirect chain, not only the
initial request - so the allowlist cannot be bypassed by an on-allowlist
page redirecting elsewhere. A refusal on a discovered link is logged at
warning level and skipped - not recorded as a breach - because sites
commonly redirect off-host for legitimate reasons (SSO,
www/apex normalisation, link-tracking wrappers), and treating every such redirect as a broken link would make this check impractical on real sites. - The one exception is the root
urlitself: if it redirects to a host outside the allowlist, the crawl fails outright rather than silently visiting nothing and reporting no broken links. Fix by pointingurlat the post-redirect host, or adding the target toinclude-domains.
This keeps a compromised or malicious link (or redirect) on the audited site from causing shipshape to fetch arbitrary third-party infrastructure during an audit run.
# Return format
A map of url to status/error, containing only non-2xx responses and transport failures:
| Key | Value |
|---|---|
| <crawled url> | The HTTP status code as a string (e.g. "404"), or "error: <message>" for a transport-level failure (DNS, connection refused, timeout). |
2xx responses - including 203 and 204 - are never recorded. A site with no
broken links emits an empty map. Pair the output with the not:empty
analyser, which breaches whenever its input is non-empty.
# Example
# Audit a site for broken links by crawling it and flagging any non-2xx
# response or transport error encountered along the way.
#
# `http:crawl` walks the link graph starting from `url`, following only
# same-host links (plus any hosts in `include-domains`) up to `max-depth`
# and a total request `limit`. It emits a map of url -> status/error for
# every non-2xx response or failed request; 2xx responses (including 203
# and 204) are not recorded. This pairs naturally with `not:empty`, which
# breaches whenever its input is non-empty.
#
# `url` here is `${CRAWL_BASE_URL}`, resolved via `resolve-env`. Note this
# reads a `.env` file in the project directory (see
# pkg/env/envresolver.go) - it does NOT read the process's OS environment.
# A real audit would point `url` directly at the site under test, e.g.
# `https://www.example.gov.au`; this example uses an env var so the same
# config can run against a disposable test server in CI (see
# tests/e2e/crawl_test.go) as well as a real site locally.
#
# allow-insecure is set to true here only because the e2e test target
# (tests/e2e/crawl_test.go) is a plain-http httptest server. A real audit
# target should always be https (ISM-1139); omit allow-insecure entirely
# when auditing production infrastructure so an accidental http:// url
# fails loudly instead of being silently crawled unencrypted.
collect:
site-links:
http:crawl:
url: ${CRAWL_BASE_URL}
resolve-env: true
allow-insecure: true
max-depth: 2
clean-links:
http:crawl:
url: ${CRAWL_BASE_URL}/clean
resolve-env: true
allow-insecure: true
max-depth: 2
analyse:
# Breaches: the site links to a page that returns 404.
site-has-no-broken-links:
not:empty:
input: site-links
severity: high
description: "Site has no broken links"
# Passes: the /clean section is a closed subgraph of 2xx pages that never
# reaches the broken /missing link reachable from the site root.
clean-section-has-no-broken-links:
not:empty:
input: clean-links
description: "All crawled pages returned a 2xx response"