API documentation · 6 of 9

Crawls

A crawl fetches the project's site, page by page, and turns what it found into a report. It runs in the background; you start it, follow it, and read the results when it is done.

POST/projects/{id}/crawlswrite

Starts a crawl and answers 202 Accepted straight away, with the project and its new crawl:

curl -X POST http://localhost:9000/api/v1/projects/1/crawls -H "Authorization: Bearer $KEY"
"last_crawl": { "id": 4, "start": "2026-09-29T12:40:00Z", "end": null, "crawling": true, "total_urls": 0, ... }

A project with basic_auth needs the site's credentials in the body. They are used for this crawl only and never stored:

curl -X POST http://localhost:9000/api/v1/projects/1/crawls -H "Authorization: Bearer $KEY" \
  -d '{"username": "staging", "password": "..."}'
  • There is no page or time limit: the crawl fetches the whole site, however big, until it finishes or you stop it.
  • Starting a crawl replaces the previous crawl's results, as the dashboard's Crawl Now does.
  • 409 crawl_in_progress if the project is already being crawled.
  • 400 credentials_required for a basic_auth project without a username.
  • Each address may start 5 crawls a minute; past that you get 429.

Following a crawl

Poll GET /projects/{id} every few seconds. While it runs, last_crawl.crawling is true and total_urls climbs. When it is done, crawling is false, end is set, and the results are ready.

POST/projects/{id}/crawls/stopwrite

Asks a running crawl to stop. What it fetched so far is still turned into a report, so it answers 202 and crawling turns false once that is done. Stopping a project that is not crawling is not an error.

curl -X POST http://localhost:9000/api/v1/projects/1/crawls/stop -H "Authorization: Bearer $KEY"

GET/projects/{id}/crawlsread

The project's recent crawls, newest first, in the shape of last_crawl:

{ "crawls": [ { "id": 4, "start": "...", "end": "...", "crawling": false, "total_urls": 87, ... } ] }