API documentation · 8 of 9

Custom search and extraction

Run your own checks on every page a crawl fetches: find the pages that contain, or lack, some text, and pull values out of each page with a CSS selector, XPath or a regular expression. The same rules are on the project's Custom tab.

A project can have up to 25 searches and 25 extractions. Rules run on every HTML page from the next crawl after they are added.

GET/projects/{id}/custom-rulesread

The project's rules, with how many pages of the latest crawl each one matched.

{ "rules": [ { "id": 3, "kind": "extract", "name": "Price", "syntax": "css", "mode": "text", "pattern": ".price", "matched_pages": 42 } ] }

POST/projects/{id}/custom-ruleswrite

Adds a rule and answers 201 with it. A pattern that does not compile answers 400 invalid_rule with the reason.

FieldSearchExtraction
kindsearchextract
nameA name to know it by, up to 100 characters.
patternThe text or regex to look for.The CSS selector, XPath or regex.
syntaxtext (default, ignores case) or regexcss (default), xpath or regex
modecontains (default) or not_containstext (default), html or attribute
scopehtml (default, the source) or text (what a reader sees)—
attribute—The attribute to take, with mode: attribute.
# Pages still showing "Out of stock"
curl -X POST http://localhost:9000/api/v1/projects/1/custom-rules -H "Authorization: Bearer $KEY" \
  -d '{"kind": "search", "name": "Out of stock", "pattern": "out of stock", "scope": "text"}'

# Every product's price
curl -X POST http://localhost:9000/api/v1/projects/1/custom-rules -H "Authorization: Bearer $KEY" \
  -d '{"kind": "extract", "name": "Price", "syntax": "css", "pattern": ".price"}'

# Every page's Open Graph image
curl -X POST http://localhost:9000/api/v1/projects/1/custom-rules -H "Authorization: Bearer $KEY" \
  -d '{"kind": "extract", "name": "OG image", "syntax": "xpath", "pattern": "//meta[@property=\"og:image\"]", "mode": "attribute", "attribute": "content"}'

A regex extraction takes the first capture group, or the whole match when there is none. Up to 50 values are kept per page.

DELETE/projects/{id}/custom-rules/{rule}write

Removes a rule and everything it found. Answers 204.

GET/projects/{id}/custom-rules/{rule}/resultsread

The pages the rule matched in the latest crawl, paginated with ?page=. For a search, count is how many times the pattern occurs; for an extraction, values are what it pulled out.

{
  "crawl_id": 7,
  "rule": { "id": 3, "kind": "extract", "name": "Price", ... },
  "pager": { "page": 1, "total_pages": 2 },
  "pages": [
    { "id": 5501, "url": "https://example.com/shoe", "title": "Red shoe", "status_code": 200, "count": 1, "values": ["€ 10.00"] }
  ]
}