API documentation · 8 of 9
Custom search and extraction
Run your own checks on every page a crawl fetches: find the pages that contain, or lack, some text, and pull values out of each page with a CSS selector, XPath or a regular expression. The same rules are on the project's Custom tab.
A project can have up to 25 searches and 25 extractions. Rules run on every HTML page from the next crawl after they are added.
GET/projects/{id}/custom-rulesread
The project's rules, with how many pages of the latest crawl each one matched.
{ "rules": [ { "id": 3, "kind": "extract", "name": "Price", "syntax": "css", "mode": "text", "pattern": ".price", "matched_pages": 42 } ] }
POST/projects/{id}/custom-ruleswrite
Adds a rule and answers 201 with it. A pattern that does not compile answers 400 invalid_rule with the reason.
| Field | Search | Extraction |
|---|---|---|
kind | search | extract |
name | A name to know it by, up to 100 characters. | |
pattern | The text or regex to look for. | The CSS selector, XPath or regex. |
syntax | text (default, ignores case) or regex | css (default), xpath or regex |
mode | contains (default) or not_contains | text (default), html or attribute |
scope | html (default, the source) or text (what a reader sees) | — |
attribute | — | The attribute to take, with mode: attribute. |
# Pages still showing "Out of stock"
curl -X POST http://localhost:9000/api/v1/projects/1/custom-rules -H "Authorization: Bearer $KEY" \
-d '{"kind": "search", "name": "Out of stock", "pattern": "out of stock", "scope": "text"}'
# Every product's price
curl -X POST http://localhost:9000/api/v1/projects/1/custom-rules -H "Authorization: Bearer $KEY" \
-d '{"kind": "extract", "name": "Price", "syntax": "css", "pattern": ".price"}'
# Every page's Open Graph image
curl -X POST http://localhost:9000/api/v1/projects/1/custom-rules -H "Authorization: Bearer $KEY" \
-d '{"kind": "extract", "name": "OG image", "syntax": "xpath", "pattern": "//meta[@property=\"og:image\"]", "mode": "attribute", "attribute": "content"}'
A regex extraction takes the first capture group, or the whole match when there is none. Up to 50 values are kept per page.
DELETE/projects/{id}/custom-rules/{rule}write
Removes a rule and everything it found. Answers 204.
GET/projects/{id}/custom-rules/{rule}/resultsread
The pages the rule matched in the latest crawl, paginated with ?page=. For a search, count is how many times the pattern occurs; for an extraction, values are what it pulled out.
{
"crawl_id": 7,
"rule": { "id": 3, "kind": "extract", "name": "Price", ... },
"pager": { "page": 1, "total_pages": 2 },
"pages": [
{ "id": 5501, "url": "https://example.com/shoe", "title": "Red shoe", "status_code": 200, "count": 1, "values": ["€ 10.00"] }
]
}