Acquire

HTML list sources

Crawl a listing page; keep or drop article URLs by path.

Type = HTML list. Feed URL is the listing page (section index), not a single article. We extract article links, then optionally enrich each publisher page.

FieldRequiredNotes
Feed URLYeshttps URL of the list page.
URL must containNoOne substring per line (e.g. /2026/). A link must match at least one line if any are set.
URL must not containNoOne substring per line (e.g. /page/). Matching links are dropped.
Language / category / poll / activeYesSame as other source types.

Use Test before save to confirm we can fetch and how many links we would take. Advanced JSON is the same filters as url_must_contain / url_must_not_contain.

Start free

32 articles on this site. Engineering notes on GitHub.