Acquire
HTML list sources
Crawl a listing page; keep or drop article URLs by path.
Type = HTML list. Feed URL is the listing page (section index), not a single article. We extract article links, then optionally enrich each publisher page.
| Field | Required | Notes |
|---|---|---|
| Feed URL | Yes | https URL of the list page. |
| URL must contain | No | One substring per line (e.g. /2026/). A link must match at least one line if any are set. |
| URL must not contain | No | One substring per line (e.g. /page/). Matching links are dropped. |
| Language / category / poll / active | Yes | Same as other source types. |
Use Test before save to confirm we can fetch and how many links we would take. Advanced JSON is the same filters as url_must_contain / url_must_not_contain.
32 articles on this site. Engineering notes on GitHub.
