Crawler / Terms

Crawler terms

The conditions under which UnzoiBot fetches your pages, what we retain, and what we will and will not do with it. Operational detail is on the crawler page.

Effective 2026-07-30. Pending legal review — not yet binding.

1. Access

We fetch publicly reachable HTML pages only. We do not authenticate, submit forms, or take any step to reach material behind a paywall, a metering wall or a consent gate. Where such a barrier returns partial content to an anonymous visitor, we index only that partial content.

We honour robots.txt for the UnzoiBot token and for *. A Crawl-delay slower than our default is applied; a faster one is ignored. We treat 429 and 503 responses as instructions to back off.

2. What we retain

Article text, headline, publication date, author line where present, the source URL and the publication name. Text is retained for as long as the article remains in the index and is deleted when it is removed.

We do not retain images, video, stylesheets or scripts, and we do not fetch them.

3. What we serve

API responses contain a headline, a short snippet, the publication name, a timestamp, the language and a link to the original article. Full article text is not served to API users, is not included in bulk exports, and is not sold or licensed onward.

Retained text is used to build the search index and to group articles covering the same event. It is not redistributed.

4. Attribution

Every result identifies the publication and links to the original article. API terms require callers to preserve that attribution and that link wherever results are displayed. Callers may not present results in a way that implies the content originated with them or with us.

5. Removal

A Disallow rule stops future crawling within 24 hours. To also remove content already indexed, or to remove specific articles, write to us through the contact form and we will action it within five working days. We do not require a legal basis for a removal request from a publisher for their own content.

Removal is permanent for the URLs specified. If the same URLs later become crawlable again and you want them re-included, tell us — we will not silently re-add them.

6. Rate and impact

We commit to a maximum of 20 requests per minute to any single domain, regardless of what your robots.txt permits, and to reducing that on request. If our traffic causes a problem, tell us and we will slow down or stop.

7. Impersonation

Traffic claiming to be UnzoiBot that fails reverse-DNS verification against crawl.unzoi.com is not ours. Blocking it will not affect your presence in our index, and we would like to be told about it.

8. Changes

Material changes to these terms — anything altering what we retain, what we serve, or the rate at which we crawl — will be published here with an updated effective date before they take effect.