Crawler / Terms
Crawler terms
The conditions under which UnzoiBot fetches your pages, what we retain, and what we will and will not do with it. Operational detail is on the crawler page.
Effective 2026-07-30. Pending legal review — not yet binding.
1. Access
We fetch publicly reachable HTML pages only. We do not authenticate, submit forms, or take any step to reach material behind a paywall, a metering wall or a consent gate. Where such a barrier returns partial content to an anonymous visitor, we index only that partial content.
We honour robots.txt for the UnzoiBot token and for
*. A Crawl-delay slower than our default is applied; a
faster one is ignored. We treat 429 and 503 responses as
instructions to back off.
2. What we retain
Article text, headline, publication date, author line where present, the source URL and the publication name. Text is retained for as long as the article remains in the index and is deleted when it is removed.
We do not retain images, video, stylesheets or scripts, and we do not fetch them.
3. What we serve
API responses contain a headline, a short snippet, the publication name, a timestamp, the language and a link to the original article. Full article text is not served to API users, is not included in bulk exports, and is not sold or licensed onward.
Retained text is used to build the search index and to group articles covering the same event. It is not redistributed.
4. Attribution
Every result identifies the publication and links to the original article. API terms require callers to preserve that attribution and that link wherever results are displayed. Callers may not present results in a way that implies the content originated with them or with us.
5. Removal
A Disallow rule stops future crawling within 24 hours. To also remove
content already indexed, or to remove specific articles, write to us through
the contact form and we will action it within five working
days. We do not require a legal basis for a removal request from a publisher for
their own content.
Removal is permanent for the URLs specified. If the same URLs later become crawlable again and you want them re-included, tell us — we will not silently re-add them.
6. Rate and impact
We commit to a maximum of 20 requests per minute to any single domain, regardless of what your robots.txt permits, and to reducing that on request. If our traffic causes a problem, tell us and we will slow down or stop.
7. Impersonation
Traffic claiming to be UnzoiBot that fails reverse-DNS verification against
crawl.unzoi.com is not ours. Blocking it will not affect your presence
in our index, and we would like to be told about it.
8. Changes
Material changes to these terms — anything altering what we retain, what we serve, or the rate at which we crawl — will be published here with an updated effective date before they take effect.