Our crawler
NearMintBot
If you are a site owner and you have found this page in your logs, this explains exactly what our crawler does, and how to stop it.
User-Agent
NearMintBot/1.0 (+https://near-mint.invalid/bot)
The domain is a placeholder. No domain is registered yet, so the crawler names a reserved address that cannot resolve rather than a real one it does not own. This line changes to the live domain, and this page moves with it, before the crawler runs against any third-party host in earnest.
What it does
- Reads public RSS and Atom feeds, twice a day, at 07:05 and 19:05 UK time.
- Takes the headline and the standfirst, and nothing else. It never fetches article bodies.
- Obeys robots.txt, checked against our own token rather than * — so if you name us specifically, that is the rule we follow.
- Obeys Crawl-delay, and never makes concurrent requests to one host.
- Deletes the text it stores after seven days. Only the outlet name, the link and the date are kept.
- Never renders your headline as ours. Everything a reader sees is written by us and links back to you.
What it will not do
- It will not change its User-Agent to get past a block. If your edge refuses an honest bot, we record the refusal and use your press office instead.
- It will not fetch a page whose robots.txt we could not read. Unknown is not permission.
- It will not go near paywalled content, archives or mirrors.
How to stop it
User-agent: NearMintBot Disallow: /
It takes effect on our next run, within twelve hours. If you would rather talk to a person, or you would like us to crawl something we currently do not, the contact address is on the About page.
Where we do not go, and why
This list is published because a refusal is a fact about the world, and hiding it would make our source list look more complete than it is. It is read by the crawler at runtime, so it is enforced rather than promised.
| Host | Status | What we do instead |
|---|---|---|
| Heritage Auctionsha.com | refuses honest botHTTP 403 on the homepage AND on /robots.txt, across www, movieposters and comics subdomains. An edge/WAF block on the User-Agent, not a robots.txt rule. | press releases and published results summaries, entered by hand |
| Bonhamsbonhams.com | refuses honest botHTTP 403 on homepage and /robots.txt. | press office; hand entry |
| Christie'schristies.com | refuses honest botTimed out again on 18 Sep 2026 — third independent confirmation after two timeouts on 16 Sep. | press releases and press coverage |
| Roseberysroseberys.co.uk | challenge interstitialHTTP 202 with a zero-byte body on homepage and /robots.txt — a bot challenge. | hand entry |
| Ewbank'sewbankauctions.co.uk | refused by robotsrobots.txt disallows our path. | hand entry |
| Invaluableinvaluable.com | terms unreadablerobots.txt allows and lot pages serve prices in the served HTML (622 price strings on a lot page, 20 Sep 2026), but the terms page will not be read: /inv/agreements/terms-of-use/ and /agreements/terms-of-use/ both return HTTP 202 with a zero-byte body to an honest bot. A terms page we cannot read is not permission. | Ask Invaluable for express written permission. The 202 is a bot challenge on the terms page, not a stated policy. |
| LiveAuctioneersliveauctioneers.com | refused by termsTerms (https://www.liveauctioneers.com/termsandconditions): the licence is 'for your personal non-commercial use', and 'You agree that you will not use any robot, spider, scraper, or other automated means to access the sites for any purpose without our express written permission.' robots.txt allows; the terms refuse, and robots.txt is not the whole permission. | Ask for express written permission. Catalogue pages do serve prices (91 in the served HTML), so a yes is worth having. |
| Artpriceartprice.com | blocks ai crawlersrobots.txt, 16 Sep 2026 | — |
| Antiques Trade Gazetteantiquestradegazette.com | refuses honest botHTTP 403 with a Cloudflare challenge page on every path tested: /rss/news/, /rss/, /rss, /feed, /feed/, /rss.xml, /news/rss/, /atom.xml, /umbraco/rss, /sitemap.xml and the homepage. Its robots.txt ALLOWS us — it disallows only /app_data, /app_plugins/, /install, /bin and /umbraco/ — so this is an edge block, not a stated policy. | ASK THEM. Their robots.txt permits crawling, so no editorial decision has been taken to exclude us; a CDN rule has. A named, honest crawler is the kind of thing a trade publisher will often allow on request. This is the highest-value permission to obtain in the whole source list. |
| Google News RSSnews.google.com | refused by robotsrobots.txt is 'User-agent: * / Disallow: /' with an allowlist covering /home, /topics/, /publications/, /stories/, /swg/, /about — not /rss/. CCBot, GPTBot, ChatGPT-User, PerplexityBot, anthropic-ai, ClaudeBot and Claude-Web are each named with Disallow: /. | None. The Record Book needs a different discovery mechanism. |
| Artnet Newsartnet.com | refuses honest botHTTP 403 with a Cloudflare challenge on every path tested, including /robots.txt itself, on both news.artnet.com and www.artnet.com (20 Sep 2026). The feed returned 10 items on 16 Sep and 'failed XML parse' on 18 Sep; that parse failure was the challenge page, not a malformed feed - the same mistake an earlier probe made about Antiques Trade Gazette. | Ask for crawler access. Unlike ATG, its robots.txt is blocked too, so no stated policy can be read. |
| the-saleroomthe-saleroom.com | refused by termsWebsite terms forbid it: the licence is 'for its own personal non-commercial use in order to view the Website only', the user may not 'copy, reproduce, modify, communicate to the public, or make derivative product from... any material available on or through the Website', and may not 'create a database in electronic or structured manual form by systematically and/or regularly downloading, caching, printing and/or storing the material... (by spidering or otherwise)'. robots.txt allows the crawl; the terms do not. Checked 20 Sep 2026. | Not read. It aggregates the UK regional houses, so this closes the cheapest route to the GBP 1k-25k band. |
| Chiswick Auctionschiswickauctions.co.uk | challenge interstitialHTTP 202 with a zero-byte body on the homepage and /robots.txt - a bot challenge, the same shape as Roseberys. Checked 20 Sep 2026. | hand entry or ask |
| Forum Auctionsforumauctions.co.uk | challenge interstitialHTTP 202 with a zero-byte body on the homepage and /robots.txt. Checked 20 Sep 2026. | hand entry or ask |
| Woolley & Walliswoolleyandwallis.co.uk | challenge interstitialHTTP 202 with a zero-byte body on the homepage and /robots.txt. Checked 20 Sep 2026. | hand entry or ask |
| Fellowsfellows.co.uk | challenge interstitialHTTP 202 with a zero-byte body on the homepage and /robots.txt. Checked 20 Sep 2026. | hand entry or ask |
| Mallamsmallams.co.uk | refuses honest botHTTP 403 on /robots.txt itself - no file to read, so no permission to infer. Checked 20 Sep 2026. | press office; hand entry |
Where we do go
Every host we may read, and what we take from it. Metadata onlymeans the page serves no price to an honest bot, so its figures come from the house’s press office rather than from the page. Open robots.txt is not permission on its own: a host whose terms forbid automated reading is on the list above, whatever its robots file says.
| Host | Status | What we read |
|---|---|---|
| Sotheby'ssothebys.com | allowedcrawl-delay 15s | metadata only — no price in the served page |
| Phillipsphillips.com | allowed | prices in the served page |
| Dreweattsdreweatts.com | allowed | metadata only — no price in the served page |
| Lyon & Turnbulllyonandturnbull.com | allowed | metadata only — no price in the served page |
| Propstorepropstore.com | allowed | metadata only — no price in the served page |
| Goldingoldin.co | allowed | metadata only — no price in the served page |
| Mandarakeorder.mandarake.co.jp | allowed | metadata only — no price in the served page |
| Friezefrieze.com | allowed | fetch with care |
| ARTnewsartnews.com | blocks ai crawlers | rss headline and standfirst only |
| ICv2icv2.com | allowed | fetch |
| Sworderssworder.co.uk | allowedcrawl-delay 10s | prices in the served page |
| Cheffinscheffins.co.uk | allowed | metadata only — no price in the served page |
| Tennantstennants.co.uk | allowed | metadata only — no price in the served page |
Policy as of 2026-09-20.