# NewOrderOnline.com — robots policy # # WHICH HOST THIS IS FOR (#1205). This file is the policy for the PUBLIC origin, # www.neworderonline.com — and as of the 2026-08-10 cutover (#17) it is what every hostname serves. # The pre-launch nginx `Disallow: /` override on `app` is GONE, `app` and the apex are now # Cloudflare-proxied and 301 to www, and the `Sitemap:` line below publishes for real (#490). # # Cloudflare PREPENDS its Managed robots.txt to this file, so what a crawler actually receives is # Cloudflare's content-signal block followed by everything below. Both declare `User-agent: *`; # crawlers merge matching groups, so the two must not contradict each other — while this file still # said `Disallow: /` and Cloudflare said `Allow: /`, the effective policy was ambiguous and Google # would have resolved the tie toward Allow. Keep them agreeing. # See docs/deployment/lightsail-production.md. # # WHY IT CHANGED. This file used to say "primary AI-scraper / bot defense is at the Cloudflare edge" # and then declare `Allow: /` with two exceptions. That edge has never been in the request path for # the host actually serving, so the only instruction with any effect was the invitation. On # 2026-08-04 crawlers took 99.1% of requests and 99.9% of bytes served, and OOM-killed the app. # robots.txt is a request, not a control: it sits under the edge and the app's own output caching, # and the crawlers that caused the outage are exactly the ones most likely to ignore it. # ── AI training / retrieval scrapers ────────────────────────────────────────────────────────── # Blocked outright: they spend bandwidth and origin CPU and return no search referrals. The # metered constraint is real — this is a 3 TB/month tier and one day of that traffic cost 5 GB. User-agent: GPTBot User-agent: OAI-SearchBot User-agent: ChatGPT-User User-agent: ClaudeBot User-agent: Claude-Web User-agent: anthropic-ai User-agent: PerplexityBot User-agent: Perplexity-User User-agent: CCBot User-agent: Bytespider User-agent: Amazonbot User-agent: Applebot-Extended User-agent: meta-externalagent User-agent: FacebookBot User-agent: Diffbot User-agent: Omgilibot User-agent: ImagesiftBot User-agent: Timpibot User-agent: YouBot User-agent: cohere-ai Disallow: / # ── SEO / marketing crawlers ────────────────────────────────────────────────────────────────── # Backlink and rank-tracking bots: same trade as above, cost without benefit to this site. User-agent: AhrefsBot User-agent: SemrushBot User-agent: DataForSeoBot User-agent: MJ12bot User-agent: DotBot User-agent: BLEXBot User-agent: Barkrowler User-agent: ZoominfoBot User-agent: serpstatbot User-agent: PetalBot Disallow: / # ── Everyone else, including search engines ─────────────────────────────────────────────────── # Social unfurlers (facebookexternalhit, Twitterbot, Slackbot, Discordbot) are deliberately NOT # blocked: they fetch once per shared link and are how a shared page renders a preview. The # pre-launch nginx shield on `app` does turn them away, but that is a property of that host, not # of this policy. Note Applebot (search) stays allowed above while Applebot-Extended (AI training) # is blocked — they are separate agents and that split is intentional. User-agent: * Allow: / # Faceted filters are an unbounded crawl space, and this is the specific trap behind the # 2026-08-04 outage. The multi-select facets combine freely AND repeat (`?bands=a&bands=b`), and # value ORDER yields distinct URLs, so the reachable set is combinatorial rather than merely large. # Free-text `search` is simply infinite. # # Canonical pages stay fully crawlable: band scope lives in the PATH # (/bands/{slug}/live/concerts, /bands/{slug}/music/releases, /bands/{artist}/news), never in these # parameters, and filtered views carry rel="canonical" back to the unfiltered page. # Free text. `search` is the catalogue/news key; `q` is the photo archive and the member directory. # Both are unbounded by definition. # # `q` is anchored on the parameter delimiter rather than written as `/*?*q=`. These patterns match a # bare substring, so `/*?*q=` would also match any parameter ENDING in q — `?faq=` and the like — # and block URLs it was never meant to. Same for `to` further down. The longer keys have no such # collision, so they keep the shorter one-line form. Disallow: /*?*search= Disallow: /*?q= Disallow: /*&q= # Multi-select and single-select facets. Note the singular/plural pairs are BOTH listed on purpose: # robots.txt matches literally, so `eras=` does not cover `era=` and `countries=` does not cover # `country=`. The photo archive uses the singular forms, the catalogue browsers the plural. Disallow: /*?*bands= Disallow: /*?*band= Disallow: /*?*artist= Disallow: /*?*decades= Disallow: /*?*countries= Disallow: /*?*country= Disallow: /*?*eras= Disallow: /*?*era= Disallow: /*?*photographer= Disallow: /*?*albumId= Disallow: /*?*tours= Disallow: /*?*tour= # #929 — the concert browser's venue facet. The venue PAGES (/live/venues/{slug}) are crawlable and in the # sitemap; it is the filtered query space that must not be, for the usual combinatorial reason. Singular # `venue=` is listed alongside for the literal-matching trap above, even though nothing emits it yet. Disallow: /*?*venues= Disallow: /*?*venue= Disallow: /*?*owners= Disallow: /*?*releaseCategories= Disallow: /*?*albumFamilies= Disallow: /*?*category= Disallow: /*?*letter= Disallow: /*?*year= Disallow: /*?*sort= Disallow: /*?*view= Disallow: /*?*pageSize= Disallow: /*?*hasSetlist= Disallow: /*?*hasPhotos= Disallow: /*?*hasVideo= Disallow: /*?*withReviews= Disallow: /*?*hasLyrics= Disallow: /*?*liveHistory= # Collaborations facets (#953). Plain handler parameters rather than [FromQuery] ones, but they are # emitted as filter links exactly like the rest. Disallow: /*?*role= Disallow: /*?*person= # Covers facets (#424). Both are emitted as filter links on /music/covers and combine with `search`, # `sort` and each other, so the query space is combinatorial in exactly the way the rest of this block # guards against. The page's canonical drops every filter, pointing each variant at the clean band URL. Disallow: /*?*source= Disallow: /*?*decade= # Screen-appearances facet (#428). `band`, `search`, `sort` and `decade` are already disallowed above; # `type` is the one new key. The page moved to /music/screen in #1350, and band now lives in the PATH # (/music/screen for all bands, /bands/{slug}/music/screen for one). Its canonical is the clean path for # whichever scope is loaded, so every filtered variant points at one of those two. Disallow: /*?*type= # Community feed (#1165). `kinds` is a comma-separated SUBSET mask over ActivityKind, so its query # space is combinatorial in the same way the multi-select facets are; `scope` requires sign-in and # renders nothing useful to a crawler. `before` is the infinite-scroll cursor and returns a fragment. Disallow: /*?*kinds= Disallow: /*?*scope= Disallow: /*?*before= # The "leaving the site" interstitial takes an arbitrary external URL, so `?to=` is an unbounded # space that renders no content of ours — and it is OutputCached varying by `to`, so crawling it # would fill the cache with one entry per outbound target as well as burning transfer. # Delimiter-anchored: a bare `/*?*to=` would also match `?photo=`. Disallow: /*?to= Disallow: /*&to= # #1567 — the canonical-guide revision surfaces. `from`/`to` select a PAIR of revisions to compare, so a # guide with N revisions offers N^2 comparison URLs, each one a bounded but real diff computation. That is # the combinatorial crawl trap the 2026-08-04 outage was made of, and it has no indexing value: the guide's # current text is on the thread page, which IS indexed. `restore` prefills the editor, which requires # authentication anyway — disallowed so crawlers do not queue redirects to the login page. # `from` takes the anchored form for the same reason `to` and `q` do: a bare `/*?*from=` substring would # also match any parameter ending in "from". Disallow: /*?from= Disallow: /*&from= Disallow: /*?*restore= # Virtual-reader variants share the canonical thread path. Chunk handlers return partial HTML; exact/cursor # entries are complete documents but duplicate the clean thread canonical and form an unbounded crawl space. Disallow: /*?*handler= Disallow: /*?*offset= Disallow: /*?*afterTicks= Disallow: /*?*afterId= Disallow: /*?*beforeTicks= Disallow: /*?*beforeId= Disallow: /*?*expectedOrdinal= Disallow: /*?*expectedTotal= Disallow: /*?*post= Disallow: /*?*position= Disallow: /*?*all=true # `?page=` is deliberately still ALLOWED. It is linear and bounded rather than combinatorial. # The original reason was that with no sitemap, paginated crawling was the ONLY route to deep # catalogue and archive pages. The sitemap has been live since #490/#17, which once looked like it # spent that reason — but #1767 withdrew the ~50,000 forum thread URLs from it, so for the largest # corpus on the site pagination is once again the ONLY route a crawler has. It is load-bearing, not # a leftover. Two further reasons it stays open: it is how a crawler reaches a page whose lastmod it # already has, and it is where the legacy YAF `m=` redirects land (`?page={n}#post-{id}`), which is # why `?post=` above is disallowed and this is not. Revisit only if pagination shows up as a # crawl-budget problem in Search Console, which needs evidence, not a guess. # # No URL count is quoted here on purpose. The figure that used to sit in this paragraph was 7x out # by the time anyone read it; the argument holds at any size, and the number was the part that rots. # Search is a crawl trap (infinite query space) and is marked noindex too. Disallow: /community/forums/search # Admin console (also auth-protected). Disallow: /admin # Auth-only areas. Nothing under these is public, so crawling them only produces redirects to the # login page — and /notifications is polled by the client, not a document. # # /members is deliberately NOT listed. It looks private and is not: /members/{slug} is the public # canonical fan profile (#46), and the member directory, news bylines, concert reviews, forum # authors and activity cards all link into it. An earlier revision of this file blocked it and # would have de-indexed the whole Community profile feature. Filtered/search variants of the # directory are already covered by the `q=` and `country=` rules above. Disallow: /account Disallow: /messages Disallow: /notifications # Problem reports (#1253). Blocked by PATH rather than per-parameter, because none of it is content: # /report is a form whose query keys (kind, subjectType, subjectId, from) only pre-fill it, and # /reports/{guid} is one person's own report reached by an unguessable link. That link is effectively # a bearer credential, so it must never end up in an index — the pages set `X-Robots-Tag: noindex` # as well, since robots.txt is a request and this one matters. The prefix covers both paths. Disallow: /report # #2153 — WHY THESE BLOCKS STAY, given Search Console reports them under "Indexed, though blocked by # robots.txt". That report is real and its mechanism is understood: a disallowed URL is never fetched, # so a `noindex` on it is never read, and Google indexes the bare URL with no content attached. The # textbook fix is to allow the crawl so the directive can be seen — which is what #1013 does for the # non-canonical host, and it is the right move THERE. # # It is the wrong move here, and the cost is asymmetric. `/report?kind=…&subjectType=…&subjectId=…` is # emitted once per catalogue entity, and `/account/login?returnUrl=…` once per page that needs sign-in, # so unblocking them would hand a crawler a combinatorial space of ~2,400 DB-touching form renders — # the same shape of trap as the faceted browsers above, on the same metered 3 TB/month tier, on a box # that crawlers OOM-killed on 2026-08-04. What we would buy is the removal of URL-only entries that # rank for nothing and display "no information is available". # # So this category is ACCEPTED, not fixed: the entries are inert and the crawl budget is not. The pages # set `X-Robots-Tag: noindex, nofollow` anyway — it costs nothing, it is what a non-compliant crawler # that fetches regardless should be told, and it means the block is a crawl-budget decision rather than # the only thing standing between these URLs and the index. Revisit only if these ever accrue real # impressions, which needs evidence rather than a guess. # # `/report/thanks?id={guid}` and `/reports/{guid}` would stay blocked under ANY future revisit. That # guid is a bearer credential the GET serves to whoever holds it, and `noindex` is an indexing # directive — the crawler must fetch the page to read it, by which point it has been handed the report. # Coverage could be perfect and the disclosure would still have happened. # Campaign unsubscribe (#1427). Same reasoning as /report, and the same shape of link: `?t=` is a # one-time capability token that suppresses an address, so it is a bearer credential and must never # be indexed, logged into a crawl frontier, or followed. Blocked by PATH rather than per-parameter # because none of it is content — the GET is a confirmation prompt and the one-click POST is a # machine endpoint. Both set `X-Robots-Tag: noindex` too, since robots.txt is only a request. # # A compliant crawler stops here; the real protection is that GET never mutates. RFC 8058 exists # precisely because scanners DO fetch header URLs, which is why the mutating half is POST-only. Disallow: /email # Be a good citizen without being slow to index. Google ignores this (crawl rate is a Search # Console setting); Bing and Yandex honour it. Crawl-delay: 2 Sitemap: https://www.neworderonline.com/sitemap.xml