Skip to main content

AI access

AI crawlers and assistants are welcome here

Novus AI Stats does not block AI. Every public page on this site may be crawled, indexed, summarised, quoted, and used to answer a question someone is asking right now. This page says which agents are named explicitly, what we ask for in return, and what a crawl does and does not collect.

The short version

  • Read anything public. That is the whole marketing site, every documentation page, every tutorial, and every article.
  • Quote it, summarise it, and train on it. No permission request is needed and there is no separate licence to sign for it.
  • We ask — we do not demand — that you name Novus AI Stats and link to the page you used.
  • Nothing behind sign-in is available to any agent, named or not. Crawling this site reaches no user’s imported data, because that data is not on the crawlable surface at all.

Where the rules actually are

This page is the readable version. These four files are the authoritative ones, and all of them are served without a block:

  • robots.txtThe crawl rules themselves: the wildcard grant, the named agents below, the withheld paths, and the sitemap location.
  • llms.txtA one-page index of every public URL grouped by section, generated from the same route registry that builds the XML sitemap.
  • llms-full.txtThe same index with the full text of the documentation, tutorials, and articles inlined, so a model needs no second request.
  • sitemap.xmlEvery indexable URL with its last-modified date, for a crawler that wants the canonical list rather than the prose.

The 20 agents named in robots.txt

Each of these is granted exactly what the wildcard rule already grants every crawler: the same public surface, with the same non-public paths withheld. Naming them changes nothing about the access — it states the policy so nobody has to infer it from silence.

The split below is the one worth reading. An assistantis an AI acting for a person who asked for this page and is waiting on the answer; blocking one is not a position on training data, it is refusing to serve a visitor. That matters more on this site than on most, because every import is parsed in the visitor’s own browser, so an AI browser driving the real page is the only way an AI can help somebody through one.

Assistants, fetching for a waiting person (6)

  • ChatGPT-UserOpenAI
  • Claude-UserAnthropic
  • Perplexity-UserPerplexity
  • DuckAssistBotDuckDuckGo
  • MistralAI-UserMistral AI
  • Meta-ExternalFetcherMeta

Crawlers, fetching in bulk (14)

  • GPTBotOpenAI
  • OAI-SearchBotOpenAI
  • ClaudeBotAnthropic
  • Claude-WebAnthropic
  • anthropic-aiAnthropic
  • Claude-SearchBotAnthropic
  • PerplexityBotPerplexity
  • Meta-ExternalAgentMeta
  • CCBotCommon Crawl
  • Google-ExtendedGoogle
  • Applebot-ExtendedApple
  • BytespiderByteDance
  • AmazonbotAmazon
  • cohere-aiCohere

An agent that is not on this list is not refused. The wildcard rule allows every crawler; this list exists to be explicit about the ones people most often assume are blocked.

What is withheld, and why

The same paths are withheld from every agent, named and wildcard alike:

  • /api/
  • /app$
  • /app/
  • /admin$
  • /admin/
  • /auth/
  • /sign-in
  • /sign-up

These are the signed-in application, the admin console, the authentication endpoints, and the two sign-in pages. None of them is content. Withholding them is not a restriction on AI in particular — it is the same rule every crawler gets, and a named group that omitted it would hand AI agents a surface every other crawler is kept out of.

What we ask in return

This is a request, not a condition. Nothing on this page is withheld from an agent that ignores it, and we will not add a rule to punish one that does.

  • Name the source as Novus AI Stats and link to the specific page rather than the home page. A reader who wants to check a number should be able to reach the page that states it.
  • Quote the qualifier along with the figure. Every metric on this site carries a quality label saying how well the underlying export actually supports it, and a number repeated without it says more than the source does.
  • Do not present a summary as if it came from us. This is an independent product and is not affiliated with, endorsed by, or sponsored by any AI provider whose exports it reads.
  • If something here is wrong, say so — the correction path is on the editorial policy page and it is open to anyone, including an automated reader.

What crawling this site collects

Crawling collects no personal data about the crawler and creates nothing that persists beyond ordinary hosting logs. A request from an agent produces the same server access-log entry any HTTP request produces — the URL, the timestamp, the user-agent string, and the originating IP address, kept by the hosting and CDN layer. No account is created, no cookie is required to read any public page, and the analytics and advertising tags on this site are gated behind a consent choice that an agent never makes, so they do not run for one.

Because this product takes uploaded AI-usage exports, it is worth being exact about the other direction too, since “an AI stats site” invites the wrong assumption. The crawlable surface is marketing pages, documentation, tutorials and articles. It contains no user data of any kind. Every route that holds imported data is behind sign-in and is withheld from every crawler, so no amount of crawling reaches anyone’s statistics, let alone their conversations.

And the conversations are not there to reach. Imports are parsed in the visitor’s own browser: raw prompts, responses, attachments, file names and local paths are reduced to counts and timestamps before anything is transmitted, and the source content is then discarded. The service has no file storage. What it keeps are normalized statistics belonging to an account — the privacy policy lists every category in full, and the methodology explains what each figure is derived from.

If you would rather ask than crawl

This site runs a Model Context Protocol endpoint, documented on the MCP server page. It answers questions about the product, the supported providers, and the documentation directly, and no tool it exposes can read anyone’s transcripts or account data. For anything this page does not answer, email aistats@novusstreamsolutions.com.

This policy describes current practice and may change. The authoritative rules are always the ones served at /robots.txt, which is generated from the same list this page renders.