Skip to content

Latest commit

 

History

3 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

llms.txt across Romanian domains cited by AI engines

Original measurement of how many Romanian websites publish an llms.txt file, what is inside those files, and whether the same sites allow the crawlers that feed AI answers.

Episode 4 of the Websem study series on AI-engine visibility. Study page: https://websem.ro/resurse/aeo/studiu-llms-txt-romania

  • Measured: 9 August 2026, one complete run
  • Version: v1.1.0 (26 August 2026) — see Changelog
  • Frame: 87 Romanian-market domains
  • Licence: CC BY 4.0
  • Author: Dan Cristian Alexandrescu (ORCID 0009-0007-1994-6700), Websem

Headline findings

Domains in frame 87
Publish llms.txt 56 (64.4%)
— plugin-generated 12 · mean 25,070 bytes
— hand-written 44 · mean 8,803 bytes
Spec-conformant (H1 + summary blockquote) 39
With no curated link at all 8
Block answer-engine crawlers 0
Block training or archive scrapers only 3

Three results are worth stating plainly.

llms.txt is already the majority practice, not a curiosity: 56 of 87 domains serve one. Public discussion still treats it as experimental.

The large files are plugin dumps, not curation. The 12 plugin-generated files are nearly three times larger than the hand-written ones and carry neither the # H1 nor the > summary line the llmstxt.org specification asks for. The single most-cited domain in the frame serves exactly such a dump — which is evidence against llms.txt being what earns citations.

Blocking a crawler is not one decision, it is two — and nobody here made the one that costs. Blocking GPTBot, ClaudeBot, Google-Extended or CCBot is a rights decision; each vendor's own documentation says it does not remove a site from AI answers. What removes a site is OAI-SearchBot, ChatGPT-User, Claude-User, Claude-SearchBot or PerplexityBot. Three domains in the frame block something. None of them block an answer-layer crawler. The intent is legible in the robots.txt files; the effect on visibility is nil.

Sampling frame

Not a list we chose. The frame is the set of Romanian-market domains that AI engines actually cite when answering Romanian-language questions about marketing, AEO and automation, taken from the sources recorded in answers monitored with LLM Pulse (25 prompts across 5 surfaces). Inclusion rule: the domain appears among the 200 most-cited sources. That yields 87 domains.

The frame is therefore biased toward the marketing niche by construction. It describes the domains AI engines cite in that niche — not Romanian websites in general, and the percentages should not be read as such.

Method

One GET to /llms.txt and one to /robots.txt per domain, in a single complete run.

A file counts as served when it answers 200, has content-type: text/plain and exceeds 80 bytes — the threshold removes error pages returned with a 200 status.

Spec conformance is checked on two objective elements: the first line is a Markdown # H1, and a > blockquote carrying the summary appears within the first eight lines.

Provenance is inferred from the first three lines: files announcing Rank Math, All in One SEO, Yoast or a generic "generated by" banner are classified as plugin-generated.

A user-agent counts as blocked only when robots.txt carries a block dedicated to it containing Disallow: /.

Agents are classified in three groups, following each vendor's current documentation (checked 26 August 2026):

Group Agents Effect of blocking
Answer layer OAI-SearchBot, ChatGPT-User, PerplexityBot, Perplexity-User, Claude-User, Claude-SearchBot removes the site from AI answers
Training / archive GPTBot, ClaudeBot, Google-Extended, CCBot, Bytespider, Applebot-Extended, Amazonbot, Meta-ExternalAgent, cohere-ai a rights decision; visibility unaffected
Legacy anthropic-ai no defined effect; no longer in the vendor's documentation

Google has no dedicated AI token for the answer layer: AI Overviews draw on the Search index, so access there is governed by Googlebot itself.

What this does not measure

It does not establish a causal relation between llms.txt and citations. It shows a suggestive negative correlation — the most-cited domain has the least conformant file — but with 87 domains and a single run, the effect of llms.txt cannot be separated from domain age, authority or content.

It also does not measure whether the bots read the file. That needs server logs, which we do not have for third-party domains.

A single run captures one day. Files change; a repeat run would produce a different snapshot, which is why the scanner is published alongside the data.

Contents

data/llmstxt-ro-2026-08.csv        the dataset, 87 rows, one per domain
data/llmstxt-ro-2026-08.raw.json   raw scanner output, including per-file detail
scripts/scan-llmstxt.py            collection — reproduces the raw JSON
scripts/build-llmstxt-study.mjs    derivation — rebuilds the CSV and the figures

Columns

Column Meaning
domeniu domain measured
llms_txt_servit 1 when /llms.txt answered 200 as text/plain, over 80 bytes
octeti file size in bytes, empty when not served
generat_de_plugin 1 when the first lines announce an SEO plugin
are_h1 1 when the first line is a Markdown # H1
are_sumar 1 when a > blockquote appears in the first eight lines
linkuri count of - [ list links
robots_accesibil 1 when /robots.txt answered 200
boti_blocati user-agents with a dedicated Disallow: /, space-separated

Reproducing

python3 scripts/scan-llmstxt.py       # rescans, writes the raw JSON
node scripts/build-llmstxt-study.mjs  # rebuilds the CSV and derived figures

No figure in the study page is typed by hand; all of them are generated by the second script from the raw data.

Citing

Alexandrescu, D. C. (2026). llms.txt across Romanian domains cited by AI engines. Websem. https://websem.ro/resurse/aeo/studiu-llms-txt-romania

Changelog

v1.1.0 — 26 August 2026. Crawler classification corrected. v1.0.0 listed GPTBot, ClaudeBot and Google-Extended as answer-engine crawlers. Each vendor's current documentation contradicts that: OpenAI states that disallowing GPTBot "indicates a site's content should not be used in training generative AI foundation models", while it is OAI-SearchBot whose exclusion means a site "will not be shown in ChatGPT search answers"; Anthropic describes ClaudeBot as collecting content that may contribute to training, with Claude-User and Claude-SearchBot serving user questions; Google states that Google-Extended "does not impact a site's inclusion in Google Search nor is it used as a ranking signal in Google Search". A third group was added for anthropic-ai, a token no longer present in Anthropic's documentation.

Consequence: domains blocking answer-engine crawlers falls from 2 to 0, and the one case of a domain publishing llms.txt while blocking answer crawlers falls to 0. The raw data is byte-identical — only the classification changed. data/llmstxt-ro-2026-08.csv and data/llmstxt-ro-2026-08.raw.json are unchanged from v1.0.0, which remains citable at its own DOI.

v1.0.0 — 9 August 2026. Initial deposit.

Licence

CC BY 4.0. Use it, republish it, disagree with it — attribution is the only condition.

About

Original measurement: how many Romanian domains cited by AI engines publish llms.txt, what is inside those files, and whether they allow the crawlers that feed AI answers. Episode 4 of the Websem study series. CC BY 4.0.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages