Original measurement of how many Romanian websites publish an llms.txt file,
what is inside those files, and whether the same sites allow the crawlers that
feed AI answers.
Episode 4 of the Websem study series on AI-engine visibility. Study page: https://websem.ro/resurse/aeo/studiu-llms-txt-romania
- Measured: 9 August 2026, one complete run
- Version: v1.1.0 (26 August 2026) — see Changelog
- Frame: 87 Romanian-market domains
- Licence: CC BY 4.0
- Author: Dan Cristian Alexandrescu (ORCID 0009-0007-1994-6700), Websem
| Domains in frame | 87 |
Publish llms.txt |
56 (64.4%) |
| — plugin-generated | 12 · mean 25,070 bytes |
| — hand-written | 44 · mean 8,803 bytes |
| Spec-conformant (H1 + summary blockquote) | 39 |
| With no curated link at all | 8 |
| Block answer-engine crawlers | 0 |
| Block training or archive scrapers only | 3 |
Three results are worth stating plainly.
llms.txt is already the majority practice, not a curiosity: 56 of 87
domains serve one. Public discussion still treats it as experimental.
The large files are plugin dumps, not curation. The 12 plugin-generated
files are nearly three times larger than the hand-written ones and carry
neither the # H1 nor the > summary line the
llmstxt.org specification asks for. The single
most-cited domain in the frame serves exactly such a dump — which is evidence
against llms.txt being what earns citations.
Blocking a crawler is not one decision, it is two — and nobody here made the
one that costs. Blocking GPTBot, ClaudeBot, Google-Extended or CCBot
is a rights decision; each vendor's own documentation says it does not remove a
site from AI answers. What removes a site is OAI-SearchBot, ChatGPT-User,
Claude-User, Claude-SearchBot or PerplexityBot. Three domains in the frame
block something. None of them block an answer-layer crawler. The intent is
legible in the robots.txt files; the effect on visibility is nil.
Not a list we chose. The frame is the set of Romanian-market domains that AI engines actually cite when answering Romanian-language questions about marketing, AEO and automation, taken from the sources recorded in answers monitored with LLM Pulse (25 prompts across 5 surfaces). Inclusion rule: the domain appears among the 200 most-cited sources. That yields 87 domains.
The frame is therefore biased toward the marketing niche by construction. It describes the domains AI engines cite in that niche — not Romanian websites in general, and the percentages should not be read as such.
One GET to /llms.txt and one to /robots.txt per domain, in a single
complete run.
A file counts as served when it answers 200, has content-type: text/plain and exceeds 80 bytes — the threshold removes error pages returned
with a 200 status.
Spec conformance is checked on two objective elements: the first line is a
Markdown # H1, and a > blockquote carrying the summary appears within the
first eight lines.
Provenance is inferred from the first three lines: files announcing Rank Math, All in One SEO, Yoast or a generic "generated by" banner are classified as plugin-generated.
A user-agent counts as blocked only when robots.txt carries a block
dedicated to it containing Disallow: /.
Agents are classified in three groups, following each vendor's current documentation (checked 26 August 2026):
| Group | Agents | Effect of blocking |
|---|---|---|
| Answer layer | OAI-SearchBot, ChatGPT-User, PerplexityBot, Perplexity-User, Claude-User, Claude-SearchBot |
removes the site from AI answers |
| Training / archive | GPTBot, ClaudeBot, Google-Extended, CCBot, Bytespider, Applebot-Extended, Amazonbot, Meta-ExternalAgent, cohere-ai |
a rights decision; visibility unaffected |
| Legacy | anthropic-ai |
no defined effect; no longer in the vendor's documentation |
Google has no dedicated AI token for the answer layer: AI Overviews draw on the
Search index, so access there is governed by Googlebot itself.
It does not establish a causal relation between llms.txt and citations. It
shows a suggestive negative correlation — the most-cited domain has the least
conformant file — but with 87 domains and a single run, the effect of
llms.txt cannot be separated from domain age, authority or content.
It also does not measure whether the bots read the file. That needs server logs, which we do not have for third-party domains.
A single run captures one day. Files change; a repeat run would produce a different snapshot, which is why the scanner is published alongside the data.
data/llmstxt-ro-2026-08.csv the dataset, 87 rows, one per domain
data/llmstxt-ro-2026-08.raw.json raw scanner output, including per-file detail
scripts/scan-llmstxt.py collection — reproduces the raw JSON
scripts/build-llmstxt-study.mjs derivation — rebuilds the CSV and the figures
| Column | Meaning |
|---|---|
domeniu |
domain measured |
llms_txt_servit |
1 when /llms.txt answered 200 as text/plain, over 80 bytes |
octeti |
file size in bytes, empty when not served |
generat_de_plugin |
1 when the first lines announce an SEO plugin |
are_h1 |
1 when the first line is a Markdown # H1 |
are_sumar |
1 when a > blockquote appears in the first eight lines |
linkuri |
count of - [ list links |
robots_accesibil |
1 when /robots.txt answered 200 |
boti_blocati |
user-agents with a dedicated Disallow: /, space-separated |
python3 scripts/scan-llmstxt.py # rescans, writes the raw JSON
node scripts/build-llmstxt-study.mjs # rebuilds the CSV and derived figuresNo figure in the study page is typed by hand; all of them are generated by the second script from the raw data.
Alexandrescu, D. C. (2026). llms.txt across Romanian domains cited by AI engines. Websem. https://websem.ro/resurse/aeo/studiu-llms-txt-romania
v1.1.0 — 26 August 2026. Crawler classification corrected. v1.0.0 listed
GPTBot, ClaudeBot and Google-Extended as answer-engine crawlers. Each
vendor's current documentation contradicts that: OpenAI states that disallowing
GPTBot "indicates a site's content should not be used in training generative
AI foundation models", while it is OAI-SearchBot whose exclusion means a site
"will not be shown in ChatGPT search answers"; Anthropic describes ClaudeBot
as collecting content that may contribute to training, with Claude-User and
Claude-SearchBot serving user questions; Google states that Google-Extended
"does not impact a site's inclusion in Google Search nor is it used as a ranking
signal in Google Search". A third group was added for anthropic-ai, a token no
longer present in Anthropic's documentation.
Consequence: domains blocking answer-engine crawlers falls from 2 to 0, and
the one case of a domain publishing llms.txt while blocking answer crawlers
falls to 0. The raw data is byte-identical — only the classification
changed. data/llmstxt-ro-2026-08.csv and data/llmstxt-ro-2026-08.raw.json
are unchanged from v1.0.0, which remains citable at its own DOI.
v1.0.0 — 9 August 2026. Initial deposit.
CC BY 4.0. Use it, republish it, disagree with it — attribution is the only condition.