Every article about AI and search contains a statistic, and almost none of them can be checked. We have avoided writing this one for a year for exactly that reason. What changed is that we now have eighteen months of a site built deliberately for machine readers, and we can simply count. So this is not a forecast. It is two days of our own access log, with the method described and the limits admitted.

What we built, briefly

Since the static-first rebuild, everything on this site is a file. Each article exists as HTML for people and as markdown for agents, served by content negotiation, which we described in Your Website Is Illiterate to AI Agents. There are hand-written llms.txt and llms-full.txt summaries, a full sitemap with hreflang twins for every page, structured data on articles, and an RSS feed. The pages render without JavaScript, because a crawler that has to execute a bundle to see your content is a crawler that may decide not to.

The measurement

Two full days, 28 and 29 September 2026, counted from the web server log for iwh.gr only. Total requests 10,457. Excluding assets, images, audio, video and feeds, that is 7,800 page requests. Then we counted requests whose user agent declares itself as an AI crawler or assistant.

  • 402 page requests from declared AI agents.
  • 53 from Googlebot, 31 from bingbot.

That is a ratio of nearly five to one in favour of the AI agents, on a corporate site in a niche professional field, in two ordinary days with no campaign running.

The distribution is heavily skewed. One assistant's crawler accounted for 287 of the 402. A large social platform's agent accounted for 120. The rest, in descending order, were a mix of model-lab crawlers, search-assistant crawlers and smaller entrants, most in single or low double digits. For comparison, commercial SEO crawlers were also present at similar volumes to the search engines, and one general web-archive crawler was busier than several of the AI agents.

Two findings surprised us.

Greek beat English. Of the agent page requests, 225 were for Greek pages and 123 for English ones. This site is English-primary by design and our English catalogue is the older one. We do not have a confident explanation, and we are not going to invent one. We are going to keep translating.

The summary files are barely touched. llms.txt and llms-full.txt together received eleven requests across two days, two of them from declared agents. We maintain them by hand and the nightly job complains when they fall behind. On this evidence they are a courtesy, not a channel. The articles themselves are what gets fetched.

The caveats, before anyone quotes this

Two days is a sample, not a trend, and a holiday week or a single crawl campaign would move every number here.

User agents are self-declared and unverified. We did not reverse-check IP ranges. Some of the requests counted above are certainly not what they claim: within the same set we found requests for a nonexistent uploads directory and for a JavaScript build tool's source-map path, which is scanner behaviour wearing a costume. Anyone publishing agent statistics without this caveat is either verifying the ranges, which is work, or guessing.

Most importantly, a request is not a citation and a citation is not a client. We can see that the machines are reading. We cannot see what they do with it, whether anything we wrote was surfaced to a person, or whether any of it produced a single enquiry. Anyone who tells you they can measure that from their own logs is describing something they cannot see.

The thing a normal log cannot tell you

Here is the most useful practical finding, and it is an embarrassing one.

We built content negotiation so that an agent sending Accept: text/markdown receives clean markdown instead of HTML. We cannot tell you how often that happens, because the standard access log format does not record request headers. Every one of those markdown responses looks, in the log, exactly like a page view.

If you are building for machine readers, log the Accept header and the negotiated response type from the start. It is one line of configuration and it is the difference between knowing and assuming. We are changing ours.

Crawlers and fetchers are different animals

Lumping all of these together, as we did above for the headline number, hides the distinction that matters commercially.

A crawler is building an index or a training set. It visits on its own schedule, nobody is waiting, and its visit may pay off in six months or never. A fetcher is retrieving your page at the moment a human being asked a question. The user-triggered agents, the ones whose names end in "-User", are this second kind, and in our two days they numbered fewer than thirty requests in total.

Thirty is a small number. It is also the only number on this page where a person was waiting for an answer at the moment the request was made. If you are going to optimise anything for this audience, optimise for that case: the page a machine fetches while somebody watches a cursor blink.

What actually follows for a business site

The unit is the answerable claim, not the keyword page. A page built to rank for a phrase has nothing to offer a system that reads the content and matches it to an intention. What gets used is a specific, checkable statement with a date and a source attached. Write the sentence that answers the question, put it near a heading that names the question, and do not bury it under four hundred words of throat-clearing.

Serve machines the same thing you serve people. Not a stripped variant, not a different set of facts. Apart from being the honest position, it is the one that survives the next policy change at any of the companies above.

Your second language may be your first audience. We would have cut Greek effort on the assumption that English carried the reach. Our own logs say otherwise.

Robots decisions are business decisions. Which of those agents you allow, and for what, is a commercial and legal question about who may use your material, not a setting for whoever maintains the server. Make it deliberately, write down why, and review it, because the list of agents changes every quarter.

The same shift, elsewhere

This is not only a web phenomenon. Working video creators describe the same movement on the large platforms: away from keyword optimisation and copied formats, towards systems that read the actual content and match it to viewer intent, with near-duplicate material suppressed rather than rewarded. Those are one practitioner's reading of the signals rather than published policy, and should be treated as such, but the direction rhymes with what we see in our own logs.

The conclusion in both places is the same and it is almost disappointingly old-fashioned. When the reader is a machine that understands content, the winning move is to have content worth understanding, said clearly, dated, sourced, and served to everyone in the same form. There is no keyword to buy. There is only whether the page says something true that somebody needed.


Figures counted from the iwh.gr web server log for 28 and 29 September 2026; method and limits as described above. Related reading: Your Website Is Illiterate to AI Agents, and a Focus note from September 2026 on the shift from keywords to intent on video platforms.