Skip to content

Machine readability index · read 2026-09-23

8 of 20 well-known sites could be read by a machine.

Each site's homepage was requested once, with one plain GET, identified as LegibilityBot, exactly as an answer engine's crawler would. No headless browser. No proxy to get around a refusal. Twenty sites is a small sample and every figure below says so.

The total is the least interesting part of it. What matters is that these are three different findings wearing one word.

Chose not to be read
8of 20
Refused the request or disallowed it in robots.txt. A decision, and often a reasonable one.
Could be read, and was not
2of 20
Reachable and readable as a page, carrying nothing a machine can use without guessing.
Inconclusive
2of 20
The read failed for reasons that are ours, not theirs. Counted against nobody.

The middle column is the one with a cost attached. 2 of 20 sites, 10 percent, are readable pages that a machine still cannot extract a fact from. That is the problem this product measures, and it is the only one of the three anybody can fix in an afternoon.

§01 · by cohort

News refuses on purpose. Retail mostly does not.

news10 sites

2 readable of 10

chose not to be read
6
could be read, and was not
1
inconclusive
1
retail10 sites

6 readable of 10

chose not to be read
2
could be read, and was not
1
inconclusive
1

§02 · every site

The whole cohort, with what was found.

Nothing is aggregated away. Each verdict links to what that finding means, what it costs and what would change it.

news

apnews.com
blocked
The site refused the request. This is a choice the site made.
bbc.co.uk
readable · jsonld
Typed facts were extracted without guessing.
bloomberg.com
blocked
The site refused the request. This is a choice the site made.
cnn.com
readable · jsonld
Typed facts were extracted without guessing.
ft.com
blocked
The site refused the request. This is a choice the site made.
nytimes.com
blocked
The site refused the request. This is a choice the site made.
reuters.com
robots_disallowed
robots.txt asks machines not to read this path. Recorded as a data point, never circumvented.
theguardian.com
no_structured_data
The page is readable as text but carries no structured data, so a machine has to guess what any of it means.
washingtonpost.com
timeout
The page did not answer in time.
wsj.com
robots_disallowed
robots.txt asks machines not to read this path. Recorded as a data point, never circumvented.

retail

allbirds.com
readable · opengraph
Typed facts were extracted without guessing.
apple.com
readable · jsonld
Typed facts were extracted without guessing.
bestbuy.com
no_structured_data
The page is readable as text but carries no structured data, so a machine has to guess what any of it means.
hm.com
timeout
The page did not answer in time.
ikea.com
readable · jsonld
Typed facts were extracted without guessing.
lego.com
blocked
The site refused the request. This is a choice the site made.
nike.com
readable · jsonld
Typed facts were extracted without guessing.
target.com
readable · jsonld
Typed facts were extracted without guessing.
walmart.com
readable · jsonld
Typed facts were extracted without guessing.
zara.com
blocked
The site refused the request. This is a choice the site made.

§03 · method and its limits

How this was measured.

One plain GET of each homepage, identified as LegibilityBot, with robots.txt checked first: a site that disallows us is recorded as having done so and its page is never requested. No headless browser, and the proxy fallback that would read sites which refuse machines exists and stays switched off, because a site being unreadable without it is the finding.

A refusal is only recorded as a refusal when a second, independent request path reproduces it. Where the two disagree the verdict is inconclusive, because an index that cannot tell must say so rather than pick the more interesting answer. Every refusal published here was additionally corroborated against this site's own live checker before it was recorded.

The honest limits. Twenty sites is a small sample and a homepage is one page, so nothing here describes a whole catalogue. The cohort is fixed rather than random, chosen for being well known. A site can change between reads, and two of these did while this index was being built.

The eight findings, and what each one costs · What the crawler does

Your own site is not in this cohort. The check on the front page reads it the same way, with the same classifier, and returns one of the same eight answers.

See what a machine sees