Machine readability index · read 2026-09-23
8 of 20 well-known sites could be read by a machine.
Each site's homepage was requested once, with one plain GET, identified as LegibilityBot, exactly as an answer engine's crawler would. No headless browser. No proxy to get around a refusal. Twenty sites is a small sample and every figure below says so.
The total is the least interesting part of it. What matters is that these are three different findings wearing one word.
- Chose not to be read
- 8of 20
- Refused the request or disallowed it in robots.txt. A decision, and often a reasonable one.
- Could be read, and was not
- 2of 20
- Reachable and readable as a page, carrying nothing a machine can use without guessing.
- Inconclusive
- 2of 20
- The read failed for reasons that are ours, not theirs. Counted against nobody.
The middle column is the one with a cost attached. 2 of 20 sites, 10 percent, are readable pages that a machine still cannot extract a fact from. That is the problem this product measures, and it is the only one of the three anybody can fix in an afternoon.
§01 · by cohort
News refuses on purpose. Retail mostly does not.
2 readable of 10
- chose not to be read
- 6
- could be read, and was not
- 1
- inconclusive
- 1
6 readable of 10
- chose not to be read
- 2
- could be read, and was not
- 1
- inconclusive
- 1
§02 · every site
The whole cohort, with what was found.
Nothing is aggregated away. Each verdict links to what that finding means, what it costs and what would change it.
news
- bbc.co.uk
- readable · jsonld
- Typed facts were extracted without guessing.
- cnn.com
- readable · jsonld
- Typed facts were extracted without guessing.
- reuters.com
- robots_disallowed
- robots.txt asks machines not to read this path. Recorded as a data point, never circumvented.
- theguardian.com
- no_structured_data
- The page is readable as text but carries no structured data, so a machine has to guess what any of it means.
- wsj.com
- robots_disallowed
- robots.txt asks machines not to read this path. Recorded as a data point, never circumvented.
retail
- allbirds.com
- readable · opengraph
- Typed facts were extracted without guessing.
- apple.com
- readable · jsonld
- Typed facts were extracted without guessing.
- bestbuy.com
- no_structured_data
- The page is readable as text but carries no structured data, so a machine has to guess what any of it means.
- ikea.com
- readable · jsonld
- Typed facts were extracted without guessing.
- nike.com
- readable · jsonld
- Typed facts were extracted without guessing.
- target.com
- readable · jsonld
- Typed facts were extracted without guessing.
- walmart.com
- readable · jsonld
- Typed facts were extracted without guessing.
§03 · method and its limits
How this was measured.
One plain GET of each homepage, identified as LegibilityBot, with robots.txt checked first: a site that disallows us is recorded as having done so and its page is never requested. No headless browser, and the proxy fallback that would read sites which refuse machines exists and stays switched off, because a site being unreadable without it is the finding.
A refusal is only recorded as a refusal when a second, independent request path reproduces it. Where the two disagree the verdict is inconclusive, because an index that cannot tell must say so rather than pick the more interesting answer. Every refusal published here was additionally corroborated against this site's own live checker before it was recorded.
The honest limits. Twenty sites is a small sample and a homepage is one page, so nothing here describes a whole catalogue. The cohort is fixed rather than random, chosen for being well known. A site can change between reads, and two of these did while this index was being built.
The eight findings, and what each one costs · What the crawler does
Your own site is not in this cohort. The check on the front page reads it the same way, with the same classifier, and returns one of the same eight answers.
See what a machine sees