If you're new here, you can read the original post from last month here, but essentially I performed a test with five AI search engines (Claude, ChatGPT, Gemini, Perplexity, and Bing Copilot) to see if they would cite any of my websites in their answer, and they did not. There are a lot of normal reasons this could be the case, including my websites being relatively small and just not having much content on them.
However, a reader (RileyCraig14) commented on the post.
He brought this to my attention because he checked two files on my sites that I did not. robots.txt and llms.txt are supposed to be tiny text files at the root of each website (e.g. yoursite.com/robots.txt) that contain information for the search engines or AI tools to know what is important, what to scrape, what to not scrape, etc. on each website. The first has been around since the 90's; the second is new. Both are supposed to be boring, simple text files.
Both, on my sites, were my homepage.
It was returning the same response as the homepage (35kb page, "Remote AI Agent Developer") whenever a request was made for that file from either humans or bots. The other site, naija-vpn.com, had an llms.txt issue too, but its robots.txt was fine.
Modern websites are often created as SPAs (single-page apps), meaning that there is a single HTML page, and different content is swapped in and out via JavaScript. When hosting these sites, it's common to configure your hosting provider (mine is Cloudflare Pages) to send the main app page if a non-existing page is requested. The idea is that the user probably clicked on a link on your website, and they should be sent to the main app page.
This is nice to have when people are visiting my site, but doesn't make sense for things like robots.txt and llms.txt. Those files need to be small plain text files, so it's not great that the process that turns my source code into the website that I upload and serve isn't generating those files, and instead is falling back to serving the homepage. This means that if a crawler from an AI company wanted to know what they can read on my website, it was getting a web page back when it asked!
I've fixed this for both of my sites, and I've confirmed that I can see the files now! Took an afternoon to do.
Now, you may say - fixing it for two of my sites is not that much of a story, you could have the same issue and not even know about it. And you would be totally correct - the issue is not apparent unless you look for it. It doesn't break your site for the visitors. It only provides wrong information for the bots.
To that end I wrote another small tool that can help you to check whether you have this issue - ai_crawlability.py - part of my open source SEO toolset. All you need to do is to provide it with your domain and it will check whether robots.txt, llms.txt and also sitemap.xml (a third, often-used text file with the list of your pages) are served as proper text files or your website instead serves a webpage.
Note: This isn't about sites that don't have an llms.txt file (most sites don't since this is a new thing). Those will return a "not found" status code, which is fine. The issue I'm describing is when a site returns a status code of 200 (which means the request was successful) and serves a webpage instead of the llms.txt file, which is a bug.
You can enter sites you want to check, but that's not the point of this tool. The point of this tool is to enter search terms you want to use to get an unbiased sample of domain names from Google (via SearchApi, which is sponsoring this article series). Then the tool will check all of the domains in the sample. If you only check your own sites (or a handful of sites you choose), you won't get a good idea of how prevalent the issue is.
Important: the 50 domains I checked were not chosen randomly, they were pulled from real Google search results for 10 completely unrelated search topics that have nothing to do with each other (and nothing to do with my own industry), in order to avoid getting domains from competitors/companies similar to my own.
Here are the 10 search topics that I used for my sample: ai agent framework, best project management software, how to learn python, best productivity apps, react vs vue, best budget laptops 2026, cheap flights to lagos, best coffee subscription, how to start a podcast, and best password manager. To get the domains for each topic I used the Google search engine from SearchApi, which returns the same results you get when you search in Google, just via API: for each of the 10 search topics I got the URLs from the organic results and added the root domain of each of the URLs to my list of 50 domains, making sure I had 50 unique root domains across all the search topics.
For each domain, I used ai_crawlability.py to make a direct request to each domain for their /robots.txt, /llms.txt, and /sitemap.xml. If the response code was a 200 (success) and the content of the file was a webpage and not plain text, I counted it as a broken file. All 50 domains were checked.
5 out of 50 domains checked had a broken file, that's 1 out of 10!
The domains with this bug are: reddit.com (llms.txt and sitemap.xml), skyscanner.com (llms.txt and sitemap.xml), paymoapp.com (llms.txt), and 2 Medium.com blogs run by different people, ilampadmanabhan.medium.com and navanathjadhav.medium.com (sitemap.xml for both, both with a large broken webpage of 41,975 bytes and 41,972 bytes).
It's pretty unlikely that two people that have nothing to do with each other would make the same mistake, and the fact that the numbers are so close is further evidence that this is indeed a problem with Medium for anyone who is using a custom domain for their Medium blog.
There are some cases where the domain names did not work for reasons unrelated to this problem. Six of the domains that did not work were w3schools.com, pcmag.com, united.com, cheapoair.com, drinktrade.com, and beanbox.com. I manually checked two of these, and w3schools.com responded with "403 Forbidden", and pcmag.com seemed to be using some kind of bot detection software. These are both unrelated to the problem that I'm looking at here, and fortunately the tool that I was using was smart enough to not consider these "broken".
It turns out this was important.
I didn't re-test using the 5 search engines on my own sites after they were fixed. Why? Because both the search engines and the AI that is doing the crawling don't come back to look at a site again until some period of time has passed. So if I ran the test again, chances are the search engines wouldn't have indexed the changes yet and would still report 0 sites citing my sites - but not for the same reason! So, I'll do that as a separate item later on down the road when it is appropriate.
However, what I can say is that this bug is very real and rather prevalent. Even the popular site of reddit.com has this bug.
Code
dannwaneri
/
seo-agent
Open-source SEO audit agent — real-browser checks, backlink scoring, and multi-engine AI citation tracking (Claude, ChatGPT, Gemini, Perplexity, Copilot) via SearchApi
seo-agent
A local SEO co-pilot built with Python, Browser Use, and the Claude API. Visits real pages in a visible browser window, extracts SEO signals, checks for broken links, scores backlinks, surfaces GSC quick wins, maps internal link clusters, and writes structured reports — resumable if interrupted.
Ran it on my own sites. Found a title cannibalising its own homepage, a position 9.5 query with 0% CTR, two missing internal links, and an orphan page with no path to it.
Everything is open source.
Table of Contents
- What It Does
- Stack
- Installation
- Configuration
- Usage
- Modules
- Output
- PASS/FAIL Rules
- Cost
- Scheduling
- Attestation Verification
- Environment Variables
- Architecture
- Contributing
- Writing
- License
- Limitations
What It Does
Core audit (runs on a URL list):
- Visits each URL in a real Chromium browser — not a headless scraper
- Extracts title, meta description, H1s, and canonical tag via Claude API
- Checks for broken same-domain links asynchronously using…
SearchApi
SearchApi's Google search engine is what powered the domain discovery for this piece.
Note: I was given some API credits from SearchApi to produce this article as part of their Developer Ambassador program. However, the integration I created and my testing of the API was done all by me.

Top comments (5)
The fact that a site can work perfectly for humans while quietly serving the wrong content to crawlers is easy to overlook. Automated crawlability checks like this can catch issues before they affect AI visibility.
Appreciate it , that's exactly why I built the tool instead of just checking my own 2 sites.
This is a good one. SPA fallback can make everything look fine because you still get a 200, while the crawler actually receives your homepage instead of llms.txt or robots.txt.
I usually check the raw response and Content-Type directly. Same for sitemap.xml.
We added similar checks into Prerender Buddy because these small crawler-facing issues are very easy to miss from a normal browser.
Ngl, that's the sort of thing that slips between the cracks... Good job taking the initiative once someone brought it up, there's alot of sites that likely need a refresh!
Yeah, the annoying part is it's invisible until someone checks. Reddit's on that 5/50 list, so "a lot of sites need a refresh" isn't an exaggeration.