A tutorial you wrote last year links to a company's documentation page. Since then, the company redesigned its site, the page moved, and the link now leads to a "404 Not Found."
Nobody tells you. Readers just click, hit a dead end, and leave.
Old posts are where this piles up. The more you publish, the harder it gets to remember what you linked to. Checking every link by hand doesn't scale, but a small script can do it for you every week.
This guide builds one. It reads your sitemap, checks every link on every page, and opens a GitHub issue when something is broken. When the links are fixed, it closes the issue on its own. It costs nothing and needs nothing to install.
What you need
- A website with a
sitemap.xmlfile (most blogs and site builders create one) - A GitHub repository to hold the script. It can be a new, empty one
- Issues turned on in that repository (Settings, then Features)
You don't need to install any Python packages. The script only uses what Python already includes.
Step 1: Add the script
Create a file called check_links.py in your repository and paste this in:
"""Weekly broken link checker.
Reads your sitemap, opens every page, and checks every link on it.
If it finds broken links, it writes a short report to report.md.
Needs nothing but Python, so there is nothing to install.
"""
import os
import urllib.error
import urllib.request
import xml.etree.ElementTree as ET
from concurrent.futures import ThreadPoolExecutor
from datetime import date
from html.parser import HTMLParser
from urllib.parse import urldefrag, urljoin
SITEMAP_URL = os.environ["SITEMAP_URL"]
HEADERS = {"User-Agent": "Mozilla/5.0 (compatible; LinkChecker/1.0)"}
TIMEOUT = 15
# Many sites use these codes to turn away bots, even when the page is fine.
# We can't tell if the link is really broken, so we don't report it.
NOT_SURE = {401, 403, 429, 999}
def open_url(url, method="GET"):
request = urllib.request.Request(url, headers=HEADERS, method=method)
return urllib.request.urlopen(request, timeout=TIMEOUT)
class LinkParser(HTMLParser):
"""Collects the href of every <a> tag on a page."""
def __init__(self):
super().__init__()
self.links = []
def handle_starttag(self, tag, attrs):
href = dict(attrs).get("href")
if tag == "a" and href:
self.links.append(href)
def get_pages():
with open_url(SITEMAP_URL) as response:
root = ET.fromstring(response.read())
return [el.text.strip() for el in root.iter() if el.tag.endswith("}loc")]
def get_links(page):
"""Return the absolute http(s) links found on one page."""
try:
with open_url(page) as response:
html = response.read(3_000_000).decode("utf-8", errors="replace")
except Exception:
return [] # the page itself gets reported when we check it below
parser = LinkParser()
parser.feed(html)
links = set()
for href in parser.links:
url, _ = urldefrag(urljoin(page, href))
if url.startswith(("http://", "https://")):
links.add(url)
return links
def check_link(url):
"""Return None if the link works, or a short reason if it doesn't."""
reason = None
for _ in range(2): # one retry, in case of a temporary hiccup
try:
open_url(url, "HEAD").close()
return None
except Exception:
pass # some servers reject HEAD requests, so try a normal GET
try:
open_url(url, "GET").close()
return None
except urllib.error.HTTPError as error:
if error.code in NOT_SURE:
return None
reason = f"HTTP {error.code}"
if error.code < 500:
return reason # a 404 won't fix itself in five seconds
except Exception as error:
reason = f"Could not connect ({type(error).__name__})"
return reason
def main():
if os.path.exists("report.md"):
os.remove("report.md")
pages = get_pages()
print(f"Found {len(pages)} pages in the sitemap")
# Map each link to the pages it appears on
found_on = {page: {"the sitemap"} for page in pages}
with ThreadPoolExecutor(max_workers=5) as pool:
for page, links in zip(pages, pool.map(get_links, pages)):
for link in links:
found_on.setdefault(link, set()).add(page)
urls = list(found_on)
print(f"Checking {len(urls)} links...")
with ThreadPoolExecutor(max_workers=8) as pool:
results = list(pool.map(check_link, urls))
broken = [(url, reason) for url, reason in zip(urls, results) if reason]
print(f"Broken links: {len(broken)}")
if not broken:
return
lines = [
f"Checked {len(pages)} pages and {len(urls)} links on {date.today()}.",
f"Found {len(broken)} broken link(s).",
"",
]
for url, reason in sorted(broken):
lines.append(f"### {url}")
lines.append(f"Problem: {reason}")
lines.append("Found on:")
lines += [f"- {source}" for source in sorted(found_on[url])]
lines.append("")
with open("report.md", "w") as report:
report.write("\n".join(lines))
main()
Here's what it does, in plain terms:
- It reads your sitemap to get a list of pages.
- It opens each page and collects every link.
- It checks each link once, even if it appears on ten pages.
- If any are broken, it writes a report that lists each broken link and where it appears.
Your own pages get checked too. That means a page in your sitemap that no longer loads shows up in the report.
Step 2: Add the workflow
Create the file .github/workflows/link-check.yml:
name: Weekly link check
on:
schedule:
- cron: '23 6 * * 1' # Mondays at 06:23 UTC
workflow_dispatch: # adds a "Run workflow" button for testing
permissions:
contents: read
issues: write
jobs:
check-links:
runs-on: ubuntu-latest
timeout-minutes: 15
steps:
- uses: actions/checkout@v4 # use the latest major version
- name: Check links
env:
SITEMAP_URL: https://example.com/sitemap.xml # change this
run: python3 check_links.py
- name: Open, update or close the issue
env:
GH_TOKEN: ${{ github.token }}
GH_REPO: ${{ github.repository }}
run: |
gh label create broken-links --color D93F0B \
--description "Found by the weekly link check" --force
existing=$(gh issue list --label broken-links --state open \
--json number --jq '.[0].number // empty')
if [ -f report.md ]; then
if [ -n "$existing" ]; then
gh issue edit "$existing" --body-file report.md
else
gh issue create --title "Broken links found" \
--label broken-links --body-file report.md
fi
elif [ -n "$existing" ]; then
gh issue close "$existing" --comment "No broken links found this week."
fi
Change SITEMAP_URL to your own sitemap address. That's the only setting you need to touch.
A few choices in there are worth explaining:
-
An odd start time.
23 6avoids the top of the hour, when GitHub is busiest and scheduled runs are most likely to be delayed. - A timeout. If something hangs, the job stops after 15 minutes instead of running for hours.
- Limited permissions. The workflow can read your code and manage issues, and nothing else.
- One issue at a time. If there's already an open issue, the workflow updates it instead of creating a new one every week. When a week comes back clean, it closes the issue.
Step 3: Test it
Don't wait until Monday. Go to the Actions tab in your repository, choose "Weekly link check", and click "Run workflow".
To see the issue appear, add a link you know is broken to one of your pages, wait for your site to update, and run it again. Then fix the link and run it once more. The issue should close itself.
What counts as "broken"?
The script reports a link when:
- The page returns a "not found" or similar error, like 404 or 410
- The server returns a server error, like 500, even after a retry
- The address can't be reached at all, for example because the domain no longer exists
It doesn't report links that return 401, 403 or 429. Many sites, LinkedIn among them, block automated visitors even when the page works fine for people. Reporting those would fill your issue with false alarms, so the script skips them. The tradeoff is that a link that's really broken behind one of those codes goes unnoticed.
Redirects are followed, so an old link that lands on a working page counts as fine.
Tweaks you might want
Ignore certain sites. If a site keeps causing false alarms, skip it. In get_links, change the if line so it also checks a list of domains to ignore:
IGNORE = ("linkedin.com", "twitter.com") # add near the top
# ...and in get_links:
if url.startswith(("http://", "https://")) and not any(d in url for d in IGNORE):
Run monthly instead of weekly. For a small blog, once a month is often enough. Change the cron line to something like '23 6 1 * *'.
Be gentler on other sites. The script checks a handful of links at a time. If you have a huge site, lower max_workers so you don't send too many requests at once.
What it can't catch
It's a simple tool, and it has limits:
- It only checks links in
<a>tags. Images, scripts and stylesheets aren't checked. - It doesn't run JavaScript, so links that only appear after a page loads in a browser are missed.
- It doesn't understand a sitemap index (a sitemap that lists other sitemaps). If yours is one, point it at each sitemap separately.
- Some sites show a "page not found" message but still return a normal success code. The script can't tell.
Will it stay running?
This is worth a thought, because the checker can fail quietly too.
In a public repository, GitHub switches off scheduled workflows after 60 days without any repository activity. If the checker repo is one you never touch, it may stop running without telling you. Put a reminder in your calendar, or glance at the Actions tab now and then.
The same idea applies to any automation you set and forget. I wrote about it in how to monitor your automations when something breaks.
Is this worth automating?
For most blogs, yes. It's a good example of a task that suits automation:
- It's repetitive and boring
- It's easy to check that it worked
- A mistake costs very little, since the worst case is a false alarm in an issue
If you want a way to judge other tasks, I wrote about when not to automate.
Quick answers
How often should I check for broken links?
Weekly is fine for an active blog. Monthly works for a small one that rarely changes.
Does this cost anything?
No. A weekly run takes a few minutes a month. Public repositories don't use up any free minutes, and private repositories only use a tiny part of the monthly allowance.
Can I check a site that isn't mine?
It works for any site with a sitemap, but be considerate. Don't run it often against other people's sites.
Why did it report a link that works in my browser?
Some servers treat scripts differently from browsers. Open the link yourself first. If it works, add that domain to the ignore list.
The short version
Broken links don't break anything loudly. They just make your older posts a little less useful, one dead end at a time.
A weekly check turns that into a short list of things to fix. If you build it once, you can leave it alone.
If you want more practical guides like this, you can find them on Procwire. I also wrote a Python uptime monitor guide that pairs well with this one.
Top comments (0)