DEV Community

Stephano kambeta
Stephano kambeta

Posted on

Build a Free Broken Link Checker with GitHub Actions

A tutorial you wrote last year links to a company's documentation page. Since then, the company redesigned its site, the page moved, and the link now leads to a "404 Not Found."

Nobody tells you. Readers just click, hit a dead end, and leave.

Old posts are where this piles up. The more you publish, the harder it gets to remember what you linked to. Checking every link by hand doesn't scale, but a small script can do it for you every week.

This guide builds one. It reads your sitemap, checks every link on every page, and opens a GitHub issue when something is broken. When the links are fixed, it closes the issue on its own. It costs nothing and needs nothing to install.

What you need

  • A website with a sitemap.xml file (most blogs and site builders create one)
  • A GitHub repository to hold the script. It can be a new, empty one
  • Issues turned on in that repository (Settings, then Features)

You don't need to install any Python packages. The script only uses what Python already includes.

Step 1: Add the script

Create a file called check_links.py in your repository and paste this in:

"""Weekly broken link checker.

Reads your sitemap, opens every page, and checks every link on it.
If it finds broken links, it writes a short report to report.md.
Needs nothing but Python, so there is nothing to install.
"""
import os
import urllib.error
import urllib.request
import xml.etree.ElementTree as ET
from concurrent.futures import ThreadPoolExecutor
from datetime import date
from html.parser import HTMLParser
from urllib.parse import urldefrag, urljoin

SITEMAP_URL = os.environ["SITEMAP_URL"]
HEADERS = {"User-Agent": "Mozilla/5.0 (compatible; LinkChecker/1.0)"}
TIMEOUT = 15

# Many sites use these codes to turn away bots, even when the page is fine.
# We can't tell if the link is really broken, so we don't report it.
NOT_SURE = {401, 403, 429, 999}


def open_url(url, method="GET"):
    request = urllib.request.Request(url, headers=HEADERS, method=method)
    return urllib.request.urlopen(request, timeout=TIMEOUT)


class LinkParser(HTMLParser):
    """Collects the href of every <a> tag on a page."""

    def __init__(self):
        super().__init__()
        self.links = []

    def handle_starttag(self, tag, attrs):
        href = dict(attrs).get("href")
        if tag == "a" and href:
            self.links.append(href)


def get_pages():
    with open_url(SITEMAP_URL) as response:
        root = ET.fromstring(response.read())
    return [el.text.strip() for el in root.iter() if el.tag.endswith("}loc")]


def get_links(page):
    """Return the absolute http(s) links found on one page."""
    try:
        with open_url(page) as response:
            html = response.read(3_000_000).decode("utf-8", errors="replace")
    except Exception:
        return []  # the page itself gets reported when we check it below
    parser = LinkParser()
    parser.feed(html)
    links = set()
    for href in parser.links:
        url, _ = urldefrag(urljoin(page, href))
        if url.startswith(("http://", "https://")):
            links.add(url)
    return links


def check_link(url):
    """Return None if the link works, or a short reason if it doesn't."""
    reason = None
    for _ in range(2):  # one retry, in case of a temporary hiccup
        try:
            open_url(url, "HEAD").close()
            return None
        except Exception:
            pass  # some servers reject HEAD requests, so try a normal GET
        try:
            open_url(url, "GET").close()
            return None
        except urllib.error.HTTPError as error:
            if error.code in NOT_SURE:
                return None
            reason = f"HTTP {error.code}"
            if error.code < 500:
                return reason  # a 404 won't fix itself in five seconds
        except Exception as error:
            reason = f"Could not connect ({type(error).__name__})"
    return reason


def main():
    if os.path.exists("report.md"):
        os.remove("report.md")

    pages = get_pages()
    print(f"Found {len(pages)} pages in the sitemap")

    # Map each link to the pages it appears on
    found_on = {page: {"the sitemap"} for page in pages}
    with ThreadPoolExecutor(max_workers=5) as pool:
        for page, links in zip(pages, pool.map(get_links, pages)):
            for link in links:
                found_on.setdefault(link, set()).add(page)

    urls = list(found_on)
    print(f"Checking {len(urls)} links...")
    with ThreadPoolExecutor(max_workers=8) as pool:
        results = list(pool.map(check_link, urls))

    broken = [(url, reason) for url, reason in zip(urls, results) if reason]
    print(f"Broken links: {len(broken)}")
    if not broken:
        return

    lines = [
        f"Checked {len(pages)} pages and {len(urls)} links on {date.today()}.",
        f"Found {len(broken)} broken link(s).",
        "",
    ]
    for url, reason in sorted(broken):
        lines.append(f"### {url}")
        lines.append(f"Problem: {reason}")
        lines.append("Found on:")
        lines += [f"- {source}" for source in sorted(found_on[url])]
        lines.append("")
    with open("report.md", "w") as report:
        report.write("\n".join(lines))


main()
Enter fullscreen mode Exit fullscreen mode

Here's what it does, in plain terms:

  1. It reads your sitemap to get a list of pages.
  2. It opens each page and collects every link.
  3. It checks each link once, even if it appears on ten pages.
  4. If any are broken, it writes a report that lists each broken link and where it appears.

Your own pages get checked too. That means a page in your sitemap that no longer loads shows up in the report.

Step 2: Add the workflow

Create the file .github/workflows/link-check.yml:

name: Weekly link check

on:
  schedule:
    - cron: '23 6 * * 1'   # Mondays at 06:23 UTC
  workflow_dispatch:        # adds a "Run workflow" button for testing

permissions:
  contents: read
  issues: write

jobs:
  check-links:
    runs-on: ubuntu-latest
    timeout-minutes: 15
    steps:
      - uses: actions/checkout@v4   # use the latest major version

      - name: Check links
        env:
          SITEMAP_URL: https://example.com/sitemap.xml   # change this
        run: python3 check_links.py

      - name: Open, update or close the issue
        env:
          GH_TOKEN: ${{ github.token }}
          GH_REPO: ${{ github.repository }}
        run: |
          gh label create broken-links --color D93F0B \
            --description "Found by the weekly link check" --force
          existing=$(gh issue list --label broken-links --state open \
            --json number --jq '.[0].number // empty')

          if [ -f report.md ]; then
            if [ -n "$existing" ]; then
              gh issue edit "$existing" --body-file report.md
            else
              gh issue create --title "Broken links found" \
                --label broken-links --body-file report.md
            fi
          elif [ -n "$existing" ]; then
            gh issue close "$existing" --comment "No broken links found this week."
          fi
Enter fullscreen mode Exit fullscreen mode

Change SITEMAP_URL to your own sitemap address. That's the only setting you need to touch.

A few choices in there are worth explaining:

  • An odd start time. 23 6 avoids the top of the hour, when GitHub is busiest and scheduled runs are most likely to be delayed.
  • A timeout. If something hangs, the job stops after 15 minutes instead of running for hours.
  • Limited permissions. The workflow can read your code and manage issues, and nothing else.
  • One issue at a time. If there's already an open issue, the workflow updates it instead of creating a new one every week. When a week comes back clean, it closes the issue.

Step 3: Test it

Don't wait until Monday. Go to the Actions tab in your repository, choose "Weekly link check", and click "Run workflow".

To see the issue appear, add a link you know is broken to one of your pages, wait for your site to update, and run it again. Then fix the link and run it once more. The issue should close itself.

What counts as "broken"?

The script reports a link when:

  • The page returns a "not found" or similar error, like 404 or 410
  • The server returns a server error, like 500, even after a retry
  • The address can't be reached at all, for example because the domain no longer exists

It doesn't report links that return 401, 403 or 429. Many sites, LinkedIn among them, block automated visitors even when the page works fine for people. Reporting those would fill your issue with false alarms, so the script skips them. The tradeoff is that a link that's really broken behind one of those codes goes unnoticed.

Redirects are followed, so an old link that lands on a working page counts as fine.

Tweaks you might want

Ignore certain sites. If a site keeps causing false alarms, skip it. In get_links, change the if line so it also checks a list of domains to ignore:

IGNORE = ("linkedin.com", "twitter.com")   # add near the top

# ...and in get_links:
if url.startswith(("http://", "https://")) and not any(d in url for d in IGNORE):
Enter fullscreen mode Exit fullscreen mode

Run monthly instead of weekly. For a small blog, once a month is often enough. Change the cron line to something like '23 6 1 * *'.

Be gentler on other sites. The script checks a handful of links at a time. If you have a huge site, lower max_workers so you don't send too many requests at once.

What it can't catch

It's a simple tool, and it has limits:

  • It only checks links in <a> tags. Images, scripts and stylesheets aren't checked.
  • It doesn't run JavaScript, so links that only appear after a page loads in a browser are missed.
  • It doesn't understand a sitemap index (a sitemap that lists other sitemaps). If yours is one, point it at each sitemap separately.
  • Some sites show a "page not found" message but still return a normal success code. The script can't tell.

Will it stay running?

This is worth a thought, because the checker can fail quietly too.

In a public repository, GitHub switches off scheduled workflows after 60 days without any repository activity. If the checker repo is one you never touch, it may stop running without telling you. Put a reminder in your calendar, or glance at the Actions tab now and then.

The same idea applies to any automation you set and forget. I wrote about it in how to monitor your automations when something breaks.

Is this worth automating?

For most blogs, yes. It's a good example of a task that suits automation:

  • It's repetitive and boring
  • It's easy to check that it worked
  • A mistake costs very little, since the worst case is a false alarm in an issue

If you want a way to judge other tasks, I wrote about when not to automate.

Quick answers

How often should I check for broken links?
Weekly is fine for an active blog. Monthly works for a small one that rarely changes.

Does this cost anything?
No. A weekly run takes a few minutes a month. Public repositories don't use up any free minutes, and private repositories only use a tiny part of the monthly allowance.

Can I check a site that isn't mine?
It works for any site with a sitemap, but be considerate. Don't run it often against other people's sites.

Why did it report a link that works in my browser?
Some servers treat scripts differently from browsers. Open the link yourself first. If it works, add that domain to the ignore list.

The short version

Broken links don't break anything loudly. They just make your older posts a little less useful, one dead end at a time.

A weekly check turns that into a short list of things to fix. If you build it once, you can leave it alone.

If you want more practical guides like this, you can find them on Procwire. I also wrote a Python uptime monitor guide that pairs well with this one.

Top comments (0)