DEV Community

Siekwie
Siekwie

Posted on AI-assisted

GitHub only keeps 14 days of repo traffic. I built an open-source tool that keeps all of it

Open Insights → Traffic on any repository you own and GitHub shows you views, unique visitors, clones and referring sites. For the last 14 days. Day 15 is gone, and there is no setting, plan or export that brings it back.

That makes some simple questions impossible to answer:

  • Did the launch post in March bring more people than the one in June?
  • Is the project growing, or was last week just a good week?
  • Which site has sent the most visitors over the whole life of the repo?

So I built RepoEasy. It signs in with GitHub, reads the traffic of your repositories every day and keeps it. The code is MIT-licensed, and there is a hosted version with a demo you can open without an account: repoeasy.wiest-lab.eu.

RepoEasy overview: lifetime and recent traffic across all repositories

What it keeps

  • Views, unique visitors, clones and unique cloners per day, private repositories included
  • Referring sites and popular pages
  • Stars, forks, open issues, pull requests and release downloads, one snapshot per day
  • Star history from before you signed up, backfilled from the timestamp GitHub keeps for every star
  • The same star, fork and release numbers for any public repository you choose to follow: a dependency, a competitor, a project you are curious about

Once the data is there, the rest follows: CSV and JSON export, an activity feed with traffic spikes and star milestones, webhook alerts for Discord or Slack, and an opt-in public stats page with README badges.

A repository page in RepoEasy: stars, forks, releases, views and clones

How the archive is built

Reading the numbers is the easy part:

gh api repos/OWNER/REPO/traffic/views
Enter fullscreen mode Exit fullscreen mode

That returns up to 14 daily entries, each with count and uniques. The interesting part is merging overlapping 14-day windows, day after day, without getting the numbers wrong. Three things stood out.

1. Today is always incomplete

The newest entry is a partial day and keeps growing until midnight UTC. So a day that is already stored is never overwritten with a lower number, only raised:

INSERT INTO traffic_daily (repo_id, day, views, uniques) VALUES (?, ?, ?, ?)
ON CONFLICT(repo_id, day) DO UPDATE SET
  views = MAX(views, excluded.views),
  uniques = MAX(uniques, excluded.uniques);
Enter fullscreen mode Exit fullscreen mode

Every sync reads all 14 days again, so a partial value is corrected by the next run, and a missed sync costs nothing as long as one succeeds within two weeks.

2. Unique visitors do not add up

GitHub reports unique visitors per day. Nothing tells you that Monday's visitor and Thursday's visitor are the same person. So "unique visitors over 90 days" in RepoEasy is the sum of the daily uniques, and the UI says so wherever that number appears. I would rather show a number with an honest label than a nicer-looking one I cannot back up.

3. Referrers only exist as 14-day totals

For referring sites and popular pages there is no per-day breakdown, only "the last 14 days". Adding up a snapshot from every day would count each visit up to 14 times. RepoEasy stores a snapshot every day and, for lifetime numbers, adds up snapshots taken 14 days apart. That tiles the timeline without overlap. It is an estimate and labelled as one.

The stack

One Node process and one SQLite file.

  • Hono for HTTP, better-sqlite3 for storage
  • A React single-page app with hand-drawn SVG charts, no chart library and no UI framework
  • GitHub tokens encrypted at rest with AES-256-GCM
  • A small VPS behind Caddy for the hosted version, docker compose up -d for your own

A sync costs about one GraphQL call per 15 repositories plus four to six REST calls per tracked repository, all against the user's own rate limit.

About the permission it asks for

GitHub has no read-only permission that includes traffic, so signing in asks for the repo scope. That is a lot to hand to a service you met five minutes ago. RepoEasy only reads, with one exception: a form that lets you edit a repository's description, homepage and topics, and only when you use it.

If you would rather not grant that to a hosted service, run it yourself. In single-user mode it works with a fine-grained token limited to Administration: read and Metadata: read, and every feature is unlocked.

Try it

  • Hosted, with a demo: repoeasy.wiest-lab.eu. Free for 3 tracked repositories and 10 followed ones, with the full history kept. Pro is €1.50 a month for unlimited repositories, sync every 6 hours, API tokens, webhooks and share pages.
  • Source: github.com/Siekwie/RepoEasy

One thing to know before you decide: the archive starts on the day you connect. GitHub cannot give back days it has already dropped, so the best time to start is before your next launch, not after it.

What would you want to learn from long-term traffic data that 14 days cannot tell you? I am collecting ideas for what to build next.

Top comments (0)