DEV Community

Cover image for I Recorded 10,847 Frames While Coding. Only 30 Were Worth Keeping.
Pavel Kazantsev
Pavel Kazantsev

Posted on Edited on Originally published at pkazantsev.com

I Recorded 10,847 Frames While Coding. Only 30 Were Worth Keeping.

I recorded a two-hour coding session and got 10,847 frames.

The final timelapse needed only 30.

So I built a pipeline to figure out which 30 were actually worth keeping.

ProgressCut compressing a coding session into a short visual story

Source session: youtube.com/watch?v=GznmPACXBlY

Most of those 10,847 frames were noise: the cursor blinked, the system clock ticked, while nothing meaningful in the code changed.

The first step isn't clever. It's just getting rid of those.

ProgressCut is a local macOS app that turns a coding session into a short visual story. No talking, no editing. You hit stop and it hands you a GIF. Here's how the algorithm actually works.

The compression pipeline

Each stage removes a different kind of noise:

Capture → dHash dedupe → SSIM novelty → Segments → 30 moments
10,847  →    1,203     →     284      →    48     →     30
Enter fullscreen mode Exit fullscreen mode

Real numbers from a 2-hour session.

Five-stage pipeline: 10,847 raw frames → dHash 1,203 → SSIM 284 → segments 48 → story 30 moments


Stage 1 — dHash: kill the identical frames

89% gone in one pass. At 2 fps, consecutive screenshots are often nearly identical: the cursor blinks, the clock changes, or a single character appears while the rest of the screen stays untouched.

dHash resizes each frame to 9×8 pixels, converts to greyscale, then compares each pixel to the one on its right (1 if brighter, 0 otherwise). 8 rows × 8 comparisons = 64 bits, stored as a BigInt. XOR two hashes and count the differing bits. Distance ≤ 8: duplicate.

dHash pipeline: source frame to 9x8 greyscale grid to 64-bit comparison bits, Hamming distance decision

Kernighan's bit trick — O(set bits) not O(64)

Counting those differing bits naively loops 64 times. Kernighan's trick loops once per set bit, so for near-duplicates that's typically 2–6 iterations:

export function hammingDistance(a: DHash, b: DHash): number {
  let diff  = a ^ b;      // XOR: 1 where bits differ
  let count = 0;
  while (diff > 0n) {
    diff &= diff - 1n;    // clears the lowest set bit
    count++;
  }
  return count;
}
Enter fullscreen mode Exit fullscreen mode

Why BigInt? JS numbers are IEEE 754 doubles: they top out at 53 safe integer bits. A dHash is 64 bits, so BigInt is the only way to get exact XOR. The n suffix (0n, 1n) is the literal syntax.

Result: 10,847 → 1,203 frames (−89%)


Stage 2 — SSIM: kill the boring frames

Unique isn't the same as interesting. After dHash you still have ~1,200 frames that each look slightly different, but most of that difference is one more character on a line. You need a score that reflects how much the visual state actually changed, not just whether the pixels moved.

SSIM (Structural Similarity Index) compares luminance, contrast, and spatial structure between consecutive frames. Score close to 1: nothing happened. Score close to 0: something worth keeping.

         (2μₓμᵧ + C₁)(2σₓᵧ + C₂)
SSIM = ─────────────────────────────────
       (μₓ² + μᵧ² + C₁)(σₓ² + σᵧ² + C₂)
Enter fullscreen mode Exit fullscreen mode

Novelty score timeline showing spikes when significant visual changes occur, with selected moments marked as purple dots

Full-frame SSIM would average out a big change in one corner against stagnation everywhere else. Instead the frame is tiled into 16 regions, SSIM computed per region, then averaged. A file-save that changes the status bar doesn't steal credit from a tab switch that changed everything.

Single-pass variance — one loop, not two

The naive path does two loops: one for the mean, one for the variance. Both collapse into one using Var(X) = E[X²] − E[X]²: accumulate Σx, Σx², Σy, Σy², Σxy in a single pass and derive everything after:

let sumA=0, sumB=0, sumAA=0, sumBB=0, sumAB=0;
const n = (x1-x0) * (y1-y0);

for (let y=y0; y<y1; y++) {
  for (let x=x0; x<x1; x++) {
    const pa = a.pixels[y*W+x] ?? 0;
    const pb = b.pixels[y*W+x] ?? 0;
    sumA  += pa;    sumB  += pb;
    sumAA += pa*pa; sumBB += pb*pb; sumAB += pa*pb;
  }
}

const muA  = sumA/n,  muB  = sumB/n;
const varA  = sumAA/n - muA*muA;   // Var(X) = E[X²] − E[X]²
const varB  = sumBB/n - muB*muB;
const covAB = sumAB/n - muA*muB;
Enter fullscreen mode Exit fullscreen mode

Result: 1,203 → 284 frames (−77%)


Stage 3 — Segmentation: budget the moments fairly

Sessions aren't uniformly active. There's 20 minutes of flow where you're actually building something, then 10 minutes reading docs, then another sprint. If you spread the 30-moment budget evenly across the timeline, the reading gaps eat slots that should go to the interesting parts.

The pipeline splits the frame sequence into activity windows by novelty density. Each window's share of the budget scales with how much happened inside it:

warm-up │ ███████████ active  │ ░ idle │ ████████ active │ ████ wrap-up
 2 mom  │   12 moments        │ 1 mom  │  9 moments      │ 6 mom
Enter fullscreen mode Exit fullscreen mode

Temporal segmentation showing timeline divided into active and idle windows with budget bars below

A 20-minute sprint might take 8 moments; a 10-minute idle stretch gets 1.

Result: 284 frames → 48 segments → 30 moments


What comes out

30 moments → a GIF or MP4, under 30 seconds, no input from you. Each moment holds for a duration proportional to the activity level of its segment.

Input:  ░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░  10,847 frames
Output: █  ██   █████  █  ████  █  ██   ████      30 moments
Enter fullscreen mode Exit fullscreen mode

Compression ratio: 10,847 grey frames in the top bar, 30 purple moments in the bottom bar

No frames leave the machine. The app calls no model. Run it twice on the same input and you get the same output. That last part matters more than it sounds: a non-deterministic story compressor is harder to trust and impossible to debug.


What's next

Better novelty scoring with embeddings

SSIM can measure structural change, but it has no idea what that change means. Opening a new file and typing a few characters may affect a similar amount of screen area even though one is a much more meaningful transition.

An embedding-based scorer might separate those cases better by measuring semantic change rather than raw pixel structure. That's the next hypothesis to test.

The port in packages/engine is already the swap point: implement a new scorer behind the same interface, run the same 187 tests, and leave the rest of the pipeline untouched.

But swapping scorers without a way to evaluate the result is just guessing. The feedback loop comes first: let users remove moments they don't think belong in the final story, then use that data to compare scoring methods.

SSIM stays the default until an embedding-based scorer demonstrates a measurable improvement.

Longer term: adaptive budget

The 30-moment budget is fixed right now. A focused 20-minute session and a messy two-hour session both get 30 moments.

The segmentation step already estimates how much activity happened in each window, so the next step is to let the total story length depend on the session itself.

A short session with lots of meaningful changes might need fewer, faster moments. A long session with large idle stretches should compress harder instead of wasting slots on inactivity.

If you were building this, would you keep SSIM or move to embeddings?

Code

MIT, pnpm monorepo, 187 tests, CI on macos-latest. The algorithm lives in packages/engine/src/ with no Electron dependency, runs in plain Node. If you want to poke at the frame selection logic without building the whole app, that's the entry point.

→ github.com/icesurf666/progress-cut

Top comments (1)

Collapse
 
entersuren profile image
paul • • Edited

SSIM is good enough for now. Embeddings would catch more but you'd need a model on device - that's a whole other problem. Curious how you're thinking about the binary size tradeoff.