DEV Community

KeyVault Edge
KeyVault Edge

Posted on Originally published at keyvaultedge.com

How to audit your codebase for exposed API keys (free tools)

Accidental credential commits are often found weeks or months after they happen, long after anyone who was watching had time to use them. This guide walks through a systematic audit of your repositories, git history, Docker images and CI pipelines using free, open-source tools.

What you're scanning and why

A credential audit is not just scanning your current working tree. Secrets committed to git are permanent in history unless you rewrite it. Deleting a file does not remove the secret: the commit with the original file is still reachable via git log and git show.

A proper audit covers:

  • Full git history: every commit, including deleted files and reverted changes
  • All branches and tags: feature branches often hold secrets that never reached main
  • Docker image layers: build-time ARG and ENV instructions that baked secrets in
  • CI/CD configs: .github/workflows, .gitlab-ci.yml, and environment variables echoed into logs
  • Lock files: rarely, npm/yarn lockfiles contain registry auth tokens
  • Build artifacts: compiled frontends that embedded env vars at build time

Scanning git history with Gitleaks

Gitleaks (MIT) is one of the most widely used open-source secret scanners, with built-in rules for OpenAI, AWS, Stripe, GitHub, Twilio and most major providers.

# Install (macOS)
brew install gitleaks

# Install (Linux)
curl -sSL https://github.com/gitleaks/gitleaks/releases/latest/download/gitleaks_linux_x64.tar.gz | tar -xz
sudo mv gitleaks /usr/local/bin/

# Scan full git history of the current repo
gitleaks detect --source . --verbose

# Write a JSON report
gitleaks detect --source . --report-path gitleaks-report.json --report-format json
Enter fullscreen mode Exit fullscreen mode

detect scans every commit. On large repos it can take a few minutes; --log-opts accepts any git log options, e.g. --log-opts="--since=2025-01-01" to scan recent history only.

Add it as a pre-commit hook so new secrets never land:

# .pre-commit-config.yaml
repos:
  - repo: https://github.com/gitleaks/gitleaks
    rev: v8.21.0
    hooks:
      - id: gitleaks
Enter fullscreen mode Exit fullscreen mode
pre-commit install
Enter fullscreen mode Exit fullscreen mode

Deep scanning with TruffleHog

TruffleHog (AGPL-3.0) goes further by verifying detected credentials against their upstream APIs: it tells you not just that a string looks like an OpenAI key, but whether that key is currently valid.

# Scan local git history, verified findings only
trufflehog git file://. --only-verified

# Scan a public GitHub repo
trufflehog github --repo https://github.com/yourusername/yourrepo

# Scan every repo in an org
trufflehog github --org yourorgname --only-verified
Enter fullscreen mode Exit fullscreen mode

--only-verified is great for triage: it filters out false positives and surfaces credentials that are exploitable right now.

Scanning Docker images

Images built with ARG or ENV instructions that reference real keys bake those keys into a layer permanently. A later instruction that overwrites the variable doesn't help: the earlier layer is still in the image.

# Scan a local or registry image
trufflehog docker --image yourimage:latest

# Quick manual check of every layer
docker save yourimage:latest | tar -xO | strings | grep -E 'sk-proj|AKIA|ghp_'
Enter fullscreen mode Exit fullscreen mode

The classic mistake:

# Bakes the real key into an image layer: don't
ARG OPENAI_API_KEY
ENV OPENAI_API_KEY=$OPENAI_API_KEY
# Inject at runtime instead (docker run -e, or your orchestrator's secrets)
Enter fullscreen mode Exit fullscreen mode

Auditing CI/CD pipelines

Build logs are often public, especially in open source. Look for:

  • env: blocks with literal values (OPENAI_KEY: sk-proj-...). References like ${{ secrets.X }} are fine.
  • Steps that print the environment: run: echo $OPENAI_API_KEY, env, printenv.
  • Test scripts that hardcode keys "for easier local testing".
  • Cached artifacts that include environment dumps from crash reporters.

GitHub's built-in scanning

GitHub secret scanning runs on public repositories and can be enabled for private ones on paid plans. For partner patterns (OpenAI, Stripe, AWS and many more) GitHub also notifies the provider, which may revoke the key.

Turn on push protection under Settings → Code security → Secret scanning to block commits containing known secret patterns before they land.

When you find something

  1. Rotate the credential immediately. Don't wait for the history rewrite; assume it's compromised.
  2. Check for unauthorised usage in the provider's logs (OpenAI usage page, AWS CloudTrail, Stripe logs).
  3. Remove it from history with git filter-repo (not filter-branch), force-push everywhere, and have contributors re-clone.
  4. Ask GitHub support to clear cached views of the old commits; rewriting history doesn't purge them.
  5. Assume forks still have it. You can't rewrite other people's forks, which is why step 1 comes first.

Making the next leak cheaper

Scanning finds the leaks you can see. The other half is making a leaked string less valuable: scoped keys, spending limits, and short-lived or origin-bound tokens.

That second half is what I'm building with KeyVault Edge: your code ships a token that the edge proxy only exchanges for the real key on requests from origins you allow. It's one layer, not a replacement for rotation and scanning.

Originally published at keyvaultedge.com.

Top comments (0)