DEV Community

Curious4AI
Curious4AI

Posted on AI-assisted

Finding a flag hidden in a .wav file with a spectrogram

Today I worked through a beginner forensics CTF challenge. The description was one line: here is a .wav file, find the flag. The hint said to try tools like Audacity or Sonic and look at the spectrogram.

I'll skip the flag so I don't spoil it for anyone. Here's what I did and what I got wrong first.

First look

The file is 10 seconds, 44.1 kHz, stereo. Both channels turned out to be identical, so I only used the left one.

from scipy.io import wavfile

rate, data = wavfile.read("audio.wav")
print(rate, data.shape, data.shape[0] / rate, "seconds")
# 44100 (441000, 2) 10.0 seconds

left = data[:, 0].astype(float)
Enter fullscreen mode Exit fullscreen mode

If you play it, it's noise. That's the clue. When a challenge says "hidden in noise" and the hint says spectrogram, the thing hiding in there is probably an image drawn in frequency over time.

What a spectrogram shows

A normal waveform plot shows loudness over time. A spectrogram splits the audio into short chunks, runs an FFT on each chunk, and plots which frequencies are loud at which moment. Time goes left to right, frequency goes bottom to top, brightness is energy.

So if someone generates sound that is loud at exactly the right frequencies at exactly the right times, the picture they want shows up. Text works fine.

First attempt: matplotlib defaults

import matplotlib.pyplot as plt

fig, ax = plt.subplots(figsize=(14, 5))
ax.specgram(left, Fs=rate)
fig.savefig("spec.png")
Enter fullscreen mode Exit fullscreen mode

Matplotlib's specgram defaults to NFFT=256. That already showed text, but it looked rough. There were thin white vertical gaps, and I got a divide by zero encountered in log10 warning.

My guess about the gaps: the plot takes a log of the energy, and a chunk that is pure silence has zero energy. I checked the samples and there are 34 runs of 256 or more zeros in a row, which fits. That's a guess about why, I didn't dig further.

Second attempt: bigger window, more overlap

ax.specgram(left, NFFT=2048, Fs=rate, noverlap=1536, cmap="viridis")
ax.set_ylim(0, 8000)
Enter fullscreen mode Exit fullscreen mode

Two things changed.

  • NFFT=2048 gives finer frequency detail, so the letters get sharper vertically. The trade-off is that you lose time detail, since each chunk is longer.
  • noverlap=1536 means each chunk shares 75% of its samples with the previous one. That smooths the picture horizontally.

Then I cut the y-axis at 8 kHz. Everything interesting sat below about 7 kHz. Zooming in made the text tall enough to read comfortably.

Now the whole line of text was clear, a short label followed by a long string of characters. Reading it out was just squinting at the picture. If two characters look alike, zoom in more.

There was also a faint line rising along the bottom edge between roughly 1 and 2 seconds. I don't know what it is. It might just be part of how the file was made.

The Audacity route

The hint mentioned Audacity. I did this one in Python and haven't tried that route, but the Audacity manual says you switch a track to spectrogram view from the track's dropdown menu. You can then adjust the view settings there if the text looks blurry. It's the same idea as changing NFFT above.

What I'd take from it

  • Noise in an audio challenge plus a spectrogram hint usually means a picture hidden in the frequencies.
  • Window size and overlap are the two knobs. If the picture is blurry, change them before you assume it's a different trick.
  • Crop the frequency range to where the signal is.
  • Always check both channels. Here they matched, but in other challenges they may not.

Sources

Top comments (0)