Today I worked through a beginner forensics CTF challenge. The description was one line: here is a .wav file, find the flag. The hint said to try tools like Audacity or Sonic and look at the spectrogram.
I'll skip the flag so I don't spoil it for anyone. Here's what I did and what I got wrong first.
First look
The file is 10 seconds, 44.1 kHz, stereo. Both channels turned out to be identical, so I only used the left one.
from scipy.io import wavfile
rate, data = wavfile.read("audio.wav")
print(rate, data.shape, data.shape[0] / rate, "seconds")
# 44100 (441000, 2) 10.0 seconds
left = data[:, 0].astype(float)
If you play it, it's noise. That's the clue. When a challenge says "hidden in noise" and the hint says spectrogram, the thing hiding in there is probably an image drawn in frequency over time.
What a spectrogram shows
A normal waveform plot shows loudness over time. A spectrogram splits the audio into short chunks, runs an FFT on each chunk, and plots which frequencies are loud at which moment. Time goes left to right, frequency goes bottom to top, brightness is energy.
So if someone generates sound that is loud at exactly the right frequencies at exactly the right times, the picture they want shows up. Text works fine.
First attempt: matplotlib defaults
import matplotlib.pyplot as plt
fig, ax = plt.subplots(figsize=(14, 5))
ax.specgram(left, Fs=rate)
fig.savefig("spec.png")
Matplotlib's specgram defaults to NFFT=256. That already showed text, but it looked rough. There were thin white vertical gaps, and I got a divide by zero encountered in log10 warning.
My guess about the gaps: the plot takes a log of the energy, and a chunk that is pure silence has zero energy. I checked the samples and there are 34 runs of 256 or more zeros in a row, which fits. That's a guess about why, I didn't dig further.
Second attempt: bigger window, more overlap
ax.specgram(left, NFFT=2048, Fs=rate, noverlap=1536, cmap="viridis")
ax.set_ylim(0, 8000)
Two things changed.
-
NFFT=2048gives finer frequency detail, so the letters get sharper vertically. The trade-off is that you lose time detail, since each chunk is longer. -
noverlap=1536means each chunk shares 75% of its samples with the previous one. That smooths the picture horizontally.
Then I cut the y-axis at 8 kHz. Everything interesting sat below about 7 kHz. Zooming in made the text tall enough to read comfortably.
Now the whole line of text was clear, a short label followed by a long string of characters. Reading it out was just squinting at the picture. If two characters look alike, zoom in more.
There was also a faint line rising along the bottom edge between roughly 1 and 2 seconds. I don't know what it is. It might just be part of how the file was made.
The Audacity route
The hint mentioned Audacity. I did this one in Python and haven't tried that route, but the Audacity manual says you switch a track to spectrogram view from the track's dropdown menu. You can then adjust the view settings there if the text looks blurry. It's the same idea as changing NFFT above.
What I'd take from it
- Noise in an audio challenge plus a spectrogram hint usually means a picture hidden in the frequencies.
- Window size and overlap are the two knobs. If the picture is blurry, change them before you assume it's a different trick.
- Crop the frequency range to where the signal is.
- Always check both channels. Here they matched, but in other challenges they may not.
Sources
- Matplotlib
specgramdocs (NFFTdefault 256,noverlap): https://matplotlib.org/stable/api/_as_gen/matplotlib.pyplot.specgram.html - Audacity manual, Spectrogram View: https://manual.audacityteam.org/man/spectrogram_view.html
Top comments (0)