DEV Community

Cover image for An LLM Heard Women in 71% of Our 'Male' AI Voice Replies
אחיה כהן
אחיה כהן

Posted on Fully Autonomous

An LLM Heard Women in 71% of Our 'Male' AI Voice Replies

On 27 September a real caller on my business line spoke to our "male" AI voice using at, the feminine Hebrew "you". The voice sits at roughly 135 Hz. We had chosen it with a pitch check, and the pitch check was the mistake.

In English that would be a matter of taste. In Hebrew it's a correctness bug, because present-tense verbs carry gender. Our assistant ends its greeting with "recording and passing it on", which is מקליט ומעביר for a man and מקליטה ומעבירה for a woman. The system prompt fixes the grammatical gender, so the audio has to match it for the whole call. If it doesn't, the sentence contradicts itself.

So we measured how often each voice actually sounds like the gender it's supposed to be.

The setup

  • The speaker was gemini-3.8-live in he-IL, one prebuilt voice per run, thinking budget 0.
  • Google's Chirp 3 HD catalogue labels 16 of its voice names male. We ran 13. The other three had been recorded with feminine grammar in an earlier run, which would have confounded the test.
  • The script was four caller turns in Hebrew, sent as text through the realtime input (that's how our speech-to-text feeds the model in production), with the production prompt built for masculine grammar. Three separate calls per voice.
  • Every reply of 1.5 seconds or more was converted to what a phone line delivers, then sent to gemini-3.5-flash at temperature 0 with a single one-word question.

The conversion and the judge call, trimmed from the harness:

import audioop, io, wave
import numpy as np, soxr
from google.genai import types

def phone_wav(pcm24: bytes) -> bytes:
    """24 kHz model output -> 8 kHz mu-law round trip, i.e. what the caller hears."""
    x = np.frombuffer(pcm24, dtype=np.int16).astype(np.float32)
    x8 = soxr.resample(x, 24000, 8000, quality="HQ")
    x8 = np.clip(np.round(x8), -32768, 32767).astype(np.int16)
    x8 = audioop.ulaw2lin(audioop.lin2ulaw(x8.tobytes(), 2), 2)
    buf = io.BytesIO()
    with wave.open(buf, "wb") as w:
        w.setnchannels(1); w.setsampwidth(2); w.setframerate(8000)
        w.writeframes(x8)
    return buf.getvalue()

Q = ("Listen to the speaker in this phone recording. Does the voice sound like "
     "a man or a woman? Answer with exactly one word: man, woman, or unclear.")

async def judge(client, pcm24: bytes) -> str:
    r = await client.aio.models.generate_content(
        model="models/gemini-3.5-flash",
        contents=[types.Part.from_bytes(data=phone_wav(pcm24), mime_type="audio/wav"), Q],
        config=types.GenerateContentConfig(temperature=0))
    a = (r.text or "").strip().lower()
    return "woman" if "woman" in a else ("man" if "man" in a else "unclear")
Enter fullscreen mode Exit fullscreen mode

Check for "woman" before "man". The second string is inside the first.

The results

One judging pass, run on 30 September over the recordings made on 27 September:

Voice man woman unclear
Zubenelgenubi 6 6 0
Algieba 5 4 2
Sadaltager 4 8 0
Algenib 3 7 0
Alnilam 3 7 0
Charon 3 8 0
Schedar 3 7 1
Umbriel 3 8 0
Enceladus 2 8 1
Sadachbia 2 9 0
Iapetus 2 10 0
Puck 2 9 1
Fenrir 0 12 0

That's 146 replies: 103 judged female (71%), 38 male (26%), 5 unclear. Only one voice was heard as a man more often than as a woman. None of the 39 calls got through without at least one reply judged female.

These voices were speaking masculine Hebrew the whole time. If the judge were leaning on grammar, that would have pushed it toward "man", so the grammar can't explain the result.

The controls

A judge that says "woman" to everything would produce the same table, so we checked it against clips where the answer is known:

  • Female voices from the same model, speaking feminine grammar: 34 of 34 replies judged female.
  • Synthetic male callers from our test set, already at 8 kHz: judged male in 10 of 12 passes. Synthetic female callers: female in 6 of 6.
  • The 135 Hz voice from the opening is called Pulcherrima, and Google's catalogue labels it female. We had it speak masculine grammar anyway. It was judged male in 0 of 11 replies.

The judge can recognise a man on a phone line. It just doesn't hear one in most of these voices.

Temperature 0 didn't make the judge deterministic

This was the part I didn't expect. The same 11 Algieba clips were judged male 73% of the time on 27 September, then 45% and 55% on two passes on 30 September. One of the male control clips, an angry caller, came back man, woman, woman across three passes.

What that means in practice:

  1. One judging pass is an anecdote. Run at least three and report the spread.
  2. Keep known-answer controls in every run, not only in the first one.
  3. Look at direction before magnitude. In both full passes, at least 12 of the 13 voices came out at or below 50% male. The exact percentages kept moving.

Telling the model to sound male made it worse

The obvious fix was a line in the system prompt. Translated from the Hebrew, it said: "Your voice: a man's voice, low and steady. The whole call at exactly the same pitch, even in a short sentence, a thank-you or a goodbye. Never rise to a woman's pitch."

Across the same 13 voices, the male share went from 26% without the line to 22% with it. Algieba, the only voice that had been winning, dropped to 11%. My guess, which I haven't tested, is that naming "a woman" in the instruction nudges the output toward one. We had shipped the line that afternoon and removed it the same evening.

Pitch isn't the variable

This is old news in speech science. Hillenbrand and Clark (Attention, Perception & Psychophysics, 2009) resynthesised sentences to flip the speaker's apparent sex. Shifting pitch alone, or formants alone, usually failed. Shifting both worked about 82% of the time. An F0 threshold checks one of two cues, and it's the cue our 135 Hz voice already passed while still sounding like a woman.

What we shipped

  • Customers now choose from 11 voices, all female. The male voices are quarantined and get re-tested whenever the model changes.
  • My own line keeps the one male voice, because I can live with the risk there.
  • Since 28 September, the greeting on customer lines is spoken by the live model itself. Before that, a separate text-to-speech model rendered it under the same voice name, and you could hear the switch between the two engines.

If you're building voice agents in a language with grammatical gender, measure the voice the way the caller will hear it, at 8 kHz rather than from the vendor's sample. Pick voices with a judge and controls, not with an F0 cutoff, and don't expect adjectives in the prompt to fix it.

I build these systems at Achiya Automation, and the service in question is Onimli, a Hebrew AI answering service for missed calls.

If you've used an LLM as a judge for audio: how many passes do you run before you believe a number, and have you seen temperature 0 flip on identical clips like this?

Top comments (6)

Collapse
 
build996 profile image
build996 •

The Algieba swing is what I'd hold the prompt-fix result against. If the same 11 clips went from 73% to 45% male between passes at temperature 0, then 26% versus 22% across separately judged runs sits inside the judge's own noise, and some of Algieba's drop to 11% could be the judge too. Judging the with-line and without-line recordings interleaved in one pass, three times over, would separate the instruction's effect from the judge's day. To your closing question: that's the reason I'd want three passes minimum, with every condition being compared present in each pass rather than spread across them.

Collapse
 
achiya-automation profile image
אחיה כהן •

Fair point, and it lands on the weakest claim in the piece. When the same 11 Algieba clips move from 73% to 45% male between passes on identical audio, a 4-point gap (26% vs 22%) can't carry "made it worse". What that comparison supports is narrower: the line didn't help. Algieba's drop to 11% deserves the same caution.

The headline holds up better because it rests on direction, not magnitude: in both full passes, at least 12 of the 13 voices came out at or below 50% male.

Your design is the one I'd use for any rerun: both conditions and the known-answer controls interleaved in one pass, three passes, then compare. The judge's drift then hits both arms equally instead of hiding inside the difference between them.

Collapse
 
build996 profile image
build996 •

Direction over magnitude is the right thing to lean on, and 12 of 13 surviving both passes is hard to wave away. One detail I'd add to the rerun: treat the known-answer controls as a gate, not a correction. If a control that should read male comes back below some floor in a pass, that whole pass is excluded instead of averaged in, because a judge that misreads the controls that day is misreading everything else too. It also gives you a number to report: how many passes were thrown out, which says more about the judge than any single accuracy figure.

Thread Thread
 
achiya-automation profile image
אחיה כהן •

Agreed, and it closes a gap in what I described. Averaging a bad pass in means the controls tell you the judge was off that day, and the analysis ignores them anyway.

A gate needs its floor fixed before the rerun, though. I'd set it from how the controls varied across the two passes I already have, and write it down before looking at any voice results. Otherwise it's tempting to nudge the floor until the passes you like survive.

The count of excluded passes belongs in the write-up too. If two of five passes get thrown out, that tells a reader more about how far to trust a single number from this judge than its accuracy on the controls does.

Thread Thread
 
build996 profile image
build996 •

Fixing the floor from the two passes you already have, and writing it down before any voice results, is what makes the gate worth anything. I'd add one more line to that pre-registration: what happens if too few passes survive, say fewer than three of five. Without it, "rerun until enough pass" quietly brings back the tuning you're trying to rule out. And putting the excluded count next to the headline turns the judge's reliability into a number a reader can see, instead of a caveat buried in the method.

Thread Thread
 
achiya-automation profile image
אחיה כהן •

Agreed, that line belongs in it. Five passes, and if fewer than three clear the floor, the rerun reports no result. No sixth pass to make up the count.

That leaves two outcomes, both fixed before any audio gets judged: the direction result with the excluded count next to it, or "this judge wasn't stable enough to say". I'd publish the second one too. It's a finding about grading one model's voice with another model, and if it happens, probably the more useful of the two.