DEV Community

orca forge
orca forge

Posted on Edited on Originally published at forge.workstyle.tech

Why "shōshō" Became "shomo": A Whitelist Was Chipping Away at Japanese Characters

📝 Originally published (in Japanese) at forge.workstyle.tech.

I received this report regarding the avatar for the customer support window.

"Shoushou omachi kudasai" (Please wait a moment) is being pronounced as "Shomo omachi kudasai" (Wait for me).

Since I had just re-baked the speech model multiple times, I first suspected the model itself. To cut it short, the model, parameters, and cache were all functioning normally; the input text being passed to the TTS was corrupted.

Debugging from downstream to upstream

I'll list the steps in the order I suspected the cause. This order itself is a reflection of where I could have improved.

The Model. I traced the database to see which model the support site was pulling. It was correctly fetching the latest trained model.

The Cache. There was a TTS cache table, so I checked if it was returning old audio. The contents were only entries from a different provider from over a month ago, so that was unrelated.

The Synthesis Parameters. The runtime was using style_weight=2.0 and sdp_ratio=0.8. Since my verification used 1.0 / 0.4, I thought this might be the cause. I ran some tests.

style_weight 1.0 / 2.0 / 3.0 / 4.0   → All ratio 1.00
sdp_ratio    0.2 / 0.4 / 0.6 / 0.8   → All ratio 1.00
12 styles × weight 2.0             → All ratio 1.00
Enter fullscreen mode Exit fullscreen mode

Emotional Styles. I synthesized "Shoushou omachi kudasai" (Please wait) with all 12 styles, and all were normal.

The Path. Since I had moved the audio pipeline to a separate service, I checked if that service was synthesizing it independently. Synthesis was happening on the backend, and the path was the same.

It took several hours by this point. Everything was clean.

Printed the output of the preprocessing once

The only thing left was the input text.

>>> _clean_tts_text('少々お待ちください。')
'少お待ちください。'
Enter fullscreen mode Exit fullscreen mode

The character "々" (iteration mark) was missing. When I synthesized the sentence with the character removed, it was pronounced like this:

Input '少々お待ちください'  → Heard '少々お待ちください'      (Normal)
Input '少お待ちください'    → Heard 'Shoo omachi kudasai'   ← This one
Enter fullscreen mode Exit fullscreen mode

"Shoo wo" (the show) was being heard as "Shomo". I left the check that should have taken 5 minutes for the very end.

Cause: The character "々" is not in the Kanji range

The preprocessing had a whitelist to strip out emoticons, emojis, and special symbols.

# TTS preprocessing: remove unnecessary symbols and emoticons for speech
_TTS_ALLOWED_RE = re.compile(
    r"[^぀-ゟ"   # Hiragana
    r"゠-ヿ"     # Katakana
    r"一-鿿"     # Kanji
    r"ヲ-゚"     # Half-width katakana
    r"a-zA-Za-Z"
    r"0-90-9"
    r"、!?,.ー"
    r"\s"
    r"]"
Enter fullscreen mode Exit fullscreen mode

The intent was clear, and the implementation was straightforward. The problem is that "々" is not included in the 一-鿿 (CJK Unified Ideographs) range. "々" belongs to the symbols/punctuation block, not kanji. To someone writing Japanese, "々" is a kanji, but Unicode classification doesn't reflect that.

Found 7 types of characters outside CJK Unified

Since I found one, there must be others. I did an exhaustive search for characters outside CJK Unified that are actually used in Japanese.

Character Example Result Damage
々 少々日々 少・日 "Shomo"
〆 〆切 切 "Kiri"
〇 〇月日 月日 Dates disappear
𠮷 (Ext A) 𠮷野 野家 Proper names break
髙﨑 (Compat) 﨑山 山 Proper names break
〜 10〜20分 1020分 Numbers merge
: 5:30 330 Time formats break

Even keeping the symbols didn't fix it. "10〜20分" becoming "1020分" is quite bad. Meanwhile, % also drops, but that's not needed for speech.

Keeping symbols didn't fix it

I added 〜 and : to the whitelist. It made things worse.

'午後30に開始します' → Heard 'Gogo ni...' (Incorrect numbers)독'午後:30に開始します' → Heard 'Gomo,30に開始します' ← Worse if kept
Enter fullscreen mode Exit fullscreen mode

TTS cannot interpret : as a time, producing noise like "Gomo". 〜 is also treated as punctuation rather than "from". Both were incorrect.

Converting to Japanese

Found by chance: The pronunciation dictionary wasn't working

While tracing the cause, I discovered that not a single pronunciation dictionary was applied to this support site. Dictionaries are scoped by project, and all 44 entries were linked to different projects.

I let the voice read them.

Notation Correct Reading Actual Reading
液冷 Eirei Digeki
主な Omono Owner
従量課金 Juryou kakin Juryou kachi
行っています Okatte imasu Itte imasu

Having "Okatte imasu" (is doing) become "Itte imasu" (is saying) is bad because it occurs frequently in service text and reverses the meaning.

Of the 44 entries, excluding 10 product names and personal names, 34 were general terms needed on any site (technical terms and Japanese with split Onyomi/Kunomi readings). I deployed them to three other projects.

What I learned

Look at the input first. I spent hours crushing models, parameters, caches, and paths, only to finish by printing the preprocessing output once.

⚠️ The wave dash has yet another pitfall. 〜 (U+301C) and ~ (U+FF5E) are different characters, and the one used varies by environment. If you only write one, the other leaks out. I included both.

Top comments (0)