📝 Originally published (in Japanese) at forge.workstyle.tech.
I received this report regarding the avatar for the customer support window.
"Shoushou omachi kudasai" (Please wait a moment) is being pronounced as "Shomo omachi kudasai" (Wait for me).
Since I had just re-baked the speech model multiple times, I first suspected the model itself. To cut it short, the model, parameters, and cache were all functioning normally; the input text being passed to the TTS was corrupted.
Debugging from downstream to upstream
I'll list the steps in the order I suspected the cause. This order itself is a reflection of where I could have improved.
The Model. I traced the database to see which model the support site was pulling. It was correctly fetching the latest trained model.
The Cache. There was a TTS cache table, so I checked if it was returning old audio. The contents were only entries from a different provider from over a month ago, so that was unrelated.
The Synthesis Parameters. The runtime was using style_weight=2.0 and sdp_ratio=0.8. Since my verification used 1.0 / 0.4, I thought this might be the cause. I ran some tests.
style_weight 1.0 / 2.0 / 3.0 / 4.0 → All ratio 1.00
sdp_ratio 0.2 / 0.4 / 0.6 / 0.8 → All ratio 1.00
12 styles × weight 2.0 → All ratio 1.00
Emotional Styles. I synthesized "Shoushou omachi kudasai" (Please wait) with all 12 styles, and all were normal.
The Path. Since I had moved the audio pipeline to a separate service, I checked if that service was synthesizing it independently. Synthesis was happening on the backend, and the path was the same.
It took several hours by this point. Everything was clean.
Printed the output of the preprocessing once
The only thing left was the input text.
>>> _clean_tts_text('少々お待ちください。')
'少お待ちください。'
The character "々" (iteration mark) was missing. When I synthesized the sentence with the character removed, it was pronounced like this:
Input '少々お待ちください' → Heard '少々お待ちください' (Normal)
Input '少お待ちください' → Heard 'Shoo omachi kudasai' ← This one
"Shoo wo" (the show) was being heard as "Shomo". I left the check that should have taken 5 minutes for the very end.
Cause: The character "々" is not in the Kanji range
The preprocessing had a whitelist to strip out emoticons, emojis, and special symbols.
# TTS preprocessing: remove unnecessary symbols and emoticons for speech
_TTS_ALLOWED_RE = re.compile(
r"[^-ゟ" # Hiragana
r"゠-ヿ" # Katakana
r"一-鿿" # Kanji
r"ヲ-゚" # Half-width katakana
r"a-zA-Za-Z"
r"0-90-9"
r"、!?,.ー"
r"\s"
r"]"
The intent was clear, and the implementation was straightforward. The problem is that "々" is not included in the 一-鿿 (CJK Unified Ideographs) range. "々" belongs to the symbols/punctuation block, not kanji. To someone writing Japanese, "々" is a kanji, but Unicode classification doesn't reflect that.
Found 7 types of characters outside CJK Unified
Since I found one, there must be others. I did an exhaustive search for characters outside CJK Unified that are actually used in Japanese.
| Character | Example | Result | Damage |
|---|---|---|---|
| 々 | 少々日々 | 少・日 | "Shomo" |
| 〆 | 〆切 | 切 | "Kiri" |
| 〇 | 〇月日 | 月日 | Dates disappear |
| 𠮷 (Ext A) | 𠮷野 | 野家 | Proper names break |
| 髙﨑 (Compat) | 﨑山 | 山 | Proper names break |
| 〜 | 10〜20分 | 1020分 | Numbers merge |
| : | 5:30 | 330 | Time formats break |
Even keeping the symbols didn't fix it. "10〜20分" becoming "1020分" is quite bad. Meanwhile, % also drops, but that's not needed for speech.
Keeping symbols didn't fix it
I added 〜 and : to the whitelist. It made things worse.
'午後30に開始します' → Heard 'Gogo ni...' (Incorrect numbers)독'午後:30に開始します' → Heard 'Gomo,30に開始します' ← Worse if kept
TTS cannot interpret : as a time, producing noise like "Gomo". 〜 is also treated as punctuation rather than "from". Both were incorrect.
Converting to Japanese
Found by chance: The pronunciation dictionary wasn't working
While tracing the cause, I discovered that not a single pronunciation dictionary was applied to this support site. Dictionaries are scoped by project, and all 44 entries were linked to different projects.
I let the voice read them.
| Notation | Correct Reading | Actual Reading |
|---|---|---|
| 液冷 | Eirei | Digeki |
| 主な | Omono | Owner |
| 従量課金 | Juryou kakin | Juryou kachi |
| 行っています | Okatte imasu | Itte imasu |
Having "Okatte imasu" (is doing) become "Itte imasu" (is saying) is bad because it occurs frequently in service text and reverses the meaning.
Of the 44 entries, excluding 10 product names and personal names, 34 were general terms needed on any site (technical terms and Japanese with split Onyomi/Kunomi readings). I deployed them to three other projects.
What I learned
Look at the input first. I spent hours crushing models, parameters, caches, and paths, only to finish by printing the preprocessing output once.
⚠️ The wave dash has yet another pitfall. 〜 (U+301C) and ~ (U+FF5E) are different characters, and the one used varies by environment. If you only write one, the other leaks out. I included both.
Top comments (0)