DEV Community

Win Aung
Win Aung

Posted on

MyanmarChemCalc-Bench: Do LLMs Do Chemistry Better in English Than in Burmese?

Kaggle Benchmarking Challenge Submission

Benchmark: https://www.kaggle.com/benchmarks/winkoaung/myanmar-chem-calc-bench

What tasks did I run?

MyanmarChemCalc-Bench — 30 tasks built from 15 chemistry calculation questions drawn from Myanmar's Grades 9–12 chemistry textbook worked examples (ground-truth answers verified against the textbooks). Each question exists in two versions: English and Burmese — the same problem, the same numbers, the same expected answer.

Topics span the full calculation curriculum: mole concept, Avogadro's number, isotopes, percentage composition, combustion analysis, stoichiometry, gas volumes, ideal gas law, molarity, dilution, titration, pH, Faraday's laws, percentage yield, and equilibrium constants.

Scoring is strict and automatic: the model must show its work, then write FINAL ANSWER: <number> on its own line. A Python checker extracts the number and compares it against the ground truth within a per-question tolerance.

The creative hook: does a model score lower when the exact same chemistry question is asked in Burmese? Burmese is a low-resource language — if the multilingual gap is real, it should show up here.

Which models did I run it against?

Five models from Kaggle's suite:

Model Score
Claude Sonnet 4.5 100.00
DeepSeek-R1 93.33
Gemini 2.5 Pro 93.33
GPT-6 Astra 93.33
Qwen 3 235B A22B Instruct 46.67

What are the main insights?

1. The most interesting finding isn't chemistry — it's run reliability

Qwen's 46.67 looks like a chemistry failure. It isn't — not mostly. Of Qwen's runs that actually produced a scored answer, the vast majority passed. The missing cells were execution errors: 429 rate-limit errors from provider overload, not wrong answers. The leaderboard score divides passes by all 30 tasks, so errors masquerade as incompetence.

Lesson: on a small benchmark, run reliability can distort the leaderboard more than model capability. I report passes, fails, and errors separately — and so should you.

2. Paired translations make failures inspectable

The EN/MM pairing paid off in a concrete case: on the mole question ("How many moles are in 8 g of NaOH?"), Qwen passed in English but wrote 23 + 16 + 1 = 50 in the Burmese version, getting 0.16 mol instead of 0.2. Same numbers, same question — a genuine arithmetic slip in the Burmese run. That's a single case, not proof of a systematic language gap, but it's exactly the kind of failure this benchmark design surfaces.

3. Small benchmarks demand restraint

With 30 tasks, one task swings the score by 3.33 points. The gap between Claude (100.00) and the 93.33 group is just two tasks. Don't over-interpret the ordering — it needs completed evaluations and repeated runs to be meaningful.

4. Burmese didn't break the top models — and the numeral fix is validated

After updating the Burmese prompts to use Arabic numerals (8 instead of ၈), every completed run on the new Burmese tasks was a PASS — 100% across all five models. The numeral change didn't degrade anything; models handle Arabic numerals in Burmese prompts reliably. Claude reached a perfect 100.00 overall. For textbook-style calculations, the current frontier models handle Burmese about as well as English.

Where can we see it?

Benchmark (public): https://www.kaggle.com/benchmarks/winkoaung/myanmar-chem-calc-bench

All 30 tasks, the leaderboard, and per-model run traces are public. The task files are generated from a single Python script with ground truths recomputed against the Myanmar textbook worked examples.

What's next?

  • Fix the remaining error cells with bounded retries and re-run for a clean comparison
  • Expand beyond 15 questions — stoichiometry and multi-step problems are where I'd expect the real gaps
  • Test whether strict Burmese explanations (not just prompts) change anything

Built for the Kaggle Benchmarking Challenge. All questions verified against Myanmar basic-education chemistry textbooks (Grades 9–12).

Top comments (0)