We run Lingvoberry, a reader where you tap a word in a graded book to save it. To place a new user on a CEFR level we built a test: show a page of a real book, ask about four words on it, and climb up or down a band depending on the answers.
The first version read one or two bands too high. Fixing it taught us that a word's frequency rank is a poor stand-in for how hard the word is. Here is what we measured, on 252 production books.
The first version: grade a word by its rank
We have a frequency list for each of 14 languages and cut it into bands, from A0 up to C1 by rank. The rarer a word, the higher the band, and a page of a C1 book was asked about its rarest words.
It looked reasonable. It was wrong in three different ways.
1. The glossary put easy words on hard pages
Every chapter of our books ends with a small vocabulary box, and we took test words from it. Measured, that box put as many top-900 words on a C1 page as on an A1 one: 29% of every word asked, and 60% of runs asked at least one. So the questions did not get harder as the walk climbed.
The fix: grade each word by its own frequency band and draw the questions out of the page's own prose. A page's prose turned out to be deep enough that the glossary is no longer the admission ticket: a C1 page carries a median of 8 distinct C1-band words.
2. A third of the "hard" words were not hard
We then looked at what a C1 probe could ask. About a third of the candidates were not C1 questions at all, and they fell into a few classes.
- Forms of an easier word (20%). Our lists count printed forms, so "blau" is A2 and "blauer" ended up in C1. The fix is to band a word by the best rank of anything it is a form of: the longest prefix that is itself a list word, with a language-specific limit on the tail. German, Dutch and Swedish glue words together, so a long tail there is a compound ("Küstenstadt" is not a form of "Küste"). Turkish stacks four endings, the Slavic languages three.
- International words (16%). "Konstruktion" is rare in German text and obvious to anyone who reads a European language. Frequency cannot see that. We compared the fourteen lists against each other: a word whose folded stem also stands in three lists outside its own language family is not a question from B1 up. The family rule matters: without it "город" matched three Slavic lists and was thrown out of the Russian ladder.
-
Names and fragments. A capitalised word that stands nowhere in lowercase is often a name, and in German every noun is capitalised, so there we ask whether the word appears several times on one page. A tokeniser also leaves
arenandisnbehind at an apostrophe, andarenhad come out of one list as a C2 word.
3. The level was the maximum, and "I know it" was free
The old level was the maximum over the pages visited, and only the first three claims of a run were ever verified. A reader could press "I know it" on everything, fail every check and still be placed at B2 in 83% of runs, and at C1 in the rest.
Now the level is the highest band that passed with no failure under it, and nothing is reported above the highest band where a claim was actually checked. On the same adversarial runs, none reach B1.
What I would do again
- Measure on your own data before tuning. The 29% and the 83% were found by counting, not by looking at a few pages.
- Treat a threshold as a calibration constant. We now write every run down (bands walked, words known and asked, checks passed) so the next calibration starts from our own numbers.
- Keep the engine in one place. The same walk runs in the Android app, on the web and on the logged-out page, pinned by one shared test fixture.
If you grade text for learners, or anything else where "rare" and "hard" get conflated, I would like to hear what you use instead of rank.
Top comments (0)