A writing agent can produce a fluent paragraph and still choose the wrong spelling convention. A word can pass a dictionary check and still conflict with a style guide. US English makes a useful example: spelling and regional usage are separate checks.
I keep those decisions outside the model in language-mcp: a small TypeScript MCP server with local spelling and simplicity checks. It currently supports US English and Dutch, with a separate Dutch word lookup and a B1 simplicity check for each language. The examples below use US English and Dutch; the same design can be extended to other languages with suitable dictionaries, rules, and tools. The agent submits text, reads the findings, and decides which changes make sense. The server does not rewrite files.
The useful part of this project is the boundary between a dictionary, a style rule, and an editorial decision. They answer different questions, even when all three produce something that looks like a correction.
A local dictionary first, a network lookup when needed
Six tools use the Hunspell executable with bundled dictionaries. Dutch uses OpenTaal; US English uses the SCOWL/LibreOffice dictionary assets. Spelling checks do not send the submitted text to a spelling API.
That privacy statement applies to this server, not to the whole agent session. If your MCP client uses a hosted language model, its conversation and tool results may still reach that provider. I explain this distinction in more detail in our guide to what stays local in a hybrid AI setup.
The one network tool, get_dutch_word_details, queries Woordenlijst.org, maintained by the Instituut voor de Nederlandse Taal for the Taalunie. It provides lexical details such as pronunciation, word forms, and hyphenation. This is a separate network operation, useful for an individual word that needs investigation.
The Dutch dictionary source is OpenTaal's Hunspell repository. Keep the dictionary attribution and license files when redistributing the assets; a server's code license does not replace them.
Why I kept the Hunspell executable
The project's architecture decision record considers replacing the executable with nspell, a JavaScript implementation. Removing a native dependency would simplify deployment. I kept the executable because changing the implementation also means proving behavioral compatibility.
A particularly easy mistake sits in Hunspell's -a pipe protocol. It does not promise one response line for every submitted token. A hyphenated input can produce a response for each part, followed by an empty separator line.
If the wrapper zips individual response lines with input words, one compound can shift the remaining results onto the wrong words. The current implementation groups output into blocks separated by empty lines, then associates each block with one unique input token. It rejects a response-count mismatch instead of returning misaligned findings.
Within a block, the wrapper accepts * and + responses, collects suggestions from & responses, and treats unknown responses as failures. It removes duplicate suggestions and returns at most eight per token. It also limits the subprocess to 20 seconds.
That explains the dependency better than a speed claim. The ADR discusses performance, but I am not presenting a Hunspell-versus-nspell benchmark here. A replacement should pass the same compound and suggestion tests before deployment gets simpler at the expense of correctness.
Install the pinned version
You need Git, Node.js 20 or newer, npm, and hunspell available on the MCP host's PATH. The Bash commands below target Linux, macOS, or WSL. Native Windows setup is not covered here; with WSL, install the dependencies and run the client inside that environment.
Install Hunspell through your operating system's package manager first. The repository README gives pacman -S hunspell for Arch Linux. Package installation may require administrator privileges. Check the actual executable before building:
node --version
npm --version
hunspell --version
git clone https://github.com/seppegadeyne/language-mcp.git
cd language-mcp
git checkout --detach e769c5d960c92a4455f1be0607cc68b873dffab2
npm ci --ignore-scripts
npm test
npm run build
The checkout pins the source discussed here rather than silently following future changes. Run the locked dependency install, tests, and build on your MCP host before configuring your client. A successful run on one machine is not proof that every Node version or operating system has been exercised.
In an existing Hermes installation, merge this entry into mcp_servers in your configuration. Replace the example path with your checkout's absolute path; do not replace unrelated server entries.
mcp_servers:
language:
command: node
args: ["/absolute/path/to/language-mcp/dist/cli.js"]
connect_timeout: 60
enabled: true
timeout: 120
Then test discovery:
hermes mcp test language
The server speaks MCP over stdio, so node dist/cli.js is not an interactive spelling prompt. Your client launches it and exchanges protocol messages. Other stdio-capable MCP clients need their own configuration syntax, and their tool names may include a server prefix.
Try the US English tools
These JSON objects are tool arguments, not shell commands. Select the named tool through your MCP client or ask your agent to call it.
For a paragraph, use check_us_english_text:
{"text":"While moving toward the center, we organized everything."}
The verified response is:
US English check (hunspell en_US + britticism scan)
8 words checked, 0 unknown, 0 British forms.
OK: no typos, no British forms found.
For one word, use validate_us_english_word:
{"word":"color"}
The verified response is "color" is correct US English.
To see the regional style check in action, replace While with the intentionally British Whilst, and toward with towards. These are test inputs, not recommended US English. The dictionary accepts them, but the separate style rules flag both and suggest the US forms. A dictionary check alone would miss that distinction.
The server also provides check_dutch_text, validate_dutch_word, and get_dutch_word_details. The first two use the local Dutch dictionary; the last uses the network for lexical details. You do not need the network lookup for the US English examples. Two simplicity checks, check_dutch_b1_text and check_us_english_b1_text, are covered next.
Checking Dutch text for B1-level simplicity
Version 0.4.0 adds check_dutch_b1_text, for a question the spelling tools cannot answer: is this Dutch text simple enough for a broad audience?
"Dutch at B1" is an editorial convention, not a validated text measurement. The Rijksoverheid communication guideline asks writers to aim for B1 (CEFR) in new public-facing material. A formula cannot certify that a text meets it, so the tool reports proxies a reviewer can check in context:
- Readability: Flesch-Douma reading ease (the Dutch-adapted Flesch formula) and ARI, with a practical B1 band around 60-70 for Flesch-Douma. Below 50 words or three sentences, the formulas are suppressed instead of reporting noise.
- Sentence and paragraph length against configurable thresholds (defaults: warn above 15 words per sentence, flag above 20; 150 words per paragraph).
- Passive voice, using a second Hunspell pass described below.
- Officialese jargon with plain alternatives (
thanstonu,betreffendetoover), filler words, and idioms that are hard for NT2 readers, each with its position. - Nominalization density (
-ing/-tie/-heid/-iteitwords per 100 words) and je/u voice consistency.
The passive detector is a second Hunspell protocol story. The -a pipe from the spelling tools returns no morphology, but hunspell -m returns st: stem and ts: tag fields from the OpenTaal affix file. Three quirks matter:
- Only affix-derived forms get the past-participle tag
ts:VBpe(gemaakte); bare participles such asverzondenare their own lemma without a tag. - Separable compounds analyze as compound stems:
uitgevoerdbecomesuit st:gevoerd. - Tags are context-free:
isplus a participle can be a perfect tense or a passive, so those hits are labeled low confidence.
The wrapper combines the tag, a bare-lemma check with Dutch participle phonology, and a compound-stem rule, then scans for auxiliaries (wordt, worden, werd, is, zijn) followed by a participle. Every hit is a review flag with its sentence and position, not a verdict.
An intentionally bureaucratic example:
{"text":"Wellicht wordt de aanvraag thans door de gemeente in behandeling genomen. De implementatie van de regeling betreffende de vergoeding wordt vervolgens door de afdeling uitgevoerd, wat betekent dat u als klant langer moet wachten op een beslissing dan u wellicht zou verwachten op basis van de eerder door ons gedane toezeggingen. Je krijgt daarna een bericht. Onder de loep genomen?"}
The verified response, abbreviated:
Dutch B1 simplicity check (B1 proxies, not a validated B1 verdict)
Readability:
- Flesch-Douma: 54 (band: below-b1; B1 target 60-70)
- ARI: 11.4 (grade-level indication)
- Words: 58, sentences: 4, avg 14.5 words/sentence
- Sentences over 15 words: 1 (1 over 20)
- Long words (>=4 syllables): 7%
Passive voice (review flags, not verdicts):
- "wordt genomen" (+ door/van agent) — "Wellicht wordt de aanvraag thans door de gemeente..." (position 0)
- "wordt uitgevoerd" (+ door/van agent) — "De implementatie van de regeling betreffende..." (position 73)
Jargon and vague wording (plain alternatives):
- "thans" → nu (position 27)
- "implementatie" → uitvoering, uitrollen (position 77)
- "betreffende" → over, bij (position 107)
Filler words:
- "Wellicht" → misschien (position 0)
Idioms (consider plain phrasing for NT2 readers):
- "Onder de loep genomen" → goed bekijken (position 359)
Nominalization density: 10.3 per 100 words (high — prefer verbs over nouns)
Voice consistency: MIXED (expected je): u-forms 2x at position 196, je-forms 1x at position 329 — pick one address form.
The same check for US English
Version 0.5.0 adds check_us_english_b1_text. It asks the same question about US English copy. The output has the same structure: readability, passive voice, rule findings with plain alternatives and positions, nominalization density, and reader address. "B1" is still shorthand for plain language, not a certified level.
The English version needs its own measurements, not translated Dutch rules:
- Readability: Flesch Reading Ease (the original English formula, with 60-70 as the practical plain-English band), Flesch-Kincaid grade level (default target 9 or lower), and ARI. A zero-dependency syllable counter feeds the formulas. It agrees with the CMU Pronouncing Dictionary on about 94% of the 10,000 most common US English words. That is accurate enough for averages over a text, not for individual words.
- Sentence length: warn above 20 words and flag above 25, following plainlanguage.gov and the GOV.UK style guide. Both thresholds are tool parameters.
- Rules from the Federal Plain Language Guidelines and the GOV.UK "words to avoid" list. They cover formal jargon (
utilizetouse,prior totobefore), wordy phrases (in order tototo), hidden verbs (make a decisiontodecide), filler words, and idioms. British spellings stay incheck_us_english_text, so nothing is reported twice. - Reader address: the tool lists references to "users" or "the customer" when the text could address the reader as "you".
Passive detection was the second Hunspell surprise. The bundled en_US affix file has no ts: tags at all. hunspell -m marks regular -ed forms with the affix flag fl:D (created st:create fl:D), while irregular participles such as written or built come back as bare lemmas. The detector therefore combines the flag with a closed list of irregular participles. It also has to skip hyphenated words, because the en_US dictionary splits data-driven into one morphology block per part. Participles that usually describe a state after "be" (is located, are required) are labeled low confidence.
An intentionally bureaucratic example:
{"text":"In order to facilitate the processing of your request, it is essential that applicants utilize the online portal prior to the deadline. Requests that are submitted after the deadline will not be considered by the committee. Users are required to make a payment before a decision can be reached. Due to the fact that the system is undergoing maintenance, additional delays may be experienced. You will get an email when we are done."}
The verified response:
US English B1 plain-language check (B1 proxies, not a validated B1 verdict)
Readability:
- Flesch Reading Ease: 51.8 (band: below-b1; B1 target 60-70)
- Flesch-Kincaid grade: 9.6 (above target 9)
- ARI: 8.9 (grade-level indication)
- Words: 71, sentences: 5, avg 14.2 words/sentence
- Sentences over 20 words: 1 (0 over 25)
- Complex words (>=3 syllables): 20%
Passive voice (review flags, not verdicts):
- "are submitted" — "Requests that are submitted after the deadline will not be considered by the committee." (position 135)
- "be considered" (+ by agent) — "Requests that are submitted after the deadline will not be considered by the committee." (position 135)
- "are required" [adjectival use likely, low confidence] — "Users are required to make a payment before a decision can be reached." (position 223)
- "be reached" — "Users are required to make a payment before a decision can be reached." (position 223)
- "be experienced" [adjectival use likely, low confidence] — "Due to the fact that the system is undergoing maintenance, additional delays may be exper…" (position 294)
- "are done" [adjectival use likely, low confidence] — "You will get an email when we are done." (position 391)
Jargon and buzzwords (plain alternatives):
- "facilitate" → help, make possible (position 12)
- "utilize" → use (position 87)
- "prior to" → before (position 113)
- "required" → need, must (position 234) — common in web copy; judge in context
- "additional" → more, extra (position 354)
Wordy phrases:
- "In order to" → to (position 0)
- "it is essential that" → (state the point directly) (position 55)
- "Due to the fact that" → because (position 295)
Hidden verbs (use the verb, not the noun):
- "make a payment" → pay (position 246)
Nominalization density: 2.8 per 100 words
Top: decision (1x), maintenance (1x)
Reader address: third-person references to the reader 2x (first at position 76): applicants (1x), users (1x) — address the reader as "you" where they are the audience; "you" used 2x.
The last hit shows why these remain review flags: "when we are done" is an adjective, not a passive, and the tool labels it low confidence instead of hiding it.
Extend the pattern to another language
MCP does not tie this design to English or Dutch. The current server implements those two languages; adding another one requires code and dictionary assets, not just a language name in a request.
Start with a compatible Hunspell dictionary from a trustworthy source. Include its dictionary and affix files, preserve its license and attribution, and add it to the dictionary resolver. Then register language-specific text-checking and word-validation tools with the same result structure.
Add regional or editorial rules only when the dictionary cannot express the distinction you need. A lexical lookup is optional and requires a suitable source for that language; the Dutch lookup is not a universal dictionary API. Finally, test valid words, misspellings, compounds, suggestions, and any language-specific rules before exposing the new tools to an agent.
The reusable idea is the separation of responsibilities: the dictionary checks spelling, explicit rules check selected conventions, and the agent reviews the findings in context. Language coverage depends on the available dictionaries and the integration work, not on MCP itself.
Treat the style rules as review suggestions
The britticism detector is a static list of regular expressions and suggested replacements. Some rules cover spelling variants. Others encode editorial preferences, such as using custom rather than bespoke.
I would not automatically apply every match. The list also includes context-sensitive vocabulary: queue is perfectly ordinary in an American programming article, and lift does not always mean an elevator. Rules can flag valid text or generate an unsuitable inflection. This implementation is not a grammar parser.
There is no measured coverage percentage behind the rule list. It is inspectable and testable, which makes it useful for enforcing a specific style contract, but it cannot certify that a paragraph is idiomatic US English.
The two checking paths also treat Markdown differently. The spelling tokenizer removes URLs, email addresses, fenced code, and inline code. The britticism scan receives the original text. A quoted example or identifier can therefore appear in style findings even when spelling ignores it. Preserve source quotations and code; review the surrounding prose instead.
The limits belong in the workflow
The tokenizer stops after 2,000 checkable tokens, skips one-character tokens, and skips short all-uppercase acronyms. For a longer document, split it into sections and keep track of what you submitted. A clean result is not evidence that every character in the original document was checked.
Reported positions are JavaScript string indexes, not editor-ready line and column locations. Since version 0.5.0, each spelling finding reports the offset of the token itself. Earlier versions searched for the first occurrence of the word, which could point inside a longer word (red inside hundred). Treat the positions as navigation aids: excluded material and later edits still make automatic patching risky.
The network lookup has its own boundaries. It uses an unofficial XML service, serializes calls with a minimum 1.2-second interval, and caches results in memory for 24 hours. This is not a bulk dictionary download API or an availability guarantee. Its parser also takes the first matching descendant fields; in the verified pizza lookup, the top-level hyphenation reflected a diminutive. Check ambiguous lexical details on the source website rather than treating every rendered field as authoritative.
My preferred workflow is simple: submit a section, inspect the findings, make deliberate edits, and check the edited section again. Keep proper nouns and technical identifiers unless there is a specific reason to change them. Read the whole paragraph afterward. A dictionary cannot tell whether a sentence says what you meant.
The Glama listing is another way to discover the server; the pinned repository is the reference for the tool names and behavior described here. A directory listing does not remove the Hunspell runtime requirement.
Top comments (0)