DEV Community

Daniel Pertu
Daniel Pertu

Posted on

Our email filter keeps two kinds of address, and a Wix subdomain is why it almost kept everything

When a Nakodo business campaign finds a cafe in a places listing, the listing usually gives us a name, an address, a category and a website. What it very often does not give us is an email, and the whole product is email. So we read the business's own website and look for one.

The naive version of this is a regex over the HTML. Run that against a few hundred real pub and salon websites and look at what you get back. A representative haul from one page:

  • hello@thecrown.co.uk, which is the right answer
  • info@webdesignleeds.co.uk, in the footer, because the agency that built the site put it there
  • support@opentable.com, from the booking widget
  • 605a7baf2e4f4f0e9b6a1b6b3c1f6d2e@o12345.ingest.sentry.io, which is an error tracker DSN that happens to be shaped like an email
  • your-email@gmail.com, left in a contact form placeholder
  • noreply@thecrown.co.uk, which is real, and useless
  • jobs@thecrown.co.uk, which is real, and a human, and the wrong human

Emailing any of the last six is somewhere between embarrassing and a spam complaint. So the rule is strict, and it is stated in two allowances rather than a pile of exclusions.

Keep an address at the site's own name, with any ending

hello@thecrown.pub is a good address for the business at thecrown.co.uk. Small businesses buy a short domain for email and a different one for the website, or move from one to the other and keep both. So the comparison is on the registrable name without its suffix, using tldts:

const nameOf = (host: string) => getDomainWithoutSuffix(host.toLowerCase())?.toLowerCase() ?? null;
Enter fullscreen mode Exit fullscreen mode
assert.equal(siteName("https://www.thecrown.co.uk/menu"), "thecrown");
assert.equal(siteName("http://shop.thecrown.pub"), "thecrown");
Enter fullscreen mode Exit fullscreen mode

Keep personal mail, because that is what half of them use

The second allowance is the one that a B2B-flavoured filter would never include. A huge number of independent businesses run on gmail.com, hotmail.co.uk, or an internet provider's mailbox from fifteen years ago:

export const isPersonalMail = (email: string) => isFreeMail(email) || ISP_MAIL.has(email.split("@")[1] ?? "");
Enter fullscreen mode Exit fullscreen mode

isFreeMail covers the global providers. ISP_MAIL is a hand-kept set of the broadband and telco mailboxes that actually turn up in the wild, with a strong UK and European bias because that is where the data is: btinternet.com, talktalk.net, virginmedia.com, sky.com, wanadoo.fr, bigpond.com, and so on. It is a boring list and it is the difference between finding an owner and finding nobody.

So: sue.landlady@btinternet.com is kept. thecrownleeds@gmail.com is kept.

And now the part that nearly broke it

Here is keepBusinessEmail, which is twelve lines and has one genuinely dangerous branch:

export function keepBusinessEmail(email: string, names: string[]): boolean {
  const [local, domain] = email.toLowerCase().split("@");
  if (!local || !domain) return false;
  if (PLACEHOLDER_LOCAL.test(local) || SYSTEM_LOCAL.test(local) || WRONG_DESK.test(local)) return false;
  // Machine-made local parts: error-tracker keys, hashes.
  if (/^[0-9a-f]{16,}$/.test(local) || /\d{6,}/.test(local)) return false;
  if (isPersonalMail(email)) return true;
  const name = nameOf(domain);
  return name !== null && names.includes(name);
}
Enter fullscreen mode Exit fullscreen mode

The dangerous branch is the last one, and the danger is site builders.

A huge number of these businesses are on a hosted subdomain: thecrown.wixsite.com, thecrown.square.site, thecrown.myshopify.com. Run getDomainWithoutSuffix on thecrown.wixsite.com and you get wixsite, not thecrown. So the site's "own name" would be wixsite, and the filter would then happily accept any address at wixsite.com, which means the platform's own support and marketing addresses sail straight through on every single customer site it has ever built. Multiply by the number of Wix sites in a places listing and you have a filter that has learned to email Wix.

The fix is a deny set consulted by siteName, which returns null rather than a wrong name:

const HOSTED = new Set([
  "wixsite", "wix", "squarespace", "square", "weebly", "wordpress", "godaddysites", "webflow",
  "shopify", "myshopify", "github", "netlify", "vercel", "framer", /* ... */
]);

export function siteName(url: string): string | null {
  const name = nameOf(new URL(url).hostname);
  return name && name.length >= 2 && !HOSTED.has(name) ? name : null;
}
Enter fullscreen mode Exit fullscreen mode

null is not an error state. It means "this site has no name of its own", and the consequence falls out of keepBusinessEmail for free: with an empty names array the own-name branch can never match, so a hosted site keeps personal mail only. The test says exactly that:

// A site we couldn't name keeps personal mail only.
assert.ok(!keepBusinessEmail("hello@thecrown.co.uk", []));
assert.ok(keepBusinessEmail("thecrown@gmail.com", []));
Enter fullscreen mode Exit fullscreen mode

That is a good shape for this kind of rule. The uncertainty is represented once, in the place that knows about it, and every consumer gets the conservative behaviour without having to know why.

Three local-part refusals that are three different ideas

They look like one list and they are not.

PLACEHOLDER_LOCAL is example, your-email, yourname, johndoe, test. These are not addresses, they are template text somebody forgot to replace. Note that example@gmail.com is rejected even though it passes the personal-mail allowance, which is why the local-part checks run before the allowances rather than after.

SYSTEM_LOCAL is no-reply, mailer-daemon, postmaster, abuse, privacy, dpo, webmaster, wpforms. Real mailboxes, no reader, or a reader who will be annoyed in a legally interesting way.

WRONG_DESK is jobs, careers, hr, invoices, accounts-payable, billing, complaints, refunds. These are real mailboxes with real humans who read them, and they are still refusals, because the message we are sending is a partnership enquiry and none of those desks handles one. This is a product rule wearing a regex costume, and it is the one I would most expect a different team to set differently.

Then the machine-made check: /^[0-9a-f]{16,}$/ for hex blobs, and /\d{6,}/ for anything with six consecutive digits. That is what kills the Sentry DSN above, along with the various analytics and CDN identifiers that are email-shaped by accident.

Fail open on DNS, fail closed on everything else

Before an address is used, the domain is checked for mail servers. The interesting part is what counts as a failure:

check = Promise.race([
  dns.resolveMx(domain).then(
    (r) => r.length > 0,
    (e: NodeJS.ErrnoException) => !["ENOTFOUND", "ENODATA", "ESERVFAIL"].includes(e.code ?? ""),
  ),
  new Promise<boolean>((r) => setTimeout(() => r(true), 5000)),
]);
Enter fullscreen mode Exit fullscreen mode

Only a definite "no such domain" or "no mail servers here" counts against an address. Any other error, and any lookup slower than five seconds, resolves to true. A resolver having a bad afternoon must not quietly delete leads, and it would, because the symptom of a strict DNS check under load is simply fewer businesses with emails, which looks exactly like a thin area.

That is the opposite of the posture everywhere else in this file. Content on a page is untrusted and refused by default. Infrastructure that we depend on is given the benefit of the doubt. Both are "do the safe thing", and the safe thing points in opposite directions.

Rank the survivors, keep three

An address that passes is then ranked by who is likely behind it:

const sorted = ["events@x.com", "bookings@x.com", "info@x.com", "manager@x.com"]
  .sort((a, b) => emailRank(a) - emailRank(b));
assert.deepEqual(sorted, ["manager@x.com", "info@x.com", "bookings@x.com", "events@x.com"]);
Enter fullscreen mode Exit fullscreen mode

Owner, landlord, licensee, director and manager rank first. The general inbox (info, hello, enquiries, reception) ranks second. Bookings and sales third, events and functions fourth. An address the ranker does not recognise, such as sarah@thecrown.co.uk, gets the same rank as the general inbox, on the reasoning that a first name at a small business is usually the person who owns it. Three are kept, best first.

Getting the address off the page at all

Addresses on these sites are frequently not in the HTML as text. The extractor undoes, in one pass: HTML entities (bookings&#64;thepub.co.uk), @ and %40 escapes inside JSON-LD, Cloudflare's email protection (a hex blob in data-cfemail, XORed against its own first byte), and mailto: hrefs cut at the ? so a subject= parameter does not end up in the address.

const cf = [...decoded.matchAll(/data-cfemail="([0-9a-f]+)"/gi)].map((m) => decodeCfEmail(m[1]));
Enter fullscreen mode Exit fullscreen mode

And then one refusal that is pure slapstick: <img src="logo@2x.png"> matches any sane email regex. There is a file-extension check for exactly that, and the test asserts no result contains logo.

Reasons are data, not log lines

If nothing survives, the business is set aside with a reason, and the reason is a column rather than a string built at render time:

const NO_EMAIL_TEXT: Record<NoEmailReason, string> = {
  no_website: "No website or email listed",
  unreachable: "Website unreachable",
  blocked: "Website blocks visits",
  form_only: "Contact form only, no email",
  none: "No email on their website",
};
Enter fullscreen mode Exit fullscreen mode

form_only is the one that earns its place. A business whose site has a working contact form and no address is a completely different situation from one with nothing, both for the user reading the list and for us deciding whether re-reading the site next month is worth a request. The crawl detects a form by looking for a <form> with an email-ish input, and sets the reason only when nothing else was found.

The crawl itself is deliberately small: at most five pages, the homepage first, then internal links ranked by what their path looks like (contact and get-in-touch, then find-us and location, then about, then bookings), and /contact plus /contact-us tried blind if nothing relevant was linked. robots.txt is checked for every URL including the homepage, and a disallow produces blocked rather than a fetch.

Read the rule we publish, then check our work

The best part of writing this up is that the rule is not a secret. It is published, in plain English, in the business campaigns section of Nakodo's methods page: nakodo.app/how-it-works#businesses.

That page says we keep an address at the website's own name with any ending, gives hello@thecrown.pub for thecrown.co.uk as the example, says we keep personal mailboxes because many small businesses use them, says everything else is thrown away including a web designer's or a booking platform's address and placeholders and no-reply and jobs and accounts, says we read at most five pages starting with the home page, says we respect robots.txt, and says a business with no email left is set aside with the reason. Go and read it against the code above. If the two ever disagree, the page is the one that is wrong.

Top comments (0)