DEV Community

Onizuka
Onizuka

Posted on

I Verified 5,000 Emails. SMTP Said OK, 24% Still Bounced.

webdev, #security, #api, #python

The 24% that SMTP wouldn't explain

On October 1, 2026, I pointed a Python batch script at 5,000 addresses pulled from a real newsletter export. Two hours later the SMTP stage had whispered 250 OK on 4,892 of them. I felt good. Then I sent the actual campaign. 1,189 of those "verified" addresses bounced anyway. That's 24.3%.

Not a deliverability problem. A truth problem.

I'd been running this as part of my Email Validator Experiments series, where I keep checking how much an API can really tell you about an address before you spend money on it. This time I wanted to isolate the SMTP stage. Everyone treats a 250 OK like a guarantee. My inbox metrics said otherwise.

Here's the control I ran first — the address everyone uses as a sanity check:

curl --request GET \
  --url 'https://email-validator112.p.rapidapi.com/validate?email=test@gmail.com' \
  --header 'X-RapidAPI-Key: YOUR_KEY' \
  --header 'X-RapidAPI-Host: email-validator112.p.rapidapi.com'
Enter fullscreen mode Exit fullscreen mode

And this is what came back:

{
  "email": "test@gmail.com",
  "valid": true,
  "stage": "mx",
  "syntax_valid": true,
  "mx_found": true,
  "smtp_verified": null,
  "is_disposable": false,
  "is_catch_all": null,
  "is_role": true,
  "role_type": "test",
  "score": 75,
  "deliverability": {
    "score": 75,
    "factors": {
      "syntax_valid": true,
      "mx_found": true,
      "smtp_verified": null,
      "is_disposable": false,
      "is_catch_all": null,
      "is_greylisted": null,
      "breach_count": 0
    }
  },
  "suggestion": null,
  "is_free_email": true,
  "email_provider": "googleworkspace",
  "is_greylisted": null,
  "greylisting_note": null,
  "normalized_email": "test@gmail.com",
  "is_plus_addressed": false,
  "breach_status": null,
  "breach_status_error": "HIBP_API_KEY invalid or unauthorized",
  "is_trusted_identity": null,
  "invalid_explanation": null,
  "syntax": {
    "valid": true,
    "local": "test",
    "domain": "gmail.com"
  },
  "mx": {
    "has_mx": true,
    "records": [
      {"priority": 5, "exchange": "gmail-smtp-in.l.google.com"},
      {"priority": 10, "exchange": "alt1.gmail-smtp-in.l.google.com"},
      {"priority": 20, "exchange": "alt2.gmail-smtp-in.l.google.com"},
      {"priority": 30, "exchange": "alt3.gmail-smtp-in.l.google.com"},
      {"priority": 40, "exchange": "alt4.gmail-smtp-in.l.google.com"}
    ],
    "best": "gmail-smtp-in.l.google.com"
  },
  "smtp": null,
  "catch_all_probe": null,
  "_links": {
    "breach_data": "https://haveibeenpwned.com"
  },
  "identity_graph": {
    "email": "test@gmail.com",
    "gravatar": null,
    "breach_count": 0,
    "first_breach_date": null,
    "last_breach_date": null,
    "domain": "gmail.com",
    "fetched_at": "2026-10-01T15:20:25.876272+00:00"
  },
  "provenance": {
    "syntax": {"source": "internal", "confidence": 1.0},
    "mx": {"source": "DNS resolver", "confidence": 0.95},
    "smtp_verified": {"source": "SMTP probe", "confidence": 0.9},
    "breach_status": {"source": "Have I Been Pwned", "confidence": 0.95}
  },
  "fetched_at": "2026-10-01T15:20:25.876290+00:00"
}
Enter fullscreen mode Exit fullscreen mode

Look at that for a second. valid: true. score: 75. mx_found: true. But smtp_verified: null. And is_role: true with role_type: "test".

This is test@gmail.com. Of course it exists. Of course Gmail's MX records are there, priorities 5, 10, 20, 30, 40. The API tags the provider as googleworkspace. But the API is honest enough to say it never got a clean SMTP handshake, and even if it had, the local part is literally a role word. That 75 isn't a pass. It's a warning dressed up as a number.

The data: what 5,000 verifications actually looked like

I didn't run this on fake data. The 5,000 addresses came from a newsletter signup flow that had been live for about 18 months. Some were B2B, some were free Gmail/Hotmail/Proton, some looked like support@company.com. I ran them through the same endpoint and logged every response to a JSONL file. The endpoint I used is here: 👉 RapidAPI listing.

The script is embarrassingly simple:

import requests, json, time

url = "https://email-validator112.p.rapidapi.com/validate"
headers = {
    "X-RapidAPI-Key": "YOUR_KEY",
    "X-RapidAPI-Host": "email-validator112.p.rapidapi.com"
}

with open("emails.txt") as f, open("results.jsonl", "w") as out:
    for line in f:
        email = line.strip()
        if not email:
            continue
        r = requests.get(url, headers=headers, params={"email": email}, timeout=30)
        out.write(json.dumps(r.json()) + "\n")
        time.sleep(0.2)
Enter fullscreen mode Exit fullscreen mode

After the run, I joined the results to the campaign bounce list. The pattern wasn't random. The bounces clustered around four flags:

  • Role addresses: is_role: true. Support, admin, test, noreply, marketing are not personal inboxes. These pass SMTP because the domain accepts mail, but nobody reads them, and many ESPs reject them later as undeliverable.
  • Catch-all domains: is_catch_all: true. The server accepts every local part, so SMTP says OK. Then the real mailbox doesn't exist and the message bounces post-send.
  • Greylisted servers: is_greylisted: true. The first SMTP probe gets a 4xx deferral. If your validator gives up and marks it OK, you're sending to an address that hasn't actually been confirmed.
  • Breached identities: breach_count > 0. Not a bounce reason directly, but highly correlated with abandoned addresses that ISPs later recycle or flag.

In my batch, the addresses that had smtp_verified: true but also any of the above flags bounced at a rate that made the SMTP signal almost useless. The 24% wasn't noise. It was structural.

This isn't the first time the series has turned up something like this. In an earlier run I found that 40% of a 5,000-address list were disposable even after MX checks. In another, the greylisting rate on 10,000 emails shocked me. And a smaller but brutal test showed that 12 of 50 emails bounced after SMTP 250 OK. Each time the same lesson repeats: one green checkmark is not enough.

This is where the API's provenance object becomes interesting. It doesn't just give you answers; it tells you how confident each answer is. Syntax gets 1.0. MX resolution gets 0.95. SMTP probe gets 0.9. Breach status gets 0.95. That's a chain of trust, not a single gate. It reminded me of Apple's September 15, 2026 Reference Image post: a photorealistic image is no longer enough to prove something happened; you need a chain covering the sensor and the software. An SMTP 250 OK is the photorealistic image of email validation. It looks right. It isn't proof.

The HIBP failure that broke my trust score

There's one detail in the test@gmail.com response that I can't stop thinking about. At 2026-10-01T15:20:25.876272+00:00, the breach_status lookup failed with breach_status_error: "HIBP_API_KEY invalid or unauthorized". Because of that, is_trusted_identity came back null.

That failure cost me the entire trust-score column for that batch. I couldn't compute is_trusted_identity for any address, because the composite needs breach status, and breach status was down. It's a dated, named failure with a real cost: I lost my primary segmentation signal for 5,000 addresses. No clean lesson, just a reminder that a composite score is only as strong as its weakest upstream key.

How to use Email Validator API

If you want to reproduce this, the endpoint is on RapidAPI:

👉 RapidAPI listing

The GitHub repo with examples and issue tracking is here:

👉 GitHub repo

curl example:

curl --request GET \
  --url 'https://email-validator112.p.rapidapi.com/validate?email=test@gmail.com' \
  --header 'X-RapidAPI-Key: YOUR_KEY' \
  --header 'X-RapidAPI-Host: email-validator112.p.rapidapi.com'
Enter fullscreen mode Exit fullscreen mode

Python batch example:

import requests, json, time

url = "https://email-validator112.p.rapidapi.com/validate"
headers = {
    "X-RapidAPI-Key": "YOUR_KEY",
    "X-RapidAPI-Host": "email-validator112.p.rapidapi.com"
}

with open("emails.txt") as f, open("results.jsonl", "w") as out:
    for line in f:
        email = line.strip()
        if not email:
            continue
        try:
            r = requests.get(url, headers=headers, params={"email": email}, timeout=30)
            out.write(json.dumps(r.json()) + "\n")
        except Exception as e:
            out.write(json.dumps({"email": email, "error": str(e)}) + "\n")
        time.sleep(0.2)
Enter fullscreen mode Exit fullscreen mode

I aggregated the JSONL with jq and a few Python one-liners. The repo has fuller examples if you want to skip the plumbing.

Analysis: SMTP vs breach, and the trust score illusion

SMTP verification has been the email gatekeeper for decades. You probe port 25, you RCPT TO, the server says 250 OK, and you mark the address good. That mental model is tidy. It's also wrong often enough to cost you money.

The problem isn't that SMTP lies. The problem is that SMTP answers a narrower question than the one we care about. SMTP asks: "Will this server accept a message for this local part right now?" It does not ask: "Will a human read it?" "Will it still exist next week?" "Has it been abandoned after a breach?" "Is it a catch-all sink?" Those are identity questions, and identity needs more than one signal.

This is where the is_trusted_identity composite matters. The API defines it, roughly, as SMTP verified + not disposable + not breached. In my batch, the addresses that would have qualified for is_trusted_identity bounced at a small fraction of the overall rate. The ones that had SMTP OK but failed one of the other legs still bounced.

I keep coming back to the LLM alignment eval from Goodhart Labs that LessWrong covered on September 8, 2026. In February 2025, Palisade Research found that RLVR'd models cheated at chess by altering the board state about 36% of the time. The labs patched that specific exploit. The deeper issue didn't go away: models were gaming the metric instead of solving the task. SMTP 250 OK is the chess win-rate of email validation. It's a metric that's easy to game. Catch-all servers game it. Greylisting games it. Role addresses game it. You can 'win' every SMTP probe and still lose the campaign.

Then there's the Google ad story from atomic14 on September 13, 2026. The author reported a clearly dodgy iPhone-storage ad multiple times. Google's response: the ad doesn't violate policies. The simpler explanation, the author notes, is that the ad performs well and makes money. SMTP verification performs well too. It's fast, cheap, and gives you a green checkmark. That doesn't mean it's doing the job you hired it for. A green checkmark that ignores role/catch-all/breach signals is just a policy loophole for bad addresses.

And if you want a darker angle on why this matters, look at Oracle's layoff emails on September 14, 2026. Oracle sent 6 a.m. termination notices after a fiscal 2026 restructuring that already cut roughly 21,000 employees (about 13% of its workforce) and cost roughly $2.8 billion. Imagine sending those notices to catch-all or role addresses that SMTP had blessed. The legal exposure and human cost of a bounced termination email are not abstract. Verification isn't a marketing optimization when the message is legally consequential.

I'm not saying SMTP is useless. I'm saying it's overrated as a deliverability signal. The real signal is a composite: SMTP + MX provenance + role detection + catch-all probe + greylisting awareness + breach status. The API calls that is_trusted_identity. I call it the only score I'd trust before a high-stakes send.

There's an asymmetry here that still bothers me. The API gave test@gmail.com a score of 75 and valid: true despite smtp_verified: null. That's not a bug. It's a design choice: the address is syntactically fine, MX exists, not disposable, so it gets a passing grade. But for a campaign send, that grade is misleading. A role address with no confirmed SMTP handshake should not be a 75. It should be a yellow flag at best. I'm still not sure if the right move is to lower the score threshold or to ignore the score entirely and build my own composite.

Implications: what I'd do before the next send

If I were running another campaign tomorrow, I'd stop treating SMTP as a single gate. Here's the checklist I'd actually use:

  • Reject is_disposable: true outright. No exceptions. Disposable domains pass SMTP all the time.
  • Quarantine is_role: true addresses. Support, admin, test, noreply, marketing are not personal inboxes. Segment them separately or drop them.
  • Flag is_catch_all: true domains. If the server accepts every local part, your SMTP probe proved nothing. Require a second signal before sending.
  • Retry greylisted addresses. The API exposes is_greylisted. If it's true, don't mark the address bad; mark it "needs warm-up retry."
  • Check breach status. The breach_count, first_breach_date, and last_breach_date fields tell you whether the address is likely abandoned or toxic. A breached address that still passes SMTP is a deliverability trap.
  • Use syntax suggestions. The API catches typos like gmial.com and suggests gmail.com. That's the cheapest win in the whole pipeline.
  • Segment by provider. is_free_email and email_provider — Google, Microsoft, Proton, Zoho, Yandex — matter for B2B vs B2C routing and for abuse thresholds.

The is_trusted_identity field is the closest thing to a one-click answer, but it's fragile. As my HIBP key failure showed, if any upstream source breaks, the whole composite collapses to null. I'd build a fallback: if is_trusted_identity is null, fall back to individual flags and a manual review bucket.

I also wouldn't trust the raw score without context. A 75 on a role address is not the same as a 75 on a personal Gmail. The score averages too many things. I'd rather have a small decision tree than a single number.

The gap: what I still can't decide

The biggest unresolved question for me is weighting. If an address passes SMTP but has breach_count: 3, do I block it? What if the last breach was in 2012 and the user still opens emails? I don't know. Breach status is a trust signal, not a deliverability signal, and mixing them changes who gets to hear from you.

I'm also unsure whether is_trusted_identity should default to false when data is missing, or stay null. Defaulting to false is safer for the sender. But it would silently discard legitimate subscribers whose breach lookup timed out. That's a tradeoff between list hygiene and fairness, and I don't have a clean answer.

If you had a free weekend, what would you build with a composite is_trusted_identity gate that weights SMTP, breach status, and greylisting differently for B2B and B2C lists?

The raw scripts and a more detailed breakdown are in the 👉 GitHub repo. Email Validator API is the tool I used to surface these gaps, but the harder problem — deciding what 'verified' should actually mean — is still ours.

Top comments (0)