DEV Community

Cover image for AI-assisted genealogy

AI-assisted genealogy

Nicolas Fränkel on September 03, 2026

My son recently came to me to brag about using AI to find our ancestors. While the results were correct, I didn't learn anything new, as it stopped...
Collapse
 
lavich profile image
Andrei Lobanov •

Great article! I've been working on tooling around GEDCOM myself, so it was really interesting to see how you're using it with AI.

I've recently developed Domorium, an Obsidian extension for working with GEDCOM files. It provides navigation between records and photos and includes a built-in GEDCOM validator.

Might be useful if you continue exploring different ways of working with GEDCOM.

Collapse
 
kenwalger profile image
Ken W Alger •

Good writeup. The GEDCOM-in-Git approach is the part I'd steal.

Three things jumped out at me from having done this by hand for my own family line.

Wrong facts are sticky. Your Switzerland example is the general case, not an edge case. Once someone infers a place, date, or relationship and commits it to a tree, it gets copied into other trees. After a few hops, the original inference can become almost indistinguishable from a sourced fact. The claim's provenance gets weaker while confidence in it somehow gets stronger.

Repeated names are worse than they look, too. I've dealt with the equivalent of John Smith having a son named John Smith who then has another son named John Smith, all in the same area with lifespans overlapping by decades. Jr. and Sr. get applied inconsistently or disappear entirely. The failure mode isn't just choosing the wrong person. You can merge two people into one or split one person into two, and either mistake can leave behind a tree that looks perfectly reasonable. The Savoy discriminant you found is a good example of the kind of evidence needed to separate them.

The third one is something I've been thinking about recently: your loop runs until no records are found. In the US, that can terminate rather dramatically at 1890 for many families because most of that census no longer survives. But "no records found" can represent very different situations. The schedule may not survive. The record may exist but still be sealed under the 72-year rule. The relevant repository may never have been searched. Or the person genuinely may not have been there.

GEDCOM has SOUR for documenting where a claim came from. What I find interesting is the inverse problem: how do we preserve the provenance of an absence?

A family tree records the evidence that allowed a branch to continue. It usually says much less about the evidence, boundaries, and failed searches that caused a branch to stop.

So the tree ends, but the reason the tree ends isn't really part of the tree.

Collapse
 
nfrankel profile image
Nicolas Fränkel •

Thanks for your feedback.

Wrong facts are sticky. Your Switzerland example is the general case, not an edge case. Once someone infers a place, date, or relationship and commits it to a tree, it gets copied into other trees. After a few hops, the original inference can become almost indistinguishable from a sourced fact. The claim's provenance gets weaker while confidence in it somehow gets stronger.

That's the reason I do use QUAY a lot. Only official records (administrative or religious) are QUAY 3 (the topmost quality). That's absolute confidence. Below that, it's an assumption. A huge part of the work, after having mined the genealogy sites is to actually find the relevant record to up the quality.

Repeated names are worse than they look, too. I've dealt with the equivalent of John Smith having a son named John Smith who then has another son named John Smith

I have the same. 3 generations of people who had the great idea to pass their first name to their child. Absolute hell.

GEDCOM has SOUR for documenting where a claim came from. What I find interesting is the inverse problem: how do we preserve the provenance of an absence?

I first used the NOTE field, but it makes the GEDCOM bloat pretty fast. I have a index of research notes, outside of the GEDCOM, plus GitHub issues.

Also, you have it pretty easy with genealogy in recent times. When you get back not that far, you get issues you didn't probably encounter:

  1. Orthography is "fluid". Sometimes, even in the same act, the priest writes the name in two different ways.
  2. Latin. As you get back in time, acts are more and more in Latin. So your ancestor Pierre is known in the acts as Petrus (and with cases, it's even more fun).
  3. Lack of last names. In some branches of my family, I went back to a time where they are just X of Y region.

Have fun!

Collapse
 
publiflow profile image
PubliFlow •

This raises some important points. In practice, I've found that the key is balancing theoretical best practices with pragmatic trade-offs — what works in a blog post doesn't always survive contact with a legacy codebase.