DEV Community

VLAD
VLAD

Posted on

Embeddings explained for people who write code

Type "how do I reset my password" into a good search box, and it finds a page titled "account recovery." Zero words in common. Nothing matched on text. So how did it know they mean the same thing?

Your app turned meaning into numbers — and once meaning is numbers, a computer can measure it. That trick is called an embedding. Get this one idea and half the AI buzzwords — "vector search," "cosine similarity," "vector database" — stop being scary.

Prefer to watch? Full walkthrough with the meaning-space animation:

An embedding is a list of numbers

Start with the machine itself. You hand it a piece of text — a word, a sentence, a whole paragraph. It hands back a list of numbers. That list is the embedding. Same text in, same numbers out, every time.

How long is the list? For a common model — OpenAI's text-embedding-3-small — it's 1,536 numbers. That sounds like a lot, until you think of each number as a coordinate.

Two numbers place a point on a map. Three place it in a room. 1,536 place it in a space you can't picture — but the math works exactly the same as the map.

Meaning becomes position

Here's the whole point of that space: the model places text so that similar meaning lands in a similar spot. "cat" and "dog" end up as neighbors. "car" ends up far away.

Nobody wrote that rule. The model learned it — it read a mountain of text and noticed which words keep the same company. Words used the same way get pushed together; words used differently get pushed apart. This is an old idea from linguistics called the distributional hypothesis: a word's meaning is shaped by the words it usually appears next to. Meaning, in this space, is just where you land relative to everything else.

"Cosine similarity" is just an angle

Now the real question in search: are these two things close? You embed the question, you embed every document, and you grab the points nearest the question. But "nearest" means one specific thing — and it's simpler than it sounds.

Draw an arrow from the center of the space out to each point.

  • Two texts with similar meaning? Their arrows point almost the same way — a small angle between them.
  • Unrelated texts? The arrows point off in different directions — a wide angle.

That angle is cosine similarity. A value near 1 means the arrows point the same way (very similar). A value near 0 means they're unrelated. That's the whole comparison — "cosine similarity" is just measuring the angle between two arrows.

In code, it's three lines

Every "AI search" feature you've used is basically this, under a nicer name:

const doc   = embed("account recovery steps")      // → [0.02, -0.91, …] · 1536 numbers
const query = embed("how do I reset my password?")

// cosine: 1 = same direction, 0 = unrelated
const score = cosine(query, doc)                   // ≈ 0.86 — close
Enter fullscreen mode Exit fullscreen mode

Embed your text, embed your query, compare them, keep the closest. (Those scores are illustrative — the real point is high = same direction.)

Trap 1: the numbers only make sense inside one model

These coordinates are only meaningful within a single model. A vector from model A and a vector from model B are not comparable — it's gibberish. Same coordinates, different maps.

So embed your query and your documents with the exact same model, always. Change the model, and you have to re-embed everything you're searching over.

Trap 2: embeddings capture meaning, not exact strings

Embeddings are great at meaning — which makes them bad at things that have no meaning. An error code. A product ID. SKU-4417. There's nothing to place on the meaning-map; it's just an exact string, and pure vector search fumbles it, because nothing is "close in meaning" to a serial number.

That's why real systems run both: keyword search to catch exact strings, vector search to catch meaning, and the results merged. If you've built RAG and watched it miss an obvious error code, this is usually why.

The takeaway

An embedding, start to finish:

  1. Text becomes a list of numbers.
  2. The numbers are coordinates in a meaning-space.
  3. Close in meaning → close in space.
  4. "Similarity" is just the angle between two arrows.

Once you see it as numbers on a map, the buzzwords fall away — and you can actually debug your search instead of trusting it. When a result looks wrong, you're not staring at magic; you're asking a concrete question about distance on a map.

What's the weirdest match your vector search ever returned — the one that made no sense at all? Drop it in the comments — I read them.


I make Vlad's Stack — how the tools you use every day actually work, for people who write code. Full video walkthrough is above.

Top comments (6)

Collapse
 
hannune profile image
Tae Kim •

Answering the question: a Korean company abbreviation matching a Unix command flag because both were four uppercase Latin letters. It's the identifier trap you mention, but the disguise was subtle enough that it took me an afternoon to spot on a knowledge-graph query. The hybrid approach you describe is right; what actually stuck for us was boosting BM25 scores for tokens shorter than six characters before the merge rather than after. We'd tried the post-merge boost and it didn't move the recall numbers enough to matter.

Collapse
 
vladut02 profile image
VLAD •

Great example. And a sneaky one, because both strings look like "identifiers" to a human, but to the embedding model they're just noise that happens to land in the same spot.

The before-vs-after merge point is really useful. My guess at why it works: a post-merge boost can only reshuffle docs that already made it into the top-k. If the exact-match doc never made it into BM25's top-k in the first place, no boost afterward can bring it back. Boosting before the merge changes which docs get in at all. Does that match what you saw?

Also curious how you landed on six characters. Did you tune it on your own data, or was it a rule of thumb? I'd guess it depends a lot on how your IDs look.

Collapse
 
hannune profile image
Tae Kim •

Yeah, top-k coverage is exactly the failure mode - KFTC scored low enough on BM25 that it just never entered the candidate set, so boosting after the merge was already too late. The six-character thing was honestly not tuned: I noticed most of the problematic identifiers in that corpus happened to be longer than common stopwords, so I tried six and it stopped the false matches, and I never really went back to validate it more rigorously. I would be suspicious of using that number anywhere else without actually looking at where your own false negatives fall by token length.

Thread Thread
 
vladut02 profile image
VLAD •

That's the more useful half of the answer, honestly. The number six isn't the takeaway, the method is: look at where your own false negatives fall by token length, instead of copying someone else's threshold. Thanks for saying the unglamorous part out loud.

Collapse
 
hannune profile image
Tae Kim •

The 'similar meaning lands in a similar spot' part works beautifully until you hit cases where the same word means completely different things - 'Apple' the fruit and 'Apple' the company end up in weirdly similar spots in most models I've tested. We ran into this building entity search over financial documents where 'Mercury' could be a car brand, a planet, a chemical element, or a startup, and the embedding alone genuinely can't tell you which one without surrounding context. Once the lookup is just 'find similar vectors to this query' you're flying blind on which sense of the word you're actually looking at. We've ended up keeping embeddings for semantic similarity but layering a graph on top to handle disambiguation when it matters.

Collapse
 
vladut02 profile image
VLAD •

Real problem, but I'd put the blame in a slightly different place. Modern embedding models are context-aware - "I ate an apple" and "Apple raised prices" land in genuinely different spots, because the vector is built for the whole piece of text, not for the word in isolation.

The trouble starts when there is no context to work with. If what you indexed is a bare entity name - just "Mercury" on its own - then the model has nothing to condition on, and you get a vector that sits somewhere in the middle of every sense of the word. Same on the query side: a one-word query gives it nothing.

So I'm curious about your setup: are you embedding the entity names by themselves, or the sentence/paragraph they appear in? If it's the name alone, giving it a short description before you embed it sometimes fixes a surprising amount of this - and it's a lot cheaper than a graph. Though for financial entities you probably want the graph anyway, for reasons that have nothing to do with search.