Building the vector store last time was only half the job. The actual point of embeddings is asking a question and getting back the right document, not just any document that happens to be sitting there. Chroma does this with distance: turn the question into a vector the same way you turned each document into one, then measure how close the question's vector is to each stored vector. Lower distance, more similar in meaning — at least in theory. The real test isn't whether it returns something, because it always will. It's whether the number actually means anything.
Reconnected to last entry's collection — collection.count() confirmed all 3 entries were still there, good — then ran two queries through the same embed-then-search pattern:
>>> q = ollama.embeddings(model="nomic-embed-text", prompt="how do I check pod status with oc")
>>> results = collection.query(query_embeddings=[q["embedding"]], n_results=2)
>>> results["ids"]
[['02-oc-cli-mentor-system-prompt.md', 'posts_03-1b-vs-3b-memory-comparison']]
>>> results["distances"]
[[437.72, 499.63]]
>>> bad_q = ollama.embeddings(model="nomic-embed-text", prompt="what's the best pizza topping")
>>> bad_results = collection.query(query_embeddings=[bad_q["embedding"]], n_results=2)
>>> bad_results["distances"]
[[542.33, 585.48]]
| Query | Top match | Best distance | Worst distance |
|---|---|---|---|
| "how do I check pod status with oc" | 02-oc-cli-mentor-system-prompt.md |
437.72 | 499.63 |
| "what's the best pizza topping" | (same 2 docs, wrong topic) | 542.33 | 585.48 |
Two things worth calling out here. The on-topic question correctly surfaced the oc-mentor file as the closest match — found the right document, not just a document, which is genuinely the thing I was hoping to see and wasn't fully sure would happen. But the more convincing result is the second one: every distance for the real question came in lower than every distance for the pizza question. Even the worst on-topic match beat the best off-topic match. Small sample, sure, but that's not noise — the number is actually tracking relevance.
One loose thread I want to flag rather than bury: the third stored document — the long draft that got truncated by 57% last entry — never showed up in either result, for either question. Can't tell yet if that's because it's genuinely unrelated to these two test questions, or because chopping away more than half its content wrecked whatever got embedded. That's a real open question, not something I'm going to pretend I've already answered.
So: the retrieval loop works. Real discrimination, not just "closest by default." But three documents and two questions is a proof of concept, not an eval — the next real thing to do is fix the chunking problem from last entry and see if it changes how that truncated document behaves here.
Top comments (0)