Find My Thingy: A Local-First AI Memory Assistant Powered by Gemma 3 4B
Have you ever remembered saving something important in a documen...
For further actions, you may consider blocking this person and/or reporting abuse
Source citations on top of a local Gemma 3 4B pipeline is the detail I like most, since that is what makes a personal RAG trustworthy instead of a confident guesser. The hard part I keep hitting is evaluating retrieval quality once it is all on-device with all-MiniLM-L6-v2; I wrote up how I test that for LlamaIndex here: dev.to/kartik-nvjk/llamaindex-make... . How are you checking that the retrieved chunk actually supports the answer it cites?
That's something I'm still improving. Right now, I mainly check this by looking at the retrieved chunks and making sure the cited chunk actually contains the information used in the answer, rather than just being semantically similar.
I haven't implemented a full automated faithfulness evaluation yet. My next step would be to build a small evaluation set with expected supporting chunks and measure things like Recall@k, context precision, and claim-level faithfulness separately. That would make it much easier to tell whether a bad answer came from retrieval or from Gemma itself.
Really like how you've approached this problem. Instead of building another generic chatbot, you've focused on personal knowledge management. That's a nice direction!
thankyou !!
How is the personal memory library different from the document-based knowledge base? Are they stored and retrieved separately?
Yes, they're kept separate. The document knowledge base is for information coming from files I upload, while the Memory Library is for personal facts I explicitly want Find My Thingy to remember.
So basically, documents are things I give it to read, while memories are things I tell it to remember. This separation also makes retrieval more predictable instead of mixing everything together.
Very good project. Good thinking
thankyou!
Could you explain how you are using ChromaDB and Sentence Transformers together for document retrieval?
Yeah! I use Sentence Transformers to turn the document chunks into embeddings, and then store those embeddings in ChromaDB. When I ask something, the query is embedded in the same way and ChromaDB finds the most relevant chunks. Those chunks are then passed to Gemma as context for the answer.
So the basic flow is: document → chunks → embeddings → ChromaDB → relevant chunks → Gemma
Thankyou!!...yeah planning to add that too
Okay
yes!
This is such a cool take on personal knowledge management! Love the local-first approach and the fact that answers are backed by source citations. Great work! 👏
Thankyou 🥰