DEV Community

Machine coding Master
Machine coding Master

Posted on

Java & AI: What Developers Need to Know

Stop Context Starvation: Implementing Parent Document Retrieval in Spring AI

Uniform 512-token chunking in production RAG pipelines is dead; it either dilutes vector search precision or starves your LLM of surrounding context. By adopting parent-child retrieval in Spring AI with pgvector, you decouple needle-in-a-haystack search granularity from generation context.

Why Most Developers Get This Wrong

  • Embedding massive chunks: Stuffing 1,000-token blocks into PgVectorStore hoping embeddings retain fine-grained facts—cosine similarity inevitably washes out.
  • Feeding micro-chunks directly to the LLM: Ingesting 128-token snippets yields high search precision, but passing isolated sentence fragments causes the LLM to hallucinate from missing context.
  • Over-engineering with extra infrastructure: Spinning up separate graph databases or external key-value caches instead of leveraging relational metadata linking already available in PostgreSQL.

The Right Way

Decouple retrieval units from synthesis units by indexing micro-chunks containing a parent_id metadata pointer, then hydrating the full parent document before building your prompt.

  • Split hierarchically: Generate large parent sections (1,024 tokens) for synthesis and split them into child micro-chunks (128 tokens) for vector indexing.
  • Link via metadata: Inject the parent's primary key into each child Document using Spring AI's doc.getMetadata().put("parent_id", parentId).
  • Query pgvector with children: Execute vector similarity searches strictly against the child embeddings for maximum retrieval sensitivity.
  • Expand context before generation: Map child hits back to unique parent IDs and fetch the full parents via standard relational queries before calling ChatClient.

Show Me The Code


java
public List<Document> retrieveWithParentContext(String query) {
    // 1. High-precision similarity search on 128-token child embeddings
    List<Document> childHits = vectorStore.similaritySearch(
        SearchRequest.builder().query(query).topK(5).similarityThreshold(0.78).build()
    );

    // 2. Extract unique parent UUIDs from Spring AI Document metadata
    Set<UUID> parentIds = childHits.stream()
        .map(doc -> UUID.fromString(doc.getMetadata().get("parent_id").toString()))
        .collect(Collectors.toSet());

    //
Enter fullscreen mode Exit fullscreen mode

Top comments (0)