DEV Community

Cover image for I Built a RAG AI Assistant That Runs in the Browser with WebGPU
Sébastien Descloux
Sébastien Descloux

Posted on

I Built a RAG AI Assistant That Runs in the Browser with WebGPU

A website may already contain the answers visitors need.

The real problem is often finding them.

So I built and deployed a first version of a RAG AI assistant for my own website, with one particular constraint:

The response model runs directly in the visitor's browser using WebGPU.

This was not meant to be another chatbot connected to an API.

I wanted to explore what happens when retrieval, local inference, source-based answers, conversation context, and a CMS are designed as one complete workflow.

How the architecture works

The system splits the work between the website server and the browser.

1. The CMS prepares the knowledge base

Selected website content is split into passages and indexed.

The CMS also lets me manage:

  • which documents are available to the assistant
  • model profiles
  • search parameters
  • assistant instructions
  • activation and validation

2. The browser understands the question

The browser can use recent conversation context to reformulate follow-up questions.

An embedding model then converts the query into a vector.

3. The server performs vector search

The vector is sent to the website API.

The server searches the document index and returns the most relevant passages together with their sources.

4. WebGPU generates the answer locally

The response model runs on the visitor's GPU.

It receives:

  • the question
  • recent conversation context
  • retrieved passages
  • CMS instructions

The first result is not displayed immediately.

A second pass checks the draft against the retrieved sources, and the system can perform one corrective search if necessary.

5. The visitor gets an answer with sources

The final response contains links back to the original website content.

If the available information is not sufficient, the assistant is expected to say so instead of inventing an answer.

Why run the model in the browser?

The main experiment was to avoid calling a hosted generation API for every message.

This changes several things.

No per-response generation API call

Once the models are loaded, generation happens on the visitor's device.

That does not make the system free — hosting, downloads, development and maintenance still have costs — but there is no token bill for every generated response in this workflow.

More control over conversation data

The conversation is not sent to an external generation API to produce the answer.

Optional conversation memory is also stored locally in the browser.

The server is still involved in document retrieval, so this is not a fully offline system.

Local memory

I also added optional semantic memory using IndexedDB.

It can keep previous exchanges locally for a limited period and retrieve an older, related question when useful.

An important design decision was to keep three concepts separate:

  • documents are the factual source
  • recent conversation maintains context
  • local memory helps recover previous intent

An old AI-generated answer never becomes a new factual source.

The limitations are just as important

Running AI models in the browser comes with real constraints.

The current response model requires roughly 2 GB of downloads, plus about 279 MB for the embedding model.

Browser support is another factor.

Having navigator.gpu available does not automatically mean that every model and inference engine will work correctly on that device.

Drivers, GPU capabilities, memory limits and browser implementation all matter.

Latency is also very different from a fast cloud API.

With the models already cached, one documented test produced a reviewed answer in:

  • 36.11 seconds on a Samsung Galaxy S25 Ultra
  • 13.06 seconds on a desktop with an NVIDIA Ampere GPU

These are measurements from my own test devices, not performance guarantees.

More complex mobile requests can take over a minute.

Why I still find this architecture interesting

The goal is not to replace every search box with an AI assistant.

The interesting use cases are situations where users have a real question and the website already has reliable information that can answer it.

For example:

  • product catalogs
  • technical documentation
  • training websites
  • tourism or accommodation
  • business software
  • real-estate search

A user can express a need naturally, while structured data and RAG keep the response connected to information that can actually be verified.

The model is only one part of the solution.

The real value comes from combining:

data + retrieval + business rules + UX + performance + validation

What I learned from this first version

Building this project reinforced something I increasingly see in applied AI work:

A better model does not automatically create a better product.

The difficult parts are often around the model:

  • preparing reliable knowledge
  • keeping retrieval consistent
  • managing context
  • handling model downloads and caching
  • keeping the interface responsive
  • validating answers
  • testing on real devices
  • deciding when the system should simply say "I don't know"

This first version is now running in production on my own website, and I am continuing to test lighter models, additional browsers, GPUs and possible business applications.


I wrote a more complete article with the architecture, limitations, benchmarks and potential use cases:

👉 RAG and WebGPU: When Your Website Becomes an AI Assistant

There is also a deeper technical case study covering the complete RAG pipeline, CMS integration, embeddings, local memory and model choices.

Top comments (3)

Collapse
 
carbonlayer profile image
CarbonLayer •

The device benchmarks and model-download size make the trade-off unusually concrete. Avoiding a per-response generation bill is compelling, but the cost shifts to download friction, uneven hardware, and potentially minute-long waits—especially on mobile. I’m also interested in how the second-pass source check behaves when retrieval is weak: does it reliably say “I don’t know,” or can a locally generated answer still sound more certain than its evidence?

Collapse
 
sdx_development profile image
Sébastien Descloux •

Thanks, that’s exactly the trade-off I wanted to make visible.

In my test on a Galaxy S25 Ultra using Samsung Internet, a complete, verified response took 36.11 seconds, with the models already downloaded. The initial setup requires around 2 GB for the response model and 279 MB for the embeddings model. More complex questions can take over a minute on mobile. Local generation avoids paying an API fee for every response, but hosting, bandwidth, and performance differences across devices are still real costs.

When information retrieval is weak, here’s what happens:

If no relevant passage is found, the assistant tries a more targeted search. If that still returns nothing, it offers a human handoff instead of generating an answer.

The first draft remains hidden. A second pass sends the model the question, the draft, and the retrieved passages to check whether the answer actually addresses the request and whether it is genuinely supported by those sources.

If the draft is rejected, the assistant can run a corrective search and produce a second draft. If that one is also rejected, neither draft is shown: the assistant offers a human handoff instead.

For example, in my tests on both the S25 and a desktop computer, this verification step rejected a fabricated price range of €15,000 to €100,000, followed by an answer that avoided mentioning the documented range of €1,500 to €5,000. It then accepted a corrected response.

That said, I wouldn’t claim that the system always knows when it doesn’t know. The verification step uses the same local model: when faced with passages that sound plausible but are actually insufficient, it could still accept an overly confident answer. So my tests show that the mechanism can catch certain specific errors, but they do not support any general conclusion about its overall reliability.

The next step is to measure this using questions that are deliberately unanswerable or based on conflicting sources, while tracking unsupported claims and the cases where a human handoff is triggered appropriately.

Overall, it’s an interesting experiment because it makes it possible to see, in practical terms, how far a fully local, in-browser RAG system can go — and where its limitations still are.

Collapse
 
carbonlayer profile image
CarbonLayer •

Thanks for laying out the full flow—and for including a concrete failure the second pass caught. The handoff path matters just as much as the verification step: when retrieval still comes up empty, the system has a defined way to avoid turning a weak search into a confident answer.

The caveat you call out is the one I’d watch most closely. Since the verifier uses the same local model, it may share the draft’s blind spots; a plausible-sounding passage could make both passes overconfident. Your proposed tests with unanswerable questions and conflicting sources seem like a good way to probe that. I’d be especially interested in how you distinguish a justified handoff from an unnecessary one, and how often unsupported answers still get through.

Are you planning to publish results across those cases—and, if so, will you compare behavior across devices or just measure answer quality? The 36-second result with models already downloaded makes that trade-off part of the evaluation too.