This is a submission for the Weekend Challenge: Generosity Edition
Don't Just Ask AI. Give the Answer Back.
AI is a real force multiplier fo...
Some comments have been hidden by the post's author - find out more
For further actions, you may consider blocking this person and/or reporting abuse
Well, technically Gemini users can still use it? While not explicitly used, it does offer a service worthwhile having. I know that finding solutions is often hit-or-miss, especially now, so having a repo to query, with a MCP in order to prioritize it, should help alot of people with common issues.
Especially if it comes to project scoped work, eg. how to set up a CLAUDE.md file for a Blazor project. It becomes a template grab system that works similar to how we used to use stack overflow.
Overall really neat project that tackles a real world frustration, getting working solutions, that would make alot of people's lives easier, especially given that it's user-fed, it can really push the boundaries.
Good point on the Stack Overflow comparison. The difference is that here the solution comes already structured and reviewed, not just a thread of answers where you have to figure out which one actually works. The project-scoped use case is interesting too. A searchable base of working setups for specific frameworks would save a lot of time compared to asking the same question fresh every time.
If you used Stack Overflow before the AI days, the 'burden of proof' was usually if the answer has more than 50 upvotes. That usually meant that people tried it and it works. Though that was a 99% case, rather than 100% verified. To get 100% verified, it'd mean that a single person needs to test every possible scenario and that's usually impossible.
Yes, that's a good comparison. 😄 Stack Overflow's votes were actually a pretty useful form of collective validation: not proof, but evidence that the answer worked for enough people to earn that level of confidence.
I'm not aiming for 100% verification either. That's usually impossible once you consider different OS versions, configurations, dependencies, and edge cases.
That's why I'd rather make the scope and known constraints explicit than pretend that a reviewed solution is universally true. Human review establishes “this is worth publishing as knowledge”; it doesn't magically prove “this works in every possible environment.”
Either way, a 90% solution still saves alot of time and can often give inspiration for getting it working for edge cases.
Exactly. 😄 The goal isn't to pretend every piece of shared knowledge is universally applicable. If it gets you 90% of the way there, that's already a huge win — and the remaining 10% can often be much easier to solve once you have a solid starting point.
That's probably a much more realistic definition of “useful” than “works in every possible scenario.”
This raises some important points. In practice, I've found that the key is balancing theoretical best practices with pragmatic trade-offs — what works in a blog post doesn't always survive contact with a legacy codebase.
Absolutely. And I think that's actually one of the reasons I'm interested in synthesizing knowledge rather than simply collecting “best practices”.
A solution can be perfectly valid in isolation and still be the wrong recommendation once you add a legacy codebase, existing constraints, or a particular environment.
That's another reason why I don't want Shared Knowledge to behave like a static cookbook. The useful part is not just “this works”, but understanding the scope and conditions under which it works — and being explicit when those conditions matter.
In the end, “pragmatic trade-offs” are often part of the knowledge, not a deviation from it.
Treating best practices as universal truths is exactly how we end up with brittle architectures. If this shared knowledge base turns into a static wiki, it will inevitably fail when developers apply rigid solutions to messy, legacy environments. How do you envision the system dynamically capturing those environmental constraints so the synthesized recommendations actually stay context-aware?
I agree with the problem, but I don't think the answer needs to be a dynamic context engine — at least not at this stage.
The environmental constraints are part of the knowledge being contributed and synthesized. If a solution only works on a particular OS, PHP version, database, deployment model, etc., that constraint shouldn't disappear during synthesis. The resulting knowledge should state its scope rather than pretending to be universal.
And if later contributions show that the scope was too broad or that there are contradictory environments, that's exactly the kind of thing the human review step should surface.
For this PoC, I'm much more interested in preserving those constraints than in trying to automatically reason about every possible environment. 🙂
Treating knowledge as a living synthesis rather than a static dump of best practices is exactly what prevents technical debt from compounding. When we ignore legacy constraints, even the most elegant isolated solution becomes a liability. I'm curious how you plan to structure the MCP to continuously ingest and adapt to those environmental shifts without turning into just another outdated wiki.
One angle I haven’t seen mentioned yet is what happens when the knowledge base itself becomes untrusted model input.
Since these articles aren’t only documentation for humans but are retrieved through MCP and consumed by agents, a merged contribution effectively enters the agent’s context boundary. That makes this a knowledge-poisoning / indirect prompt injection problem as well as a content review problem.
Human PR review is a strong publication gate, but “this technical solution is correct” and “this Markdown is safe for an AI agent to consume as data rather than instructions” are slightly different checks.
I’d be interested in how you see that boundary evolving as contributions become less trusted: whether retrieved articles should be structurally treated as quoted/untrusted data, whether certain content patterns should be rejected at publication time, or whether the MCP should enforce that distinction independently of whatever client/model is calling it.
The model agnostic design actually makes that question more interesting, because you can’t assume every downstream client will interpret retrieved content with the same trust semantics.
Maybe I just missed this being mentioned but there are a lot of comments in this post. Good work
I think there may be a slight category error here.
The knowledge base isn't intended to be an untrusted collection of arbitrary model-generated content. A contribution goes through the MCP, becomes a PR, and is explicitly reviewed by a human before it becomes part of the knowledge base.
So if a reviewer validates the solution, the resulting Markdown is simply validated knowledge. The fact that it can subsequently be consumed either by a human or by an AI doesn't fundamentally change the trust model of the knowledge itself.
Of course an agent can be vulnerable to prompt injection from external content in general, but that isn't specific to this knowledge base. The same question applies to any documentation, API response, web page, repository, etc. that an agent is allowed to consume.
I would therefore distinguish “untrusted contribution” from “published knowledge”. The former is exactly why the PR/review boundary exists. Once merged, the content has passed that boundary.
If we started treating every published knowledge artifact as inherently untrusted model input, we'd effectively be saying that human validation doesn't establish anything about the trustworthiness of the knowledge — which isn't really the model I'm building.
That said, the distinction between knowledge and instructions to the agent is an interesting one. But I'd see that as a property of how a particular MCP client consumes the returned data, rather than something that turns a validated knowledge base into a poisoning vector by itself.
I think I may have used “untrusted” too broadly in my first sentence, because I’m not arguing that a merged article should remain epistemically untrusted after human review.
What I’m distinguishing is trust in the knowledge from authority over the consuming agent.
A reviewer can establish “this is a technically valid solution” without necessarily establishing “this content is authorized to influence downstream agent control flow.” To me those are separate properties.
So I’m not sure the distinction is really untrusted contribution vs. published knowledge. I think it’s more like publication trust vs. execution authority.
And I agree this isn’t unique to this knowledge base. Documentation, API responses, repositories, and web pages all create the same boundary when agents consume them. That’s actually why I’d expect the MCP client to preserve a strong data/instruction distinction regardless of how trusted the upstream source is.
In that sense I think we may be closer than my use of “untrusted” initially made it sound.
Yes, I think I see the distinction you're making, but I'd probably stop the boundary there.
My responsibility is to make the application robust and ensure that what gets published has been properly validated. Whether a human reads that knowledge, an agent consumes it, or someone tries to use it in a way I never intended isn't something the knowledge base itself can really control.
So for me the important boundary remains the publication gate: is this knowledge valid and safe to publish? Once published, it's a knowledge resource. What a downstream system does with it is a separate problem.
That makes sense. I think we’re aligned on the distinction then we’re mostly drawing the system boundary at different places. Your publication gate establishes the properties of the knowledge resource, while the consuming client owns the instruction/authority boundary downstream. Thanks for engaging .
Exactly. I think that's where we landed. Thanks for the thoughtful discussion! 👍
The architectural pivot is the part that stuck with me most, deliberately taking Gemini out of the critical path and giving up the prize category eligibility rather than forcing a dependency that would have contradicted the whole point of the project. That's a genuinely principled call and it shows in the architecture.
The "conversation stays private, knowledge extracted from it can be shared" boundary is also really cleanly drawn. A lot of knowledge-sharing tools blur that line or make sharing the default, and the explicit human review gate before anything goes public is the kind of design decision that feels obvious in hindsight but actually requires discipline to hold to when you're building fast.
The ElevenLabs Voice Library bug is a useful one too, the empty environment variable silently overriding the default because os.environ.get() only falls back when the key is missing rather than when it's empty is exactly the kind of thing that takes way too long to find the first time you hit it.
Thank you — especially for taking the time to read the article at that level of depth.
What you picked up on is actually very close to what I was trying to do with the article itself. I deliberately didn't want to turn it into a polished “here's what I built and why it works” story. I wanted to document the reasoning, the architectural decisions, the things that didn't work, and the solutions that emerged from those problems.
In a way, I tried to apply the same approach to the article as to the project: don't hide the messy parts behind the final result, because that's often where the useful knowledge is.
So I'm particularly glad you noticed the Gemini decision, the privacy/knowledge boundary, and even the rather mundane ElevenLabs environment-variable bug. Those details are part of the story precisely because they shaped the system that eventually emerged.
And yes, that
os.environ.get()behavior is one of those wonderfully small bugs that can consume a disproportionate amount of time the first time you encounter it. 😄"Don't hide the messy parts behind the final result, because that's often where the useful knowledge is" is exactly the kind of writing philosophy that makes a technical post worth reading twice. The polished version of a project teaches you what someone built. The honest version teaches you how they thought, and those are completely different things to learn from.
The Gemini decision is the clearest example of that in the article, because it would have been very easy to just not mention it, ship the Google AI integration, qualify for the prize category, and write a cleaner narrative. Documenting the reasoning behind not doing that, and what it cost, is the part that actually transfers to someone else's next decision. Really appreciate you writing it that way.
Thank you — that's exactly the distinction I was trying to make.
A polished project can show you the destination, but the messy parts often show you which roads are actually worth taking — and which ones aren't. 😄
The Gemini decision was particularly important to document for that reason. The “cleaner” story would have been easy to write, but it wouldn't have explained the architectural trade-off. If someone else faces a similar choice later, the reasoning and the cost of the decision are probably more useful than the fact that the project worked.
That's really what I want my technical writing to preserve: not just what I built, but enough of the thinking behind it that someone else can reuse the reasoning.
This is a really great idea. I love the thought of taking those useful little things we figure out while working with AI and turning them into something other people can actually find and use later, instead of letting them disappear into chat history.
I’ve run into a few things that I think could be useful to share, so I might have to contribute something. The human review part is especially nice too. AI can help put the knowledge together, but that doesn’t mean everything it says should automatically become trusted documentation.
Really cool project. I’m interested to see what people add to it.
Thank you! And if you already have a few things in mind that you've learned the hard way, that's exactly the kind of material I'm hoping Shared Knowledge can capture. 😄
The contributor shouldn't have to turn that experience into polished documentation first — that's part of what the MCP and the synthesis workflow are supposed to help with.
And I completely agree about the review boundary. The AI can help extract, structure and generalize something useful, but “the model produced a plausible explanation” is very different from “this is now trusted shared knowledge.”
I'm really curious to see what people end up contributing. The interesting part of the project will probably be what emerges from those small, otherwise-forgotten discoveries once they start accumulating.
I know I definitely have a few “well, that took way longer to figure out than it should have” discoveries sitting around.
I’ve been experimenting with running and converting smaller models lately, and some of the compatibility issues and workarounds might be useful to other people. I’ll have to look through what I’ve got and see what would make a good first contribution.
And yeah, I think that distinction between a plausible explanation and trusted knowledge is really important. It’s easy for something to sound right, especially when an AI explains it confidently, so it's definitely a good idea to have that review step.
I’m really curious to see what people contribute too. I think the little things people figure out along the way could end up being some of the most useful stuff in there.
Absolutely — the smaller-model stuff sounds like exactly the kind of thing that could be valuable there.
Compatibility issues are often a perfect example of knowledge that is painful to discover once, but incredibly useful when someone else can find the answer instead of repeating the same experiment. And the fact that you had to work around them is probably more interesting than a generic “this model works with X” statement — the constraints and the actual workaround are what make the knowledge reusable.
So if you dig through your notes and find a couple of those “why does this work that way?” moments, I'd definitely be interested in seeing them. 😄
And yes, I think that's the core distinction: the AI can help turn those experiences into something structured and reusable, but someone still needs to decide whether the resulting explanation actually deserves to become shared knowledge.
Those little discoveries may well end up being the most valuable part of the corpus precisely because nobody thinks they're worth writing down at the time.
Yeah, exactly! I think the constraints and workarounds are usually the most useful part too. A generic “this model works” doesn’t tell you much when you’ve spent three hours trying to figure out why it doesn’t work on your setup.
I’ve got some notes from a recent experiment that could be a good fit. I might dig through them and see what I can pull together. I like that the MCP can help with the structuring part, because that’s usually the bit that makes me put off writing things down.
Really looking forward to seeing how this grows!
Exactly — and I think that's one of the less obvious problems Shared Knowledge is trying to solve.
The frustrating part is that the useful information is often already there in someone's notes, terminal history, or “three hours of figuring out why this doesn't work” — but turning that into a proper piece of documentation feels like a second, completely different task.
If the MCP can take the messy experience and help structure and generalize it, without requiring the contributor to become a technical writer, then the barrier to sharing becomes much lower.
So definitely dig through those notes when you get a chance. I'd much rather see a rough contribution that can be refined through the workflow than have another useful workaround disappear into someone's project folder. 😄
And I'm curious to see what comes out of your experiment too!
The duplicate check has a race that only shows up once the base is worth searching. search_knowledge reads merged articles, so two people who hit the same wall in the same week both search, both find nothing because the first article is still an open PR, and both open one.
Your Copilot run is the good case and it shows the limit: the assistant did check before publishing, and it could still only see main.
Having publish_knowledge search open PRs as well, or deriving the branch name from a slug of the title so a second attempt collides on the branch instead of on the reviewer, keeps that off the human's desk.
That's a fair observation, although I wouldn't actually consider this a race condition that the architecture needs to eliminate.
Two people can independently encounter the same problem within a short enough window that neither can see the other's contribution yet. That's not necessarily a failure of the system — it's simply a property of collaborative knowledge creation.
The important distinction for me is between preventing duplicate published knowledge and preventing duplicate proposals. The former is something the system should handle automatically; the latter is perfectly legitimate.
That's also why the human review gate is there. If two PRs describe essentially the same solution, the reviewer can reject the second one. But they can also merge the useful parts of both, or even decide that the second solution is better and replace the first one.
So in this particular case, I actually want the architecture to tolerate the race rather than trying to eliminate it. The PR is a proposal, not yet knowledge.
And I think your observation about the Copilot run is still valuable: it makes that boundary very explicit — the assistant can verify against published knowledge, but it cannot know what another contributor is currently proposing. That's a feature of the workflow rather than something I necessarily want to hide from it.
The "AI-generated bug that wasn't there" section hit home — I just published a piece where my own skill read my Zapier test posts as a promise I'd made to my audience and confidently offered to "finish the thread." Same shape as your Copilot phantom bug: the model wasn't wrong in a way that looks wrong, it was wrong in a way that looks completely reasonable, which is the only kind that actually costs you.
And the human-review gate as the one non-negotiable boundary is exactly where I keep landing too. Not because the model is dumb — because "structure a contribution" and "decide it deserves to be public" are different jobs, and only one of them is safe to hand over. Nice to see someone draw that line so cleanly, and hold it even when it cost you a prize category.
Yes — “wrong in a way that looks completely reasonable” is exactly the dangerous part.
A completely absurd answer is easy to catch. The difficult cases are the ones where the model constructs a perfectly coherent explanation from something that was never actually there. That's what made the phantom bug in my Copilot experiment interesting: the failure wasn't a lack of reasoning, it was reasoning applied to a false premise.
And I really like your distinction between “structure a contribution” and “decide it deserves to be public”. That's very close to how I see the human-review gate.
The model can help turn messy experience into a structured, potentially useful contribution. But whether that contribution is accurate, sufficiently general, non-sensitive, and actually worth adding to the shared knowledge base is a different kind of decision.
That's also why I was willing to lose the prize category. If removing that boundary made the project more eligible but less trustworthy, it would have defeated the point of the experiment.
And your Zapier example is a great illustration of the same problem from another angle: the model didn't invent something random — it inferred a plausible intention from context that didn't actually imply it. Those are exactly the errors that deserve a human in the loop.
"Reasoning applied to a false premise" is the cleanest way anyone's put it, and it reframes the whole thing for me. We keep building guardrails against models that reason badly, when the expensive failures come from models that reason well from a wrong starting point. You can't catch that by checking the logic — the logic is fine. You have to check the premise, and the premise is the one thing the model can't verify on its own.
Which loops right back to your review gate: the human isn't there to grade the reasoning, they're there to confirm the premise was real. Good conversation to have stumbled into — I'll be keeping an eye on where Shared Knowledge goes.
Exactly. And I think that's the part I hadn't fully articulated when I started the project.
We tend to think of AI guardrails in terms of how the model reasons: constrain the output, validate the logic, add more checks around the reasoning process. But if the model is reasoning correctly from a premise that never existed, all of those checks can happily pass.
That's why I increasingly see the human review as an epistemic boundary rather than just a quality-control step. The question isn't only “is this reasoning sound?” but “is this actually grounded in something that happened, and does the resulting knowledge deserve to become shared knowledge?”
And there is an interesting consequence for Shared Knowledge: the MCP can help structure and generalize an experience, but it can't be the final authority on whether that experience is real or whether the generalization is valid.
So yes — “reasoning applied to a false premise” may end up being one of the concepts I carry forward from this experiment. 😄
And thank you for following the project. I'm genuinely curious to see where it leads too.
"Epistemic boundary rather than quality-control step" — that's the upgrade. Quality control asks if the work is good; an epistemic boundary asks if it's true, and those come apart exactly where a fluent model is most convincing. I think that's the sentence I'll be stealing back from you. 😄
This was one of the better threads I've fallen into here. I'll be watching Shared Knowledge — and if that premise-checking layer ever turns into something concrete, I'd love to read the writeup.
Haha, go ahead — consider “epistemic boundary” officially open source. 😄
And now that I think about it, there’s a pretty interesting parallel with the problems you’re dealing with in Publora.
Your “right shape, wrong target” failure is almost the operational version of the same problem: the system can validate that something is technically correct without being able to establish that it is actually what was intended.
For Shared Knowledge, the equivalent is: the model can produce a perfectly coherent explanation from a premise that was never true.
In both cases, the dangerous failure isn't a broken call or obviously bad reasoning. It's a successful operation built on a false assumption.
That makes me think the “epistemic boundary” may actually be more useful as a design principle than I initially expected. The interesting challenge now is turning it into something concrete without pretending that truth itself can be automated.
And yes — if I manage to turn that premise-checking layer into more than an architectural box on a diagram, there will definitely be a writeup. 😉
Thanks again for the conversation. This has been one of the more interesting threads I've had around the project so far.
The next thing I'd want as a reader is a small "tested under" block: dependency versions, a minimal input, expected output, and the date someone actually reran it. That can stay in Markdown alongside the solution, without retaining the private conversation.
Your optional-dependency example is a good candidate: can a second person reproduce the import failure and then verify the fix in a clean environment? A merged article and an independently reproduced fix are different kinds of evidence; both are useful, but I'd label them separately.
Would you consider letting the next reader contribute a reproduction result to the existing article rather than another solution? That could make the smallest useful contribution "I checked this on version X; here's what changed," not writing a whole new post.
I like the idea of recording independent reproduction, but I think there is an important constraint here: we can't simply accumulate individual experiences alongside the original solution.
The goal is to synthesize knowledge, not to build a collection of personalized installation stories. If several people independently solve the same problem, the useful result should become something like “Installing n8n on Fedora 18+” or “Installing n8n on Debian / Ubuntu”, rather than a collection of “how Pascal did it”, “how X did it”, etc.
So I agree that an independent reproduction is valuable evidence, but I think that evidence should feed back into the synthesis rather than become another public artifact on its own.
That also means the MCP would need to distinguish between an experience, the synthesized knowledge extracted from multiple experiences, and new evidence that can confirm or challenge that synthesis.
That's actually a rather interesting consequence of the model: the contribution doesn't necessarily produce another article. It can improve an existing piece of shared knowledge.
Yes, that distinction improves what I was suggesting: reproduction should support a shared procedure, not generate another installation article.
I'd keep a compact evidence record attached to the revision it tested: environment, result, and date. The synthesis can then state its supported environments and any unresolved contradictions. A failure on a different OS shouldn't automatically invalidate the procedure for environments where it still works.
The interesting editorial decision is whether new evidence changes the procedure, narrows its scope, or leaves an open question. Would you have a maintainer review those changes, or let the MCP propose a revision with the conflicting evidence attached?
Yes, exactly. I’d have the MCP propose the revision, with the conflicting evidence attached, but the maintainer would still review and decide what actually changes.
The MCP can identify that new evidence contradicts the current scope or procedure. It shouldn’t decide whether that means changing the procedure, narrowing its scope, or leaving the contradiction open.
That keeps the same principle as the initial contribution: AI can prepare the change, but a human decides what becomes shared knowledge.
Great project, Pascal, as usual! And I really appreciate the practical, common-sense approach to AI, especially in the middle of all the current AI paranoia. 😄
Thank you, Sylwia! 😄
And coming from you, “as usual” probably means more than you realize. We’ve been around quite a few AI discussions together by now, so I’m glad that this practical, common-sense approach still comes through.
I think there’s a middle ground that often gets lost: neither “AI will solve everything” nor “AI is the source of everything that’s wrong with software and society.” Just build things, see what actually works, understand where it fails, and make sensible decisions from there.
And yes, there is quite enough AI paranoia around at the moment to keep us busy. 😄
Brilliant idea, Pascal! 😄Sharing knowledge is a great form of generosity. It's part of developer culture.
That said, it had never occurred to me that sharing knowledge fits the challenge topic until I saw your solution.
I suppose it's like a technical wiki. I wonder how others would find and use it. Should we select the website as a knowledge base in an AI agent?
Thanks Julie! 😄 I think that's one of the interesting things about the challenge: the idea isn't necessarily obvious until you look at “knowledge” as something that can be shared back with the community.
It's wiki-like in the sense that the result is a searchable knowledge base, but I don't expect people to manually browse the website all the time.
That's partly why I built the MCP interface: an AI agent can call
search_knowledgewhen it needs that knowledge, rather than relying on the user to know that the website exists. Of course, the client/agent still decides when and how to use that capability.So yes, you could treat the knowledge base as a resource for an AI agent — but through the MCP, rather than simply telling the agent “use this website as your knowledge base.” 🙂
Wonderful! 😄So the tool is agent-oriented. We don't fetch knowledge by searching manually; the agent searches for us.
I've seen people online use tools like AI Agent + Obsidian to save chat history to their knowledge base for review and future use. Your solution pushes this forward, since others can reuse these solutions and save a lot of time. Many of solutions that LLMs provide are summaries of solutions already online, some of which may be outdated. That's why it's helpful when someone verifies a solution and shares it with others.
Pascal, Thank you so much — I've learned a lot from your solution. You designed the knowledge search as an MCP server. I have used and built skills, but I haven't used MCP before. I thought MCP servers were mainly provided by big companies; now I know we can build our own. Awesome!
That's exactly the kind of takeaway I was hoping someone would get from the project. 😄
MCP can look like something that belongs to the big AI providers when you mostly see their servers and integrations, but there's nothing particularly “big company” about building one. It's just a protocol for exposing capabilities to an AI client — you can build a small server for your own use case just as well.
And I like your comparison with Obsidian too. The interesting shift here is from “my AI remembers what I've done” to “the useful part of what I've learned can become reusable knowledge for other people and other agents.”
If this project made you want to build your first MCP server, then I'd say the experiment has already achieved something. 🙂
@pascal_cescato_692b7a8a20 Great write-up! The decision to move Gemini out of the critical path to prioritize MCP interoperability over prize qualification was a sharp call. Standardizing through an MCP prompt guidelines contract makes this far more practical for real-world workflows.
I was particularly interested in your note about Markdown not translating cleanly into speech. When generating the narration script for ElevenLabs, are you planning to use a lightweight LLM pass to rewrite the Markdown into natural spoken prose, or are you aiming for a rule-based parser that inserts pauses and transition cues based on heading levels?
Thanks! For the ElevenLabs part, I'm actually keeping it much simpler for the PoC. 🙂
The validated Markdown remains the single source of truth, and the workflow extracts the content for TTS rather than asking another LLM to rewrite it. I don't want a second model silently changing the knowledge between the human-reviewed article and the audio.
There may be room for a dedicated speech representation later, because Markdown and spoken prose obviously aren't the same thing. But for this PoC, I'm deliberately avoiding another transformation layer.
Makes total sense! Keeping the validated Markdown as the single source of truth without another LLM pass definitely protects against hallucinated changes or altered meaning. It's a clean, safe choice for a PoC.
Are you planning to handle structural pauses using basic regex/rule-based rules in the script extraction step, or will you just pass the raw text string directly to ElevenLabs for now?
For now, just the raw text string. 🙂
No regex layer or structural pause processing in the PoC. ElevenLabs can handle the text as-is, and I'd rather keep the pipeline boring until there's an actual problem worth solving.
If the audio turns out to need a proper speech representation later, that's a separate problem — and probably a good candidate for its own experiment.
An MCP for turning "answer → PR" is such a clean idea 🔥 We ask the same question 50 times a week on different channels and the good answer just evaporates into Slack history. Curious how it handles conflicting answers though, if two people PR different "correct" answers to the same question, does the MCP flag the collision or does last-write-win?
No last-write-wins here. 😄
Both contributions can become PRs, and the maintainer decides whether to merge one, combine them, reject one, or leave the contradiction open. The PR is the proposal; the merged article is the knowledge.
That's also why I don't see duplicate or conflicting contributions as something the MCP needs to automatically “solve”.
That actually makes sense , treating the contradiction itself as signal instead of noise to auto-resolve. Feels closer to how a good tech lead handles conflicting PRs anyway: "the proposal" vs "the accepted knowledge" is a clean split. Bookmarking this pattern, might steal it for how I think about doc review at work.
Exactly. 😄 And honestly, if that distinction is useful outside this PoC, then the experiment has already paid off.
Feel free to steal it — that's kind of the point of the whole thing. 😉
I like the explicit privacy and human review boundary, but I would like to see the operational side of shared memory.
In my experience, memory is not free. There is ingestion and indexing cost, storage, retrieval latency, additional context tokens, and maintenance when entries become stale or contradictory. Memory can save work, but poor retrieval can also add noise or lead an agent confidently in the wrong direction.
Search_knowledge is available, and the article shows Copilot choosing to search for duplicates, but how do you make that behaviour reliable across clients? Is it a system prompt, client policy, hook, or simply a learned convention? If the agent skips the search, the knowledge base is invisible.
I would love to see an evaluation comparing sessions with and without retrieval. Do the results improve correctness, reduce time and token usage, or mainly avoid duplicate work? Measuring retrieval quality, whether the result was actually used, and whether the final solution improved would make the value of the shared knowledge loop much clearer.
That's a fair set of questions, but I think you're looking quite a bit further down the road than this PoC is meant to go. 😄
This was a weekend challenge PoC, not an attempt to design a production-grade shared-memory infrastructure. I wanted to demonstrate the basic loop: contribute → review → publish → retrieve.
search_knowledgeis exposed by the MCP, but I'm not trying to force every client to use it. That's partly the point of making this an MCP: the client decides how it wants to use the capability.From here, there are plenty of possible directions: retrieval strategies, evaluations, client-specific workflows, freshness, scaling, etc. But none of those are requirements for the PoC to demonstrate the underlying idea.
And it's deliberately open: people can participate through development PRs, fork it, implement their own workflows, or simply use the knowledge base. I'm much more interested in seeing what others do with the idea than pretending I already designed the final system. 🙂
That makes sense, and the basic loop is already interesting. I use a similar approach in my own setup, with both code, memory and project documentation available to the agent.
It has been very useful, older problems do not automatically become new problems in later sessions. Past project decisions remain available, so I do not have to rediscover or repeat work that has been settled. Even old bugs have value as historical knowledge because they explain what failed, why it failed, and which paths should not be repeated.
That is why I find your idea compelling even at the PoC stage. The contribution and retrieval loop creates a practical way to preserve experience instead of letting every solved problem disappear with the session. I am curious to see what workflows people build around it.
Yes — I think the “old bugs” point is particularly interesting. A failed approach can be just as valuable as a successful one, provided the knowledge captures the context and the reason it failed.
That's really what I'm hoping the contribution/retrieval loop can preserve: not just answers, but the reasoning and experience that would otherwise disappear with the session.
And at this stage, I'm deliberately leaving the workflows open. I'm curious to see what other people build around the same basic loop rather than trying to prescribe it upfront. 🙂
Respect for dropping Gemini eligibility to keep the architecture model-agnostic. That's the same kind of scope call I had to make this weekend (cut ElevenLabs entirely rather than force it in for a prize category). The os.environ.get() empty-string-vs-missing-key bug is a good catch too, that exact gotcha gets everyone eventually.
Thank you — and having read your Ajo Chain submission, I think I understand the parallel much better now.
We ended up making almost the same kind of scope decision, for different technologies and for slightly different reasons, but with the same underlying principle: don't let the prize category dictate the architecture.
In my case, keeping Gemini in the critical path would have contradicted the model-agnostic design I was aiming for. In yours, keeping ElevenLabs would have meant spending more time forcing an integration that wasn't essential to the actual product. In both cases, the category became less important than the integrity of the project.
And I particularly appreciate the comparison because that's exactly what I tried to preserve in my article: not a polished account where every technology ends up working perfectly, but the actual reasoning, failures and scope decisions that shaped the final architecture.
Also, apparently the empty-string
os.environ.get()trap is now officially a shared experience. 😄the human-review gate is the right call, but curious how it holds up once volume goes up. if ten people hit the same import-chain bug and all ten PRs land in the queue same week, does the reviewer dedupe by re-reading each one, or is there a search-before-submit step that stops the duplicate PR before it's even opened? sounds like search_knowledge already does this in agent mode, so maybe that's the real answer.
Yes, I think that's the important distinction. I don't really want to prevent ten people from proposing knowledge about the same problem — they may actually bring useful differences in environments or solutions.
search_knowledgecan help an agent discover that the problem is already covered before contributing, but I wouldn't make that a hard prerequisite either. The PRs are proposals, not published knowledge.The maintainer can then merge, reject, or use a new contribution to improve an existing article. So the thing I want to avoid duplicating is the published knowledge, not necessarily the contributions themselves.
"Don't just ask AI, give back to it" is a great framing for the challenge. Turning personal troubleshooting notes into shared MCP tools is exactly the kind of knowledge compounding the ecosystem needs. We took a similar angle on the server side: making an MCP server something a developer can build in one annotated Java class, so the cost of "sharing" drops to near zero. Nice work.
Thank you! And I like the parallel with what you did on the server side.
I think we're attacking the same problem from two different directions: if contributing knowledge requires a significant amount of extra work, people will naturally keep it to themselves. Your annotated Java class makes the cost of creating an MCP server almost disappear; my goal is to make the cost of turning an already-solved problem into shared knowledge similarly small.
That's really what I meant by “Don't just ask AI, give back to it”. The useful contribution doesn't have to be another piece of work created from scratch — ideally, it should be something that emerges naturally from the work you're already doing.
The interesting part, for me, is what happens once those small contributions start compounding. That's where the MCP becomes more than just another way of calling an AI model.
Really interesting approach to MCP as a shared, governed knowledge layer, not just a tool interface.
The explicit contribution + PR review flow is especially important—it creates a trust boundary around AI-generated knowledge and helps mitigate hallucination/knowledge poisoning.
GitHub as the initial persistence layer is also a smart MVP choice: versioning, provenance, review, and rollback come almost for free. 🚀
Curious to see how this evolves around knowledge freshness, semantic deduplication, provenance, and retrieval ranking at scale. 🔥
Thanks! That's pretty much the idea: GitHub gives me a surprisingly strong foundation for the MVP — versioning, review, provenance and rollback without having to build another persistence layer.
For the trust boundary, I see the PR review less as “quality control” and more as the point where a contribution becomes published knowledge. The model can structure and synthesize a contribution, but it doesn't get to decide that the result is true or worth publishing.
The scale questions you mention are definitely interesting, especially semantic deduplication and retrieval ranking. But I’m deliberately keeping those simple for now: I’d rather see where the actual bottlenecks appear before adding another layer of AI to solve hypothetical ones. 🙂
PR with human review is what makes this trustworthy. AI can structure the knowledge, but someone (real human) still has to own and judge it.
Exactly. AI can help turn an experience into structured knowledge, but it doesn't get to decide that the knowledge is true. That's the human's job at the PR stage.
Great concept. AI can solve problems quickly, but preserving and sharing those solutions is where the real community value comes in. Combining AI automation with human review creates a much stronger foundation for trusted knowledge sharing.
Thanks! That's really the idea: let AI do the tedious structuring work, while keeping the decision about what becomes shared knowledge with humans.
The Git-based approach has another interesting advantage beyond keeping the MVP simple: it gives shared knowledge a natural history. Technical solutions don't stay correct forever, so being able to see when an article was created, what changed, and why it was updated could become just as valuable as the original solution. As the knowledge base grows, I could see “staleness” becoming an important part of the MCP itself,especially for solutions tied to specific library versions, APIs, or tooling. A solution being reviewed once doesn't necessarily mean it should remain trusted indefinitely. Version-aware retrieval could help an agent distinguish between knowledge that was validated recently and knowledge that may need another review.
Yes, I think that's one of the interesting consequences of using Git as the source of truth. The history is already there, so you don't have to invent a separate provenance system just to know what changed and when.
And I agree that “reviewed once” shouldn't mean “true forever”. Version-specific knowledge will eventually need some notion of freshness or scope, especially around APIs and fast-moving tooling.
That's probably a very natural direction for the project, but definitely not something I wanted to solve in a weekend PoC. 😄 For now, I'd rather have a small system with real history than a sophisticated freshness model with no real corpus to test it against.
How can communities effectively measure the success of a Collaborative Problem-Solving Model (MCP) in fostering shared knowledge?
I think the simplest answer is: measure whether the knowledge is actually useful, not just how much the system produces.
For this PoC, I'd look at things like whether people can find an existing solution instead of solving the same problem again, whether published knowledge gets reused, and whether human review catches contributions that shouldn't become shared knowledge.
Beyond that, I'd rather let real usage reveal what needs measuring. I don't want to invent a sophisticated evaluation framework before there's a real corpus and real usage to evaluate. 🙂
What metrics could best capture "usefulness" in early MCP adoption—e.g., reuse rate, PR feedback, or time saved?
For an early PoC, I'd probably start with reuse rate and time saved, because they measure whether the knowledge actually prevents duplicated work.
PR feedback is useful too, but more as a signal about the quality of the contribution/review process than about the usefulness of the resulting knowledge.
So if I had to pick only two: “Was this knowledge reused?” and “Did it save someone from solving the same problem again?” Those seem much more meaningful to me than raw contribution or PR counts.
Thanks for info :-)
The Gemini part stood out to me. You gave up the prize just so the thing still works with any AI instead of being locked to one company. Most people would've taken the prize and quietly ended up stuck to one. You took the loss to keep it open instead. That's genuinely cool.
And that Copilot review flagging a bug with a line number for code that doesn't even exist, yep, I know that feeling. I run a local model over my own files and it gives me confident broken stuff all the time, looks totally fine until you actually check it. That's why a human having the final say is the part that makes the whole thing work. That's what I personally want to teach practically to others.
Thanks! 😄 The Gemini decision was really less about “giving up the prize” and more about not wanting to make the architecture serve the competition. If the project is supposed to be model-agnostic, putting one model in the critical path just to qualify for a category would have been a strange trade-off.
And yes, that Copilot example is exactly why I keep coming back to the distinction between producing something plausible and establishing that it's actually true. The output looked perfectly reasonable until you checked the premise — there wasn't even a line of code at that location.
I think that's one of the most practical things we can teach people about AI: don't just ask whether the answer looks right. Check what it's actually based on. 🙂
That insight about raw Markdown collapsing in TTS engines due to missing prosodic breaks is a goldmine. Most people assume hooking an LLM pipeline to ElevenLabs is just a simple string-pass, but speech engines require narrative cadence that Git-formatted text completely lacks.
The real strength of this architecture, though, is the strict human-in-the-loop gate before publication. In a world currently being flooded with automated AI-generated documentation garbage, using an agent to structure knowledge without giving it auto-merge authority is the only way to maintain a high signal-to-noise ratio.
Great pragmatic build!
Thanks! 😄
You're right about one small limitation of the TTS output: the Markdown structure isn't translated into prosodic breaks. For example, with an H2, ElevenLabs reads the heading and immediately continues with the paragraph, so it can sound as if the heading is part of the next sentence.
It's noticeable, but for a weekend PoC I didn't consider it worth adding another transformation layer just for that. The validated Markdown remains the source of truth, and the audio is simply generated from it.
A proper speech representation could certainly improve the result later. But for now, I'd rather have slightly imperfect audio than introduce another processing step that could alter the approved content. 🙂
Really like that you kept Gemini out of the core flow — choosing openness over an easy prize category says a lot. Curious how you'll know when keyword search isn't enough anymore. Great weekend project!
Thanks! And for the search question, I think the answer is simply: when keyword search starts failing in practice. 😄
For a PoC, I'd rather have a boring search that works than add semantic retrieval because it might be needed someday. If the corpus grows enough that recall or relevance becomes a real problem, that's a good reason to revisit it and measure what actually helps.
And yes, keeping Gemini out of the critical path was a deliberate trade-off. The architecture mattered more to me than the prize category.
The decision to keep GitHub as the source of truth makes a lot of sense. I’ve built small tools where adding a database early created more work than value.
The boring stack is often the one that lets you prove the actual idea first.
Exactly. 😄 The point of the PoC was to prove the knowledge loop, not to build a database around it.
If GitHub already gives me versioning, review, provenance and rollback, adding another persistence layer would mostly mean adding another problem to solve before knowing whether I actually have one.
The MCP abstraction is the right call — one protocol for tools instead of N SDKs. The bottleneck I keep hitting is discovery and hosting (who runs the server, how you handle auth), not the protocol itself. Nice to see a community-focused take on it.
Thanks! I think that's an important distinction. MCP solves the protocol/interface side; it doesn't magically solve the question of who hosts a server or how authentication should work.
For this PoC I deliberately kept that part simple: it's a community project that can be run, forked or extended by whoever wants to use it. Hosting and auth become much more interesting once you turn it into a shared service rather than a proof of concept.
That's probably one of the areas where the community will end up taking it in different directions.
Wow! Bloody well done!
The ElevenLabs Voice Library trap got me too, same weekend. Mine surfaced as a
402 "Free users cannot use library voices" rather than your invalid_uid, which
was at least readable, but the cause was identical: a voice ID copied out of the
web interface that a free account cannot reach through the API.
Worth flagging for your self-hosted path specifically. "Add the voice to My
Voices" is a manual setup step, and anyone cloning your repo with their own key
hits that wall before they get their first MP3. I gave up on hardcoding a voice
in the end. On a 400, 402 or 404, the code calls /v2/voices, picks one the account
actually has, prefers a British one, retries and remembers the answer. One extra
request, once, and a setup instruction nobody reads disappears.
On Markdown not being built for speech, I came at that from the other end, and it
might save you a step. I stopped sending the document to the TTS engine at all.
The generated text is full of bullet lists, and they read aloud terribly, so the
audio path picks five fields and drops everything else: what the gap is, what the
role exists to change, the time commitment, and who they want to hear from. No
responsibilities list, no skills list. The useful part was that it forced a
decision about what somebody listening actually needs, and it turned out not to
be the whole article. Your narration script step might end up being less about
inserting pauses than about deciding what to leave out.
That's a really useful data point, especially because it confirms the exact distinction I was documenting.
I actually don't think the self-hosted path needs a dynamic voice fallback, though. The intended setup is simpler: go to ElevenLabs → Voice Lab → My Voices, choose a voice that is available to your account/API, then configure the three values explicitly: API key, voice ID and model ID.
The important trap is that a voice being visible and usable in the Voice Library doesn't necessarily mean that its ID is available through the API on a Free plan. That's exactly what caused my misleading
invalid_uiderror. Adding the voice to My Voices changes that entitlement, which is why the same API call starts working. That's also what I documented in the knowledge article.So I wouldn't want the code to silently pick another voice if the configured one isn't available. That would make the setup appear to work while quietly changing one of the user's configuration choices.
Your point does reinforce something useful, though: the setup instructions need to make the Voice Library → My Voices distinction very explicit. That's probably more valuable than trying to make the code guess what the user intended.
And I really like your second point about the audio representation. I think that's going to require a more fundamental rethink than just making Markdown TTS-friendly. The canonical Markdown and the spoken representation probably need to be treated as two different artifacts with two different purposes.
That's actually one of the cases where I don't want the MCP to make the decision automatically.
If two solutions work in different environments, I'd prefer the synthesis to capture those constraints in one article rather than pretending there's a single universal answer. For example: “works this way on X, while Y requires this other approach.”
If the contributions genuinely contradict each other, the contradiction can remain explicit and the human reviewer decides whether to merge them, keep one, or leave the question open.
And yes, the published article is what
search_knowledgesees, not the individual PRs. So ideally a query would surface the synthesized knowledge with its environmental scope, rather than two competing drafts.The important distinction for me is that PRs are proposals; the merged article is the knowledge. 🙂
The explicit human-review gate is the part I’d keep even as the workflow grows. Id also record the source conversation and the final diff with each PR, so a useful extraction stays easy to audit later.
I definitely agree about keeping the human-review gate as the workflow grows. For me, that's a structural part of the architecture, not something to remove once the process becomes more mature.
The audit trail is an interesting point, but I would be very reluctant to retain the source conversation itself. One of the explicit boundaries of the project is that the conversation remains private while the extracted knowledge may become shareable.
Anonymising the conversation before storing it could reduce the risk, but it doesn't eliminate it. There is always a possibility of re-identification through context or information that the anonymisation process failed to catch. And in some cases, the sensitive information isn't even something that can simply be anonymised.
So I'd rather keep the provenance and the approved diff where possible, without turning the private conversation into a persistent artifact. The audit trail should help us understand and verify the knowledge, without creating a back door to the conversation that produced it.
That's a constraint I'll probably explore further as the project evolves.
great job
Thanks! Really appreciate it! 🙏
Appreciate the practical approach here. The real test of any pattern is how it holds up over time — would be interesting to see a follow-up covering how this has scaled as the project grew.
Thanks! That's actually the interesting part of doing it as a PoC: now we can see what problems appear when people actually use it.
If it grows, I'll be happy to write the follow-up — but I'd rather let the project tell me what needs to change than predict the scaling problems upfront. 🙂
What a good idea 💡
great project pascal
Thank you! Really appreciate it! 🙏
This is a really thoughtful take on turning AI-assisted problem solving into lasting community knowledge. 💜
Thanks! That's exactly what I was hoping to explore with the PoC: making the useful part of an AI-assisted solution survive beyond the conversation.
Cover image looks really good!!
Nice write-up!! :D
Thanks! Cover is chatGPT generated on a Claude prompt.
I can see this being really useful inside companies too. Half the useful engineering knowledge usually lives in Slack threads, chats, and people’s heads instead of somewhere searchable.
Absolutely. In some ways, the company use case may be even more obvious: the knowledge already exists, it's just scattered across conversations, tickets, commits, and people's memory.
The interesting part is turning those individual experiences into something the whole team can actually reuse, without requiring everyone to become a documentation specialist. That's a big part of what motivated the PoC.
I lost nearly $150,000 to a crypto scam. Mrs. Roberts Lee ( Roberts Lee 6 1 8 Gmail C 0 m ) ( +1 (856) 549‑7469 ) provided clear, professional guidance and support that helped me recover my funds. I’m truly grateful. Always verify recovery services independently before making any payments.