What 1.9.9 fixed, how we tested, and what we are building next.
14 min read · vektormemory.com
The pace is changing
The field keeps renaming itself: assistants, agents, and now SI “super intelligence."
I always liked “the everything machine” better, because it says what we want from the thing. If it is going to merge into the one "thing," it has to be built from parts that all work in unison, accurately with speed.
A few months ago, VEKTOR was a tiny 4-layer memory store with a lot of lofty ideas around it, unbuilt modules on a Kanban board, and saved to memory as future planning projects.
Today we can ask an assistant to run a workflow, watch it ask for permission (Hitl), and read afterwards exactly what it did, saved to the memory layers graph, with full transparency.
A search that took half a second now takes a few milliseconds, shaving off time from the dozens of iterations we made. A memory that used to pile up now notices when a fact has changed. None of the ideas are ever finished, and we list what is still rough around the edges below. But the pace has changed, and I think I know why.
Building the foundations was the hardest part.
Memory, context, skills, MCP tools, and the connectors all had to be built and work properly and agree with each other before anything built on top could be trusted. For a long time that meant slow, careful back-end plumbing that nobody sees, wiring, event handlers, and smoke testing 100’s of LLMs constantly changing and improving.
Now the majority of those pieces hold, and each new feature starts with more already in place. The assistant can search your notes, read the custom Vektor guide in the library we made for users and LLMs, see which panel you have open, and use the same 60+ MCP tools and memory Claude via sees. We stopped re-explaining things to it, so new work lands faster. We expect the pace to keep climbing, as the Frontier models improve in coding and the open-source models as well.
There is a second reason. Somewhere in the last few weeks we reviewed our design and research methods and started measuring it more efficiently, primarily because we are inpatient, with a small framework we describe below and will release as a skill once refined.
It borrows from John von Neumann and adds a few habits of our own that we learned from watching LLMs struggle with benchmarking and looping. It’s not claiming new science or a white paper, just a remix of old and new ideas into the agentic tools we use now. It culls the bad ideas early and cuts our test time from days to hours via microtests in a matrix.
**The real problem
**The field has moved so fast in the last two months from Openclaw and Hermes agents.
Now with OpenAI launching “dots” on 29 September: always-on personal agents with their own cloud computer, connected to more than 4,000 apps, working read-only in the background and asking before they send or change anything. xAI shipped Grok Bot in August, an agent that operates apps by looking at the screen and remembers your corrections, amazing. Meta’s Muse Spark plans a job and hands the pieces to parallel sub-agents, with cute characters and payments.
All three live in the vendor’s cloud data centers, highly polished with millions spent on devs, and all three keep your memory there, and all have agent credit card payments built in, some even work in your Tesla EV, controlling real vehicle settings, the future promised now.
We look at where the market is going, and we are building the other trade-off: assistant and memory on your own machine, any model you like, with a human in the loop by default, security, hitl, privacy, soveriegn data.
Those choices cost something, complete agent autonomy mode. Currently a small local model is worse than a frontier model at deciding when to use a tool, the cloud computer option to lean on, or built-in credit card payments, which we can see as being an absolute cluster if an agent conducts incorrect transactions, loop errors, and customer support issues in the future.
Now a billion-dollar company will have thousands of offshore support staff agents that can handle 1000s of tickets for insurance and cashback claims when the agents go rogue and order 200 pairs of shoes instead of 2.
So the question for this release was not “which payment feature is next" because that is where the market is at. but “which parts of the foundations work well for us based on our own ethos."
It's all a sacrifice to maintain our local philosophy, but then again, it avoids a lot of the foreseeable payment headaches. Maybe we are right or wrong; sitting on sidelines and watching it all play out, we can learn from the mistakes and implement and revise as needed at our own pace.
Vek Panel
How we test: grids, gates, worst-case rules
Von Neumann minmax matrix
Every “which method is best” question now goes through the same three steps.
The grid. Rows are methods, and the shipped method is always the first row. Columns are accuracy, speed, and accuracy on a holdout set. Each cell is one small experiment. Because each cell takes seconds, we can try thirty ideas in an afternoon.
The gates. Each gate is a yes or no. We tune on 10 questions and keep a change only if the ranking score rises and the top-10 hits do not fall.
We then re-check the kept change on a separate 20-question form, and it must not be worse there. The remaining 70 questions are never used while tuning, and they are the only numbers we report.
Ten questions are noisy: one question moves the score by 10 points. The second gate exists to catch luck. Several window-size changes passed the first gate and failed the second.
The worst-case rule. For each row we compute two regrets: how much accuracy it gives up against the best row, and how much speed it gives up against the fastest row. We take the worse of the two and pick the row with the smallest worst case.
Here is a real example from the library search matrix:
Matrix
The 30-candidate row is the most accurate, but it pays for that in speed. The 20-candidate row gives up 5% of the accuracy and loses far less time, so it wins. A fast method that is wrong and an accurate method that is slow both lose under this rule. Timing is noisy, so the milliseconds moved by a few between runs. The ordering held.
Where the idea comes from. In 1928 von Neumann proved the minimax theorem: in a two-player zero-sum game, you can pick the strategy that limits your worst-case loss. We borrowed the decision rule, not the game.
Nobody is playing against the test set. The “opponent” is our own habit of picking the method that wins on the metric we happen to like, whilst refining and improving on our ideas, reiterate quickly, fail fast, and win fast, stacking upwards into the matrix.
Why do this? Because the market is moving that fast, there is little point watching the LLMs hammering tokens and performing benchmark tests over and over to find out the original theory and ideas were wrong in the first place and bear no tangible results…
It is also worth knowing what the research says about people and minimax.
In 2009 Levitt, List and Reiley tested whether professional poker and soccer players actually play minimax in laboratory games (NBER working paper 15609). They mostly did not, and the skill did not transfer from the field to the lab.
Our reading is modest: do not assume theory describes the behavior, and use worst-case regret as a simple tie-breaker when no single metric can be trusted.
What the research says about the pieces
A few results shaped what we built and what we left alone.
Keyword search stays the base. A 2026 scaling study, “BM25 Wins at Scale”, ran 28 corpus sizes up to about 600 million tokens. Keyword search overtook dense retrieval around 10 million tokens and led at every larger size. So the Library keeps keyword search as its foundation and adds meaning on top.
Fusion by rank. To combine two ranked lists we use reciprocal rank fusion (Cormack, Clarke and Büttcher, 2009). It adds up 1 divided by each list’s rank, so a result that is high in both lists wins, and no one has to compare raw scores from different systems.
A model-written search, only when needed. HyDE (Gao and colleagues, 2022) has a model write a hypothetical answer and searches with that. It costs a model call, so VEKTOR runs it only when the quick searches disagree, and waits about a second at most.
The embedding window. The small embedding model reads about 128 tokens, which is roughly 512 characters of English. Our library passages average 2,600. That one fact explains why plain re-ranking by meaning hurt at first, and why scoring overlapping windows fixed it.
Dates at search time. Memory tools describe handling time questions with rule-based pattern matching and no extra model calls, using dates stored with each memory. We store when each fact was replaced, and we plan to test our own version of that for the ledger described near the end.
The initial assumption is not always correct
The Library holds about 154,000 passages from our test books. Search took 545 ms on our test queries, which you notice when an assistant searches three times before it answers.
We assumed the ranking code was the cost. A profile said otherwise: one line counted every passage in the index before each search, and that count alone took about 470 ms.
Reading the total the index already stores brought it under a millisecond. Searching first for passages that contain all the meaningful words, and widening only when needed, took ten short queries to 5.6 ms on average. One hundred full-sentence questions average 96 ms.
Speed was the easy half. To measure accuracy, we had a local model write 100 natural questions from random passages, each tied to the passage that answers it. Plain keyword search put the right passage in the top five for 63% of them.
We tried five ideas to improve that. Dropping search terms for speed cut it to somewhere between 23% and 52%. Splitting passages into smaller pieces cut it to 53%. Letting a local 9-billion-parameter model pick the best passage made first place worse on a 30-question check, from 47% to 27%.
One thing helped: re-ranking the top 20 by meaning. Because of the window limit above, we score overlapping 480-character windows and keep the best one.
On 70 questions we never tuned on, the ranking score went from 0.41 to 0.50, about 22% better, at the same speed once the vectors are stored. They fill in quietly as you search.
Limits: a local model wrote the questions, 70 questions leaves roughly five points of noise, and your first search on a new topic is ordinary keyword search.
A memory that says “that changed”
Save “the meeting moved to Thursday” and the old “Tuesday” memory should stop answering.
VEKTOR now compares each new memory with the closest older ones. A different date, number, owner or address, or a “no longer”, links the pair with a “contradicts” edge. It uses wording and the stored embeddings only, with no model call, so it is private and free.
By default nothing is replaced or deleted. An optional setting retires the older fact, which stays in your file.
The first version was wrong in a useful way. On a test set of 46 pairs it caught 19 of 20 changes with no false alarms. On our real test file of 2,256 memories it flagged 37% of everything, because the file is full of long chat transcripts and test timestamps whose numbers differ constantly.
We restricted it to short plain statements, ignored long numbers and anything shaped like code, and required clear change wording before a bare number counts. It now flags 10 memories, 0.4%.
One gap we cannot close yet: the test file has no recorded “this replaced that” pairs, so we cannot tell you how many real changes it misses. That is why it flags and does not rewrite.
An assistant must not claim what it did not do
Vek the mini assistant can now use tools: search your notes, memory, the guide and the Library, open a panel, save a note, run a Workflow.
You choose how much it may do. Read means it only looks, Ask (the default) shows each change on a card and waits for Approve or Deny, and Auto just does it. Every action lands in a list called “What Vek did”.
No tool can buy anything in cyberspace with a credit card, that's the compromise we made so far.
Our test model is a local 9B model through Ollama, and it does not call tools reliably, yet.
So the loop does three things:
It tells the model which tools exist and gives an example. If you asked for an action and the answer only says it was done, the loop replies that nothing happened and asks once more. Plain chat answers get a rule that they cannot run anything and must not say they did.
Over 21 requests across seven kinds of task, the model used the right tool 20 times. The miss was a claim that a workflow did not exist.
Privacy you can switch on or off
An internal auditor run we did critiqued our privacy chapter and asked what the “few small automatic requests” were, and noted that weather and currency widgets also call out.
The chapter lists four: the update check, licence activation, first-time model downloads and the model list. The widgets were not on the list.
They are now, and there is a switch. In Config, under Privacy, “Block automatic network requests” stops everything VEKTOR does on its own. A test records every outside address and shows zero requests with the switch on.
The app sends no telemetry in either state. We searched the code again for analytics and error-reporting addresses and found none.
What is in 1.9.9
Memory:
Keeps itself up to date in the background: backup first, never deletes a memory, resumes if you close it
Claude, the app and the terminal share one memory
A richer Graph, linked as you save: by names, by time order, by meaning
Searches follow names, and the slower model-written search step runs only when the quick searches disagree
Replaced facts stay replaced, and the new contradiction check
Vaults work everywhere: Claude and the terminal follow the vault you pick
Fading memories now covers every memory, with safeguards, and vektor cold shows and undoes it
New terminal commands: vektor recall, backup, export, tags. db-merge imports into your normal file
Vek assistant:
Tools with Read, Ask and Auto levels, an approval card and an activity list
Reads public links and your notes, runs Workflows, answers how-to questions from the guide
Two-part answers, personalities that change how it answers, an Ask box under it, spoken answers, orb avatars and a larger face
Library, guide, workflows:
The VEKTOR guide ships as a 55-chapter book in your Library, with matching MCP tools (60+ in all)
Library search much faster (545 ms to under 6 ms on short queries) and re-ordered by meaning. EPUB pictures and a wider reader
Workflows cards fill the page, with All, Scheduled and On demand views, search, a details window and last-run time. Results can go to Slack or email
Setup and privacy:
The privacy switch described above
Both setup wizards offer all 18 providers, and your question goes only to the provider you chose
API keys set from the terminal go into the encrypted vault
A 15-tool profile for apps that cap how many tools they accept
Terminal commands start in 0.15 seconds instead of 0.9
The road to v2.0
These are our plans, in the order we expect them being built now over the next month:
A decision ledger. Ask “what did we decide about the Quokka plan, and what changed since?” or “what did I believe in March?” The memory already records when a fact was replaced. The ledger would tie decisions to their sources.
Memory that works while you sleep. Merging scattered notes into short insights that link back to their sources. We will test it on a copy first, because a model rewriting your memory is where trust gets lost.
One front door. One assistant decides who handles a request. The Desk modes stay as manual overrides.
The benchmark rerun. We are re-running our memory benchmarks and plan to publish full results with 2.0, including the categories where we need most improvements.
Phone and text. Only after we find the most private method to do it. Local-only means your computer has to be on.
A simpler start. Opening the app without a terminal command.
On the shelf: letting Vek see your screen through a small local vision model, strictly opt-in. We want the privacy rules written before the we start coding and testing how it works.
Where VEKTOR fits
None of this needed a bigger model. It needed measuring each step, keeping what held up on questions we had not tuned against, and dropping what did not.
Two of our best-sounding ideas lost to the baseline in the first test, leading to quicker iterations and pivots.
That is the method we plan to keep through 2.0: a small test, a yes or no, then move on. Memory, context, skills, MCP and the connectors give each test something solid to stand on, and we think that is why this stretch moved so fast and wil continue this way into the future.
If you run local models, we would like to hear what you would let an assistant do without asking, and what you would never want it to touch.
Those answers determine the default levels and where and what we build next.
Upgrading
Drop-in from any earlier version. Your memory database is untouched, and v1.9.9 includes everything from v1.9.8.
npm install -g ./vektor-slipstream-1.9.9.tgz
The full changelog is at vektormemory.com/docs/changelog. The forum is the fastest way to reach us with questions or feedback.
VEKTOR Memory builds local-first persistent memory infrastructure for AI agents. Documentation and downloads are at vektormemory.com.
Sources:
OpenAI dots launch coverage (androidheadlines.com/2026/09/openai-launches-dots-always-on-ai-agents.html), xAI Grok Bot (layer3labs.io/guides/what-is-grok-bot), Meta Muse Spark 1.1 (siliconangle.com/2026/07/09/meta-launches-flagship-muse-spark-1–1-model-multi-agent-upgrades), “BM25 Wins at Scale” (arxiv.org/abs/2607.26497), Levitt, List and Reiley, “What Happens in the Field Stays in the Field”, NBER working paper 15609 (nber.org/papers/w15609), Mem0 on temporal reasoning (mem0.ai/blog/introducing-temporal-reasoning-in-mem0), Cormack, Clarke and Büttcher, “Reciprocal Rank Fusion outperforms Condorcet and individual rank learning methods”, SIGIR 2009, Gao et al., “Precise Zero-Shot Dense Retrieval without Relevance Labels”, 2022.
AI Agent
LLM
Memory Improvement
Minmax


Top comments (0)