Summary
From the article:
Over the past few months I have been playing around quite a bit with LLM agents, particularly for vulnerability research.
[...]
If an observation changes, I don’t want the model to reconstruct the entire investigation from a transcript and hopefully notice all of the consequences. I want the affected conclusions to become invalid automatically.
When looking at the problem from this perspective, I started wondering why we were making the LLM reconstruct its entire state over and over again.
[...]
This eventually turned into Lemmalog.
[...]
This means that the LLM is still responsible for understanding natural language, source code, debugger output and all the other messy information that appears during an investigation.
LLMs happen to be quite good at this.
But once that information has been converted into structured facts, we no longer need the model to repeatedly determine all of its consequences. The database can do that instead.
[...]
Because Lemmalog already tracks the dependencies of derived facts, we can ask it for the provenance of a conclusion.
[...]
This was originally mostly necessary to make incremental evaluation work correctly, but it turns out that being able to ask an AI agent why it believes something is quite useful as well :)
[...]
At some point it became fairly obvious that I had approached the problem like a static analysis engine without intentionally meaning to.
[...]
The engine itself now supports incremental evaluation, retractions, provenance, temporal facts, aggregations, entity reconciliation, hybrid retrieval, demand-driven queries and a bunch of other things that I probably added because implementing Datalog features is more fun than I expected.
[...]
So Lemmalog currently sits third among the dedicated memory systems in this comparison, behind PropMem and OpenClaw.
If we count throwing the entire conversation into the prompt as a memory system, it is fourth.
[...]
The second result I like is adversarial questions.
[...]
These questions deliberately contain false or misattributed premises.
For example, the conversation may contain a story about somebody receiving a gift, followed by a question which attributes the same gift to somebody else.
A language model with a giant transcript is rather tempted to find the semantically similar story and answer anyway. A structured memory can instead notice that there is simply no supporting fact about the person in the question.
[...]
At some point the full-context version doesn’t merely become expensive.
It stops fitting in the context window.
[...]
Lemmalog is already competitive with dedicated LLM memory systems, substantially outperforms full context on some of the tasks it was designed for, and does so while giving the reader a tiny fraction of the original history.
[...]
The source code for Lemmalog is available here.