punnerud/Local_Knowledge_Graph
Overview
Local_Knowledge_Graph builds a graph out of a local model’s reasoning trace rather than out of a document corpus. You ask a question, the model reasons step by step, and each step is drawn as a node with typed edges between them. Everything runs on the operator’s machine through Ollama. Published to PyPI as mpe-lkg by Morten Punnerud-Engelstad.
This inverts the direction every other graph tool in this vault runs in. lightrag, hipporag, and cognee all extract a graph from documents; this extracts one from an inference trace. The corpus is the model’s own thinking.
Key Facts
- Stars/forks: 547 stars, 48 forks, 26 commits, 1 open issue at capture (2026-08-12). Popularity is running well ahead of code history.
- Install and run:
pip install mpe-lkg, thenmpe-lkg; serves onhttp://localhost:5100. On startup it reports what is missing and the single command that fixes it. - Backend: Ollama for both chat and embeddings; SQLite for storage; Flask with server-sent events for the UI.
- Model sensitivity: on the project’s arithmetic battery,
qwen3:4b-instruct-2507answers 82.5% againstllama3.2:3b’s 40%. The README states plainly that “the model matters more than anything else here.” - Export: any run is retrievable as RDF via
GET /jobs/<id>/rdf, which makes runs portable into other graph tooling. - Licence: “The mpedb License 1.0” — not OSI-approved. Free for every person and organisation except groups above five billion dollars in revenue or valuation, who owe a one-time fee of seven US cents per device. This is deliberately not tagged
open-sourcein this vault; the tag would misdescribe it.
Two Edge Colours, Two Epistemic Statuses
The graph distinguishes how strongly it believes each connection, and shows the difference:
- Blue edges — embedding similarity. Association. Two steps landed near each other in vector space.
- Green edges — settled by an exact evaluator. Sums and unit conversions computed in fractions via the author’s companion library
mpeqs.
A reader can therefore see at a glance which links are merely plausible and which were checked by arithmetic.
Implications: This is the same distinction attested-computation draws in the OKF v0.2 spec — a value that was computed the sanctioned way versus one that merely looks right — rendered as a visual property of a graph rather than as a verification verdict in a bundle. It is the cheapest expression of that idea this vault has seen: no attester, no receipt, just a colour that means “an evaluator settled this.”
Reasoning Modes
Three modes escalate how much independent agreement is required before an answer is accepted:
reason— a single step-by-step pass.explore— each sub-question gets its own run.settle— the question is explored twice and finished only when two independent runs agree.
Implications: settle is self-consistency applied as a stopping condition rather than as a post-hoc vote — the run does not end until agreement exists. That makes verification part of the control flow, not a filter applied afterwards, which is the same move verifiability identifies as what makes a task automatable at all.
Strongest Path
The strongest path between two nodes maximises the product of similarities along it. That is equivalent to minimising a sum of -log(similarity); because those costs are non-negative, Dijkstra returns the exactly optimal path rather than an approximation. The score reported for a path is the geometric mean of its edges.
Implications: A named, exactly-solvable path metric is a stronger primitive than “related pages” — it gives a defensible answer to how two ideas connect and how confidently, not just that they do. This vault’s graph_neighbors and community detection answer proximity questions; nothing here answers path questions with a score attached.
The Refusal to Add an ANN Index
The project measured the case for an approximate-nearest-neighbour index and published the numbers rather than the conclusion alone. At 768 dimensions, via scripts/bench_search.py:
| Vectors | Exact scan (numpy) | Annoy query | Annoy build, per insert |
|---|---|---|---|
| 100 | 0.007 ms | 0.031 ms | 3.5 ms |
| 1 000 | 0.017 ms | 0.032 ms | 36 ms |
| 10 000 | 0.30 ms | 0.031 ms | 366 ms |
| 100 000 | 3.3 ms | 0.032 ms | 4 020 ms |
Three conclusions follow. The exact scan is already fast enough at any plausible size — 3.3 ms at a hundred thousand vectors, against an LLM call measured in seconds. An Annoy index cannot be appended to: it is immutable once built, and this application inserts after every reasoning step, so the entire index must be rebuilt each time, which is worse than the exact scan at every size measured. And on annoy 1.17.3 with numpy 2.5.2 on Python 3.12, the index returns wrong answers — get_nns_by_item(7, 5) returns [1], one arbitrary row instead of five, omitting even the query vector itself, which must always be its own nearest neighbour at distance zero. The consequence was that the “Related Questions” panel had been silently returning a single arbitrary row.
The refusal is pinned as a test that fails if a future build ever starts behaving correctly, so the decision gets revisited rather than inherited. tests/test_store.py asserts exactness directly: a vector is its own nearest neighbour, and the ranking matches a full brute-force sort. The author also names what would change the decision first — at a few hundred thousand vectors the binding constraint becomes memory (100,000 × 768 × 4 bytes ≈ 300 MB resident), and the answer then is a memory-mapped index, not a faster query.
Implications: This is a model for how a rejected optimisation should be recorded. The vault’s own LESSONS.md and DECISIONS.md carry decisions without the measurements that justified them; a decision pinned to a test that fires when its premise expires cannot rot into cargo cult the way a prose rationale can. The concrete finding also matters directly — this vault has embedding infrastructure in scripts/hybrid_query_helper.py, and at its page count an exact scan is unambiguously the right choice.
Relationships
- “implements” knowledge-graph — builds one from a reasoning trace rather than from documents
- “uses” embeddings — similarity edges and exact nearest-neighbour search over 768-dimensional vectors
- “differs-from” lightrag — LightRAG graphs a corpus; this graphs an inference trace
- “differs-from” rag — nothing is retrieved at query time; the graph is the reasoning itself
- “references” attested-computation — its green edges express the same verified-versus-associative distinction
Open Questions
- The 82.5%-versus-40% arithmetic figure and the search benchmark are author-reported and single-source. The search bench at least names its script (
scripts/bench_search.py); the arithmetic battery is not described anywhere in the captured sources. - The
annoy+ numpy 2.5.2 defect is a specific, plausible claim about third-party software from one environment combination, recorded here as reported and not independently verified. - Does the RDF export carry the blue/green edge distinction, or does it flatten both into untyped similarity edges? The captured sources do not say, and the answer decides whether the verified/associative split survives leaving the tool.
- 547 stars against 26 commits is an unusual ratio. Whether the project is early or simply compact is not determinable from these sources.
External Links
- https://github.com/punnerud/Local_Knowledge_Graph
- https://raw.githubusercontent.com/punnerud/Local_Knowledge_Graph/main/docs/design.md
- https://github.com/punnerud/MPEqs
- https://pypi.org/project/mpe-lkg/
Sources
^[raw/external/github-com-local-knowledge-graph-0793437a.md]