Definition

Embeddings are dense vector representations of text that capture semantic meaning. They are central to RAG systems, where documents are converted to vectors and compared by cosine similarity. In the LLM Wiki pattern, embeddings play a secondary role — the primary mechanism is structured markdown with explicit wikilinks rather than vector similarity.

Key Points

  • In traditional RAG, embeddings are the core retrieval mechanism: documents are chunked, embedded, and searched by similarity
  • In the LLM Wiki pattern, knowledge is organized through explicit wikilinks and structured pages — not vector proximity
  • Embeddings can still be useful in LLM Wiki for: routing queries to the right wiki, discovering potential new links, and initial clustering of sources
  • The Level Up Coding deep dive mentions embeddings as part of step 4 of the ingest pipeline (“Embed — updated page is re-embedded”)

Exact Search Beats Approximate Search at Vault Scale

An approximate-nearest-neighbour (ANN) index is usually treated as the obvious upgrade once a vector store grows. local-knowledge-graph measured that assumption instead of inheriting it, at 768 dimensions:

VectorsExact scan (numpy)Annoy queryAnnoy build, per insert
1000.007 ms0.031 ms3.5 ms
1 0000.017 ms0.032 ms36 ms
10 0000.30 ms0.031 ms366 ms
100 0003.3 ms0.032 ms4 020 ms

Three findings. An exact numpy scan costs 3.3 ms at a hundred thousand vectors, against an LLM call measured in seconds — the search is never the bottleneck. An Annoy index is immutable once built, so any application that inserts incrementally must rebuild the whole index per insert, which is worse than the exact scan at every size measured. And on annoy 1.17.3 with numpy 2.5.2 under Python 3.12, the index returned wrong results outright: get_nns_by_item(7, 5) gave one arbitrary row instead of five, omitting even the query vector itself. The binding constraint past a few hundred thousand vectors is reported to be memory rather than speed — roughly 300 MB resident at 100,000 × 768 × 4 bytes — for which the answer is a memory-mapped index, not a faster query.

These figures are author-reported and single-source, though the benchmark script is named (scripts/bench_search.py).

Implications: For this vault the conclusion is unambiguous — at ~140 pages an exact scan costs microseconds, and any ANN index in scripts/hybrid_query_helper.py would be pure complexity with a correctness risk attached. More usefully, the project pinned its refusal as a test that fails if a future build starts behaving, so the decision is revisited rather than inherited. That is a pattern worth copying for this vault’s own rejected optimisations, which currently live in DECISIONS.md as prose that cannot detect when its premise expires.

  • local-knowledge-graph — measures exact scan against ANN and rejects the index, with the numbers published
  • rag — RAG relies on embeddings for retrieval; LLM Wiki does not
  • retrieval-augmented-generation — Embeddings are the foundation of RAG retrieval
  • llm-wiki — LLM Wiki uses structured markdown instead of vector embeddings

Sources

^[raw/articles/karpathy-llm-wiki-gist.md] ^[raw/articles/levelup-llm-wiki-deep-dive.md]