Overview

Graphify is an open-source (MIT) skill that maps any folder — code, docs, PDFs, images, videos, SQL schemas, scripts — into a queryable knowledge graph you traverse instead of grepping files. You type /graphify . in an AI coding assistant (Claude Code, Codex, OpenCode, Cursor, Gemini CLI, 15+ hosts) and it builds the graph locally. It is the most-starred project in the vault’s ecosystem (85.8k⭐ on GitHub; the site reports 100 repos graphified with 854k nodes / 1.9M edges).

Key Facts

  • License: MIT (free for personal + commercial)
  • Install: pip install graphifyy (note: two y’s) or uv tool install graphifyy; then graphify install registers the skill
  • Stack: 100% Python; tree-sitter AST parsing (deterministic, no LLM) for code; NetworkX for the graph; Leiden community detection (no embeddings, no vector store)
  • Code parsing: 36 tree-sitter grammars covering Python, TypeScript, JavaScript, Go, Rust, Java, C/C++, CUDA, Metal, Ruby, C#, Kotlin, Scala, PHP, Swift, Lua, Elixir, shell, JSON, Scala, and many more
  • Output: three files — graph.html (interactive, clickable), GRAPH_REPORT.md (highlights + surprising connections + suggested questions), graph.json (full graph, queryable)
  • Graph features: god nodes (most-connected concepts), communities (Leiden clustering), cross-file links (calls/imports/inherits resolved), query/path/explain capabilities, rationale + doc refs as first-class nodes
    • graphify explain <x> — explain a node, listing all edges with EXTRACTED/INFERRED tags
    • graphify path A B — shortest path between two nodes
    • graphify query "<question>" — scoped subgraph for a plain-language question
  • Edge transparency: every edge tagged EXTRACTED (explicit in source), INFERRED (resolved by graphify), or AMBIGUOUS — you can tell read-directly from inferred; graphify explain shows the confidence per edge
  • MCP server: python -m graphify.serve graph.json — stdio by default, optional Streamable HTTP transport for team sharing with --transport http, API key support, and per-session state management; 10 tools (query_graph, get_node, get_neighbors, shortest_path, get_community, god_nodes, graph_stats, list_prs, get_pr_impact, triage_prs), flags --transport {stdio,http}, --host (default 127.0.0.1), --port 8080, --api-key
  • Local-first / privacy: code parsed locally with tree-sitter (deterministic AST, 0 LLM credits, nothing leaves machine during extraction); only the optional semantic pass over docs/media calls a backend, and only if you configure one (ANTHROPIC_API_KEY/OPENAI_API_KEY/etc). Supports multiple backends: OpenAI, Anthropic, Gemini, DeepSeek, Bedrock, Azure, Ollama
  • Beyond code: docs (.md wikilinks become references edges), PDFs, Office, Google Workspace, images, video/audio, YouTube/URLs; rationale markers # NOTE: / # WHY: / # HACK: and ADR/RFC citations become first-class nodes
  • Benchmarks: evaluated as long-term memory (LOCOMO recall@10: 0.497, QA accuracy: 45.3% vs supermemory 49.7%/mem0 27.3%) and as code-intelligence layer
  • Always-on integration: strict mode (graphify claude install --strict, toggled via GRAPHIFY_HOOK_STRICT=1/0) blocks first raw source read and redirects to graph; soft nudge mode (graphify install) fires once per session; hook platforms (Claude Code, Gemini CLI) vs instruction-file platforms (Codex, OpenCode, Cursor via AGENTS.md/.cursor/rules/graphify.mdc)
  • Automation: graphify hook install embeds interpreter path and sets up git-aware merge driver so graph.json never has conflict markers
  • Install notes: PyPI package is graphifyy (double-y; avoid graphify* impostors); prefer uv tool install graphifyy (isolated env) over plain pip; skill lands at ~/.claude/skills/graphify/SKILL.md (user) or .agents/skills/graphify/SKILL.md (project); PowerShell: use graphify . not /graphify .; optional extras [pdf], [office], [google], [video]/[youtube], [terraform], [dm]

Commands

/graphify .                       build graph for current folder
/graphify . --update              re-extract only changed files
/graphify . --cluster-only        rerun clustering without re-extracting
/graphify . --cluster-only --resolution 1.5   # more granular communities
/graphify . --no-viz              skip HTML, just report + JSON
/graphify . --wiki                build a markdown wiki from the graph
graphify query "what connects auth to the database?"
graphify path "UserService" "DatabasePool"
graphify explain "RateLimiter"
graphify hook install             auto-rebuild on git commit
graphify prs                      PR dashboard: CI state, review status, impact
graphify prs 42 --triage          AI ranks review queue; --conflicts flags merge risk
/graphify add <paper-url>         fetch a paper and add it

graphify prs maps open PRs onto the graph, highlights overlapping nodes and pairs carrying merge risk.

Relationships

  • “belongs-to” knowledge-graph — Graphify is a production knowledge-graph builder; closest in spirit to cognee and swarmvault for the explicit-graph layer
  • “operates-with” llm-wiki — it can run over a wiki/ folder, but builds ITS OWN graph from the files; it does not read OKF frontmatter or wikilinks as native source
  • “operates-with” model-context-protocol — exposes an MCP server (stdio + optional HTTP)
  • “sibling-of” cognee — sibling explicit-graph option; Cognee is memory-engine (Neo4j/Postgres + 14-tool MCP), Graphify is codebase→graph skill with edge transparency
  • “sibling-of” rightmemory — sibling typed-edge approach for coding agents
  • “sibling-of” understand-anything — sibling codebase-comprehension graph with business-domain explanations
  • “belongs-to” agent-memory-systems — appears in benchmarks as a long-term-memory layer
  • “sibling-of” wikilinks — Graphify’s EXTRACTED/INFERRED edge tags mirror this vault’s rule that inferred links must be visually distinct from manual ones

Implications

Graphify is the strongest explicit-graph / navigation option for the vault’s Fase C (gap #3 explicit graph, gap #1 navigation at scale). Two caveats for OUR vault:

  1. Format: it builds its own graph.json from raw files — it does not import our OKF frontmatter or wikilinks. So it would be a parallel graph layer, not an extension of the OKF artifact. (Same trade-off we noted for cognee and nashsu-llm-wiki.)
  2. Best fit: over a CODEBASE or a large wiki/ corpus (100+ pages) where index.md grep breaks. For our current 75-page vault, wikilinks still suffice (per connection-methods and the comprehensive report’s “no graph before index.md breaks” rule).

The EXTRACTED/INFERRED distinction is exactly the discipline this vault already enforces for inferred connections — Graphify makes that visible at the edge level, which is why it is worth ingesting despite the format mismatch.

Sources

^[https://github.com/Graphify-Labs/graphify] ^[https://graphify.net/]