From 669187d1483be74ded2d114f8fd1f35e080a52a4 Mon Sep 17 00:00:00 2001 From: Alex Clarke Date: Mon, 13 Jul 2026 17:00:22 -0600 Subject: [PATCH] docs: Updated docs to mention features about graph-based RAG --- Environment-Variables.md | 3 + Graph-Agents.md | 30 +++++++--- RAG.md | 117 +++++++++++++++++++++++++++++++++++---- 3 files changed, 133 insertions(+), 17 deletions(-) diff --git a/Environment-Variables.md b/Environment-Variables.md index 4de4f90..9f7546d 100644 --- a/Environment-Variables.md +++ b/Environment-Variables.md @@ -37,6 +37,9 @@ Below are the most commonly used configuration settings and their corresponding | `rag_top_k` | `COYOTE_RAG_TOP_K` | | `rag_chunk_size` | `COYOTE_RAG_CHUNK_SIZE` | | `rag_chunk_overlap` | `COYOTE_RAG_CHUNK_OVERLAP` | +| `rag_extractor_model` | `COYOTE_RAG_EXTRACTOR_MODEL` | +| `rag_extractor_prompt` | `COYOTE_RAG_EXTRACTOR_PROMPT` | +| `rag_graph_hops` | `COYOTE_RAG_GRAPH_HOPS` | | `highlight` | `COYOTE_HIGHLIGHT` | | `theme` | `COYOTE_THEME` | | `serve_addr` | `COYOTE_SERVE_ADDR` | diff --git a/Graph-Agents.md b/Graph-Agents.md index 85d6e78..9cbf6b2 100644 --- a/Graph-Agents.md +++ b/Graph-Agents.md @@ -521,10 +521,11 @@ Agent nodes (which spawn full sub-agents) intentionally have no ## rag -Runs a hybrid (vector + keyword) retrieval against a per-node knowledge base -and writes the result into state. This is how a graph agent does -Retrieval-Augmented Generation: the `rag` node retrieves context, downstream -`llm`/`agent` nodes inject it into their prompts via normal templating. +Runs a hybrid (vector + full-text + optional graph) retrieval against a +per-node knowledge base and writes the result into state. This is how a +graph agent does Retrieval-Augmented Generation: the `rag` node retrieves +context, downstream `llm`/`agent` nodes inject it into their prompts via +normal templating. ```yaml research_context: @@ -569,6 +570,18 @@ base is first built): - **`chunk_size`:** Document chunk size. - **`chunk_overlap`:** Overlap between chunks. - **`reranker_model`:** Reranker applied to hybrid-search results. +- **`extractor_model`:** Chat model for graph-based entity/relationship extraction. + When set, a knowledge graph is built at index time and used as an additional + retrieval signal alongside vector and BM25. Falls back to the global + `rag_extractor_model` config when omitted. See [RAG > Graph-Based RAG](RAG#graph-based-rag) + for guidance on model selection. +- **`extractor_prompt`:** Custom extraction prompt template. Must contain a `__CHUNK__` + placeholder. Falls back to `rag_extractor_prompt` then the built-in prompt. Useful for + domain-specific entity types (e.g. legal, medical, code). See + [RAG > Custom extraction prompt](RAG#custom-extraction-prompt). +- **`graph_hops`:** Number of graph hops to expand from matched entities at query time + (default: `1`). `0` returns only documents directly linked to matched entities with no traversal. + Falls back to `rag_graph_hops`. Higher values surface more loosely related documents; keep at `1` for most corpora. - **`batch_size`:** Embedding-request batch size. Each falls back to the app-level `rag_*` config when omitted. **When @@ -615,9 +628,12 @@ inspected, not run, so knowledge-base building is skipped entirely.) ### Retrieval -Retrieval at execution time is fast (no re-embedding of the corpus). It's -the same hybrid vector + keyword search normal Coyote RAG uses. The corpus -embedding/chunking cost is paid once, at load time. +Retrieval at execution time is fast (no re-embedding of the corpus). The corpus +embedding/chunking cost is paid once, at load time. The retrieval strategy +matches normal Coyote RAG: vector + full-text (BM25) signals are always active, +and a third graph-based signal is added when an `extractor_model` is set at +build time. All active signals are fused via RRF; if a `reranker_model` is set, +it replaces RRF and graph search is not applied. --- diff --git a/RAG.md b/RAG.md index 54c89aa..601817f 100644 --- a/RAG.md +++ b/RAG.md @@ -2,7 +2,8 @@ Retrieval Augmented Generation (RAG) is a method of minimizing LLM hallucination without consuming a significant portion of the context length. It uses documents and other additional resources that you provide to give the model more context for all of your prompts. -Coyote has a built-in vector database and full-text search engine to support RAG knowledge bases for your queries. +Coyote has a built-in vector database, full-text search engine, and optional knowledge graph to support RAG knowledge +bases for your queries. At query time these signals are fused together to maximize retrieval quality. The generated knowledge bases are stored in the `rag` subdirectory of your Coyote configuration directory. The location of this directory varies by system, so you can use the following command to find your RAG directory: @@ -68,14 +69,34 @@ When you define RAG, Coyote will first "build" the RAG. This means that Coyote w generate [embeddings](https://huggingface.co/spaces/hesamation/primer-llm-embedding) for that text. This essentially just means that Coyote translates the document into a language the LLM can understand. -These embeddings are then stored in an in-memory vector database. +These embeddings are stored in an in-memory vector database. Coyote also indexes every document chunk in a full-text +search index (BM25) for keyword-based retrieval. + +If an [extractor model is configured](#graph-based-rag), Coyote additionally runs an LLM-based entity and relationship +extraction pass over each document chunk and builds a **knowledge graph** of named entities and their connections. This +graph is saved alongside the vector and BM25 indexes and enables a third, graph-based retrieval signal at query time. ### 2. Lookup -Coyote sits between you and the model. So when you submit a prompt to the model, before Coyote ever sends it, it will first -convert your prompt into embeddings (LLM language), and look for relevant snippets of text in the vector database. +Coyote sits between you and the model. So when you submit a prompt to the model, before Coyote ever sends it, it runs a +**hybrid search** over your knowledge base using two (or three) complementary signals: -Coyote then passes the top `n`-snippets of text that it finds in the vector database as additional context to the model -before your prompt. +- **Vector search**: Your query is converted to embeddings and matched against the document embeddings using cosine + similarity (HNSW index). +- **Full-text search (BM25)**: A keyword-based search that finds documents containing the same terms as your query. +- **Graph search** *(only when [graph-based RAG](#graph-based-rag) is enabled)*: Entities mentioned in your query are + looked up in the knowledge graph. Coyote expands up to `rag_graph_hops` hops (default: 1) from matched entities, + scoring neighbors by edge weight and query relevance, then returns documents linked to the highest-scoring entities. + +The results from all active signals are merged via [Reciprocal Rank Fusion (RRF)](https://www.elastic.co/docs/reference/elasticsearch/rest-apis/reciprocal-rank-fusion), +giving slightly more weight to semantic similarity: + +| Signal | RRF Weight | +|------------------|------------| +| Vector (HNSW) | 1.125 | +| Full-text (BM25) | 1.0 | +| Graph | 0.9 | + +Coyote then passes the top `n` merged results as additional context to the model before your prompt. ### 2a. Reranking (Optional) The lookup for relevant snippets of texts uses embeddings to find text that is semantically similar to your prompt, and @@ -163,15 +184,91 @@ rag_embedding_model: null # Specifies the embedding model used for contex ``` ## Reranker -By default, Coyote uses [Reciprocal Rank Fusion (RRF)](https://www.elastic.co/docs/reference/elasticsearch/rest-apis/reciprocal-rank-fusion) to merge vector and keyword search results. +By default, Coyote uses [Reciprocal Rank Fusion (RRF)](https://www.elastic.co/docs/reference/elasticsearch/rest-apis/reciprocal-rank-fusion) +to merge results from all active retrieval signals (vector, BM25, and graph when enabled). See [How It Works](#how-it-works) for the exact weights. -You can change the default reranker model to any other reranking model in your configured clients. To change the default -reranker model, simply change the value of the `rag_reranker_model` setting in your global configuration file: +When a reranker model is set it replaces RRF: the union of vector and BM25 candidates is re-scored by the reranker for +direct query relevance, and graph-based search is not applied. You can change the default reranker model to any +reranking model in your configured clients: ```yaml -rag_reranker_model: null # By default, +rag_reranker_model: null # Reranker model; when set, replaces RRF (graph search is not applied) ``` +## Graph-Based RAG + +When Coyote builds a knowledge base, it can optionally run an LLM-based entity and relationship extraction pass +over each document chunk to construct a **knowledge graph**. At query time, this graph becomes a third retrieval +signal alongside vector and full-text search. + +To enable graph-based RAG, set `rag_extractor_model` to any chat model in your configured clients: + +```yaml +rag_extractor_model: null # Chat model for entity/relationship extraction; enables graph RAG when set +``` + +When set, Coyote extracts named entities and their relationships from each chunk at build time and stores a +knowledge graph alongside the vector and BM25 indexes. At query time, entities in your query are matched +against the graph and neighbors are expanded up to `rag_graph_hops` hops to surface documents linked to those +entities. This graph signal is then fused into the hybrid search result via [Reciprocal Rank Fusion](#2-lookup). + +> **Graph-based RAG and reranking are mutually exclusive.** If a `rag_reranker_model` is also set, the +> reranker replaces RRF entirely and the graph signal is not applied. See [Reranker](#reranker) for details. + +### Graph expansion depth + +`rag_graph_hops` controls how many hops to expand from matched entities at query time (default: `1`): + +```yaml +rag_graph_hops: 1 # 0 = seed nodes only; 1 = direct neighbors; 2 = neighbors of neighbors; etc. +``` + +- **`0`:** Returns only documents directly linked to query-matched entities; no graph traversal. +- **`1` (default):** Expands to entities directly connected to query matches. Good for most corpora. +- **`2+`:** Traverses further into the graph, surfacing more loosely related documents. Useful for dense, + highly interconnected knowledge bases (e.g. ontologies, large technical wikis). May increase noise on + sparse corpora. + +### Custom extraction prompt + +By default, Coyote uses a built-in prompt that extracts entities of types `PERSON`, `ORGANIZATION`, `CONCEPT`, +`TECHNOLOGY`, `LOCATION`, `EVENT`, and `OTHER`. For domain-specific corpora you can override this with +`rag_extractor_prompt`: + +```yaml +rag_extractor_prompt: null # Custom extraction prompt; must contain __CHUNK__ placeholder +``` + +The prompt must contain the literal string `__CHUNK__`, which Coyote replaces with the document chunk at +extraction time. The response must be a JSON object with `entities` and `relationships` arrays in the same +structure the built-in prompt produces. Example custom prompt for a legal corpus: + +``` +Extract legal entities and relationships from the following text. + +Return JSON: +{ + "entities": [{"name": "...", "type": "STATUTE|CASE|PARTY|COURT|CONCEPT", "description": "..."}], + "relationships": [{"from": "...", "to": "...", "type": "cites|governs|decided_by", "weight": 0.9}] +} + +Only extract what is clearly stated. Return ONLY the JSON object. + +Text: +__CHUNK__ +``` + +### Extractor model guidance + +- **Use a fast, cheap chat model.** e.g. `anthropic:claude-haiku-4-5` or `openai:gpt-4o-mini`. Extraction + runs once per chunk at build time, so speed and cost matter more than raw capability. +- **Graph-based RAG is most useful** for knowledge bases with rich entity relationships: technical documentation, + research papers, wikis. For small corpora or plain prose, vector + BM25 alone is usually sufficient. +- **Individual `rag` nodes in graph agents can override this** with their own `extractor_model`, `extractor_prompt`, + and `graph_hops` fields. See [Graph-Agents](Graph-Agents#rag) for details. +- The extractor model is prompted interactively when you create a new RAG knowledge base via `.rag`. If you skip + it, the knowledge base uses vector + full-text search only (no graph). + ## Chunk Size In the context of RAG, the chunk size is the maximum length of each text chunk (measured in characters) that is created when splitting documents. In Coyote, this defaults to `2000` characters.