diff --git a/Graph-Agents.md b/Graph-Agents.md index ebe13ee..040e3e9 100644 --- a/Graph-Agents.md +++ b/Graph-Agents.md @@ -585,6 +585,14 @@ base is first built): (default: `1`). `0` returns only documents directly linked to matched entities with no traversal. Falls back to `rag_graph_hops`. Higher values surface more loosely related documents; keep at `1` for most corpora. - **`batch_size`:** Embedding-request batch size. +- **`driver`:** [Storage driver](RAG#storage-drivers) for this node's knowledge base — `yaml` + (default) or `duckdb`. Use `duckdb` when the node's corpus is large enough that rebuilding + the in-memory index on every agent start is noticeable; it keeps vectors and text on disk in + a `.duckdb` file beside the generated knowledge base. An invalid driver is + rejected when the graph is validated, naming the offending node. Like every field in this + section it only applies when the knowledge base is *first* built — changing it later has no + effect until the knowledge base is deleted and rebuilt. The `qdrant` driver is not available + here: attaching a remote collection is an interactive flow (`.rag attach`). Each falls back to the app-level `rag_*` config when omitted. **When `embedding_model`, `chunk_size`, and `chunk_overlap` are all set, the diff --git a/RAG.md b/RAG.md index 601817f..304a02b 100644 --- a/RAG.md +++ b/RAG.md @@ -69,9 +69,15 @@ When you define RAG, Coyote will first "build" the RAG. This means that Coyote w generate [embeddings](https://huggingface.co/spaces/hesamation/primer-llm-embedding) for that text. This essentially just means that Coyote translates the document into a language the LLM can understand. -These embeddings are stored in an in-memory vector database. Coyote also indexes every document chunk in a full-text +These embeddings are stored in a vector database — in memory, on disk, or in a remote Qdrant collection, depending on +the RAG's [storage driver](#storage-drivers). Coyote also indexes every document chunk in a full-text search index (BM25) for keyword-based retrieval. +Building costs real time and real money: every chunk is sent to an embedding model, and every chunk is sent again to a +chat model if graph RAG is enabled. Adding or editing documents with `.edit rag-docs` only processes what actually +changed, but `.rebuild rag` deliberately re-embeds the entire corpus from scratch — reach for it when a document's +contents changed underneath Coyote or the knowledge base looks wrong, not as routine maintenance. + If an [extractor model is configured](#graph-based-rag), Coyote additionally runs an LLM-based entity and relationship extraction pass over each document chunk and builds a **knowledge graph** of named entities and their connections. This graph is saved alongside the vector and BM25 indexes and enables a third, graph-based retrieval signal at query time. @@ -114,6 +120,95 @@ below for more details on how to customize this. Finally, the text snippets that were looked up in RAG are passed to the model as additional context to your prompt, giving the model query-specific context to answer your question. +# Storage Drivers +Every RAG picks a **storage driver** when it's created. The driver decides where your vectors and document text +actually live: inside the RAG's own file, in a local database beside it, or in a Qdrant server you already run. + +| Driver | Where the vectors live | Reach for it when | +|--------|------------------------|-------------------| +| `yaml` *(default)* | In the RAG's own YAML file, loaded into memory on use | You want the simplest thing that works, or several Coyote processes need the same RAG at once | +| `duckdb` | A `.duckdb` file beside the RAG | The corpus is large enough that re-loading it into memory on every start is annoying | +| `qdrant` | A collection on a Qdrant server you already run | Something else already built and maintains the collection | + +Coyote asks which one you want while creating a RAG with `.rag`: + +```text +? RAG storage driver: +> yaml — portable, in-memory HNSW; usable from several Coyote processes at once (default) + duckdb — persistent on-disk store; vectors and content survive restarts; HNSW approximate search. +``` + +If you don't care, take the default. `yaml` is what Coyote has always done. + +> **A RAG's driver is fixed when it's created.** There is no conversion step — switching means deleting the RAG and +> building it again, which re-embeds every document. Worth a moment's thought before you pick, but not worth agonizing +> over: for small corpora the difference is not something you'll notice. + +## yaml +The default. Vectors and document text are stored in the RAG's YAML file and read into memory when the RAG is loaded, +where the HNSW index is built on the fly. + +Its useful property is that nothing holds a lock: as many Coyote processes as you like can use the same RAG +simultaneously. Its cost is startup — a large corpus means a large file to parse every time the RAG is loaded. + +## duckdb +Vectors, document text and the search indexes live in a `.duckdb` file next to the RAG, and stay there between +runs. Nothing is rebuilt at startup, so loading a large RAG is quick regardless of size. + +Two things are worth knowing before you choose it: + +* **Concurrency is readers-or-writer.** Several Coyote processes can query the same duckdb RAG at once. While one + process is *ingesting* or *rebuilding* it, though, the others can't read it until that finishes. If you routinely + run several Coyotes against one RAG and rebuild it often, `yaml` will annoy you less. +* **First use needs network access.** DuckDB's vector (`vss`) and full-text (`fts`) extensions are downloaded the + first time a duckdb RAG is created, then cached and reused. On a machine that's offline from the start, create one + duckdb RAG while connected and every later one works offline. Pre-installing them from a DuckDB shell is fiddlier + than it sounds — the cache is keyed by DuckDB version, and Coyote bundles its own — so connecting once is much the + easier route. + +Deleting the RAG removes the `.duckdb` file and its write-ahead log along with it. + +## qdrant — attaching an existing collection +The `qdrant` driver connects a RAG to a collection on a Qdrant server you already run. It is **read-only**: Coyote +queries the collection and never writes to it, so whatever populated it stays the sole owner of its contents. This is +the driver to use when another system — a LangChain pipeline, an ingestion job, a colleague's tooling — already +maintains a collection you'd like to ask questions about. + +Because the collection already exists, you don't create this kind of RAG with `.rag`; you *attach* it: + +```shell +.rag attach my-docs +``` + +Coyote then walks you through it: + +1. **Host** — for example `qdrant.company.com:6333`. +2. **API key** — Coyote asks whether the instance needs one, then which [Vault](Vault) secret holds it (defaulting to + `QDRANT_API_KEY`). If that secret doesn't exist yet you can create it right there, without leaving the wizard. The + key itself is never written into the RAG file; only the secret's name is. +3. **Collection** — Coyote connects, lists what's available, and lets you pick. + +It then checks the collection is actually usable and tells you what it found — the vector dimension, and a warning if +the collection is empty (you can attach anyway, but every query will return nothing until something writes to it). + +A collection has to meet two requirements: + +* **A single unnamed vector.** Named or multi-vector collections are rejected, because Coyote queries with one unnamed + vector and Qdrant would reject every request. +* **Document text in a `page_content` payload field.** This is the field Coyote reads answers out of. Collections + built by LangChain use this name by default. + +Point IDs may be integers or UUIDs; both work. + +> **Use the same embedding model that built the collection.** Coyote can't detect a mismatch — the dimensions may line +> up perfectly while the vectors mean entirely different things. The symptom is retrieval that returns confident +> nonsense, so it's worth checking rather than discovering later. + +### Attached RAGs inside a sandbox +If you run Coyote in a [sandbox](Sandboxes), an attached RAG's host is allow-listed and its API key is handed to the +sbx proxy automatically, so queries work inside the sandbox exactly as they do outside. See +[Sandboxes > Credentials & Secrets](Sandboxes#credentials--secrets). + # Supported Document Sources Coyote supports a number of document sources that can be used for RAG: diff --git a/REPL.md b/REPL.md index eab7766..030b9e5 100644 --- a/REPL.md +++ b/REPL.md @@ -148,8 +148,9 @@ complete tasks using the documents as additional context. | Command | Description | |------------------|------------------------------------------------------------------------------| | `.rag` | Initialize or access a RAG | +| `.rag attach ` | Attach a RAG to an existing remote [Qdrant](RAG#qdrant--attaching-an-existing-collection) collection (read-only) | | `.edit rag-docs` | Add or remove documents from the active RAG using your preferred text editor | -| `.rebuild rag` | Rebuild the active RAG to accommodate document changes | +| `.rebuild rag` | Rebuild the active RAG from scratch, re-embedding every document | | `.sources rag` | Show a works-cited of the sources used in the last query | | `.info rag` | Display information about the active RAG | | `.exit rag` | Exit the active RAG | diff --git a/Sandboxes.md b/Sandboxes.md index 584df75..d803bf3 100644 --- a/Sandboxes.md +++ b/Sandboxes.md @@ -180,6 +180,22 @@ Secrets already registered with `sbx` are silently skipped to keep re-attaches f If your LLM client uses **OAuth** instead of an API key (`auth: oauth` in your config), no secret is injected. The sbx proxy handles OAuth natively without a stored key. +### Attached RAG collections + +A RAG [attached to a remote Qdrant collection](RAG#qdrant--attaching-an-existing-collection) needs two things to work +inside a sandbox, and Coyote arranges both without being asked: the collection's host has to be reachable through the +network policy, and the API key has to reach the server without ever entering the VM. + +When you attach a RAG, Coyote writes a small mixin next to it declaring the host as an allowed domain and the API key +as a proxy-managed credential. On launch that mixin is discovered alongside the generated `coyote-mcp` one, the key is +registered with the sbx secret store on your host, and the proxy writes it into the outgoing request header at the +network edge — the same treatment LLM provider keys get. Queries then behave identically inside and outside the +sandbox. + +This is also why a RAG's `api_key` must be a `{{SECRET_NAME}}` placeholder rather than the key itself: Coyote reads +that placeholder back out to learn which vault secret to hand to sbx. A literal key can't be provisioned, so Coyote +refuses to load a RAG configured that way and tells you which secret to create. + ### The generated `coyote-mcp` mixin Alongside secret registration, every non-`--fresh` launch renders a mixin kit named `coyote-mcp` and diff --git a/_Sidebar.md b/_Sidebar.md index e52d82b..962b411 100644 --- a/_Sidebar.md +++ b/_Sidebar.md @@ -57,6 +57,9 @@ - [File Discovery](Workspace-Instructions#file-discovery) - [Configuration](Workspace-Instructions#configuration) - [RAG](RAG) + - [Storage Drivers](RAG#storage-drivers) + - [Attaching a Qdrant Collection](RAG#qdrant--attaching-an-existing-collection) + - [Graph-Based RAG](RAG#graph-based-rag) - [Macros](Macros) - [Roles](Roles) - [Skills](Skills)