Skip to content

Knowledge Bases

A KB is a curated corpus the agent can search and read at runtime via kb_search_pages, kb_list_pages and kb_read_page. Where a skill teaches the agent how to do something, a KB tells it what it knows.

For the operator’s view, see Work mode → Knowledge bases.

Creating a knowledge base: name, slug, description and the curator model

Source files (markdown, PDF, CSV)
│
▼ uploaded via POST /api/knowledge-bases/{id}/sources
┌─────────────────────────────┐
│ Curator (an LLM) │
│ Reads sources, extracts │
│ structured knowledge, │
│ deduplicates │
└─────────────┬───────────────┘
│ COMPILE (POST /api/knowledge-bases/{id}/compile)
▼
┌─────────────────────────────┐
│ Wiki — searchable index │
│ (Hub-versioned) │
└─────────────┬───────────────┘
│
▼ kb_list_pages / kb_read_page at runtime
Agent's session context grows by relevant pages

The curator distills and links the material into a wiki. Hybrid search covers passages from both original sources and generated pages. Agents can read exact source passages, expand to the parent document, or follow a PDF section tree to its page ranges.

Field Notes
Name Unique within your org
Description Free-form
Curator model Locked at creation — can’t be changed later. Default works for most cases; pick a smarter model (Claude Sonnet 4.6, GPT-5) at creation for highly technical content.
Sources List of files / S3 prefixes / GitHub repos / Notion pages / URLs (only File works in current builds)
Status error (no successful compile yet) → compiling → active
Counters sources, wiki pages, attached agents

Published artifacts are versioned in Hub. Search results and reads refer to the same published content. A failed compile preserves the previous publication; changing a Hub branch alone does not roll back the search registry.

Five types are listed in ADD SOURCE. Only File is enabled in current builds — the rest are scaffolded but not yet wired.

Type Available Notes
File ✅ Upload .md, .pdf, .txt, .docx, .csv
S3 🚧 Recursive pull from a bucket/prefix
GitHub 🚧 Auto-pull from a repo’s docs folder with branch tracking
Notion 🚧 OAuth into a Notion workspace
URL 🚧 Crawl a domain or specific URLs

When the other source types ship, they’ll have continuous sync (re-compile on detected changes).

Locked at KB creation. Can’t be changed later. Common choices:

  • Surogate default — fine for most cases
  • Claude Sonnet 4.6 — better for highly technical content
  • GPT-5 — alternative

The curator runs once per source on COMPILE. The cost is dominated by the size of the source corpus.

Terminal window
curl -X POST http://localhost:8000/api/knowledge-bases/{id}/compile \
-H "Authorization: Bearer $TOKEN"

What happens:

  1. For each source: read, parse, chunk
  2. For each chunk: curator extracts structured knowledge
  3. Per-source wiki page assembled
  4. Cross-source synthesis index built
  5. Wiki committed to the Hub

Compile time: 30s to several minutes per source.

A failed source is marked error with its recorded message in the SOURCES list. Source, concept-generation, upload or publication failures leave the previous published knowledge available. Successfully prepared summaries are checkpointed for a retry; a PDF retry also recovers its extracted source and section tree.

  • Retry the failing file — transient failures (curator LLM rate limits) are common, and files a dead compile left locked are requeued automatically on every finalize path, so retry never reports “no files to compile”
  • Reduce source size — splitting large documents into chapters can shorten individual retries. Long PDFs use the section-tree indexer; oversized documents and sections are summarized in parts, then combined so the summary process covers the full text.
  • Check format — corrupt PDFs and non-UTF-8 markdown choke the parser
  • Read the per-file error — it names the actual reason

Document identity and ingestion diagnostics

Section titled “Document identity and ingestion diagnostics”

Expand Document identity and ingestion under a source file to see extracted titles, identifiers, explicit editions and aliases with their source quotations. This works across document types and languages. Filename and content claims are kept separately; disagreements are flagged, and matching identifiers never automatically merge documents. Edition strings are preserved without inferring dates from codes.

Identity extraction uses a bounded sample from the beginning and end of the source and separate optional model calls for filename and content. Quotes are checked against extracted source text, but identity interpretation can still be incomplete or incorrect. Agents should read the source before selecting a document or edition.

Identity extraction uses the configured base model by default; bulk summaries keep using the summary model. Operators can set kb_identity_llm_profile: summary to use the summary model for both. The base profile can cost more per source. An identifier that exactly equals the complete filename stem is also confirmed in code, preserving punctuation and version suffixes.

The same disclosure shows the latest ingestion attempt’s warnings and failures, including pages with no extracted text, failed or limited image captioning, section-index fallback, rejected identity claims and missing embeddings at publication. Diagnostic counts include a few examples. A blank PDF page can be intentional, so inspect the original before treating it as missing content.

Published identity remains available when a later compile fails. The ingestion report records the latest attempt, not a full history or current health status; an embedding backfill may have repaired an earlier warning. Existing sources receive this metadata when compiled again; rebuild the KB to populate all files.

After attaching a KB, agents use kb_search_pages for hybrid lexical/vector search across their permitted knowledge bases, then kb_read_page to read the matching pages. Validate retrieval by asking the attached agent representative questions and inspecting its search and read tool results.

For a known document, use kb_search_pages(query="DOC-003", mode="documents") to look up an exact filename, title, identifier or explicit alias. Results make ambiguity and filename/content conflicts explicit. Case and whitespace are normalized; punctuation, accents and edition differences remain meaningful. If identity lookup finds nothing, search content with mode="passages" (the default). Both modes use the same permitted KB scope.

Search hits include a passage_id when an exact evidence passage is available. Pass it to kb_read_page with the returned KB ID and path. Omit it to expand the document; use pages for a PDF range. Reads report the published version and location.

For table headings or clauses continued around a passage, pass context="surrounding" together with passage_id. The reader keeps the passage and expands its PDF page and immediate neighbors within a budget. Markdown sources receive a window around the passage. The default is 8,000 source characters, with a maximum of 24,000; limit sets the budget. Each slice reports its original location and marks omitted text. Use a parent read or PDF pages range for further context; surrounding cannot be combined with pages or offset.

Deleting a source file retires its dependent searchable evidence. Existing KBs need a full rebuild to populate source passages and provenance.

The wiki browser’s “Search wiki” field filters page titles locally; it does not run the agent’s retrieval pipeline.

Terminal window
curl -X POST http://localhost:8000/api/agents/{agent_id}/knowledge-bases \
-H "Authorization: Bearer $TOKEN" \
-d '{"kb_id": "..."}'

The agent’s system prompt gains:

## Available Knowledge Bases
- acme-product-docs: Product documentation, runbooks, known-issues database.
Browse via kb_list_pages("acme-product-docs"), then kb_read_page(path).

Plus the agent’s tool catalogue grows by kb_search_pages, kb_list_pages and kb_read_page.

Method Endpoint Purpose
GET /api/knowledge-bases List
GET /api/knowledge-bases/{id} Detail
POST /api/knowledge-bases Create. Body: {name, description, curator_model}
DELETE /api/knowledge-bases/{id} Delete (pre-flight: warns if attached to agents)
GET /api/knowledge-bases/{id}/sources List sources
POST /api/knowledge-bases/{id}/sources Add source (multipart for File, JSON for others)
DELETE /api/knowledge-bases/{id}/sources/{source_id} Remove source
POST /api/knowledge-bases/{id}/compile Trigger compile
GET /api/knowledge-bases/{id}/wiki Browse wiki

MCP & Vault for external system integration. Datasets for the other corpus type used in training.