How the OpenDerisk Knowledge CRUD API Manages Domain Knowledge

The knowledge CRUD API exposes FastAPI endpoints under the /knowledge prefix that perform create-read-update-delete operations on knowledge spaces and documents, using Pydantic models for validation, a service layer for business logic, and DAOs for persistence, while delegating chunking and embedding to an integrated RAG service.

The derisk-ai/openderisk repository implements a comprehensive knowledge management system for LLM-driven agents. The knowledge CRUD API provides a type-safe, RESTful interface for managing domain-specific knowledge spaces, documents, and vector embeddings through a layered architecture that separates HTTP handling, business logic, and data persistence.

Architecture of the Knowledge CRUD API

The implementation follows a strict layered pattern defined in packages/derisk-app/src/derisk_app/knowledge/api.py and its supporting modules.

API Router Layer

The FastAPI router exposes typed endpoints under the /knowledge prefix. Each route accepts Pydantic request models, invokes the service layer, and returns standardized Result wrappers. Key routes include @router.post("/knowledge/space/add") for space creation and @router.post("/knowledge/{space_name}/document/add") for document ingestion.

Service Layer

Located in packages/derisk-app/src/derisk_app/knowledge/service.py, the service layer implements business logic through methods like create_knowledge_space(), create_knowledge_document(), and delete_document(). It validates inputs, constructs database entities, and coordinates between DAOs and the RAG Service for vector operations.

Data Access Objects (DAOs)

DAOs provide low-level ORM access to MySQL tables including knowledge_space, knowledge_document, and knowledge_chunk. Instantiated at module load time (e.g., knowledge_space_dao = KnowledgeSpaceDao()), these objects handle row-level CRUD operations in derisk_serve/rag/models/.

RAG and Storage Integration

Document synchronization delegates to the RAG Service in derisk_serve/rag/service/service.py, which manages chunking, embedding generation, and vector store insertion. File uploads utilize the FileStorageClient abstraction from derisk/core/interface/file.py to persist binaries to object storage.

Managing Knowledge Spaces

Knowledge spaces are logical containers for domain-specific documents that support full CRUD operations.

Create a Knowledge Space

To create a space, send a POST request to /knowledge/space/add with a SpaceServeRequest payload containing name, storage_type, and description. The endpoint invokes KnowledgeService.create_knowledge_space(), which persists the entity via knowledge_space_dao.create_knowledge_space().

import requests

payload = {
    "name": "finance_reports",
    "desc": "Annual financial statements",
    "owner": "alice",
    "storage_type": "VectorStore"
}
r = requests.post(
    "http://localhost:8000/knowledge/space/add",
    json=payload,
)
print(r.json())   # {"code":0,"msg":"success","data":[123]}

List and Query Spaces

Retrieve spaces via POST /knowledge/space/list using KnowledgeSpaceRequest filters such as user_id or name. The service layer enriches results with document counts before returning a list of SpaceQueryResponse objects.

payload = {"user_id": 42}
r = requests.post(
    "http://localhost:8000/knowledge/space/list",
    json=payload,
)
print(r.json()["data"])

Update Space Configuration

Space metadata and arguments are managed through POST /knowledge/{space_id}/arguments (read) and POST /knowledge/{space_id}/argument/save (write). These endpoints read from and write to the context column via knowledge_space_dao.update_knowledge_space().

Delete a Knowledge Space

The POST /knowledge/space/delete endpoint removes database rows and the corresponding upload folder under KNOWLEDGE_UPLOAD_ROOT_PATH.

Document Lifecycle Operations

Documents exist within spaces and support multiple ingestion paths including direct text input, file upload, and Yuque URL import.

Add a Document

Create document entries via POST /knowledge/{space_name}/document/add. The endpoint validates uniqueness, builds a KnowledgeDocumentEntity, and persists via knowledge_document_dao. For Yuque URLs, use the /knowledge/{space_name}/document/yuque/add variant.

payload = {
    "doc_name": "2024_Q1_Statement",
    "doc_type": "TEXT",
    "content": "Revenue ...",
    "labels": ["finance", "Q1"],
}
r = requests.post(
    "http://localhost:8000/knowledge/finance_reports/document/add",
    json=payload,
)
print(r.json())   # {"code":0,"msg":"success","data": 987}

Upload File Attachments

Upload binary files via POST /knowledge/{space_name}/document/upload using multipart form data. The API sanitizes filenames before calling FileStorageClient.save_file().

curl -X POST "http://localhost:8000/knowledge/finance_reports/document/upload" \
  -F "doc_name=2024_Q1.pdf" \
  -F "doc_type=PDF" \
  -F "doc_file=@2024_Q1.pdf"

List and Filter Documents

Query documents via POST /knowledge/{space_name}/document/list, supporting pagination and filters for doc_type, status, and doc_name. The service layer returns dictionaries formatted for front-end consumption.

Edit and Delete Documents

Update documents through POST /knowledge/{space_name}/document/edit, which loads the space and forwards a DocumentServeRequest to the RAG service (service.update_document()). Delete via POST /knowledge/{space_name}/document/delete, which calls knowledge_space_service.delete_document() to remove DB rows and associated chunks.

Vector Store Synchronization and Chunk Management

After ingestion, documents must be synchronized to generate embeddings for semantic search.

Synchronize Documents for RAG

The POST /knowledge/{space_name}/document/sync endpoint transforms documents into chunks, generates embeddings using the specified model_name, and pushes vectors to the selected store. Batch operations are supported via /knowledge/{space_name}/document/sync_batch.

payload = {
    "doc_ids": [987],
    "model_name": "gpt-4o-mini",
    "chunk_size": 512,
    "chunk_overlap": 50,
}
r = requests.post(
    "http://localhost:8000/knowledge/finance_reports/document/sync",
    json=payload,
)
print(r.json())

Manage Document Chunks

Retrieve chunks via POST /knowledge/{space_name}/chunk/list using ChunkQueryRequest parameters including document_id, page, and page_size. Edit chunk content via POST /knowledge/{space_name}/chunk/edit.

payload = {
    "document_id": 987,
    "page": 1,
    "page_size": 20,
}
r = requests.post(
    "http://localhost:8000/knowledge/finance_reports/chunk/list",
    json=payload,
)
print(r.json()["data"])

Request Models and Error Handling

All endpoints utilize Pydantic BaseModel classes defined in packages/derisk-app/src/derisk_app/knowledge/request/request.py, including KnowledgeSpaceRequest, KnowledgeDocumentRequest, and ChunkQueryRequest. These models provide automatic validation, serialization, and OpenAPI schema documentation. Every endpoint wraps execution in try … except Exception blocks to catch errors, log them, and return standardized Result.failed payloads.

Summary

  • The knowledge CRUD API in derisk-ai/openderisk provides FastAPI endpoints under /knowledge for full lifecycle management of spaces and documents.
  • The architecture separates concerns across the API router (api.py), service layer (service.py), DAOs (derisk_serve/rag/models/), and RAG integration (derisk_serve/rag/service/service.py).
  • Operations include creating spaces, uploading files via FileStorageClient, synchronizing documents for vector search with configurable chunk_size and chunk_overlap, and managing chunks with pagination support.
  • Pydantic models in request/request.py ensure type safety and automatic validation for all CRUD operations against the MySQL schema.

Frequently Asked Questions

What is the base URL path for all knowledge CRUD operations?

All endpoints are prefixed with /knowledge and are defined in packages/derisk-app/src/derisk_app/knowledge/api.py. Space management uses paths like POST /knowledge/space/add, while document operations use parameterized routes such as /knowledge/{space_name}/document/add.

How does the API handle file storage for uploaded documents?

The API uses the FileStorageClient abstraction from derisk/core/interface/file.py. When you call the document_upload endpoint, it sanitizes filenames and persists binaries to the configured object storage or local path under KNOWLEDGE_UPLOAD_ROOT_PATH before creating the database record.

What happens when I synchronize a document via the sync endpoint?

Calling POST /knowledge/{space_name}/document/sync triggers the RAG service in derisk_serve/rag/service/service.py. This process chunks the document content according to the provided chunk_size and chunk_overlap parameters, generates embeddings using the specified model_name (e.g., gpt-4o-mini), and inserts the vectors into the configured vector store for semantic search.

Can I update individual chunks of a document after synchronization?

Yes. The API exposes POST /knowledge/{space_name}/chunk/edit for modifying chunk content, and POST /knowledge/{space_name}/chunk/list for retrieving paginated chunks using ChunkQueryRequest filters including document_id and page_size. These operations allow fine-grained management of the vector store content without re-processing entire documents.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →