How the OpenDerisk Knowledge CRUD API Manages Domain Knowledge
The knowledge CRUD API exposes FastAPI endpoints under the /knowledge prefix that perform create-read-update-delete operations on knowledge spaces and documents, using Pydantic models for validation, a service layer for business logic, and DAOs for persistence, while delegating chunking and embedding to an integrated RAG service.
The derisk-ai/openderisk repository implements a comprehensive knowledge management system for LLM-driven agents. The knowledge CRUD API provides a type-safe, RESTful interface for managing domain-specific knowledge spaces, documents, and vector embeddings through a layered architecture that separates HTTP handling, business logic, and data persistence.
Architecture of the Knowledge CRUD API
The implementation follows a strict layered pattern defined in packages/derisk-app/src/derisk_app/knowledge/api.py and its supporting modules.
API Router Layer
The FastAPI router exposes typed endpoints under the /knowledge prefix. Each route accepts Pydantic request models, invokes the service layer, and returns standardized Result wrappers. Key routes include @router.post("/knowledge/space/add") for space creation and @router.post("/knowledge/{space_name}/document/add") for document ingestion.
Service Layer
Located in packages/derisk-app/src/derisk_app/knowledge/service.py, the service layer implements business logic through methods like create_knowledge_space(), create_knowledge_document(), and delete_document(). It validates inputs, constructs database entities, and coordinates between DAOs and the RAG Service for vector operations.
Data Access Objects (DAOs)
DAOs provide low-level ORM access to MySQL tables including knowledge_space, knowledge_document, and knowledge_chunk. Instantiated at module load time (e.g., knowledge_space_dao = KnowledgeSpaceDao()), these objects handle row-level CRUD operations in derisk_serve/rag/models/.
RAG and Storage Integration
Document synchronization delegates to the RAG Service in derisk_serve/rag/service/service.py, which manages chunking, embedding generation, and vector store insertion. File uploads utilize the FileStorageClient abstraction from derisk/core/interface/file.py to persist binaries to object storage.
Managing Knowledge Spaces
Knowledge spaces are logical containers for domain-specific documents that support full CRUD operations.
Create a Knowledge Space
To create a space, send a POST request to /knowledge/space/add with a SpaceServeRequest payload containing name, storage_type, and description. The endpoint invokes KnowledgeService.create_knowledge_space(), which persists the entity via knowledge_space_dao.create_knowledge_space().
import requests
payload = {
"name": "finance_reports",
"desc": "Annual financial statements",
"owner": "alice",
"storage_type": "VectorStore"
}
r = requests.post(
"http://localhost:8000/knowledge/space/add",
json=payload,
)
print(r.json()) # {"code":0,"msg":"success","data":[123]}
List and Query Spaces
Retrieve spaces via POST /knowledge/space/list using KnowledgeSpaceRequest filters such as user_id or name. The service layer enriches results with document counts before returning a list of SpaceQueryResponse objects.
payload = {"user_id": 42}
r = requests.post(
"http://localhost:8000/knowledge/space/list",
json=payload,
)
print(r.json()["data"])
Update Space Configuration
Space metadata and arguments are managed through POST /knowledge/{space_id}/arguments (read) and POST /knowledge/{space_id}/argument/save (write). These endpoints read from and write to the context column via knowledge_space_dao.update_knowledge_space().
Delete a Knowledge Space
The POST /knowledge/space/delete endpoint removes database rows and the corresponding upload folder under KNOWLEDGE_UPLOAD_ROOT_PATH.
Document Lifecycle Operations
Documents exist within spaces and support multiple ingestion paths including direct text input, file upload, and Yuque URL import.
Add a Document
Create document entries via POST /knowledge/{space_name}/document/add. The endpoint validates uniqueness, builds a KnowledgeDocumentEntity, and persists via knowledge_document_dao. For Yuque URLs, use the /knowledge/{space_name}/document/yuque/add variant.
payload = {
"doc_name": "2024_Q1_Statement",
"doc_type": "TEXT",
"content": "Revenue ...",
"labels": ["finance", "Q1"],
}
r = requests.post(
"http://localhost:8000/knowledge/finance_reports/document/add",
json=payload,
)
print(r.json()) # {"code":0,"msg":"success","data": 987}
Upload File Attachments
Upload binary files via POST /knowledge/{space_name}/document/upload using multipart form data. The API sanitizes filenames before calling FileStorageClient.save_file().
curl -X POST "http://localhost:8000/knowledge/finance_reports/document/upload" \
-F "doc_name=2024_Q1.pdf" \
-F "doc_type=PDF" \
-F "doc_file=@2024_Q1.pdf"
List and Filter Documents
Query documents via POST /knowledge/{space_name}/document/list, supporting pagination and filters for doc_type, status, and doc_name. The service layer returns dictionaries formatted for front-end consumption.
Edit and Delete Documents
Update documents through POST /knowledge/{space_name}/document/edit, which loads the space and forwards a DocumentServeRequest to the RAG service (service.update_document()). Delete via POST /knowledge/{space_name}/document/delete, which calls knowledge_space_service.delete_document() to remove DB rows and associated chunks.
Vector Store Synchronization and Chunk Management
After ingestion, documents must be synchronized to generate embeddings for semantic search.
Synchronize Documents for RAG
The POST /knowledge/{space_name}/document/sync endpoint transforms documents into chunks, generates embeddings using the specified model_name, and pushes vectors to the selected store. Batch operations are supported via /knowledge/{space_name}/document/sync_batch.
payload = {
"doc_ids": [987],
"model_name": "gpt-4o-mini",
"chunk_size": 512,
"chunk_overlap": 50,
}
r = requests.post(
"http://localhost:8000/knowledge/finance_reports/document/sync",
json=payload,
)
print(r.json())
Manage Document Chunks
Retrieve chunks via POST /knowledge/{space_name}/chunk/list using ChunkQueryRequest parameters including document_id, page, and page_size. Edit chunk content via POST /knowledge/{space_name}/chunk/edit.
payload = {
"document_id": 987,
"page": 1,
"page_size": 20,
}
r = requests.post(
"http://localhost:8000/knowledge/finance_reports/chunk/list",
json=payload,
)
print(r.json()["data"])
Request Models and Error Handling
All endpoints utilize Pydantic BaseModel classes defined in packages/derisk-app/src/derisk_app/knowledge/request/request.py, including KnowledgeSpaceRequest, KnowledgeDocumentRequest, and ChunkQueryRequest. These models provide automatic validation, serialization, and OpenAPI schema documentation. Every endpoint wraps execution in try … except Exception blocks to catch errors, log them, and return standardized Result.failed payloads.
Summary
- The knowledge CRUD API in
derisk-ai/openderiskprovides FastAPI endpoints under/knowledgefor full lifecycle management of spaces and documents. - The architecture separates concerns across the API router (
api.py), service layer (service.py), DAOs (derisk_serve/rag/models/), and RAG integration (derisk_serve/rag/service/service.py). - Operations include creating spaces, uploading files via
FileStorageClient, synchronizing documents for vector search with configurablechunk_sizeandchunk_overlap, and managing chunks with pagination support. - Pydantic models in
request/request.pyensure type safety and automatic validation for all CRUD operations against the MySQL schema.
Frequently Asked Questions
What is the base URL path for all knowledge CRUD operations?
All endpoints are prefixed with /knowledge and are defined in packages/derisk-app/src/derisk_app/knowledge/api.py. Space management uses paths like POST /knowledge/space/add, while document operations use parameterized routes such as /knowledge/{space_name}/document/add.
How does the API handle file storage for uploaded documents?
The API uses the FileStorageClient abstraction from derisk/core/interface/file.py. When you call the document_upload endpoint, it sanitizes filenames and persists binaries to the configured object storage or local path under KNOWLEDGE_UPLOAD_ROOT_PATH before creating the database record.
What happens when I synchronize a document via the sync endpoint?
Calling POST /knowledge/{space_name}/document/sync triggers the RAG service in derisk_serve/rag/service/service.py. This process chunks the document content according to the provided chunk_size and chunk_overlap parameters, generates embeddings using the specified model_name (e.g., gpt-4o-mini), and inserts the vectors into the configured vector store for semantic search.
Can I update individual chunks of a document after synchronization?
Yes. The API exposes POST /knowledge/{space_name}/chunk/edit for modifying chunk content, and POST /knowledge/{space_name}/chunk/list for retrieving paginated chunks using ChunkQueryRequest filters including document_id and page_size. These operations allow fine-grained management of the vector store content without re-processing entire documents.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →