# How the OpenDerisk Knowledge CRUD API Manages Domain Knowledge

> Discover how the OpenDerisk Knowledge CRUD API effectively manages domain knowledge through create read update delete operations. Learn about its service layer, DAOs, and RAG integration.

- Repository: [derisk-ai/openderisk](https://github.com/derisk-ai/openderisk)
- Tags: how-to-guide
- Published: 2026-02-28

---

**The knowledge CRUD API exposes FastAPI endpoints under the `/knowledge` prefix that perform create-read-update-delete operations on knowledge spaces and documents, using Pydantic models for validation, a service layer for business logic, and DAOs for persistence, while delegating chunking and embedding to an integrated RAG service.**

The `derisk-ai/openderisk` repository implements a comprehensive knowledge management system for LLM-driven agents. The **knowledge CRUD API** provides a type-safe, RESTful interface for managing domain-specific knowledge spaces, documents, and vector embeddings through a layered architecture that separates HTTP handling, business logic, and data persistence.

## Architecture of the Knowledge CRUD API

The implementation follows a strict layered pattern defined in [`packages/derisk-app/src/derisk_app/knowledge/api.py`](https://github.com/derisk-ai/openderisk/blob/main/packages/derisk-app/src/derisk_app/knowledge/api.py) and its supporting modules.

### API Router Layer

The FastAPI router exposes typed endpoints under the `/knowledge` prefix. Each route accepts Pydantic request models, invokes the service layer, and returns standardized `Result` wrappers. Key routes include `@router.post("/knowledge/space/add")` for space creation and `@router.post("/knowledge/{space_name}/document/add")` for document ingestion.

### Service Layer

Located in [`packages/derisk-app/src/derisk_app/knowledge/service.py`](https://github.com/derisk-ai/openderisk/blob/main/packages/derisk-app/src/derisk_app/knowledge/service.py), the service layer implements business logic through methods like `create_knowledge_space()`, `create_knowledge_document()`, and `delete_document()`. It validates inputs, constructs database entities, and coordinates between DAOs and the RAG `Service` for vector operations.

### Data Access Objects (DAOs)

DAOs provide low-level ORM access to MySQL tables including `knowledge_space`, `knowledge_document`, and `knowledge_chunk`. Instantiated at module load time (e.g., `knowledge_space_dao = KnowledgeSpaceDao()`), these objects handle row-level CRUD operations in `derisk_serve/rag/models/`.

### RAG and Storage Integration

Document synchronization delegates to the RAG `Service` in [`derisk_serve/rag/service/service.py`](https://github.com/derisk-ai/openderisk/blob/main/derisk_serve/rag/service/service.py), which manages chunking, embedding generation, and vector store insertion. File uploads utilize the `FileStorageClient` abstraction from [`derisk/core/interface/file.py`](https://github.com/derisk-ai/openderisk/blob/main/derisk/core/interface/file.py) to persist binaries to object storage.

## Managing Knowledge Spaces

Knowledge spaces are logical containers for domain-specific documents that support full CRUD operations.

### Create a Knowledge Space

To create a space, send a POST request to `/knowledge/space/add` with a `SpaceServeRequest` payload containing `name`, `storage_type`, and `description`. The endpoint invokes `KnowledgeService.create_knowledge_space()`, which persists the entity via `knowledge_space_dao.create_knowledge_space()`.

```python
import requests

payload = {
    "name": "finance_reports",
    "desc": "Annual financial statements",
    "owner": "alice",
    "storage_type": "VectorStore"
}
r = requests.post(
    "http://localhost:8000/knowledge/space/add",
    json=payload,
)
print(r.json())   # {"code":0,"msg":"success","data":[123]}

```

### List and Query Spaces

Retrieve spaces via `POST /knowledge/space/list` using `KnowledgeSpaceRequest` filters such as `user_id` or `name`. The service layer enriches results with document counts before returning a list of `SpaceQueryResponse` objects.

```python
payload = {"user_id": 42}
r = requests.post(
    "http://localhost:8000/knowledge/space/list",
    json=payload,
)
print(r.json()["data"])

```

### Update Space Configuration

Space metadata and arguments are managed through `POST /knowledge/{space_id}/arguments` (read) and `POST /knowledge/{space_id}/argument/save` (write). These endpoints read from and write to the `context` column via `knowledge_space_dao.update_knowledge_space()`.

### Delete a Knowledge Space

The `POST /knowledge/space/delete` endpoint removes database rows and the corresponding upload folder under `KNOWLEDGE_UPLOAD_ROOT_PATH`.

## Document Lifecycle Operations

Documents exist within spaces and support multiple ingestion paths including direct text input, file upload, and Yuque URL import.

### Add a Document

Create document entries via `POST /knowledge/{space_name}/document/add`. The endpoint validates uniqueness, builds a `KnowledgeDocumentEntity`, and persists via `knowledge_document_dao`. For Yuque URLs, use the `/knowledge/{space_name}/document/yuque/add` variant.

```python
payload = {
    "doc_name": "2024_Q1_Statement",
    "doc_type": "TEXT",
    "content": "Revenue ...",
    "labels": ["finance", "Q1"],
}
r = requests.post(
    "http://localhost:8000/knowledge/finance_reports/document/add",
    json=payload,
)
print(r.json())   # {"code":0,"msg":"success","data": 987}

```

### Upload File Attachments

Upload binary files via `POST /knowledge/{space_name}/document/upload` using multipart form data. The API sanitizes filenames before calling `FileStorageClient.save_file()`.

```bash
curl -X POST "http://localhost:8000/knowledge/finance_reports/document/upload" \
  -F "doc_name=2024_Q1.pdf" \
  -F "doc_type=PDF" \
  -F "doc_file=@2024_Q1.pdf"

```

### List and Filter Documents

Query documents via `POST /knowledge/{space_name}/document/list`, supporting pagination and filters for `doc_type`, `status`, and `doc_name`. The service layer returns dictionaries formatted for front-end consumption.

### Edit and Delete Documents

Update documents through `POST /knowledge/{space_name}/document/edit`, which loads the space and forwards a `DocumentServeRequest` to the RAG service (`service.update_document()`). Delete via `POST /knowledge/{space_name}/document/delete`, which calls `knowledge_space_service.delete_document()` to remove DB rows and associated chunks.

## Vector Store Synchronization and Chunk Management

After ingestion, documents must be synchronized to generate embeddings for semantic search.

### Synchronize Documents for RAG

The `POST /knowledge/{space_name}/document/sync` endpoint transforms documents into chunks, generates embeddings using the specified `model_name`, and pushes vectors to the selected store. Batch operations are supported via `/knowledge/{space_name}/document/sync_batch`.

```python
payload = {
    "doc_ids": [987],
    "model_name": "gpt-4o-mini",
    "chunk_size": 512,
    "chunk_overlap": 50,
}
r = requests.post(
    "http://localhost:8000/knowledge/finance_reports/document/sync",
    json=payload,
)
print(r.json())

```

### Manage Document Chunks

Retrieve chunks via `POST /knowledge/{space_name}/chunk/list` using `ChunkQueryRequest` parameters including `document_id`, `page`, and `page_size`. Edit chunk content via `POST /knowledge/{space_name}/chunk/edit`.

```python
payload = {
    "document_id": 987,
    "page": 1,
    "page_size": 20,
}
r = requests.post(
    "http://localhost:8000/knowledge/finance_reports/chunk/list",
    json=payload,
)
print(r.json()["data"])

```

## Request Models and Error Handling

All endpoints utilize Pydantic `BaseModel` classes defined in [`packages/derisk-app/src/derisk_app/knowledge/request/request.py`](https://github.com/derisk-ai/openderisk/blob/main/packages/derisk-app/src/derisk_app/knowledge/request/request.py), including `KnowledgeSpaceRequest`, `KnowledgeDocumentRequest`, and `ChunkQueryRequest`. These models provide automatic validation, serialization, and OpenAPI schema documentation. Every endpoint wraps execution in `try … except Exception` blocks to catch errors, log them, and return standardized `Result.failed` payloads.

## Summary

- The **knowledge CRUD API** in `derisk-ai/openderisk` provides FastAPI endpoints under `/knowledge` for full lifecycle management of spaces and documents.
- The architecture separates concerns across the API router ([`api.py`](https://github.com/derisk-ai/openderisk/blob/main/api.py)), service layer ([`service.py`](https://github.com/derisk-ai/openderisk/blob/main/service.py)), DAOs (`derisk_serve/rag/models/`), and RAG integration ([`derisk_serve/rag/service/service.py`](https://github.com/derisk-ai/openderisk/blob/main/derisk_serve/rag/service/service.py)).
- Operations include creating spaces, uploading files via `FileStorageClient`, synchronizing documents for vector search with configurable `chunk_size` and `chunk_overlap`, and managing chunks with pagination support.
- Pydantic models in [`request/request.py`](https://github.com/derisk-ai/openderisk/blob/main/request/request.py) ensure type safety and automatic validation for all CRUD operations against the MySQL schema.

## Frequently Asked Questions

### What is the base URL path for all knowledge CRUD operations?

All endpoints are prefixed with `/knowledge` and are defined in [`packages/derisk-app/src/derisk_app/knowledge/api.py`](https://github.com/derisk-ai/openderisk/blob/main/packages/derisk-app/src/derisk_app/knowledge/api.py). Space management uses paths like `POST /knowledge/space/add`, while document operations use parameterized routes such as `/knowledge/{space_name}/document/add`.

### How does the API handle file storage for uploaded documents?

The API uses the `FileStorageClient` abstraction from [`derisk/core/interface/file.py`](https://github.com/derisk-ai/openderisk/blob/main/derisk/core/interface/file.py). When you call the `document_upload` endpoint, it sanitizes filenames and persists binaries to the configured object storage or local path under `KNOWLEDGE_UPLOAD_ROOT_PATH` before creating the database record.

### What happens when I synchronize a document via the sync endpoint?

Calling `POST /knowledge/{space_name}/document/sync` triggers the RAG service in [`derisk_serve/rag/service/service.py`](https://github.com/derisk-ai/openderisk/blob/main/derisk_serve/rag/service/service.py). This process chunks the document content according to the provided `chunk_size` and `chunk_overlap` parameters, generates embeddings using the specified `model_name` (e.g., `gpt-4o-mini`), and inserts the vectors into the configured vector store for semantic search.

### Can I update individual chunks of a document after synchronization?

Yes. The API exposes `POST /knowledge/{space_name}/chunk/edit` for modifying chunk content, and `POST /knowledge/{space_name}/chunk/list` for retrieving paginated chunks using `ChunkQueryRequest` filters including `document_id` and `page_size`. These operations allow fine-grained management of the vector store content without re-processing entire documents.