# How to Configure Chunking Strategies in WeKnora: A Complete API and SDK Guide

> Learn to configure WeKnora chunking strategies via the API and SDK. Choose from heading, heuristic, auto, or recursive tiers for optimal knowledge base organization.

- Repository: [Tencent/WeKnora](https://github.com/tencent/WeKnora)
- Tags: how-to-guide
- Published: 2026-09-13

---

**To configure chunking strategies in WeKnora, set the `chunking_config` JSONB field when creating or updating a knowledge base, selecting from tiers like `heading`, `heuristic`, `auto`, or `recursive`, with optional parent-child hierarchy support.**

WeKnora turns raw documents into searchable units through a pluggable, tier-based chunking pipeline defined in the Tencent/WeKnora repository. The system interprets **ChunkingConfig** parameters stored as JSONB in the `knowledge_bases` table, processing them through the `internal/infrastructure/chunker` package to produce semantically coherent chunks. Understanding how to manipulate these configurations allows you to optimize retrieval accuracy for different document types ranging from structured Markdown to unstructured plain text.

## ChunkingConfig Schema and Available Parameters

The chunking behavior is controlled by a JSONB column attached to each knowledge base definition. According to the API documentation in [`website-docs/04-api/02-api-chunks.md`](https://github.com/Tencent/WeKnora/blob/main/website-docs/04-api/02-api-chunks.md), the configuration accepts the following fields:

### Core Configuration Fields

- **`strategy`** (string): The splitting algorithm to apply. Valid values are `legacy`, `auto`, `heading`, `heuristic`, or `recursive`.
- **`chunk_size`** (int): Desired character length for each chunk, defaulting to 1000 characters.
- **`chunk_overlap`** (int): Number of characters that overlap between successive chunks to preserve context.
- **`separators`** ([]string): Custom delimiter array (e.g., `["\n\n", "."]`) that overrides default splitting boundaries.
- **`token_limit`** (int): Maximum token count per chunk when using token-based language models.
- **`languages`** ([]string): Language hints (e.g., `["en", "zh"]`) that inform language-specific heuristics.

### Parent-Child Hierarchical Options

When `enable_parent_child` is set to `true`, WeKnora implements a two-tier splitting strategy:

- **`enable_parent_child`** (bool): Activates hierarchical chunking where large parent windows contain smaller child chunks.
- **`parent_chunk_size`** (int): Size of the large container chunk (default 4096 characters).
- **`child_chunk_size`** (int): Size of the granular retrieval units (default 384 characters).

In this mode, the chunker first creates parent chunks using [`internal/infrastructure/chunker/splitter.go`](https://github.com/Tencent/WeKnora/blob/main/internal/infrastructure/chunker/splitter.go), then subdivides them into child chunks while preserving breadcrumb context through `ContextHeader` linkage.

## Available Chunking Strategies in WeKnora

The chunker package implements distinct **tiers** in [`internal/infrastructure/chunker/strategy.go`](https://github.com/Tencent/WeKnora/blob/main/internal/infrastructure/chunker/strategy.go), each optimized for specific document structures. The tier selector executes candidates in order, falling back automatically if a tier rejects the document format.

### Legacy and Recursive Tiers

**`TierLegacy`** (aliased as `recursive`) resides in [`internal/infrastructure/chunker/splitter.go`](https://github.com/Tencent/WeKnora/blob/main/internal/infrastructure/chunker/splitter.go) and implements the original recursive character splitting using a fixed hierarchy of separators. This strategy works best for plain text documents without semantic markup.

### Heading-Based Splitting

**`TierHeading`** in [`internal/infrastructure/chunker/heading_hierarchy.go`](https://github.com/Tencent/WeKnora/blob/main/internal/infrastructure/chunker/heading_hierarchy.go) parses Markdown documents according to header levels (H1, H2, H3), preserving document outline structure in the chunk boundaries. This tier activates automatically when the `strategy` field is set to `"heading"`.

### Heuristic and Auto Strategies

**`TierHeuristic`** from [`internal/infrastructure/chunker/heuristic_splitter.go`](https://github.com/Tencent/WeKnora/blob/main/internal/infrastructure/chunker/heuristic_splitter.go) employs multi-language, sentence-aware splitting that respects punctuation and linguistic boundaries. **`TierAuto`** runs a profiling step to dynamically select the most appropriate tier (heading vs. heuristic vs. legacy), retrying with the next candidate if validation fails.

### Tier Selection and Fallback Logic

The `SplitWithDiagnostics` function in [`strategy.go`](https://github.com/Tencent/WeKnora/blob/main/strategy.go) chains candidate tiers (typically `auto → heading → heuristic → legacy`) and validates output via `ValidateChunks`. If a tier returns an error or produces invalid chunks, the selector immediately attempts the next strategy in the chain, ensuring robust processing across heterogeneous document collections.

## How Configurations Are Applied

When you create or update a knowledge base via the REST API or Go SDK, WeKnora persists the configuration through a specific pipeline:

1. **Persistence**: The API writes the `chunking_config` JSONB to the `knowledge_bases` table.
2. **Configuration Building**: In [`internal/application/service/knowledge.go`](https://github.com/Tencent/WeKnora/blob/main/internal/application/service/knowledge.go), the `buildSplitterConfig` function unmarshals the JSONB into a `SplitterConfig` struct.
3. **Execution**: The server calls `SplitWithDiagnostics` from the chunker package, passing the populated `SplitterConfig` to the selected tier.
4. **Enrichment**: Each resulting `Chunk` struct receives a `ContextHeader` containing heading breadcrumbs and optional parent-child links if hierarchical mode is enabled.

Already-stored chunks remain immutable unless you trigger a full re-index after updating the configuration.

## Practical Configuration Examples

### REST API Configuration via cURL

Create a knowledge base using the **heading** strategy with parent-child hierarchy:

```bash
curl -X POST https://weknora.example.com/api/v1/knowledge-bases \
  -H "Authorization: Bearer $TOKEN" \
  -H "Content-Type: application/json" \
  -d '{
        "name": "ProductGuide",
        "type": "document",
        "chunking_config": {
          "chunk_size": 900,
          "chunk_overlap": 120,
          "separators": ["\n\n", "."],
          "strategy": "heading",
          "enable_parent_child": true,
          "parent_chunk_size": 4096,
          "child_chunk_size": 384
        }
      }'

```

### Go SDK Implementation

Configure chunking programmatically using the client defined in [`client/chunk.go`](https://github.com/Tencent/WeKnora/blob/main/client/chunk.go):

```go
import (
    "context"
    "github.com/Tencent/WeKnora/main/client"
)

func configureChunking(ctx context.Context, api *client.Client) error {
    cfg := client.ChunkingConfig{
        ChunkSize:         800,
        ChunkOverlap:      100,
        Separators:        []string{"\n\n", "."},
        Strategy:          "heuristic",
        Languages:         []string{"en", "zh"},
        EnableParentChild: true,
        ParentChunkSize:   4096,
        ChildChunkSize:    384,
    }

    kb := client.KnowledgeBase{
        Name:           "TechnicalDocs",
        Type:           "document",
        ChunkingConfig: cfg,
    }

    return api.CreateKnowledgeBase(ctx, kb)
}

```

### Updating Existing Knowledge Bases

Modify chunking parameters without recreating the knowledge base:

```bash
curl -X PATCH https://weknora.example.com/api/v1/knowledge-bases/kb-12345 \
  -H "Authorization: Bearer $TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"chunking_config":{"strategy":"heuristic","chunk_size":600}}'

```

The server applies the new configuration to subsequent document ingestions; existing chunks retain their previous structure until re-indexed.

## Debugging and Tuning Chunking Behavior

WeKnora provides diagnostic tools to validate configuration choices before committing them to production documents.

**Preview Endpoint**: Send text to `GET /api/v1/chunker/preview?text=...` to see which tier the auto-selector chooses and view intermediate chunk boundaries. The response includes `selected_tier`, `chunks` array, and `diagnostics` showing rejected tiers and profiling metadata.

**Debug Logging**: Set `LOG_LEVEL=debug` to observe tier rejection messages (e.g., `chunker: tier heading rejected`) in the server logs, helping identify why the fallback chain activated.

**Performance Tuning**: When using parent-child mode, ensure `parent_chunk_size` remains significantly larger than `child_chunk_size` (recommended ratio 10:1 or greater) to maintain hierarchical coherence without fragmenting context excessively.

## Summary

- **ChunkingConfig** is stored as JSONB in the `knowledge_bases` table and controls all splitting behavior through fields like `strategy`, `chunk_size`, and `enable_parent_child`.
- **Five strategies** are available: `legacy`/`recursive` (character-based), `heading` (Markdown-aware), `heuristic` (language-aware), and `auto` (dynamic profiler with automatic fallback).
- **Configuration propagation** flows from API/Go SDK → database → `buildSplitterConfig` in [`knowledge.go`](https://github.com/Tencent/WeKnora/blob/main/knowledge.go) → `SplitWithDiagnostics` in the chunker package.
- **Parent-child hierarchy** creates large parent windows (4096 chars) containing smaller child chunks (384 chars), enabling granular retrieval with broad context preservation.
- **Debugging tools** include the `/api/v1/chunker/preview` endpoint and debug-level logging to trace tier selection decisions.

## Frequently Asked Questions

### What is the difference between auto and heuristic chunking strategies?

**`auto`** runs a document profiling step that attempts heading-based splitting first, then falls back to heuristic or legacy methods if validation fails, while **`heuristic`** immediately applies language-aware sentence boundaries without attempting Markdown parsing. Use `auto` for mixed document collections where structure varies, and `heuristic` when you know the content lacks Markdown headers but requires linguistic precision.

### How do I enable hierarchical parent-child chunking?

Set `enable_parent_child` to `true` in your `chunking_config` and define `parent_chunk_size` (default 4096) and `child_chunk_size` (default 384). According to [`internal/infrastructure/chunker/splitter.go`](https://github.com/Tencent/WeKnora/blob/main/internal/infrastructure/chunker/splitter.go), this creates large parent chunks that are subsequently split into retrievable child units, with each child maintaining a reference to its parent for context enrichment.

### Can I update chunking configuration without re-indexing existing documents?

Yes, but with limitations. As implemented in [`internal/application/service/knowledge.go`](https://github.com/Tencent/WeKnora/blob/main/internal/application/service/knowledge.go), updating `chunking_config` via PATCH immediately affects new document ingestions, while previously chunked documents retain their original structure. To apply new strategies to existing content, you must trigger a full re-index operation through the knowledge base management API.

### Which chunking strategy works best for Markdown documentation?

**`heading`** (tier defined in [`internal/infrastructure/chunker/heading_hierarchy.go`](https://github.com/Tencent/WeKnora/blob/main/internal/infrastructure/chunker/heading_hierarchy.go)) is optimized for Markdown files, splitting precisely at header boundaries to preserve document outline structure. For documentation containing mixed content types, **`auto`** provides robust fallback protection by attempting heading detection first, then degrading gracefully to heuristic or recursive splitting if the document lacks sufficient header hierarchy.