# How to Configure Chunk Size for Cognee: Complete API and CLI Guide

> Learn to configure Cognee chunk size easily. Optimize token limits using environment variables, Python API, or CLI commands for your documents. Get the complete guide now.

- Repository: [Topoteretes/cognee](https://github.com/topoteretes/cognee)
- Tags: how-to-guide
- Published: 2026-03-16

---

**Cognee defaults to 1500 tokens per chunk, but you can override this via environment variables, Python API calls, or CLI commands to optimize token limits for your specific documents.**

Cognee is an open-source framework for building knowledge graphs from unstructured data. Configuring the chunk size is critical for balancing context preservation with processing efficiency in the `topoteretes/cognee` repository. This guide covers the three supported methods to adjust chunking parameters at runtime.

## Understanding the ChunkConfig Architecture

Cognee's chunking pipeline is governed by the `ChunkConfig` class located in [`cognee/infrastructure/data/chunking/config.py`](https://github.com/topoteretes/cognee/blob/main/cognee/infrastructure/data/chunking/config.py). This Pydantic `BaseSettings` model handles chunking parameters and implements a singleton caching pattern.

Key implementation details:
- **Default value**: `chunk_size` initializes to **1500** tokens (lines 13-14)
- **Environment integration**: Uses `SettingsConfigDict(env_file=".env", extra="allow")` to load `.env` files automatically (lines 18-19)
- **Cached accessor**: The `get_chunk_config()` function returns a singleton instance (lines 37-53) ensuring all pipeline components reference the same configuration object

## Method 1: Environment Variable Configuration

The simplest approach uses a `.env` file to set the `CHUNK_SIZE` variable without modifying code. Because `ChunkConfig` inherits from Pydantic's `BaseSettings`, it automatically reads environment variables at startup.

Create a `.env` file in your project root:

```bash
CHUNK_SIZE=1024

```

When Cognee initializes, this value replaces the default 1500 tokens. This method is ideal for Docker containers and CI/CD pipelines where code changes are undesirable.

## Method 2: Python API Configuration

For runtime adjustments within your application, import the configuration module from `cognee.api.v1.config.config`.

### Using set_chunk_size()

The dedicated setter modifies the cached `ChunkConfig` singleton directly:

```python
from cognee import config
from cognee.infrastructure.data.chunking.config import get_chunk_config

# Update chunk size to 1024 tokens

config.set_chunk_size(1024)

# Verify the change

print(get_chunk_config().chunk_size)  # Output: 1024

```

Implementation reference: The `set_chunk_size` method retrieves the singleton via `get_chunk_config()` and assigns the new value (lines 40-43 in [`cognee/api/v1/config/config.py`](https://github.com/topoteretes/cognee/blob/main/cognee/api/v1/config/config.py)).

### Using the Generic Setter

For dynamic configuration keys, use the `set()` method:

```python
from cognee import config

config.set("chunk_size", 2048)

```

Both approaches immediately affect **all subsequent chunking operations** in the same process because they modify the cached singleton returned by `get_chunk_config()`.

## Method 3: Command Line Interface

Cognee exposes chunk configuration through the CLI, mapping commands to the same Python API setters. The mapping logic resides in [`cognee/cli/commands/config_command.py`](https://github.com/topoteretes/cognee/blob/main/cognee/cli/commands/config_command.py) (lines 172-174).

Set chunk size via terminal:

```bash
cognee config set chunk_size 1024

```

This executes the same underlying logic as the Python API, ensuring configuration consistency across interfaces.

## How Configuration Propagates Through the Pipeline

All downstream chunking tasks—including `cognee.tasks.*` modules and `chunk_by_*` utilities—read the current value from `get_chunk_config()`. Because the configuration object is cached at the module level, changes propagate instantly to:

- Document ingestion pipelines
- Vector embedding generation  
- Knowledge graph construction

The configuration flow follows this hierarchy:

```

Process Start → get_chunk_config() returns ChunkConfig (default: 1500)
   │
   ├─ .env CHUNK_SIZE overrides default (Pydantic BaseSettings)
   ├─ config.set_chunk_size(value) updates cached instance
   └─ CLI cognee config set chunk_size <value> updates cached instance
   ↓
Chunking tasks reference chunk_config.chunk_size for token limits

```

## Practical Configuration Examples

### Environment-Based Setup

```bash

# .env file

CHUNK_SIZE=2048

```

```python

# application.py - automatically loads from environment

from cognee import cognify
cognify("./documents/")

```

### Runtime Script Configuration

```python
from cognee import config

# Optimize for shorter documents

config.set_chunk_size(512)

# Execute pipeline with new chunk size

from cognee.api.v1.cognify import cognify
cognify("./technical_docs/")

```

### CLI Workflow

```bash

# Configure for current session

cognee config set chunk_size 1800

# Process data

cognee cognify --data ./my_documents/

```

## Summary

- **Default behavior**: Cognee initializes with 1500 tokens per chunk as defined in [`cognee/infrastructure/data/chunking/config.py`](https://github.com/topoteretes/cognee/blob/main/cognee/infrastructure/data/chunking/config.py)
- **Environment override**: Set `CHUNK_SIZE` in a `.env` file for containerized deployments using Pydantic `BaseSettings` integration
- **Runtime API**: Call `config.set_chunk_size()` or `config.set()` for programmatic control within Python scripts
- **Command line**: Use `cognee config set chunk_size <value>` for operational or shell-scripted adjustments
- **Propagation**: All methods modify the cached `ChunkConfig` singleton accessed via `get_chunk_config()`, affecting subsequent operations immediately without pipeline restarts

## Frequently Asked Questions

### What is the default chunk size in Cognee?

Cognee defaults to **1500 tokens** per chunk. This value is defined in the `ChunkConfig` class within [`cognee/infrastructure/data/chunking/config.py`](https://github.com/topoteretes/cognee/blob/main/cognee/infrastructure/data/chunking/config.py) (lines 13-14) and applies to all text chunking operations unless overridden via environment variables, API calls, or CLI commands.

### Can I change the chunk size without modifying my Python code?

Yes. Create a `.env` file containing `CHUNK_SIZE=<value>`, or use the CLI command `cognee config set chunk_size <value>`. Both methods update the cached configuration singleton without requiring code changes, making them suitable for DevOps workflows, Docker environments, and production deployments.

### How does chunk size configuration affect Cognee's memory usage?

Larger chunk sizes increase memory consumption during embedding generation and knowledge graph construction because each chunk requires separate vector processing. Smaller chunks reduce per-operation memory footprints but increase the total number of operations. The 1500-token default balances these concerns for general-purpose documents processing.

### Where is the active chunk configuration stored during runtime?

Cognee stores the active configuration in a cached singleton instance managed by `get_chunk_config()` in [`cognee/infrastructure/data/chunking/config.py`](https://github.com/topoteretes/cognee/blob/main/cognee/infrastructure/data/chunking/config.py). When you call `config.set_chunk_size()` or use the CLI, you modify this in-memory object. All chunking tasks reference this same cached instance, ensuring configuration changes apply immediately to subsequent operations in the same process.