How to Set Classification and Summarization Models in Cognee

Use cognee.config.set_classification_model() and cognee.config.set_summarization_model() to replace the default Pydantic models that control how Cognee tags content and generates summaries.

Cognee is an open-source knowledge graph engine that relies on configurable data models to classify and summarize content during the cognification pipeline. Learning how to set classification and summarization models enables you to customize Cognee's behavior without forking the repository or modifying environment variables. This guide explains the configuration architecture based on the actual source code in the topoteretes/cognee repository.

Understanding the Default Models

Cognee's classification and summarization behavior is governed by two default Pydantic models defined in cognee/shared/data_models.py:

  • DefaultContentPrediction – A Pydantic model that represents a single class-label prediction for content classification
  • SummarizedContent – A Pydantic model that holds a summary string and description for generated summaries

These models are referenced by the CognifyConfig class in cognee/modules/cognify/config.py, which stores the active configuration in a singleton pattern. The runtime values determine how Cognee processes raw text into structured knowledge graph nodes.

The Configuration API Pathway

Three files form the complete configuration pathway:

  1. cognee/shared/data_models.py – Defines the base Pydantic models (DefaultContentPrediction and SummarizedContent)
  2. cognee/modules/cognify/config.py – Contains the CognifyConfig class that holds the active model references
  3. cognee/api/v1/config/config.py – Exposes the public configuration API with static setter methods

According to the source code, the setter methods are implemented at lines 64-73 in cognee/api/v1/config/config.py, while a generic mapping table for string-based configuration exists at lines 28-33.

Setting Custom Classification and Summarization Models

Direct Setter Methods (Python API)

The most explicit way to configure models is through the dedicated setter methods. Your custom classes should inherit from the default models to ensure field compatibility.

import cognee
from cognee.shared.data_models import DefaultContentPrediction, SummarizedContent

# Define custom models inheriting from defaults

class MyClassificationModel(DefaultContentPrediction):
    confidence_threshold: float = 0.8  # Add custom validation fields

class MySummarizationModel(SummarizedContent):
    metadata: dict = {}  # Extend with additional metadata

# Apply custom models at runtime

cognee.config.set_classification_model(MyClassificationModel)
cognee.config.set_summarization_model(MySummarizationModel)

# Verify the configuration

from cognee.modules.cognify.config import get_cognify_config
cfg = get_cognify_config()
print(cfg.classification_model)  # → <class '__main__.MyClassificationModel'>

print(cfg.summarization_model)   # → <class '__main__.MySummarizationModel'>

Generic Configuration Setter

For dynamic configuration or when working with string-based settings, use the generic set() method:

import cognee

cognee.config.set("classification_model", MyClassificationModel)
cognee.config.set("summarization_model", MySummarizationModel)

The set() method maps string keys to the appropriate dedicated setters via the configuration mapping table in cognee/api/v1/config/config.py.

Command Line Interface

You can also configure models via the CLI without writing Python code, assuming your custom models are importable:

cognee config set classification_model my_package.models.CustomClassifier
cognee config set summarization_model my_package.models.CustomSummarizer

The CLI command forwards the string path to config.set(), which resolves the class and updates the in-memory CognifyConfig singleton.

Persisting Configuration Across Sessions

Because CognifyConfig exists only as an in-memory singleton, changes do not persist between Python process restarts. To ensure your custom models are always active, execute the configuration during application startup:


# app/startup.py

import cognee
from my_models import CustomClassification, CustomSummarization

def initialize_cognee():
    cognee.config.set_classification_model(CustomClassification)
    cognee.config.set_summarization_model(CustomSummarization)
    # Proceed with cognify operations

Place this initialization in your application's entry point or __init__.py to maintain consistent behavior across runs.

Summary

  • Default models (DefaultContentPrediction and SummarizedContent) are defined in cognee/shared/data_models.py and stored in CognifyConfig
  • Configuration API methods live in cognee/api/v1/config/config.py, providing both dedicated setters and a generic set() interface
  • Direct setters (set_classification_model and set_summarization_model) offer type-safe model replacement at lines 64-73
  • CLI configuration uses string key mapping defined at lines 28-33 to resolve model paths dynamically
  • In-memory state requires initialization on every application startup for persistent behavior

Frequently Asked Questions

What fields must custom classification and summarization models implement?

Custom models must conform to the Pydantic BaseModel structure with fields matching DefaultContentPrediction (for classification) and SummarizedContent (for summarization) as defined in cognee/shared/data_models.py. While inheriting from these classes ensures compatibility, any custom implementation must provide the same core attributes that Cognee's internal logic expects during the cognification process.

Can I use different models for different datasets within the same application?

No, the current implementation uses a singleton CognifyConfig pattern where get_cognify_config() returns a single global instance. All classification and summarization operations within the same Python process will use the same model classes. To use different models for different datasets, you must reconfigure the models between operations or run separate processes.

Where are the default model classes actually instantiated during processing?

The default models are referenced within the cognification pipeline defined in cognee/modules/cognify/config.py. When content is processed, Cognee retrieves the active configuration via get_cognify_config() and uses the classification_model and summarization_model attributes stored in that singleton to validate and structure output data.

Why does the CLI use string paths instead of Python imports?

The CLI accepts dot-notation string paths (e.g., my_package.models.CustomModel) because the generic set() method in cognee/api/v1/config/config.py dynamically imports and resolves these classes at runtime. This allows configuration without requiring users to write Python bootstrap code, though the referenced modules must be available in the Python path.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →