# How LogSentinelAI Uses Pydantic Models for Declarative Log Extraction

> Discover how LogSentinelAI uses Pydantic models for declarative log extraction. Ensure LLM extractions are validated, typed, and enforced for reliable security event storage.

- Repository: [JungJungIn/logsentinelai](https://github.com/call518/logsentinelai)
- Tags: deep-dive
- Published: 2026-02-26

---

**LogSentinelAI leverages Pydantic `BaseModel` classes to define strict schemas for security events, ensuring LLM-generated log extractions are automatically validated, typed, and enforced before storage.**

The LogSentinelAI framework transforms raw log lines into structured security intelligence by adopting a declarative approach to data extraction. Instead of imperative parsing logic, the system uses **Pydantic models for declarative log extraction** to specify exactly what constitutes a valid security event. This architecture lives in the `call518/logsentinelai` repository, where analyzer modules define expected output shapes and the core processing pipeline enforces them.

## Declarative Schema Architecture

The foundation of LogSentinelAI's extraction logic rests on Pydantic models that act as contracts between the LLM and the processing pipeline. Each analyzer module declares its own hierarchy of models to capture domain-specific security events.

### Defining Security Event Models

Analyzer modules in `src/logsentinelai/analyzers/` define Pydantic classes that specify required fields, validation constraints, and type conversions. For example, the Linux system analyzer defines enums and models to constrain severity levels and event categories:

```python
class SeverityLevel(str, Enum):
    CRITICAL = "CRITICAL"
    HIGH = "HIGH"
    MEDIUM = "MEDIUM"
    LOW = "LOW"
    INFO = "INFO"


class SecurityEvent(BaseModel):
    event_type: EventType
    severity: SeverityLevel
    related_logs: list[str] = Field(
        min_length=1,
        description="Original log lines that triggered this event"
    )
    description: str = Field(description="Detailed event description")
    confidence_score: float = Field(ge=0.0, le=1.0)

```

These definitions enforce that every security event must include at least one related log entry (`min_length=1`) and that confidence scores fall between 0.0 and 1.0. Similar models exist in [`httpd_server.py`](https://github.com/call518/logsentinelai/blob/main/httpd_server.py), [`httpd_access.py`](https://github.com/call518/logsentinelai/blob/main/httpd_access.py), and [`general_log.py`](https://github.com/call518/logsentinelai/blob/main/general_log.py) for their respective log formats.

### Schema Generation for LLM Prompts

The framework converts these Pydantic models into JSON schemas that guide the LLM's output format. Before invoking the model, the analyzer calls `model_json_schema()` to embed the exact structure into the prompt:

```python
model_schema = LogAnalysis.model_json_schema()
prompt = prompt_template.format(
    logs="\n".join(chunk),
    model_schema=model_schema,
    response_language="en"
)

```

This ensures the LLM knows precisely which fields are required, what data types to return, and any constraints that apply to the extraction.

## Validation Pipeline in Action

Once the LLM returns a response, the system uses Pydantic's validation mechanisms to guarantee the output matches the declared schema before any data reaches storage.

### The process_log_chunk Workflow

The central processing routine in [`src/logsentinelai/core/commons.py`](https://github.com/call518/logsentinelai/blob/main/src/logsentinelai/core/commons.py) orchestrates the validation step. The `process_log_chunk` function receives the raw LLM response and validates it against the analyzer's Pydantic model:

```python

# `parsed` is the JSON object returned by the LLM

parsed = json.loads(llm_response)

# Enforce the schema – raises ValidationError on mismatch

LogAnalysis.model_validate(parsed)

```

This call performs several critical functions:
- **Type coercion**: Converts JSON primitives to Python objects, including enum values like `SeverityLevel.CRITICAL`
- **Constraint checking**: Validates that `related_logs` contains at least one entry and that `confidence_score` falls within the 0.0-1.0 range
- **Field presence**: Ensures all required fields defined in the `BaseModel` exist in the response

### Error Handling and Type Safety

When validation fails, Pydantic raises a `ValidationError` that the pipeline catches to prevent malformed data from contaminating downstream systems. The error handling pathway in `process_log_chunk` logs detailed diagnostics about which fields failed validation and why, then ships a failure document to Elasticsearch for audit purposes.

This approach provides **declarative type safety**: the schema definition in the analyzer module serves as the single source of truth for both the LLM's output format and the validation logic that enforces it.

## Extending Log Coverage with Modular Analyzers

The Pydantic-based architecture enables rapid extension to new log sources. Each analyzer module in `src/logsentinelai/analyzers/` operates independently with its own model hierarchy:

- **[`linux_system.py`](https://github.com/call518/logsentinelai/blob/main/linux_system.py)**: Models for authentication failures, privilege escalation, and system anomalies
- **[`httpd_server.py`](https://github.com/call518/logsentinelai/blob/main/httpd_server.py)**: Schemas for Apache error logs and server-level issues
- **[`httpd_access.py`](https://github.com/call518/logsentinelai/blob/main/httpd_access.py)**: Structures for web access patterns and potential web attacks
- **[`general_log.py`](https://github.com/call518/logsentinelai/blob/main/general_log.py)**: Flexible models for arbitrary log formats without predefined structure

Because each analyzer declares its own Pydantic models, adding support for a new log type requires only creating a new module with appropriate `BaseModel` definitions. The existing validation infrastructure in [`core/commons.py`](https://github.com/call518/logsentinelai/blob/main/core/commons.py) automatically handles the new schemas without modification.

## Summary

LogSentinelAI implements **Pydantic models for declarative log extraction** to ensure LLM-generated security events are structurally sound and type-safe:

- **Schema Definition**: Analyzer modules in `src/logsentinelai/analyzers/` declare strict `BaseModel` classes with validation constraints like `min_length` and field ranges
- **LLM Guidance**: The `model_json_schema()` method embeds the exact JSON schema into prompts, directing the LLM to produce compliant output
- **Runtime Validation**: The `process_log_chunk` function in [`core/commons.py`](https://github.com/call518/logsentinelai/blob/main/core/commons.py) uses `model_validate()` to enforce type safety and catch malformed responses before storage
- **Modular Extensibility**: New log sources require only new Pydantic model definitions, with the validation pipeline handling them automatically

## Frequently Asked Questions

### How does LogSentinelAI ensure the LLM returns valid JSON matching the expected structure?

LogSentinelAI embeds the JSON schema generated by `model_json_schema()` directly into the prompt template. This schema defines required fields, data types, and constraints. After the LLM returns a response, the system calls `model_validate()` on the parsed JSON to enforce compliance, raising a `ValidationError` if the structure deviates from the declared Pydantic model.

### What happens when the LLM output fails Pydantic validation?

When validation fails in `process_log_chunk` (located in [`src/logsentinelai/core/commons.py`](https://github.com/call518/logsentinelai/blob/main/src/logsentinelai/core/commons.py)), the system catches the `ValidationError` and logs detailed diagnostics about which specific fields violated constraints. Rather than discarding the data, it ships a failure document to Elasticsearch for audit purposes, ensuring operators can inspect malformed outputs while preventing bad data from contaminating the security event pipeline.

### Can I add custom fields to the log extraction schema without breaking existing functionality?

Yes. The modular analyzer architecture allows you to extend Pydantic models by adding new fields with appropriate type annotations and `Field` constraints in the relevant analyzer module (e.g., [`src/logsentinelai/analyzers/linux_system.py`](https://github.com/call518/logsentinelai/blob/main/src/logsentinelai/analyzers/linux_system.py)). Existing validation logic in [`core/commons.py`](https://github.com/call518/logsentinelai/blob/main/core/commons.py) automatically handles the updated schema. If you add required fields, ensure the LLM prompt includes examples of the new structure; optional fields with defaults maintain backward compatibility automatically.

### Which analyzer modules define the Pydantic models for different log types?

LogSentinelAI organizes Pydantic models into specialized analyzer modules within `src/logsentinelai/analyzers/`:
- [`linux_system.py`](https://github.com/call518/logsentinelai/blob/main/linux_system.py) defines models for authentication and system events
- [`httpd_server.py`](https://github.com/call518/logsentinelai/blob/main/httpd_server.py) handles Apache error log structures
- [`httpd_access.py`](https://github.com/call518/logsentinelai/blob/main/httpd_access.py) manages web access log schemas
- [`general_log.py`](https://github.com/call518/logsentinelai/blob/main/general_log.py) provides flexible models for unstructured log formats

Each module declares its own `SecurityEvent` and `LogAnalysis` hierarchies tailored to the specific security indicators present in that log source.