# Shannon's Security Validation Rules for Configuration Files: A Deep Dive into Dangerous Pattern Blocking

> Discover Shannon's security validation rules for configuration files. Learn how this five-layer defense strategy blocks dangerous patterns in YAML files to prevent code injection and path traversal.

- Repository: [KeygraphHQ/shannon](https://github.com/keygraphhq/shannon)
- Tags: deep-dive
- Published: 2026-02-16

---

**Shannon's `parseConfig` function implements a five-layer defense-in-depth strategy that validates YAML configuration files against file-size limits, restrictive schemas, JSON-Schema constraints, regex-based dangerous pattern detection, and logical rule consistency to prevent code injection and path traversal attacks.**

Shannon, an open-source agent framework from KeygraphHQ/shannon, processes user-supplied YAML configurations to control web automation behavior. To prevent malicious inputs from compromising the agent or underlying system, Shannon's security validation rules for configuration files employ a hardened parsing pipeline that blocks dangerous patterns before they reach the execution engine.

## The Five-Layer Validation Pipeline in [`src/config-parser.ts`](https://github.com/KeygraphHQ/shannon/blob/main/src/config-parser.ts)

The core validation logic resides in [`src/config-parser.ts`](https://github.com/KeygraphHQ/shannon/blob/main/src/config-parser.ts), where the `parseConfig` function orchestrates sequential security checks. Each layer acts as a gate that must be passed before the next validation stage executes.

### Layer 1: File-Level Safeguards (Lines 60-73)

Before parsing begins, Shannon enforces basic file integrity constraints. The loader verifies that the file exists, is non-empty, and does not exceed **1 MiB** in size. This prevents denial-of-service attacks via massive file uploads or empty file processing loops.

### Layer 2: Restrictive YAML Parsing with FAILSAFE_SCHEMA (Lines 82-89)

Shannon parses YAML using the `yaml` library with `FAILSAFE_SCHEMA` explicitly enabled. This schema restriction is critical: it disables JavaScript evaluation and complex type reconstruction, ensuring the parser produces only primitive types (strings, numbers, booleans, null) and basic collections. This neutralizes YAML-based code injection attacks such as the infamous `!!python/object` or `!!js/function` exploits.

### Layer 3: JSON-Schema Validation via AJV (Lines 34-38)

After YAML parsing, the resulting object is validated against [`configs/config-schema.json`](https://github.com/KeygraphHQ/shannon/blob/main/configs/config-schema.json) using the AJV validator. This schema enforces structural constraints including required fields, type checking, string length limits, and URL format validation. Any configuration missing mandatory authentication fields or containing malformed URLs is rejected at this stage.

### Layer 4: Dangerous Pattern Detection with Regex Blacklists (Lines 48-55, 70-96)

Following schema validation, Shannon applies custom security checks using a regex blacklist defined in the `DANGEROUS_PATTERNS` array:

```typescript
const DANGEROUS_PATTERNS: RegExp[] = [
  /\.\.\//,        // path traversal
  /[<>]/,          // HTML/XML injection
  /javascript:/i,  // JavaScript URLs
  /data:/i,        // Data URLs
  /file:/i,        // File URLs
];

```

These patterns are tested against high-risk configuration fields:
- **Authentication credentials**: `username` and `password` values
- **Login flow steps**: Every string in `authentication.login_flow`
- **Rule definitions**: `url_path` and `description` fields within both `avoid` and `focus` rules

If any field matches a dangerous pattern, `parseConfig` throws an error specifying the exact field location (e.g., "authentication.credentials.password contains potentially dangerous pattern").

### Layer 5: Logical Rule Consistency Checks (Lines 92-113)

The final validation layer detects logical conflicts within rule definitions. Shannon identifies:
- **Duplicate rules**: Configurations containing multiple rules with identical `type` and `url_path` combinations
- **Conflicting rules**: Instances where the same pattern appears in both `focus` (include) and `avoid` (exclude) lists

These checks prevent ambiguous agent behavior that could result from contradictory configuration directives.

## Dangerous Patterns Blocked by Shannon's Security Rules

Shannon's regex blacklist specifically targets five attack vectors commonly exploited through configuration injection:

| Pattern | Regex | Attack Vector Prevented |
|---------|-------|------------------------|
| Path Traversal | `/\.\.\//` | Directory traversal attacks attempting to access files outside intended paths (e.g., `../../../etc/passwd`) |
| HTML/XML Injection | `/[<>]/` | Cross-site scripting (XSS) and XML external entity (XXE) injection via angle brackets |
| JavaScript URLs | `/javascript:/i` | Execution of arbitrary JavaScript through `javascript:` protocol URLs |
| Data URLs | `/data:/i` | Data exfiltration or code execution via base64-encoded data URLs |
| File URLs | `/file:/i` | Unauthorized local file system access through `file://` protocol references |

These patterns are applied case-insensitively where appropriate (`/i` flag) to catch obfuscated attacks using mixed-case spellings like `JaVaScRiPt:`.

## Fields Subject to Dangerous Pattern Validation

Shannon applies the dangerous pattern blacklist to specific high-risk fields within the configuration structure:

**Authentication Section**
- `authentication.credentials.username`
- `authentication.credentials.password`
- `authentication.login_flow` (array of step descriptions)

**Rules Sections**
- `rules.focus[].url_path`
- `rules.focus[].description`
- `rules.avoid[].url_path`
- `rules.avoid[].description`

This targeted approach ensures that user-controlled strings which might influence agent behavior or authentication flows are sanitized, while purely internal configuration metadata (such as boolean flags or numeric timeouts) bypasses the expensive regex checks.

## Practical Examples: Safe Configurations vs. Blocked Attacks

### Example 1: Minimal Safe Configuration

The following YAML configuration passes all five validation layers because it contains no dangerous patterns and satisfies the JSON schema:

```yaml
authentication:
  login_type: form
  login_url: https://example.com/login
  credentials:
    username: alice
    password: secret123
  login_flow:
    - "Enter username"
    - "Enter password"
    - "Click submit"
  success_condition:
    type: url_contains
    value: "/dashboard"

rules:
  focus:
    - description: "Test all POST endpoints"
      type: method
      url_path: "POST"
  avoid:
    - description: "Skip health checks"
      type: path
      url_path: "/health"

```

Loading this with `parseConfig('/path/to/config.yml')` returns a validated configuration object ready for agent initialization.

### Example 2: Blocked HTML Injection Attack

This configuration attempts to inject a script tag through the password field:

```yaml
authentication:
  login_type: form
  login_url: https://example.com/login
  credentials:
    username: "admin"
    password: "p@ssw0rd<script>"
  login_flow:
    - "Enter username"
    - "Enter password"
  success_condition:
    type: url_contains
    value: "/dashboard"

```

When `parseConfig` processes this file, the dangerous pattern check (lines 70-84 in [`src/config-parser.ts`](https://github.com/KeygraphHQ/shannon/blob/main/src/config-parser.ts)) matches the `<` character in the password against the `/[<>]/` regex. Shannon immediately throws:

```

Error: authentication.credentials.password contains potentially dangerous pattern

```

The configuration is rejected before any agent initialization occurs.

### Example 3: Duplicate Rule Detection

Shannon also enforces logical consistency. This configuration contains redundant focus rules:

```yaml
rules:
  focus:
    - description: "Test GET"
      type: method
      url_path: "GET"
    - description: "Again GET"
      type: method
      url_path: "GET"

```

During the logical rule check (lines 92-100), Shannon identifies that both entries share the same `type` (`method`) and `url_path` (`GET`). It throws:

```

Error: Duplicate rule found in rules.focus[1]: method 'GET'

```

## Key Files Implementing Configuration Security

Shannon's validation pipeline spans three critical files in the repository:

| File | Purpose |
|------|---------|
| [`src/config-parser.ts`](https://github.com/KeygraphHQ/shannon/blob/main/src/config-parser.ts) | Core parsing logic, file-size enforcement, dangerous pattern regex matching, and logical consistency checks. |
| [`configs/config-schema.json`](https://github.com/KeygraphHQ/shannon/blob/main/configs/config-schema.json) | JSON-Schema definition specifying required fields, types, formats, and structural constraints. |
| [`src/types/config.ts`](https://github.com/KeygraphHQ/shannon/blob/main/src/types/config.ts) | TypeScript interfaces defining the configuration object structure used throughout the validation pipeline. |

These files work sequentially: the parser loads the schema, validates file constraints, parses YAML safely, validates against the schema, checks for dangerous patterns, and finally verifies logical rule consistency.

## Summary

Shannon's security validation rules for configuration files employ a defense-in-depth strategy that neutralizes common attack vectors before agent execution:

- **File-level guards** enforce size limits (≤1 MiB) and existence checks to prevent resource exhaustion.
- **Restrictive YAML parsing** using `FAILSAFE_SCHEMA` eliminates JavaScript evaluation and complex type attacks.
- **JSON-Schema validation** via AJV ensures structural integrity and type safety across all configuration fields.
- **Regex-based dangerous pattern detection** blocks path traversal (`../`), HTML injection (`<>`), and malicious protocols (`javascript:`, `data:`, `file:`) in user-controlled strings.
- **Logical consistency checks** prevent ambiguous agent behavior by detecting duplicate or conflicting focus/avoid rules.

This layered approach ensures that malformed, malicious, or contradictory configurations are rejected with clear error messages before the Shannon agent processes any web automation tasks.

## Frequently Asked Questions

### What happens when Shannon detects a dangerous pattern in a configuration file?

When the `parseConfig` function identifies a string matching any entry in the `DANGEROUS_PATTERNS` regex array, it immediately throws a descriptive error specifying the exact field location (e.g., "authentication.credentials.password contains potentially dangerous pattern"). The configuration loading process halts entirely, preventing the agent from initializing with compromised credentials or malicious rule definitions.

### Which specific attack vectors does Shannon's regex blacklist prevent?

Shannon's dangerous pattern detection specifically blocks five attack vectors: **path traversal** (via `../` sequences), **HTML/XML injection** (via angle brackets), **JavaScript execution** (via `javascript:` URLs), **data exfiltration** (via `data:` URLs), and **local file access** (via `file:` URLs). These patterns are tested against authentication credentials, login flow descriptions, and rule definitions to ensure user-controlled strings cannot manipulate the agent into unsafe operations.

### How does Shannon prevent JavaScript code execution during YAML parsing?

Shannon explicitly configures the YAML parser to use `FAILSAFE_SCHEMA`, which restricts the document to basic primitive types and standard collections while disabling any JavaScript evaluation or complex type reconstruction. This schema restriction neutralizes YAML-based code injection attacks (such as `!!js/function` or `!!python/object` exploits) by ensuring the parser outputs only safe, literal values that cannot execute arbitrary code.

### Can Shannon detect duplicate or conflicting rules in configuration files?

Yes, after passing dangerous pattern checks, Shannon performs logical consistency validation to detect **duplicate rules** (identical combinations of `type` and `url_path` within the same rule category) and **conflicting rules** (where a pattern appears in both `focus` and `avoid` lists simultaneously). These checks prevent ambiguous agent behavior and configuration errors that could cause unpredictable web automation results.