# RTK Docker and kubectl Log Deduplication: How Line Counts Work

> RTK deduplicates Docker and kubectl logs with regex normalization and severity classification. See how unique messages aggregate with line counts displayed using the ×N prefix.

- Repository: [rtk-ai/rtk](https://github.com/rtk-ai/rtk)
- Tags: how-to-guide
- Published: 2026-04-24

---

**RTK deduplicates Docker and kubectl logs by normalizing each line with compiled regex patterns, classifying by severity level, and aggregating occurrences into HashMaps, displaying unique messages with an `×N` prefix indicating the total line count.**

The rtk-ai/rtk repository provides a unified log analysis CLI that reduces noise in container orchestration output. RTK processes raw streams from `docker logs` and `kubectl logs` through a shared deduplication pipeline defined in [`src/cmds/system/log_cmd.rs`](https://github.com/rtk-ai/rtk/blob/main/src/cmds/system/log_cmd.rs), collapsing repetitive entries while preserving the original text of unique events.

## Core Log Deduplication Engine

The deduplication logic resides in [`src/cmds/system/log_cmd.rs`](https://github.com/rtk-ai/rtk/blob/main/src/cmds/system/log_cmd.rs), specifically within the `run_stdin_str` function and its supporting structures. This engine ingests raw log strings and performs a four-stage analysis to produce condensed, countable output.

### 1. Line Normalization via Regex

First, the engine normalizes each line to create a canonical key for comparison. Using `lazy_static!` compiled regex constants defined at lines 12-21, the system replaces variable content with standardized placeholders:

- `<TIMESTAMP>` for date/time strings
- `<UUID>` for UUID patterns
- `<HEX>` for hexadecimal numbers
- `<NUM>` for long numeric sequences
- `<PATH>` for file system paths

This normalization ensures that timestamps or unique IDs do not prevent otherwise identical log messages from being recognized as duplicates.

### 2. Severity Classification

After normalization, each line is lower-cased and scanned for severity keywords. The engine categorizes entries into three buckets based on the presence of **error**, **fatal**, **panic**, **warn**, or **info** in the message text. This classification determines which counting HashMap will track the occurrence.

### 3. Counting Occurrences with HashMaps

The engine maintains three `HashMap<String, usize>` structures: `error_counts`, `warn_counts`, and `info_counts`. As implemented in lines 69-100 of [`log_cmd.rs`](https://github.com/rtk-ai/rtk/blob/main/log_cmd.rs), the normalized line serves as the HashMap key, while the value tracks the number of occurrences.

When a normalized key is encountered for the first time (counter equals zero), the engine stores the original (non-normalized) line text in a separate "unique" vector. This preserves the original formatting for final display while the normalized version handles deduplication logic.

### 4. Summary Generation with Line Counts

After processing the complete input stream, the engine constructs the final output (lines 105-166). The summary displays:

- Total counts per severity category
- Unique message counts
- Individual log lines prefixed with `[×N]` when `N > 1`

The `[×N]` syntax represents the line count stored in the HashMap counter, immediately showing how many times that specific log pattern appeared in the original stream.

## Docker and kubectl Integration

Both container runtimes invoke the same deduplication engine through thin wrappers located in [`src/cmds/cloud/container.rs`](https://github.com/rtk-ai/rtk/blob/main/src/cmds/cloud/container.rs). These functions execute the underlying CLI commands, capture stdout, and delegate processing to `crate::log_cmd::run_stdin_str`.

### Docker Logs Implementation

The `docker_logs()` function constructs a `docker logs` command with a 100-line tail limit and processes the output through the runner module:

```rust
// src/cmds/cloud/container.rs – docker_logs()
let mut cmd = resolved_command("docker");
cmd.args(["logs", "--tail", "100", pod]);
…
runner::run_filtered(
    cmd,
    "docker",
    &label,
    |stdout| {
        format!(
            "Logs for {}:\n{}",
            pod,
            crate::log_cmd::run_stdin_str(stdout)
        )
    },
    RunOptions::stdout_only().early_exit_on_failure(),
)

```

This wrapper prepends `[docker]` branding to the output before printing the deduplicated results.

### kubectl Logs Implementation

The `kubectl_logs()` function follows an identical pattern, invoking `kubectl logs` and passing the captured stdout to the same `run_stdin_str` engine:

```rust
// src/cmds/cloud/container.rs – kubectl_logs()
let mut cmd = resolved_command("kubectl");
cmd.args(["logs", "--tail", "100", pod]);
…
runner::run_filtered(
    cmd,
    "kubectl",
    &label,
    |stdout| {
        format!(
            "Logs for {}:\n{}",
            pod,
            crate::log_cmd::run_stdin_str(stdout)
        )
    },
    RunOptions::stdout_only().early_exit_on_failure(),
)

```

Both implementations rely on [`src/core/runner.rs`](https://github.com/rtk-ai/rtk/blob/main/src/core/runner.rs) to execute the external binary and capture output, while [`src/core/tracking.rs`](https://github.com/rtk-ai/rtk/blob/main/src/core/tracking.rs) monitors execution metrics.

## Usage Examples

When running RTK commands, the deduplication engine automatically collapses repetitive container logs:

```bash

# Docker container log analysis

$ rtk docker logs my_container
[docker] Logs for my_container:
Log Summary
   [error] 12 errors (3 unique)
   [warn]  5 warnings (2 unique)
   [info]  0 info messages

[ERRORS]
   [×8] Connection timed out to <PATH>/api
   [×3] Failed to start service <UUID>

[WARNINGS]
   [×5] Retrying connection (attempt 1 of 5)

```

```bash

# Kubernetes pod log analysis

$ rtk kubectl logs my-pod
[kubectl] Logs for my-pod:
Log Summary
   [error] 7 errors (2 unique)
   [warn]  4 warnings (1 unique)
   [info]  13 info messages

[ERRORS]
   [×5] CrashLoopBackOff: container <UUID> terminated
   [×2] Failed to pull image <PATH>/my-image:latest

[WARNINGS]
   [×4] ImagePullBackOff: retrying

```

The `×N` prefix indicates the deduplication count recorded by the engine's HashMap counters.

## Summary

- **RTK deduplicates logs** from Docker and kubectl through a shared engine in [`src/cmds/system/log_cmd.rs`](https://github.com/rtk-ai/rtk/blob/main/src/cmds/system/log_cmd.rs) that normalizes lines using compiled regex patterns for timestamps, UUIDs, hex values, numbers, and paths.
- **Line counting** is implemented via `HashMap<String, usize>` structures that track occurrences while preserving the original text of unique messages.
- **Wrappers** in [`src/cmds/cloud/container.rs`](https://github.com/rtk-ai/rtk/blob/main/src/cmds/cloud/container.rs) execute `docker logs` and `kubectl logs`, capturing stdout and passing it to `run_stdin_str` for processing.
- **Output formatting** displays unique messages with an `[×N]` prefix showing the total count of collapsed duplicate lines.
- **Runner integration** in [`src/core/runner.rs`](https://github.com/rtk-ai/rtk/blob/main/src/core/runner.rs) handles external command execution and output capture, ensuring consistent processing across container runtimes.

## Frequently Asked Questions

### How does RTK identify duplicate log lines?

RTK identifies duplicates by first normalizing each line using regex patterns that replace variable content like timestamps, UUIDs, and paths with placeholders such as `<TIMESTAMP>` and `<UUID>`. This normalized string serves as the key in a HashMap; if the same normalized pattern appears again, the engine increments the counter rather than storing a new entry.

### What placeholders does RTK use for log normalization?

According to the `lazy_static!` definitions in [`src/cmds/system/log_cmd.rs`](https://github.com/rtk-ai/rtk/blob/main/src/cmds/system/log_cmd.rs) (lines 12-21), RTK replaces matched patterns with `<TIMESTAMP>`, `<UUID>`, `<HEX>`, `<NUM>`, and `<PATH>`. These placeholders ensure that log messages differing only in transient values are recognized as identical during deduplication.

### Can I see the original log text after deduplication?

Yes. The engine stores the original (non-normalized) line text in a "unique" vector the first time a normalized pattern is encountered (when the HashMap counter equals zero). This preserved text is what appears in the final output, ensuring you see the actual log message rather than the placeholder version used for matching.

### Where is the deduplication logic implemented in the RTK source code?

The core deduplication logic is implemented in [`src/cmds/system/log_cmd.rs`](https://github.com/rtk-ai/rtk/blob/main/src/cmds/system/log_cmd.rs), specifically within the `run_stdin_str` function and its associated counting logic (lines 62-166). Docker and kubectl specific integrations are located in [`src/cmds/cloud/container.rs`](https://github.com/rtk-ai/rtk/blob/main/src/cmds/cloud/container.rs), where the `docker_logs()` and `kubectl_logs()` functions invoke the shared engine via `crate::log_cmd::run_stdin_str`.