How RTK Optimizes Output from Python Linters: Ruff and mypy Integration
RTK applies a three-stage filter pipeline to Python linter output, converting verbose diagnostics into concise summaries that reduce LLM token consumption by 60–90%.
The rtk-ai/rtk repository implements an intelligent interception layer to optimize output from Python linters like Ruff and mypy. By capturing stdout before it reaches the terminal and rebuilding diagnostics into compact reports, RTK eliminates noise while preserving essential file paths, error codes, and fixable issue counts.
The Filter Pipeline Architecture
RTK processes linter output through a detect‑run‑filter pipeline implemented in Rust. This architecture ensures raw console output never prints directly to the terminal; instead, it passes through specialized filters that extract only actionable intelligence.
Command Preparation
RTK first detects whether the user invoked a check or format operation. For Ruff, it automatically injects --output-format=json when running ruff check, guaranteeing deterministic parsing. This structured output enables reliable extraction of error codes, line numbers, and fixability metadata without fragile text parsing.
Execution via runner.rs
The core::runner::run_filtered function in src/core/runner.rs executes the linter as a child process and captures its complete stdout stream. This function holds the raw output in memory and immediately passes it to a tool-specific filter function, preventing unprocessed text from reaching the developer's screen.
Summarization and Compression
Each linter implements a dedicated filter that rebuilds output into a concise format:
- Total issue and file counts
- Top offending rules or error codes with frequency indicators
- Compact file paths for the most problematic sources
- Count of automatically fixable issues with contextual hints
- Single-line success messages when no issues exist
This compression typically reduces output from hundreds of lines to fewer than ten, achieving 60–90% token savings.
Ruff Optimization in ruff_cmd.rs
RTK handles Ruff through src/cmds/python/ruff_cmd.rs, with distinct implementations for check and format operations.
Check Command JSON Filtering
For ruff check, RTK forces JSON output (lines 37–94) and deserializes the results in the JSON filter (lines 96–150). This filter groups diagnostics by rule code (e.g., "F401") and file path, counts fixable items, and generates a "Top rules" frequency table. It appends a hint to run ruff check --fix when applicable, transforming verbose per-line errors into a scannable summary.
Format Command Handling
The format filter (lines 198–270) processes ruff format output to extract only the list of files requiring reformatting. Instead of displaying diff details, it emits a count of files needing changes versus those already formatted, plus a hint to apply fixes.
# Ruff check – full output collapsed to a 4‑line summary
$ rtk ruff check .
Ruff: 12 issues in 5 files (3 fixable)
═══════════════════════════════════════
Top rules:
F401 (6x)
E501 (3x)
Top files:
src/main.py (5 issues)
src/utils.py (3 issues)
[hint] Run `ruff check --fix` to auto-fix 3 issues
# Ruff format – tells you only the files that need reformatting
$ rtk ruff format src/
Ruff format: 2 files need formatting
═══════════════════════════════════════
1. src/main.py
2. tests/test_utils.py
3 files already formatted
[hint] Run `ruff format` to format these files
mypy Optimization in mypy_cmd.rs
For mypy, RTK uses src/cmds/python/mypy_cmd.rs (lines 9–33) to resolve the executable and invoke the runner with filter_mypy_output.
Regex-Based Parsing
Since mypy lacks JSON output, the filter (lines 43–130) compiles a regex to parse each diagnostic line, extracting file paths, line numbers, error codes, and messages. It groups errors by source file and generates a "Top codes" summary when multiple error types appear, surfacing the most frequent categories (e.g., "return-value (3x)").
No-Issue Handling
When mypy reports success, the filter returns mypy: No issues found (lines 28–34), replacing verbose or empty success output with an explicit minimal confirmation.
# mypy – groups errors by file and shows top error codes
$ rtk mypy src/
mypy: 7 errors in 3 files
═══════════════════════════════════════
Top codes: return-value (3x), name-defined (2x)
src/models/user.py (3 errors)
L8: [return-value] Incompatible return value type
L10: [name-defined] Name `foo` is not defined
L20: [return] Missing return statement
src/server/auth.py (2 errors)
L12: [return-value] Incompatible return value type
L15: [arg-type] Argument 1 has incompatible type
src/utils/helpers.py (2 errors)
L5: [assignment] Incompatible types in assignment
L9: [arg-type] Argument 2 has incompatible type
Token Savings and Performance Impact
Raw linter output often spans hundreds of lines averaging 80 characters each. RTK's filtered output compresses this to under 300 characters—typically fewer than ten lines. This yields approximately 70% token reduction on average for both tools.
The filtering process adds negligible runtime overhead. Since run_filtered processes the complete output buffer after the child process terminates, the overhead is insignificant compared to the linting operation itself. The reduction in output size often improves subsequent processing latency in LLM pipelines.
Summary
- RTK implements a three-stage pipeline (command preparation, execution via
run_filtered, and structured filtering) to process Python linter output. - Ruff integration in
src/cmds/python/ruff_cmd.rsforces JSON output for check commands (lines 37–94) and applies specialized filters for check (lines 96–150) and format (lines 198–270) operations. - mypy integration in
src/cmds/python/mypy_cmd.rsuses regex-based parsing (lines 43–130) to group type errors by file and highlight top error codes. - The token reduction averages 60–90%, collapsing verbose diagnostics into concise summaries under 300 characters.
- When no issues exist, RTK outputs a minimal confirmation (e.g.,
mypy: No issues found) rather than verbose success messages.
Frequently Asked Questions
Why does RTK force JSON output for Ruff but use regex for mypy?
Ruff provides a stable --output-format=json flag that emits structured diagnostics, allowing RTK to parse error codes, line numbers, and fixability metadata without fragile text manipulation (implemented in src/cmds/python/ruff_cmd.rs lines 37–94). mypy lacks native JSON output support, so RTK implements regex-based parsing in src/cmds/python/mypy_cmd.rs (lines 43–130) to extract structured information from plain text.
How does RTK determine which are the "Top files" and "Top rules"?
The filters aggregate diagnostics into hash maps keyed by file path and error code. They count occurrences of each key, sort by frequency, and display the top entries with compact relative paths. For Ruff, this logic appears in lines 96–150 of ruff_cmd.rs, while mypy implements similar grouping in lines 43–130 of mypy_cmd.rs.
What happens when a linter finds no issues?
When Ruff or mypy exits cleanly, RTK intercepts the empty or minimal success output and replaces it with a single explicit line, such as mypy: No issues found (implemented in mypy_cmd.rs lines 28–34). This confirms successful execution while consuming the minimum possible tokens.
Does the filtering process slow down linter execution?
No. RTK executes the linter normally via core::runner::run_filtered and processes the complete output buffer only after the child process terminates. The filtering overhead is negligible compared to the linting operation itself, and the resulting reduction in output size often improves subsequent processing latency in LLM pipelines.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →