How ContentRouter Detects Content Type for Compression in Headroom
ContentRouter uses a two-stage detection pipeline—first invoking a Rust-based native detector and then falling back to a Python regex classifier if the content is identified as plain text—to accurately select the optimal compression strategy.
In the chopratejas/headroom repository, the ContentRouter class automatically analyzes input text to determine how ContentRouter detects content type for compression before applying specialized algorithms. This detection mechanism balances performance with accuracy by combining native Rust code with regex-based heuristics.
The Two-Stage Detection Pipeline
The content classification logic resides in headroom/transforms/content_router.py within the _detect_content method. This function implements a hierarchical detection strategy that prioritizes speed while maintaining fallback safety.
Stage 1: Rust-Based Native Detection
The primary detection path leverages the native Rust implementation exposed through headroom._core.detect_content_type. This function performs fast, low-level analysis of the input string to categorize it into a ContentType enum value.
# headroom/transforms/content_router.py
def _detect_content(content: str) -> DetectionResult:
from headroom._core import detect_content_type as _rust_detect
rust_result = _rust_detect(content) # ← Rust detection
content_type = ContentType(rust_result.content_type) # map to Python enum
# ...
The Rust detector handles common formats including source code, JSON arrays, HTML, and git diffs. When the native code identifies a specific content type, the function immediately returns the result without additional overhead.
Stage 2: Python Regex Fallback
If the Rust detector classifies the content as PLAIN_TEXT, the system invokes a secondary Python regex-based detector (_regex_detect_content_type) to catch non-text types that the native implementation may have missed, such as code fences, JSON arrays, or search results.
if content_type is ContentType.PLAIN_TEXT:
regex_result = _regex_detect_content_type(content)
if regex_result.content_type is not ContentType.PLAIN_TEXT:
return regex_result
This fallback ensures edge cases—like malformed JSON or mixed-format documents—receive proper classification without sacrificing the performance benefits of the Rust fast path.
Mapping Content Types to Compression Strategies
Once detection completes, the DetectionResult.content_type is translated into a CompressionStrategy via the static _strategy_from_detection mapping in headroom/transforms/content_router.py:
mapping = {
ContentType.SOURCE_CODE: CompressionStrategy.CODE_AWARE,
ContentType.JSON_ARRAY: CompressionStrategy.SMART_CRUSHER,
ContentType.SEARCH_RESULTS: CompressionStrategy.SEARCH,
ContentType.BUILD_OUTPUT: CompressionStrategy.LOG,
ContentType.GIT_DIFF: CompressionStrategy.DIFF,
ContentType.HTML: CompressionStrategy.HTML,
ContentType.PLAIN_TEXT: CompressionStrategy.TEXT,
}
Each strategy implements format-specific optimizations. For example, CODE_AWARE preserves indentation and syntax, while SMART_CRUSHER aggressively compresses JSON arrays by removing structural whitespace.
Mixed-Content Short-Circuit
Before initiating the detection pipeline, the router checks for mixed content using the is_mixed_content helper function. If the input contains at least two distinct indicators—such as code fences, JSON blocks, prose sections, or search results—the router immediately selects the MIXED strategy.
if is_mixed_content(content):
return CompressionStrategy.MIXED
When MIXED is selected, the router splits the input into sections and routes each section individually through the detection pipeline, ensuring each content type receives appropriate compression.
Implementation in the compress() Method
The complete decision flow inside the compress() method orchestrates these components:
mixed = is_mixed_content(content) # mixed-content check
detection = _detect_content(content) # Rust → Python fallback
strategy = (
CompressionStrategy.KOMPRESS
if self._runtime_force_kompress
else self._determine_strategy(content) # mapping + mixed-content guard
)
This implementation ensures that content type detection occurs only when necessary, respecting runtime configuration flags that may force specific compression modes.
Practical Usage Examples
Detecting a JSON Array
When the router receives valid JSON, it automatically selects the SMART_CRUSHER strategy:
from headroom.transforms import ContentRouter
router = ContentRouter()
content = '[{"id":1,"name":"Alice"},{"id":2,"name":"Bob"}]'
result = router.compress(content) # detection → SMART_CRUSHER
print(result.strategy_used) # → smart_crusher
Handling Mixed Content
Documents containing multiple formats trigger section-based routing:
content = """\
```python
def hello(): print("hi")
{ "users": [{"id": 1}, {"id": 2}] } """ router = ContentRouter() result = router.compress(content) # mixed → section routing
print(result.strategy_used) # → mixed
print([d.strategy.value for d in result.routing_log])
### Disabling Smart Detection
You can force plain-text compression by disabling specific detection paths:
```python
router = ContentRouter()
router.config.enable_smart_crusher = False # disable JSON path
content = '[1,2,3,4]' # would be JSON_ARRAY
result = router.compress(content)
print(result.strategy_used) # → text (fallback to KOMPRESS)
Summary
- Two-stage detection: ContentRouter first runs
headroom._core.detect_content_type(Rust) and falls back to_regex_detect_content_type(Python) if the result is plain text. - Strategy mapping: Detected content types map to specific strategies like
CODE_AWARE,SMART_CRUSHER, orLOGvia_strategy_from_detection. - Mixed-content handling: The
is_mixed_contentcheck short-circuits detection to enable section-by-section routing when multiple formats are present. - Source location: All detection logic resides in
headroom/transforms/content_router.py, with the Rust bindings exposed inheadroom/_core/__init__.py.
Frequently Asked Questions
What happens if the Rust detector fails to identify the content type?
If the Rust implementation returns ContentType.PLAIN_TEXT, the system automatically triggers _regex_detect_content_type as a safety net. This Python-based regex fallback scans for patterns like code fences, JSON brackets, or HTML tags that the native detector might have missed, ensuring accurate classification of edge cases.
How does ContentRouter handle documents containing multiple content types?
Before running detection, the router checks is_mixed_content to identify inputs containing at least two distinct format indicators (code, JSON, prose, or search results). When mixed content is detected, the router returns CompressionStrategy.MIXED and splits the document into sections, processing each section individually through the detection pipeline to apply format-specific compression.
Can I disable smart detection and force a specific compression strategy?
Yes. Setting configuration flags like enable_smart_crusher = False or using the _runtime_force_kompress flag bypasses the detection logic. When forced, the router skips content type detection entirely and uses the specified strategy or defaults to CompressionStrategy.KOMPRESS regardless of the input format.
Where is the content type detection logic implemented in the source code?
The primary detection workflow resides in headroom/transforms/content_router.py, specifically within the _detect_content method (lines 12-36) and the _strategy_from_detection mapping (lines 134-142). The Rust native detection is implemented in the headroom._core module, exposed via detect_content_type in headroom/_core/__init__.py.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →