How to Use the JSON Output Format in betterhtmlchunking for Programmatic Processing with jq
Activate JSON mode by passing --format json (or -f json) to the CLI, then pipe the clean stdout stream to jq for filtering and extraction.
The betterhtmlchunking library splits HTML documents into manageable chunks for LLM processing. When you need to automate workflows or extract specific metadata, the JSON output format provides a structured, machine-readable alternative to plain text output.
Enabling JSON Output Mode
To switch from the default human-readable output to machine-readable JSON, append the --format json flag to any CLI invocation. The short form -f json is also accepted.
When this flag is present, the CLI bypasses standard output formatting and constructs a Python dictionary named output. This dictionary is serialized with json.dumps(..., indent=2, ensure_ascii=False) and printed to stdout. All diagnostic logs are emitted to stderr, ensuring the JSON stream remains pure and safe to pipe directly into tools like jq without filtering noise.
Understanding the JSON Schema Structure
The JSON document returned by betterhtmlchunking contains top-level metadata about the chunking operation and an array of individual chunk objects.
Top-Level Metadata Fields
According to the implementation in betterhtmlchunking/cli.py (lines 112-114), the root object includes:
total_chunks: Integer count of chunks produced for the input document.max_length: The--max-lengthvalue supplied by the user, indicating the size constraint applied during chunking.compared_by: String value of"html"if chunk size was measured on raw HTML, or"text"if measured on rendered text (triggered by the--textflag).
The Chunks Array
The chunks field contains an array where each element describes a single chunk. As implemented in betterhtmlchunking/cli.py (lines 104-108), each chunk object provides:
index: Integer position in the DOM traversal order.html: The raw HTML markup for the chunk.text: The plain-text rendering of the same chunk with tags stripped.html_length: Character count of thehtmlstring.text_length: Character count of thetextstring.
Practical jq Examples for betterhtmlchunking JSON
The following examples demonstrate how to process betterhtmlchunking JSON output using jq in Unix pipelines.
Extracting Total Chunk Count
To retrieve just the number of chunks generated:
cat my_page.html | \
betterhtmlchunking --max-length 5000 --format json | \
jq '.total_chunks'
This outputs a single integer (e.g., 12) suitable for scripting conditionals.
Filtering Chunks by Index
To extract the HTML content of a specific chunk (e.g., index 2):
cat my_page.html | \
betterhtmlchunking --max-length 4000 --format json | \
jq -r '.chunks[] | select(.index == 2) | .html'
The -r flag returns raw text, removing JSON string quotes and escaping, delivering clean HTML markup.
Exporting HTML to Individual Files
To save each chunk to its own file:
cat my_page.html | \
betterhtmlchunking --max-length 2500 --format json | \
jq -r '.chunks[] | "chunk_\(.index).html \(.html)"' | \
while read -r filename content; do
printf '%s' "$content" > "$filename"
done
This creates chunk_0.html, chunk_1.html, etc., each containing the exact HTML chunk.
Processing Text Content Streams
To stream each chunk's plain text with custom prefixes:
cat my_page.html | \
betterhtmlchunking --max-length 2000 --format json | \
jq -c '.chunks[]' | \
while read -r chunk; do
echo "$chunk" | jq -r '.text' | sed -e 's/^/>>> /'
done
This outputs each chunk's text content prefixed with >>> , useful for debugging or feeding into text-processing pipelines.
Key Implementation Details
The JSON output functionality is implemented in betterhtmlchunking/cli.py. The code constructs the output dictionary during the chunking process and serializes it using Python's standard library:
json.dumps(output, indent=2, ensure_ascii=False)
This ensures human-readable indentation while preserving Unicode characters without escaping. The separation of concerns—JSON to stdout and logs to stderr—is verified in tests/test_cli.py, which confirms that the JSON schema contains all required fields and that diagnostic output does not contaminate the data stream.
Summary
- Activate JSON output format with
--format jsonor-f jsonwhen running thebetterhtmlchunkingCLI. - The JSON schema includes metadata (
total_chunks,max_length,compared_by) and a chunks array withhtml,text, and length metrics for each chunk. - Output is sent to stdout while logs go to stderr, enabling safe piping to
jqand other JSON processors. - Use
jqselectors like.chunks[],select(.index == N), and-rfor raw string output to extract specific chunks or metadata.
Frequently Asked Questions
How do I activate JSON output format in betterhtmlchunking?
Pass the --format json flag (or -f json shorthand) to the CLI command. This instructs the tool to emit a structured JSON document instead of the default human-readable text format.
What fields are included in the JSON output?
The root object contains total_chunks (integer), max_length (integer), and compared_by (string). The chunks array contains objects with index, html, text, html_length, and text_length fields, as implemented in betterhtmlchunking/cli.py.
Can I pipe betterhtmlchunking JSON output directly to jq?
Yes. Because the tool writes the JSON payload to stdout and diagnostic logs to stderr, you can safely pipe the output directly into jq without filtering. For example: betterhtmlchunking --format json < page.html | jq '.total_chunks'.
How does betterhtmlchunking handle logging when outputting JSON?
All log messages are emitted to stderr, while the JSON data stream is written to stdout. This separation ensures that JSON parsers receive a clean, valid document even when verbose logging is enabled, as verified in tests/test_cli.py.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →