How to Use the Format Command to Clean XiaoHongShu API Output in Agent-Reach

The format command in Agent-Reach reads raw XiaoHongShu JSON from stdin and outputs a sanitized, LLM-ready structure by stripping unnecessary nesting and redundant fields.

Agent-Reach (available at Panniantong/Agent-Reach) provides a dedicated CLI sub-command specifically designed to process noisy XiaoHongShu API responses. This tool extracts only the essential fields needed for downstream LLM processing, dramatically reducing token usage and eliminating manual data wrangling.

How the Format Command Works

The cleaning workflow is encapsulated in two main components: the CLI entry point and the platform-specific formatter.

CLI Entry Point and Stdin Handling

In agent_reach/cli.py, the format sub-command is registered under p_format = sub.add_parser("format", …) and delegated to the _cmd_format function (lines 88-107). This handler reads the entire payload from stdin, verifies that it is non-empty JSON, and aborts with a clear error if the input is missing or malformed.

Platform-Specific Formatter

For the xhs platform, the command imports format_xhs_result from agent_reach/channels/xiaohongshu.py (lines 40-59). This function recognizes three distinct shapes of XiaoHongShu data:

  • A single note dictionary
  • A list of note objects
  • Wrapper objects such as {"items": …} or {"data": {"items": …}}

Each note is then delegated to _clean_note, which performs aggressive pruning to reduce token usage.

Understanding the XiaoHongShu Data Cleaner

The cleaning logic in agent_reach/channels/xiaohongshu.py focuses on extracting only LLM-relevant fields while normalizing nested structures.

Extracted Fields

The _clean_note function preserves the following data points:

  • Core identifiers: id, note_id, xsec_token
  • Human-readable content: title, desc, content
  • Author information: nickname and associated IDs
  • Engagement metrics: liked_count, collected_count, comment_count, share_count
  • Media and metadata: Image URLs, tags, and top-level comments (each normalized via _clean_comment (lines 33-46))

Output Format

The cleaned structure is emitted as pretty-printed JSON to stdout, ready to be piped into other tools or saved to a file. This format eliminates the deep nesting and redundant wrapper objects common in raw XiaoHongShu API responses.

Practical Usage Examples

You can invoke the format command through python -m agent_reach.cli or directly if the package is installed as agent-reach.

Pipe Live API Responses

Stream raw API output directly into the formatter:

curl -s "https://api.xiaohongshu.com/fe_api/bill/detail/notes?keyword=travel" \
  | python -m agent_reach.cli format xhs > cleaned_xhs.json

Process Local Files

Clean existing JSON files without writing additional Python code:

cat raw_xhs_response.json | python -m agent_reach.cli format xhs > cleaned.json

Test with Inline JSON

For quick validation or testing:

printf '{"items":[{"note_id":"123","title":"Demo","desc":"Sample","user":{"nickname":"Alice"}}]}' \
  | python -m agent_reach.cli format xhs | jq .

Output:

[
  {
    "note_id": "123",
    "title": "Demo",
    "desc": "Sample",
    "user": {
      "nickname": "Alice"
    }
  }
]

Summary

  • The format command in Agent-Reach provides a CLI-native way to sanitize XiaoHongShu API output without writing custom parsing scripts.
  • Input is read from stdin and must be valid JSON; the command aborts with a clear error if validation fails.
  • The format_xhs_result function in agent_reach/channels/xiaohongshu.py handles multiple response shapes (single notes, lists, or nested wrappers) and delegates cleaning to _clean_note.
  • Only essential LLM fields are preserved, including identifiers, content, author info, engagement metrics, and media URLs, significantly reducing token consumption.
  • Cleaned output is emitted as pretty-printed JSON to stdout, compatible with pipes and redirects.

Frequently Asked Questions

What input formats does the XiaoHongShu format command support?

The format_xhs_result function in agent_reach/channels/xiaohongshu.py recognizes three input variations: a single note dictionary, a flat list of notes, or wrapper objects containing an items or data.items key. This flexibility allows the command to handle both direct API responses and pre-saved JSON files without manual preprocessing.

Which fields are preserved when cleaning XiaoHongShu API data?

The _clean_note function extracts core identifiers (id, note_id, xsec_token), human-readable content (title, desc, content), author metadata (nickname and IDs), engagement statistics (likes, collections, comments, shares), and media elements (image URLs, tags, and top-level comments). All other nested metadata and redundant wrapper fields are stripped to minimize token usage.

Can I use the format command with local JSON files instead of live API calls?

Yes. The _cmd_format function reads exclusively from stdin, so you can pipe any local file containing XiaoHongShu JSON into the command using cat raw_xhs_response.json | python -m agent_reach.cli format xhs. This approach works for saved API responses, exported data dumps, or test fixtures.

How does the format command handle invalid JSON input?

The _cmd_format implementation in agent_reach/cli.py (lines 88-107) validates that stdin contains non-empty, well-formed JSON before processing. If the input is missing, empty, or malformed, the command aborts immediately with a clear error message rather than passing corrupt data to the formatter.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →