How the Format Command Cleans XiaoHongShu API Output in Agent-Reach

The format command in Agent-Reach sanitizes raw XiaoHongShu JSON responses by stripping unnecessary nesting and redundant fields, outputting a lightweight structure optimized for LLM processing.

Agent-Reach provides a dedicated CLI sub-command that processes raw API responses from XiaoHongShu (XHS), transforming bloated payloads into clean, token-efficient JSON. This command eliminates the need for manual data wrangling when preparing social media content for downstream language model workflows.

CLI Entry Point and stdin Handling

The formatting pipeline begins in agent_reach/cli.py, where the format sub-command is registered under p_format = sub.add_parser("format", …) and delegated to the _cmd_format function (lines 88-107).

The _cmd_format implementation reads the entire payload from stdin, verifies that it contains valid JSON, and aborts with a descriptive error if the input is missing or malformed. This design follows Unix philosophy, allowing users to pipe API responses directly into the tool without intermediate files.

Platform-Specific Cleaning Logic

For the xhs platform, the command imports format_xhs_result from agent_reach/channels/xiaohongshu.py (lines 40-59). This function serves as the primary dispatcher, recognizing three distinct shapes of XiaoHongShu data:

  • A single note dictionary
  • A list of note dictionaries
  • Wrapper objects such as {"items": …} or {"data": {"items": …}}

Once normalized, each note delegates to _clean_note, which performs aggressive field pruning to reduce token usage dramatically (addressing issue #134).

The _clean_note Extraction Strategy

The _clean_note function extracts only fields essential for LLM consumption:

  • Core identifiers: id, note_id, xsec_token
  • Human-readable content: title, desc, content
  • Author metadata: nickname and user IDs
  • Engagement metrics: liked_count, collected_count, comment_count, share_count
  • Media and taxonomy: image URLs, tags, and top-level comments (each normalized via _clean_comment)

Auxiliary cleaning occurs in _clean_comment (defined in the same file), which ensures comment objects maintain consistent structure without extraneous metadata.

Practical Usage Examples

The command accepts raw JSON via stdin and emits pretty-printed JSON to stdout, ready for piping or redirection.


# Pipe raw API response directly from curl

curl -s "https://api.xiaohongshu.com/fe_api/bill/detail/notes?keyword=travel" \
  | python -m agent_reach.cli format xhs > cleaned_xhs.json

# Process a local file containing raw XHS output

cat raw_xhs_response.json | python -m agent_reach.cli format xhs > cleaned.json

# Quick test with inline JSON

printf '{"items":[{"note_id":"123","title":"Demo","desc":"Sample","user":{"nickname":"Alice"}}]}' \
  | python -m agent_reach.cli format xhs | jq .

Example output:

[
  {
    "note_id": "123",
    "title": "Demo",
    "desc": "Sample",
    "user": {
      "nickname": "Alice"
    }
  }
]

Summary

  • The format command in agent_reach/cli.py provides a CLI interface that reads raw JSON from stdin and validates input before processing.
  • Platform-specific logic resides in agent_reach/channels/xiaohongshu.py, where format_xhs_result handles multiple response wrappers and delegates to _clean_note.
  • The cleaning process preserves only LLM-relevant fields (identifiers, content, author info, engagement stats, media URLs) while stripping redundant nesting.
  • Output is pretty-printed JSON sent to stdout, enabling seamless integration with Unix pipelines and file redirection.

Frequently Asked Questions

What input format does the format command expect for XiaoHongShu data?

The command expects raw JSON from stdin, which can be a single note object, a list of notes, or a wrapped response containing an items or data.items structure. If the input is empty or malformed JSON, _cmd_format exits with a clear error message.

Which specific fields are preserved when cleaning XiaoHongShu API output?

The _clean_note function retains core identifiers (id, note_id, xsec_token), textual content (title, desc, content), author information (nickname and IDs), engagement metrics (likes, collections, comments, shares), image URLs, tags, and normalized top-level comments. All other metadata is discarded to minimize token consumption.

How does the format command handle different XiaoHongShu response structures?

The format_xhs_result function in agent_reach/channels/xiaohongshu.py detects three patterns: individual note dictionaries, lists of notes, and wrapper objects with nested items arrays. It normalizes these into a uniform list before applying the cleaning logic, ensuring consistent output regardless of API endpoint variations.

Can the format command process local files instead of stdin?

While the command specifically reads from stdin, you can redirect file contents into stdin using standard shell operators like cat file.json | python -m agent_reach.cli format xhs. This approach maintains the tool's pipeline-friendly design while supporting file-based workflows.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →