How the `format xhs` Command Cleans Xiaohongshu API Output

The format xhs command normalizes raw Xiaohongshu API responses into compact, deterministic JSON by extracting only essential fields like id, title, desc, and engagement metrics while discarding metadata and nested wrappers.

The format xhs command is a critical component of the Agent-Reach repository designed to reduce token bloat when feeding Xiaohongshu (XHS) content into LLM pipelines. Located in agent_reach/channels/xiaohongshu.py, this functionality transforms verbose CLI/API outputs into minimal, predictable structures that preserve semantic value while eliminating noise.

Core Processing Pipeline

Input Type Discrimination

The process begins in format_xhs_result (lines 46‑58) where the function discriminates between list payloads, dictionary wrappers, and plain values. When encountering lists, it processes items iteratively; for dictionaries, it inspects for known wrapper keys including items, data.items, and data.notes to locate the actual content.

Note Extraction and Field Whitelisting

Once inside _clean_note (lines 67‑69), the extractor navigates nested structures by checking both note_card and note keys to retrieve the inner note object. The function then applies a strict field whitelist (lines 73‑76), retaining only id, note_id, xsec_token, title, desc, type, and time while stripping all other metadata.

Content Fallback and User Simplification

To handle inconsistent API schemas, the implementation includes a content fallback mechanism (lines 78‑80): if desc is absent but content exists, the latter is promoted to a top-level key. Author data is reduced to a minimal object containing just nickname and user_id (or nick_name as fallback) as implemented in lines 82‑87.

Engagement Metrics and Image Reduction

Engagement counters—likes, collects, comments, and shares—are aggregated from potential locations including interact_info, note_interact_info, or the top-level note (lines 89‑98). For media, the function normalizes image_list or images_list arrays into simple URL strings, discarding dimensions, trace IDs, and other metadata (lines 100‑112).

Tag Simplification and Comment Handling

Tag arrays are flattened from complex objects to plain strings containing only tag names (lines 114‑124). When comments are present, each entry is processed by _clean_comment (lines 33‑46) to retain only the content, author nickname, like_count, and sub_comment_count.

Implementation Architecture

The format_xhs_result Entry Point

This function serves as the primary entry point in agent_reach/channels/xiaohongshu.py, handling input validation and dispatching. Non-dictionary and non-list inputs (such as strings or None) pass through unchanged (lines 59‑60), ensuring robustness against malformed inputs.

Helper Functions

The _clean_note function contains the bulk of the cleaning logic (lines 62‑130), orchestrating field extraction, normalization, and metric aggregation. For comment threads, _clean_comment provides analogous reduction capabilities, focusing exclusively on textual content and engagement signals.

Usage Examples

Command-Line Interface

Integrate with xhs-cli to process raw API output directly:


# Fetch and format a single note

xhs read 123456 | agent-reach format xhs

The resulting JSON contains only whitelisted fields:

{
  "id": "abc123",
  "title": "测试笔记",
  "desc": "这是正文内容",
  "type": "normal",
  "xsec_token": "tok_xxx",
  "user": {
    "nickname": "小红",
    "user_id": "u123"
  },
  "liked_count": "100",
  "collected_count": "50",
  "comment_count": "20",
  "share_count": "10",
  "images": [
    "https://img.example.com/1.jpg"
  ],
  "tags": ["旅行", "美食"]
}

Programmatic Integration

Import the cleaner directly into Python workflows to process search results or individual notes:

from agent_reach.channels.xiaohongshu import format_xhs_result
import json
import subprocess

# Example: Processing search results

raw_data = json.loads(subprocess.check_output(["xhs", "search", "AI", "-n", "2"]))
cleaned_results = format_xhs_result(raw_data)

# cleaned_results is now a list of compact dictionaries suitable for LLM prompts

print(json.dumps(cleaned_results, ensure_ascii=False, indent=2))

For single-note processing:

raw_note = {...}  # Dict from xhs-cli

clean_note = format_xhs_result(raw_note)

Key Source Files

Summary

  • The format xhs command reduces Xiaohongshu API payloads by extracting only essential fields and discarding metadata.
  • Input discrimination handles lists, wrapper dictionaries, and plain values transparently.
  • Field whitelisting retains id, title, desc, type, and authentication tokens while removing noise.
  • Content fallbacks ensure content is preserved when desc is missing.
  • Image normalization converts complex image objects to simple URL arrays.
  • Comment reduction extracts only text, author, and engagement metrics from nested comment trees.
  • The implementation significantly reduces token usage for downstream LLM consumption (see issue #134).

Frequently Asked Questions

What fields does the format xhs command preserve from the API?

The command whitelists specific top-level fields: id, note_id, xsec_token, title, desc (or content as fallback), type, and time. It also extracts simplified user objects (nickname, user_id), engagement metrics (likes, collects, comments, shares), image URLs, and tag names.

How does the command handle different Xiaohongshu API response formats?

The format_xhs_result function in agent_reach/channels/xiaohongshu.py detects whether the input is a list, a dictionary with wrapper keys like items or data.notes, or a plain value. Lists are processed iteratively, dictionaries are unwrapped to find the actual note data, and other types pass through unchanged.

Can I use the formatting logic outside the CLI?

Yes. The cleaning functionality is available as a Python import from agent_reach.channels.xiaohongshu. Simply call format_xhs_result(raw_json) with any dictionary or list returned by the XHS API to receive the normalized output.

How are comments processed by the format xhs command?

When present, comments are processed by the _clean_comment helper function (lines 33‑46 in xiaohongshu.py). Each comment is reduced to its text content, author nickname, like count, and sub-comment count, removing all other metadata and nested structures.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →