How the Agent-Reach Format Command Cleans Xiaohongshu API Output

The Agent-Reach format command sanitizes verbose Xiaohongshu API responses through the format_xhs_result function, which extracts only essential fields like identifiers, engagement metrics, and media URLs while discarding structural redundancy to minimize LLM token consumption.

The Agent-Reach repository provides specialized tooling for normalizing social media API data into LLM-friendly formats. When working with Xiaohongshu (XHS) API responses, the format command leverages a dedicated sanitization pipeline implemented in agent_reach/channels/xiaohongshu.py that dramatically reduces payload size while preserving critical content and metadata.

The Three-Stage Sanitization Pipeline

The core logic resides in format_xhs_result (lines 40-60 of xiaohongshu.py). This helper implements a three-stage normalization process capable of handling both single notes and batch search results without configuration changes.

Stage 1: Envelope Type Detection

First, the function inspects whether the raw data arrives as a list (typical for search feeds) or a dict (single note or wrapper object). This determination dictates how the extractor navigates the nested JSON structure, ensuring the formatter adapts to various Xiaohongshu API endpoints automatically.

Stage 2: Note List Extraction

For dictionary payloads, the code searches for common container keys including items, data.items, or data.notes. If found, each entry within these arrays is queued for processing; otherwise, the entire dictionary is treated as a solitary note. This flexibility allows the Agent-Reach format command to normalize diverse response shapes from different XHS endpoints.

Stage 3: Field Sanitization via _clean_note

Every extracted note is delegated to the private helper _clean_note (lines 62-130), which performs aggressive pruning to retain only fields relevant for downstream processing. According to the source code documentation, this routine "drastically reduces token usage by stripping structural redundancy." The specific fields preserved include:

  • Core identifiers: id, note_id, xcs_token, title, desc, type, and time
  • Content: The content field (when distinct from desc)
  • Author metadata: Limited to nickname, user_id, and nick_name
  • Engagement metrics: liked_count, collected_count, comment_count, and share_count
  • Media assets: Image URLs are flattened from nested objects into a simple string list
  • Taxonomy: Tag names are extracted as plain strings, discarding tag IDs and metadata
  • Comment threads: Each comment is recursively processed through _clean_comment

Normalizing Comments with _clean_comment

Nested within the cleaning pipeline, the _clean_comment helper (lines 32-46) handles comment sanitization by preserving only the comment text, author nickname, and basic counters (like_count, sub_comment_count). This prevents bloated user objects, avatar URLs, and tracking parameters present in raw XHS responses from inflating token counts or cluttering LLM context windows.

Practical Implementation Example

Here is a runnable example demonstrating how to clean Xiaohongshu API output from a typical search feed response:

from agent_reach.channels.xiaohongshu import format_xhs_result

# Raw response from the XHS search_feeds endpoint

raw = {
    "data": {
        "items": [
            {
                "note_id": "12345",
                "title": "My travel diary",
                "desc": "Beautiful mountains",
                "user": {"nickname": "Alice", "user_id": "u123"},
                "interact_info": {"liked_count": 42, "comment_count": 5},
                "image_list": [{"url": "https://img.xhs.com/1.jpg"}],
                "tag_list": [{"name": "travel"}, {"name": "mountains"}],
                "comments": [
                    {"content": "Nice!", "user_info": {"nickname": "Bob"}}
                ],
            }
        ]
    }
}

cleaned = format_xhs_result(raw)
print(cleaned)

Output:

[
  {
    "note_id": "12345",
    "title": "My travel diary",
    "desc": "Beautiful mountains",
    "user": {"nickname": "Alice", "user_id": "u123"},
    "liked_count": 42,
    "comment_count": 5,
    "images": ["https://img.xhs.com/1.jpg"],
    "tags": ["travel", "mountains"],
    "comments": [
      {"content": "Nice!", "user": "Bob"}
    ]
  }
]

The function automatically adapts to single note objects (pass the dict directly) or lists of notes returned from batch operations.

Key Source Files and Architecture

Understanding the implementation requires examining three critical files in the Agent-Reach codebase:

  • agent_reach/channels/xiaohongshu.py: Contains format_xhs_result, _clean_note, and _clean_comment—the complete sanitization logic for Xiaohongshu data.
  • tests/test_xhs_format.py: Unit tests verifying correct normalization of both list-based search results and dict-based single note payloads.
  • agent_reach/channels/base.py: Defines the generic Channel interface that XiaoHongShuChannel inherits from, ensuring consistent formatting behavior across social platforms.

Summary

  • Agent-Reach provides specialized tooling to clean Xiaohongshu API output through the format_xhs_result function in xiaohongshu.py.
  • The sanitization pipeline operates in three stages: envelope detection, note list extraction, and field-level cleaning via _clean_note.
  • Token optimization is achieved by stripping structural redundancy while preserving critical fields like identifiers, engagement metrics, and media URLs.
  • Comment threads are recursively cleaned through _clean_comment to retain only text content and basic metadata.
  • The formatter handles both single notes and batch search results without requiring configuration changes.

Frequently Asked Questions

What specific fields does the Agent-Reach format command preserve when cleaning Xiaohongshu API output?

The formatter retains core identifiers (note_id, xsec_token), content metadata (title, desc), author information (nickname, user_id), engagement statistics (liked_count, comment_count, share_count), flattened image URLs, tag names, and cleaned comment text. All tracking parameters, deeply nested structural wrappers, and redundant metadata are discarded to optimize the payload.

How does the format command reduce LLM token consumption?

By delegating to _clean_note and _clean_comment, the command removes the deep nesting and repetitive structural elements present in raw Xiaohongshu payloads. As implemented in the Agent-Reach source code, this approach "drastically reduces token usage by stripping structural redundancy," delivering compact JSON-compatible objects that minimize context window usage during LLM inference.

Can the format command handle both single notes and search result lists?

Yes. The format_xhs_result function automatically detects whether the input is a list (search feeds) or dictionary (single note or wrapper). It searches common container keys like items, data.items, or data.notes to locate note arrays, processing each entry uniformly while gracefully handling standalone note objects passed directly.

Where is the Xiaohongshu formatting logic implemented in the Agent-Reach codebase?

The implementation resides in agent_reach/channels/xiaohongshu.py, specifically lines 40-60 for the entry point format_xhs_result, lines 62-130 for the _clean_note helper, and lines 32-46 for the _clean_comment utility. The XiaoHongShuChannel class inherits from the base interface defined in agent_reach/channels/base.py, while validation tests live in tests/test_xhs_format.py.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →