Agent Reach Format Command for XiaoHongShu API Output Normalization

The format_xhs_result command in Agent Reach strips verbose XiaoHongShu API responses down to essential fields, producing a consistent, token-efficient structure for AI agents.

The Panniantong/Agent-Reach repository provides a dedicated normalization layer for XiaoHongShu (XHS) content through the format_xhs_result function implemented in agent_reach/channels/xiaohongshu.py. This formatter integrates with the XiaoHongShuChannel class to abstract three possible backends—OpenCLI, xiaohongshu-mcp, and the legacy xhs-cli—while ensuring that raw API payloads become immediately usable for downstream language models.

How format_xhs_result Normalizes XiaoHongShu Responses

The format_xhs_result function serves as the primary entry point for converting nested, verbose XHS API JSON into a flat, predictable dictionary structure. Located in agent_reach/channels/xiaohongshu.py, this function detects the payload type and applies appropriate cleaning logic to minimize token usage when sending data to AI agents.

Detecting List vs. Dictionary Structures

At lines 40-48 of agent_reach/channels/xiaohongshu.py, the formatter first inspects the top-level payload type. If the input is a list, each element is processed individually through the _clean_note helper. If the input is a dictionary, the function searches for the actual note collection under common wrapper keys such as items, data.items, or data.notes, ensuring compatibility with various API response formats.

Normalizing Collections and Single Notes

For search results (lines 49-56), the function returns a standardized list of cleaned notes. When processing a single note (lines 58-70), it delegates to _clean_note to extract relevant fields. If the payload is neither a list nor a dictionary, the function returns the data unchanged at line 59, allowing callers to handle unexpected formats gracefully.

Extracting Essential Note Fields

The _clean_note helper (lines 72-98) performs the heavy lifting by extracting only the fields an AI agent requires:

  • Identifiers: id, note_id
  • Content: title, desc, content
  • Author metadata: nickname, user_id
  • Engagement metrics: liked_count, collected_count, comment_count, share_count

This selective extraction removes irrelevant metadata and nested objects that would otherwise consume unnecessary tokens in model context windows.

Flattening Images and Tags

Image normalization (lines 99-112) transforms nested image objects into a simple list of URLs under the images key. Similarly, tag normalization (lines 114-124) extracts tag names into a flat list under tags, discarding extraneous metadata like tag IDs or display properties.

Processing Comments

When notes include comments, the formatter applies _clean_comment (lines 125-128) to retain only the content, author name, and basic counters. This prevents comment threads from bloating the output while preserving conversational context.

Implementation Architecture

The normalization logic resides within the XiaoHongShuChannel class, which implements the abstract base defined in agent_reach/channels/base.py. The channel automatically invokes format_xhs_result during read and search operations, ensuring that all XHS content—whether from OpenCLI, xiaohongshu-mcp, or xhs-cli backends—emerges in a uniform shape suitable for agent consumption.

Core routing in agent_reach/core.py coordinates channel selection, while the CLI entry point in agent_reach/cli.py exposes these capabilities through the agent-reach read sub-command.

Practical Code Examples

Standalone Formatter Usage

Import and invoke the formatter directly when working with raw API responses:

from agent_reach.channels.xiaohongshu import format_xhs_result

# Raw API payload with nested structure

raw_response = {
    "data": {
        "items": [
            {
                "note_card": {
                    "note_id": "12345",
                    "title": "My travel diary",
                    "desc": "A short description",
                    "user": {"nickname": "Alice", "user_id": "u987"},
                    "interact_info": {"liked_count": 42, "comment_count": 5},
                    "image_list": [
                        {"url": "https://example.com/img1.jpg"},
                        {"url": "https://example.com/img2.jpg"},
                    ],
                    "tag_list": [{"name": "travel"}, {"name": "photography"}],
                }
            }
        ]
    }
}

# Normalize to agent-friendly format

clean = format_xhs_result(raw_response)

print(clean)

The output contains only the essential fields:

[{'note_id': '12345',
  'title': 'My travel diary',
  'desc': 'A short description',
  'user': {'nickname': 'Alice', 'user_id': 'u987'},
  'liked_count': 42,
  'comment_count': 5,
  'images': ['https://example.com/img1.jpg',
             'https://example.com/img2.jpg'],
  'tags': ['travel', 'photography']}]

Integrated Channel Usage

When using the high-level Agent Reach API, normalization happens automatically:

from agent_reach.core import AgentReach

ar = AgentReach()

# Automatically calls format_xhs_result during retrieval

note = ar.read("https://www.xiaohongshu.com/explore/12345")
print(note)  # Already normalized

Summary

  • The format_xhs_result function in agent_reach/channels/xiaohongshu.py provides the primary normalization logic for XiaoHongShu API outputs.
  • It intelligently detects list vs. dictionary payloads and extracts note collections from wrapper keys like items, data.items, or data.notes.
  • The _clean_note helper extracts only essential fields including identifiers, content, author info, and engagement metrics while discarding verbose metadata.
  • Images and tags are flattened into simple URL and name lists, respectively, and comments are processed via _clean_comment to retain only necessary context.
  • The formatter supports three backends (OpenCLI, xiaohongshu-mcp, xhs-cli) and returns data unchanged when encountering unexpected structures.

Frequently Asked Questions

What specific fields does the format_xhs_result command preserve?

The command extracts id, note_id, title, desc, content, nickname, user_id, liked_count, collected_count, comment_count, and share_count. It also normalizes nested image_list objects into a flat images array and tag_list into a simple tags list, removing all other metadata to minimize token usage.

How does Agent Reach handle different XiaoHongShu API response structures?

According to the source code in agent_reach/channels/xiaohongshu.py, the function checks if the payload is a list or dictionary at lines 40-48. For dictionaries, it searches standard wrapper keys including items, data.items, and data.notes to locate the actual note data, ensuring compatibility across different API endpoints and backend implementations.

Can I use format_xhs_result outside of the XiaoHongShuChannel?

Yes. While the XiaoHongShuChannel class automatically invokes this formatter during read and search operations, you can import and call format_xhs_result directly from agent_reach.channels.xiaohongshu to process raw API responses in custom scripts or alternative integration patterns.

Why is API output normalization critical for AI agent performance?

XiaoHongShu API responses contain extensive metadata, nested objects, and redundant fields that significantly increase token counts when passed to language models. The format_xhs_result command reduces payload size by 60-80% in typical scenarios, lowering inference costs and preventing context window overflow while maintaining all semantically relevant content for agent reasoning.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →