What Does the Format Command Do for Xiaohongshu Platform Output in Agent Reach

The format_xhs_result command in Agent-Reach strips away unnecessary fields from Xiaohongshu API responses and normalizes the data structure to create a compact, AI-friendly schema that reduces token costs and ensures consistent downstream processing.

The format command for Xiaohongshu (XHS) platform output is implemented in the agent_reach/channels/xiaohongshu.py module of the Panniantong/Agent-Reach repository. This function serves as a critical data sanitation layer that transforms raw API responses from various XHS backends—such as OpenCLI, xiaohongshu-mcp, and xhs-cli—into a standardized, minimal format suitable for large language model (LLM) consumption.

How the Format Command Works

Input Detection and Wrapper Normalization

The format_xhs_result function begins by detecting the input type at lines 40-59. It handles three primary scenarios:

  • List inputs: Treated as collections of notes to be processed iteratively
  • Dictionary wrappers: Inspected for common container keys like items, data.items, or data.notes
  • Raw dictionaries: Passed through directly when no wrapper is detected

If the input structure does not match expected patterns, the function returns the raw value unchanged to prevent data loss.

Note Extraction via _clean_note

For every note extracted from the input, the helper function _clean_note (lines 62-71) performs the core sanitization logic. This function locates the actual note payload by searching for nested structures under note_card, note, or using the dictionary directly as the inner payload.

Core Metadata Preservation

The formatter explicitly copies only essential fields that provide value to AI agents. According to lines 73-76, the following top-level keys are retained if present:

  • id or note_id
  • xsec_token
  • title
  • desc (with fallback to content if description is missing, lines 78-80)
  • type
  • time

Author and Interaction Metrics

Author information is normalized under a unified user key containing nickname, user_id, and nick_name (lines 81-87). Interaction metrics are extracted from either interact_info, note_interact_info, or directly from the note root, capturing:

  • liked_count
  • collected_count
  • comment_count
  • share_count

Image and Tag Processing

The formatter flattens complex image objects into a simple array of URLs (result["images"]). As implemented in lines 100-112, the code handles both dictionary-based image objects (extracting URL fields) and plain string URLs. Tags are reduced to a simple list of tag names by processing the tag_list or tags arrays (lines 114-124).

Comment Sanitization

When comments are present in the input, the _clean_comment helper function processes each entry to retain only content, a shortened user identifier, and engagement counts (like_count, sub_comment_count). This prevents comment threads from overwhelming the token budget while preserving conversational context.

Why Data Normalization Matters for AI Agents

Token Economy: By removing platform-specific noise such as avatar, extra_field, and geo_info, the formatter significantly reduces payload size. This directly lowers API costs for LLM calls that charge per token.

Consistency: Different XHS backends return slightly varying schemas. The format_xhs_result function normalizes these disparities, allowing downstream agents to treat all Xiaohongshu notes as a single, predictable shape.

Safety: Limiting output to known fields reduces the risk of exposing sensitive data such as raw cookies or internal platform identifiers in log files or prompt traces.

Code Examples

The following examples demonstrate practical usage of the format_xhs_result function:

from agent_reach.channels.xiaohongshu import format_xhs_result

# Single note cleaning

raw_note = {
    "id": "abc123",
    "title": "旅行日记",
    "desc": "美丽的山川",
    "type": "normal",
    "user": {"nickname": "小红", "user_id": "u001", "avatar": "https://..."},
    "interact_info": {"liked_count": "42", "collected_count": "10"},
    "image_list": [{"url": "https://img.example.com/1.jpg"}],
    "tag_list": [{"name": "旅行"}, {"name": "摄影"}],
    "extra_field": "should disappear"
}

clean = format_xhs_result(raw_note)
print(clean)

# Result: {'id': 'abc123', 'title': '旅行日记', 'desc': '美丽的山川', 

#          'type': 'normal', 'user': {'nickname': '小红', 'user_id': 'u001'},

#          'liked_count': '42', 'collected_count': '10', 

#          'images': ['https://img.example.com/1.jpg'], 'tags': ['旅行', '摄影']}

# Handling wrapped search results

wrapped_data = {"items": [raw_note, raw_note]}
clean_list = format_xhs_result(wrapped_data)

assert isinstance(clean_list, list)
assert len(clean_list) == 2

# Each item in the list is a sanitized note dictionary

# Processing comments

raw_with_comments = dict(raw_note)
raw_with_comments["comments"] = [
    {
        "content": "写得好!",
        "user_info": {"nickname": "路人甲"},
        "like_count": 5,
        "sub_comment_count": 1,
    }
]

clean = format_xhs_result(raw_with_comments)
print(clean["comments"])

# Result: [{'content': '写得好!', 'user': '路人甲', 'like_count': 5, 'sub_comment_count': 1}]

Source Files and Implementation Details

The Xiaohongshu formatting logic resides in two primary locations within the Panniantong/Agent-Reach repository:

  • agent_reach/channels/xiaohongshu.py: Contains the format_xhs_result function (lines 40-59), _clean_note helper (lines 62-131), and _clean_comment helper. This module implements the complete data transformation pipeline for the XHS channel.

  • tests/test_xhs_format.py: Provides comprehensive unit test coverage validating the retention of useful fields, removal of irrelevant data, and correct handling of wrapper objects and comment arrays.

Summary

  • The format_xhs_result command in agent_reach/channels/xiaohongshu.py sanitizes raw Xiaohongshu API responses by removing fields irrelevant to AI processing.
  • It normalizes varying response structures from different backends (OpenCLI, xiaohongshu-mcp, xhs-cli) into a consistent schema.
  • Key preserved fields include id, title, desc, user metadata, interaction counts, image URLs, and tag names.
  • The formatter handles nested wrappers (items, note_card), lists, and optional comments through dedicated helper functions.
  • Comprehensive test coverage exists in tests/test_xhs_format.py to ensure data integrity across different input formats.

Frequently Asked Questions

What fields does the format command remove from Xiaohongshu data?

The format_xhs_result function explicitly strips platform-specific noise such as avatar, extra_field, geo_info, and any other fields not explicitly whitelisted in the _clean_note helper. This selective retention ensures only data points valuable to AI agents—such as content, engagement metrics, and identifiers—are preserved, significantly reducing token counts for downstream LLM processing.

How does format_xhs_result handle different API response structures?

The function detects input types automatically at lines 40-59, handling raw dictionaries, lists of notes, and common wrapper objects containing items, data.items, or data.notes keys. For each note found, it extracts the payload from nested locations like note_card or note before applying sanitization, ensuring consistent output regardless of which XHS backend (OpenCLI, xiaohongshu-mcp, or xhs-cli) generated the response.

Where is the format command tested in the Agent-Reach repository?

The formatter is fully exercised by tests/test_xhs_format.py, which validates retention of critical fields (title, ID, user nickname, likes), removal of irrelevant data, and proper handling of list inputs, wrapper objects, and comment arrays. These tests ensure the sanitization logic remains stable across updates to the Xiaohongshu channel implementation.

Does the formatter modify the original data or create a copy?

The format_xhs_result function and its helper _clean_note construct new dictionary objects during processing rather than mutating the original input. This immutability ensures that raw API responses remain intact for logging or debugging purposes while the sanitized copy is passed to AI agents for processing.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →