# What Does the Format Command Do for Xiaohongshu Platform Output in Agent Reach

> Understand the format command for Xiaohongshu output in Agent Reach. It optimizes API responses into a compact, AI-friendly schema, reducing costs and improving processing.

- Repository: [Pnant/Agent-Reach](https://github.com/Panniantong/Agent-Reach)
- Tags: how-to-guide
- Published: 2026-06-24

---

**The `format_xhs_result` command in Agent-Reach strips away unnecessary fields from Xiaohongshu API responses and normalizes the data structure to create a compact, AI-friendly schema that reduces token costs and ensures consistent downstream processing.**

The `format` command for Xiaohongshu (XHS) platform output is implemented in the [`agent_reach/channels/xiaohongshu.py`](https://github.com/Panniantong/Agent-Reach/blob/main/agent_reach/channels/xiaohongshu.py) module of the Panniantong/Agent-Reach repository. This function serves as a critical data sanitation layer that transforms raw API responses from various XHS backends—such as OpenCLI, xiaohongshu-mcp, and xhs-cli—into a standardized, minimal format suitable for large language model (LLM) consumption.

## How the Format Command Works

### Input Detection and Wrapper Normalization

The `format_xhs_result` function begins by detecting the input type at lines 40-59. It handles three primary scenarios:

- **List inputs**: Treated as collections of notes to be processed iteratively
- **Dictionary wrappers**: Inspected for common container keys like `items`, `data.items`, or `data.notes`
- **Raw dictionaries**: Passed through directly when no wrapper is detected

If the input structure does not match expected patterns, the function returns the raw value unchanged to prevent data loss.

### Note Extraction via `_clean_note`

For every note extracted from the input, the helper function `_clean_note` (lines 62-71) performs the core sanitization logic. This function locates the actual note payload by searching for nested structures under `note_card`, `note`, or using the dictionary directly as the inner payload.

### Core Metadata Preservation

The formatter explicitly copies only essential fields that provide value to AI agents. According to lines 73-76, the following top-level keys are retained if present:

- `id` or `note_id`
- `xsec_token`
- `title`
- `desc` (with fallback to `content` if description is missing, lines 78-80)
- `type`
- `time`

### Author and Interaction Metrics

Author information is normalized under a unified `user` key containing `nickname`, `user_id`, and `nick_name` (lines 81-87). Interaction metrics are extracted from either `interact_info`, `note_interact_info`, or directly from the note root, capturing:

- `liked_count`
- `collected_count`
- `comment_count`
- `share_count`

### Image and Tag Processing

The formatter flattens complex image objects into a simple array of URLs (`result["images"]`). As implemented in lines 100-112, the code handles both dictionary-based image objects (extracting URL fields) and plain string URLs. Tags are reduced to a simple list of tag names by processing the `tag_list` or `tags` arrays (lines 114-124).

### Comment Sanitization

When comments are present in the input, the `_clean_comment` helper function processes each entry to retain only `content`, a shortened `user` identifier, and engagement counts (`like_count`, `sub_comment_count`). This prevents comment threads from overwhelming the token budget while preserving conversational context.

## Why Data Normalization Matters for AI Agents

**Token Economy**: By removing platform-specific noise such as `avatar`, `extra_field`, and `geo_info`, the formatter significantly reduces payload size. This directly lowers API costs for LLM calls that charge per token.

**Consistency**: Different XHS backends return slightly varying schemas. The `format_xhs_result` function normalizes these disparities, allowing downstream agents to treat all Xiaohongshu notes as a single, predictable shape.

**Safety**: Limiting output to known fields reduces the risk of exposing sensitive data such as raw cookies or internal platform identifiers in log files or prompt traces.

## Code Examples

The following examples demonstrate practical usage of the `format_xhs_result` function:

```python
from agent_reach.channels.xiaohongshu import format_xhs_result

# Single note cleaning

raw_note = {
    "id": "abc123",
    "title": "旅行日记",
    "desc": "美丽的山川",
    "type": "normal",
    "user": {"nickname": "小红", "user_id": "u001", "avatar": "https://..."},
    "interact_info": {"liked_count": "42", "collected_count": "10"},
    "image_list": [{"url": "https://img.example.com/1.jpg"}],
    "tag_list": [{"name": "旅行"}, {"name": "摄影"}],
    "extra_field": "should disappear"
}

clean = format_xhs_result(raw_note)
print(clean)

# Result: {'id': 'abc123', 'title': '旅行日记', 'desc': '美丽的山川', 

#          'type': 'normal', 'user': {'nickname': '小红', 'user_id': 'u001'},

#          'liked_count': '42', 'collected_count': '10', 

#          'images': ['https://img.example.com/1.jpg'], 'tags': ['旅行', '摄影']}

```

```python

# Handling wrapped search results

wrapped_data = {"items": [raw_note, raw_note]}
clean_list = format_xhs_result(wrapped_data)

assert isinstance(clean_list, list)
assert len(clean_list) == 2

# Each item in the list is a sanitized note dictionary

```

```python

# Processing comments

raw_with_comments = dict(raw_note)
raw_with_comments["comments"] = [
    {
        "content": "写得好！",
        "user_info": {"nickname": "路人甲"},
        "like_count": 5,
        "sub_comment_count": 1,
    }
]

clean = format_xhs_result(raw_with_comments)
print(clean["comments"])

# Result: [{'content': '写得好！', 'user': '路人甲', 'like_count': 5, 'sub_comment_count': 1}]

```

## Source Files and Implementation Details

The Xiaohongshu formatting logic resides in two primary locations within the Panniantong/Agent-Reach repository:

- **[`agent_reach/channels/xiaohongshu.py`](https://github.com/Panniantong/Agent-Reach/blob/main/agent_reach/channels/xiaohongshu.py)**: Contains the `format_xhs_result` function (lines 40-59), `_clean_note` helper (lines 62-131), and `_clean_comment` helper. This module implements the complete data transformation pipeline for the XHS channel.

- **[`tests/test_xhs_format.py`](https://github.com/Panniantong/Agent-Reach/blob/main/tests/test_xhs_format.py)**: Provides comprehensive unit test coverage validating the retention of useful fields, removal of irrelevant data, and correct handling of wrapper objects and comment arrays.

## Summary

- The `format_xhs_result` command in [`agent_reach/channels/xiaohongshu.py`](https://github.com/Panniantong/Agent-Reach/blob/main/agent_reach/channels/xiaohongshu.py) sanitizes raw Xiaohongshu API responses by removing fields irrelevant to AI processing.
- It normalizes varying response structures from different backends (OpenCLI, xiaohongshu-mcp, xhs-cli) into a consistent schema.
- Key preserved fields include `id`, `title`, `desc`, `user` metadata, interaction counts, image URLs, and tag names.
- The formatter handles nested wrappers (`items`, `note_card`), lists, and optional comments through dedicated helper functions.
- Comprehensive test coverage exists in [`tests/test_xhs_format.py`](https://github.com/Panniantong/Agent-Reach/blob/main/tests/test_xhs_format.py) to ensure data integrity across different input formats.

## Frequently Asked Questions

### What fields does the format command remove from Xiaohongshu data?

The `format_xhs_result` function explicitly strips platform-specific noise such as `avatar`, `extra_field`, `geo_info`, and any other fields not explicitly whitelisted in the `_clean_note` helper. This selective retention ensures only data points valuable to AI agents—such as content, engagement metrics, and identifiers—are preserved, significantly reducing token counts for downstream LLM processing.

### How does format_xhs_result handle different API response structures?

The function detects input types automatically at lines 40-59, handling raw dictionaries, lists of notes, and common wrapper objects containing `items`, `data.items`, or `data.notes` keys. For each note found, it extracts the payload from nested locations like `note_card` or `note` before applying sanitization, ensuring consistent output regardless of which XHS backend (OpenCLI, xiaohongshu-mcp, or xhs-cli) generated the response.

### Where is the format command tested in the Agent-Reach repository?

The formatter is fully exercised by [`tests/test_xhs_format.py`](https://github.com/Panniantong/Agent-Reach/blob/main/tests/test_xhs_format.py), which validates retention of critical fields (title, ID, user nickname, likes), removal of irrelevant data, and proper handling of list inputs, wrapper objects, and comment arrays. These tests ensure the sanitization logic remains stable across updates to the Xiaohongshu channel implementation.

### Does the formatter modify the original data or create a copy?

The `format_xhs_result` function and its helper `_clean_note` construct new dictionary objects during processing rather than mutating the original input. This immutability ensures that raw API responses remain intact for logging or debugging purposes while the sanitized copy is passed to AI agents for processing.