# How the Format Command Cleans and Processes Xiaohongshu API Output in Agent Reach

> Learn how Agent Reach's format command efficiently cleans Xiaohongshu API output, reducing data by over 80% for LLMs while preserving essential info. Discover streamlined data processing.

- Repository: [Pnant/Agent-Reach](https://github.com/Panniantong/Agent-Reach)
- Tags: how-to-guide
- Published: 2026-06-21

---

**The `format_xhs_result` command in Agent Reach strips redundant Xiaohongshu API JSON down to essential fields—reducing payload size by over 80% while preserving critical data for LLM consumption.**

Agent Reach, an open-source automation framework, includes a dedicated formatter for Xiaohongshu (XHS) API responses that transforms bloated JSON payloads into streamlined, token-efficient dictionaries. The `format_xhs_result` function, implemented in the XiaoHongShu channel, automatically detects input structures and applies intelligent field extraction to prepare data for downstream language model processing.

## Input Type Detection and Routing

The formatter begins by inspecting the incoming payload structure in [`agent_reach/channels/xiaohongshu.py`](https://github.com/Panniantong/Agent-Reach/blob/main/agent_reach/channels/xiaohongshu.py). When `format_xhs_result` receives data, it first determines whether the input is a list, dictionary, or unexpected type.

- **List inputs**: The function maps `_clean_note` across each element individually.
- **Dictionary inputs**: It searches for container keys such as `"items"` or `"data"`—common wrappers in search feed responses—and extracts the inner list for processing. If neither container key exists, the dictionary is treated as a single note.
- **Invalid types**: Any non-list, non-dictionary input returns unchanged, preventing exceptions from malformed payloads.

This routing logic ensures that both single notes and batch search results process correctly without manual intervention.

## Note-Level Normalization and Field Extraction

Each note passes through the `_clean_note` helper function (lines 62-130), which performs surgical field extraction while discarding structural noise. The cleaning process targets specific data categories:

### Core Identifiers and Content Fallback

The formatter preserves essential identity fields including `id`, `note_id`, `xsec_token`, `title`, `desc`, `type`, and `time` when present. If a note contains `content` but lacks `desc`, the function stores the content under the `"content"` key to ensure text availability for LLM prompts.

### Author and Engagement Metrics

Author information merges from either `user` or `author` dictionaries into a minimal `user` object containing only `nickname`, `user_id`, and `nick_name`. Engagement statistics such as likes, collects, comments, and shares are gathered flexibly from `interact_info`, `note_interact_info`, or top-level keys, eliminating duplication while ensuring metrics are captured regardless of API endpoint variations.

### Media and Tags

Image URLs extract from `image_list` or `images_list` arrays, preserving only the `url`, `url_default`, or `original` fields and discarding metadata like dimensions or processing flags. Tags flatten into simple strings, handling both dictionary-based tags with `"name"` keys and plain string entries.

### Comment Processing

When notes include a `comments` array, each comment undergoes `_clean_comment` processing (lines 132-146). The function retains only the `content`, a simplified `user` string representing the nickname, and engagement counters (`like_count`, `sub_comment_count`), creating compact comment objects suitable for context windows.

## Graceful Pass-Through for Invalid Inputs

The formatter implements defensive handling for edge cases including empty dictionaries, empty lists, plain strings, or `None` values. Rather than raising exceptions, `format_xhs_result` returns these inputs unchanged, ensuring CLI stability and library flexibility when encountering unexpected API responses or network timeouts.

## Usage Examples

The formatter handles diverse Xiaohongshu API response patterns:

```python

# Single note processing

from agent_reach.channels.xiaohongshu import format_xhs_result

raw_note = {
    "id": "abc123",
    "title": "旅行日记",
    "desc": "在欧洲的美好瞬间",
    "type": "normal",
    "xsec_token": "tok_xyz",
    "user": {"nickname": "小红", "user_id": "u123", "avatar": "..."},
    "interact_info": {"liked_count": 10, "collected_count": 2},
    "image_list": [{"url": "https://img.example.com/1.jpg"}],
    "tag_list": [{"name": "旅行"}, {"name": "摄影"}],
}
clean = format_xhs_result(raw_note)
print(clean)

# → {

#     "id": "abc123",

#     "title": "旅行日记",

#     "desc": "在欧洲的美好瞬间",

#     "type": "normal",

#     "xsec_token": "tok_xyz",

#     "user": {"nickname": "小红", "user_id": "u123"},

#     "liked_count": 10,

#     "collected_count": 2,

#     "images": ["https://img.example.com/1.jpg"],

#     "tags": ["旅行", "摄影"],

# }

```

```python

# Search feed wrapper (batch processing)

raw_feed = {"items": [raw_note, raw_note]}
clean_list = format_xhs_result(raw_feed)
print(type(clean_list), len(clean_list))

# → <class 'list'> 2

```

```python

# Note with comments

raw_note_with_comments = dict(raw_note)
raw_note_with_comments["comments"] = [
    {
        "content": "写得真好!",
        "user_info": {"nickname": "路人甲"},
        "like_count": 5,
        "sub_comment_count": 1,
    }
]
clean = format_xhs_result(raw_note_with_comments)
print(clean["comments"][0])

# → {"content": "写得真好!", "user": "路人甲", "like_count": 5, "sub_comment_count": 1}

```

## Summary

- The `format_xhs_result` command in [`agent_reach/channels/xiaohongshu.py`](https://github.com/Panniantong/Agent-Reach/blob/main/agent_reach/channels/xiaohongshu.py) reduces Xiaohongshu API payloads by over 80% through selective field retention.
- Input type detection handles both single notes and batch search results by checking for `"items"` or `"data"` wrapper keys.
- The `_clean_note` helper extracts core identifiers, author data, engagement metrics, images, and tags while removing redundant metadata.
- Comments process through `_clean_comment` to retain only content, user nicknames, and engagement counts.
- Invalid inputs pass through unchanged, ensuring robust error handling for malformed API responses.

## Frequently Asked Questions

### What fields does the format command preserve from Xiaohongshu API output?

The formatter retains `id`, `note_id`, `xsec_token`, `title`, `desc` (or `content` as fallback), `type`, `time`, simplified `user` objects with `nickname` and `user_id`, engagement counts (likes, collects, comments, shares), image URLs, tags, and condensed comment data. All other metadata is discarded to optimize token usage.

### How does Agent Reach handle nested author information in XHS notes?

The `_clean_note` function checks both `user` and `author` dictionary keys, then constructs a minimal user object containing only `nickname`, `user_id`, and `nick_name`. This approach strips avatar URLs and other profile metadata that consume tokens without adding analytical value for LLM processing.

### Can the format command process batch search results or single notes?

Yes. The `format_xhs_result` function automatically detects list inputs and processes each element individually. For dictionary inputs, it looks for `"items"` or `"data"` keys (common in search feeds) to extract note arrays, or treats the dict as a single note if no wrapper keys exist.

### Where is the format_xhs_result function implemented in the Agent Reach codebase?

The core formatting logic resides in [`agent_reach/channels/xiaohongshu.py`](https://github.com/Panniantong/Agent-Reach/blob/main/agent_reach/channels/xiaohongshu.py) (lines 40-59 for the main entry point, lines 62-130 for `_clean_note`, and lines 132-146 for `_clean_comment`), with unit tests available in [`tests/test_xhs_format.py`](https://github.com/Panniantong/Agent-Reach/blob/main/tests/test_xhs_format.py).