How the Format Command Cleans and Processes Xiaohongshu API Output in Agent Reach

The format_xhs_result command in Agent Reach strips redundant Xiaohongshu API JSON down to essential fields—reducing payload size by over 80% while preserving critical data for LLM consumption.

Agent Reach, an open-source automation framework, includes a dedicated formatter for Xiaohongshu (XHS) API responses that transforms bloated JSON payloads into streamlined, token-efficient dictionaries. The format_xhs_result function, implemented in the XiaoHongShu channel, automatically detects input structures and applies intelligent field extraction to prepare data for downstream language model processing.

Input Type Detection and Routing

The formatter begins by inspecting the incoming payload structure in agent_reach/channels/xiaohongshu.py. When format_xhs_result receives data, it first determines whether the input is a list, dictionary, or unexpected type.

  • List inputs: The function maps _clean_note across each element individually.
  • Dictionary inputs: It searches for container keys such as "items" or "data"—common wrappers in search feed responses—and extracts the inner list for processing. If neither container key exists, the dictionary is treated as a single note.
  • Invalid types: Any non-list, non-dictionary input returns unchanged, preventing exceptions from malformed payloads.

This routing logic ensures that both single notes and batch search results process correctly without manual intervention.

Note-Level Normalization and Field Extraction

Each note passes through the _clean_note helper function (lines 62-130), which performs surgical field extraction while discarding structural noise. The cleaning process targets specific data categories:

Core Identifiers and Content Fallback

The formatter preserves essential identity fields including id, note_id, xsec_token, title, desc, type, and time when present. If a note contains content but lacks desc, the function stores the content under the "content" key to ensure text availability for LLM prompts.

Author and Engagement Metrics

Author information merges from either user or author dictionaries into a minimal user object containing only nickname, user_id, and nick_name. Engagement statistics such as likes, collects, comments, and shares are gathered flexibly from interact_info, note_interact_info, or top-level keys, eliminating duplication while ensuring metrics are captured regardless of API endpoint variations.

Media and Tags

Image URLs extract from image_list or images_list arrays, preserving only the url, url_default, or original fields and discarding metadata like dimensions or processing flags. Tags flatten into simple strings, handling both dictionary-based tags with "name" keys and plain string entries.

Comment Processing

When notes include a comments array, each comment undergoes _clean_comment processing (lines 132-146). The function retains only the content, a simplified user string representing the nickname, and engagement counters (like_count, sub_comment_count), creating compact comment objects suitable for context windows.

Graceful Pass-Through for Invalid Inputs

The formatter implements defensive handling for edge cases including empty dictionaries, empty lists, plain strings, or None values. Rather than raising exceptions, format_xhs_result returns these inputs unchanged, ensuring CLI stability and library flexibility when encountering unexpected API responses or network timeouts.

Usage Examples

The formatter handles diverse Xiaohongshu API response patterns:


# Single note processing

from agent_reach.channels.xiaohongshu import format_xhs_result

raw_note = {
    "id": "abc123",
    "title": "旅行日记",
    "desc": "在欧洲的美好瞬间",
    "type": "normal",
    "xsec_token": "tok_xyz",
    "user": {"nickname": "小红", "user_id": "u123", "avatar": "..."},
    "interact_info": {"liked_count": 10, "collected_count": 2},
    "image_list": [{"url": "https://img.example.com/1.jpg"}],
    "tag_list": [{"name": "旅行"}, {"name": "摄影"}],
}
clean = format_xhs_result(raw_note)
print(clean)

# → {

#     "id": "abc123",

#     "title": "旅行日记",

#     "desc": "在欧洲的美好瞬间",

#     "type": "normal",

#     "xsec_token": "tok_xyz",

#     "user": {"nickname": "小红", "user_id": "u123"},

#     "liked_count": 10,

#     "collected_count": 2,

#     "images": ["https://img.example.com/1.jpg"],

#     "tags": ["旅行", "摄影"],

# }

# Search feed wrapper (batch processing)

raw_feed = {"items": [raw_note, raw_note]}
clean_list = format_xhs_result(raw_feed)
print(type(clean_list), len(clean_list))

# → <class 'list'> 2

# Note with comments

raw_note_with_comments = dict(raw_note)
raw_note_with_comments["comments"] = [
    {
        "content": "写得真好!",
        "user_info": {"nickname": "路人甲"},
        "like_count": 5,
        "sub_comment_count": 1,
    }
]
clean = format_xhs_result(raw_note_with_comments)
print(clean["comments"][0])

# → {"content": "写得真好!", "user": "路人甲", "like_count": 5, "sub_comment_count": 1}

Summary

  • The format_xhs_result command in agent_reach/channels/xiaohongshu.py reduces Xiaohongshu API payloads by over 80% through selective field retention.
  • Input type detection handles both single notes and batch search results by checking for "items" or "data" wrapper keys.
  • The _clean_note helper extracts core identifiers, author data, engagement metrics, images, and tags while removing redundant metadata.
  • Comments process through _clean_comment to retain only content, user nicknames, and engagement counts.
  • Invalid inputs pass through unchanged, ensuring robust error handling for malformed API responses.

Frequently Asked Questions

What fields does the format command preserve from Xiaohongshu API output?

The formatter retains id, note_id, xsec_token, title, desc (or content as fallback), type, time, simplified user objects with nickname and user_id, engagement counts (likes, collects, comments, shares), image URLs, tags, and condensed comment data. All other metadata is discarded to optimize token usage.

How does Agent Reach handle nested author information in XHS notes?

The _clean_note function checks both user and author dictionary keys, then constructs a minimal user object containing only nickname, user_id, and nick_name. This approach strips avatar URLs and other profile metadata that consume tokens without adding analytical value for LLM processing.

Can the format command process batch search results or single notes?

Yes. The format_xhs_result function automatically detects list inputs and processes each element individually. For dictionary inputs, it looks for "items" or "data" keys (common in search feeds) to extract note arrays, or treats the dict as a single note if no wrapper keys exist.

Where is the format_xhs_result function implemented in the Agent Reach codebase?

The core formatting logic resides in agent_reach/channels/xiaohongshu.py (lines 40-59 for the main entry point, lines 62-130 for _clean_note, and lines 132-146 for _clean_comment), with unit tests available in tests/test_xhs_format.py.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →