# How the `format xhs` Command Cleans Xiaohongshu API Output

> Learn how the format xhs command cleans Xiaohongshu API output. Extract essential fields for compact, deterministic JSON, simplifying your data. Optimize your app

- Repository: [Pnant/Agent-Reach](https://github.com/Panniantong/Agent-Reach)
- Tags: how-to-guide
- Published: 2026-06-17

---

**The `format xhs` command normalizes raw Xiaohongshu API responses into compact, deterministic JSON by extracting only essential fields like `id`, `title`, `desc`, and engagement metrics while discarding metadata and nested wrappers.**

The `format xhs` command is a critical component of the **Agent-Reach** repository designed to reduce token bloat when feeding Xiaohongshu (XHS) content into LLM pipelines. Located in [`agent_reach/channels/xiaohongshu.py`](https://github.com/Panniantong/Agent-Reach/blob/main/agent_reach/channels/xiaohongshu.py), this functionality transforms verbose CLI/API outputs into minimal, predictable structures that preserve semantic value while eliminating noise.

## Core Processing Pipeline

### Input Type Discrimination

The process begins in `format_xhs_result` (lines 46‑58) where the function discriminates between list payloads, dictionary wrappers, and plain values. When encountering lists, it processes items iteratively; for dictionaries, it inspects for known wrapper keys including `items`, `data.items`, and `data.notes` to locate the actual content.

### Note Extraction and Field Whitelisting

Once inside `_clean_note` (lines 67‑69), the extractor navigates nested structures by checking both `note_card` and `note` keys to retrieve the inner note object. The function then applies a strict **field whitelist** (lines 73‑76), retaining only `id`, `note_id`, `xsec_token`, `title`, `desc`, `type`, and `time` while stripping all other metadata.

### Content Fallback and User Simplification

To handle inconsistent API schemas, the implementation includes a **content fallback** mechanism (lines 78‑80): if `desc` is absent but `content` exists, the latter is promoted to a top-level key. Author data is reduced to a minimal object containing just `nickname` and `user_id` (or `nick_name` as fallback) as implemented in lines 82‑87.

### Engagement Metrics and Image Reduction

Engagement counters—likes, collects, comments, and shares—are aggregated from potential locations including `interact_info`, `note_interact_info`, or the top-level note (lines 89‑98). For media, the function normalizes `image_list` or `images_list` arrays into simple URL strings, discarding dimensions, trace IDs, and other metadata (lines 100‑112).

### Tag Simplification and Comment Handling

Tag arrays are flattened from complex objects to plain strings containing only tag names (lines 114‑124). When comments are present, each entry is processed by `_clean_comment` (lines 33‑46) to retain only the `content`, author nickname, `like_count`, and `sub_comment_count`.

## Implementation Architecture

### The `format_xhs_result` Entry Point

This function serves as the primary entry point in [`agent_reach/channels/xiaohongshu.py`](https://github.com/Panniantong/Agent-Reach/blob/main/agent_reach/channels/xiaohongshu.py), handling input validation and dispatching. Non-dictionary and non-list inputs (such as strings or `None`) pass through unchanged (lines 59‑60), ensuring robustness against malformed inputs.

### Helper Functions

The `_clean_note` function contains the bulk of the cleaning logic (lines 62‑130), orchestrating field extraction, normalization, and metric aggregation. For comment threads, `_clean_comment` provides analogous reduction capabilities, focusing exclusively on textual content and engagement signals.

## Usage Examples

### Command-Line Interface

Integrate with `xhs-cli` to process raw API output directly:

```bash

# Fetch and format a single note

xhs read 123456 | agent-reach format xhs

```

The resulting JSON contains only whitelisted fields:

```json
{
  "id": "abc123",
  "title": "测试笔记",
  "desc": "这是正文内容",
  "type": "normal",
  "xsec_token": "tok_xxx",
  "user": {
    "nickname": "小红",
    "user_id": "u123"
  },
  "liked_count": "100",
  "collected_count": "50",
  "comment_count": "20",
  "share_count": "10",
  "images": [
    "https://img.example.com/1.jpg"
  ],
  "tags": ["旅行", "美食"]
}

```

### Programmatic Integration

Import the cleaner directly into Python workflows to process search results or individual notes:

```python
from agent_reach.channels.xiaohongshu import format_xhs_result
import json
import subprocess

# Example: Processing search results

raw_data = json.loads(subprocess.check_output(["xhs", "search", "AI", "-n", "2"]))
cleaned_results = format_xhs_result(raw_data)

# cleaned_results is now a list of compact dictionaries suitable for LLM prompts

print(json.dumps(cleaned_results, ensure_ascii=False, indent=2))

```

For single-note processing:

```python
raw_note = {...}  # Dict from xhs-cli

clean_note = format_xhs_result(raw_note)

```

## Key Source Files

- **[`agent_reach/channels/xiaohongshu.py`](https://github.com/Panniantong/Agent-Reach/blob/main/agent_reach/channels/xiaohongshu.py)**: Contains `format_xhs_result`, `_clean_note`, and `_clean_comment` implementations.
- **[`agent_reach/cli.py`](https://github.com/Panniantong/Agent-Reach/blob/main/agent_reach/cli.py)**: Provides the CLI wiring that reads stdin and invokes the formatter.
- **[`tests/test_xhs_format.py`](https://github.com/Panniantong/Agent-Reach/blob/main/tests/test_xhs_format.py)**: Documents expected behavior through unit tests covering wrapper handling, field whitelisting, and comment extraction.

## Summary

- The `format xhs` command reduces Xiaohongshu API payloads by extracting only essential fields and discarding metadata.
- **Input discrimination** handles lists, wrapper dictionaries, and plain values transparently.
- **Field whitelisting** retains `id`, `title`, `desc`, `type`, and authentication tokens while removing noise.
- **Content fallbacks** ensure `content` is preserved when `desc` is missing.
- **Image normalization** converts complex image objects to simple URL arrays.
- **Comment reduction** extracts only text, author, and engagement metrics from nested comment trees.
- The implementation significantly reduces token usage for downstream LLM consumption (see issue #134).

## Frequently Asked Questions

### What fields does the `format xhs` command preserve from the API?

The command whitelists specific top-level fields: `id`, `note_id`, `xsec_token`, `title`, `desc` (or `content` as fallback), `type`, and `time`. It also extracts simplified user objects (`nickname`, `user_id`), engagement metrics (likes, collects, comments, shares), image URLs, and tag names.

### How does the command handle different Xiaohongshu API response formats?

The `format_xhs_result` function in [`agent_reach/channels/xiaohongshu.py`](https://github.com/Panniantong/Agent-Reach/blob/main/agent_reach/channels/xiaohongshu.py) detects whether the input is a list, a dictionary with wrapper keys like `items` or `data.notes`, or a plain value. Lists are processed iteratively, dictionaries are unwrapped to find the actual note data, and other types pass through unchanged.

### Can I use the formatting logic outside the CLI?

Yes. The cleaning functionality is available as a Python import from `agent_reach.channels.xiaohongshu`. Simply call `format_xhs_result(raw_json)` with any dictionary or list returned by the XHS API to receive the normalized output.

### How are comments processed by the `format xhs` command?

When present, comments are processed by the `_clean_comment` helper function (lines 33‑46 in [`xiaohongshu.py`](https://github.com/Panniantong/Agent-Reach/blob/main/xiaohongshu.py)). Each comment is reduced to its text content, author nickname, like count, and sub-comment count, removing all other metadata and nested structures.