# Distilly Built-In Material Parsers: Supported Formats and Implementation Guide

> Discover Distilly's built-in material parsers for Feishu Lark Email Slack and DingTalk. Learn how to normalize diverse exports into a unified schema for persona generation.

- Repository: [Tianyi Zhou/distilly](https://github.com/titanwings/distilly)
- Tags: how-to-guide
- Published: 2026-09-10

---

**Distilly includes four built-in material parsers for Feishu/Lark, Email, Slack, and DingTalk that normalize diverse communication exports into a unified message schema for persona generation.**

Distilly provides a dedicated ingestion layer that transforms raw enterprise communication logs into structured training data. Located in the repository's `tools/` directory, these built-in material parsers handle JSON exports, email archives, and chat logs from major platforms. Each parser implements a normalization layer that converts disparate formats into a canonical dictionary structure suitable for the downstream research and skill-writing pipeline.

## Feishu and Lark Message Parsers

The [`tools/feishu_parser.py`](https://github.com/titanwings/distilly/blob/main/tools/feishu_parser.py) module provides comprehensive support for Feishu (Lark) communication exports, handling both official and handcrafted formats.

### Supported Formats and Functions

This parser offers two primary entry points for different export types:

- **`parse_feishu_json()`**: Processes official JSON exports containing arrays of message objects
- **`parse_feishu_txt()`**: Parses TXT exports formatted as `<timestamp> <sender>: <content>`

Both functions read the source file, normalize fields into `sender`, `content`, and `timestamp` keys, filter out system artifacts and non-target messages, and return a list of standardized dictionaries.

### Generic Content Classification

The module exposes **`extract_key_content()`**, a versatile classifier that buckets parsed messages into long-form content, decision-type exchanges, and daily chat categories. Located in [`tools/feishu_parser.py`](https://github.com/titanwings/distilly/blob/main/tools/feishu_parser.py), this function serves as a generic extractor used across multiple collectors to bridge raw parsing and Distilly's research pipeline.

## Email Archive Parsers

Located in [`tools/email_parser.py`](https://github.com/titanwings/distilly/blob/main/tools/email_parser.py), the email ingestion module supports three distinct archive formats common in enterprise environments.

### EML, MBOX, and Plain Text Processing

The implementation provides three specialized functions:

- **`parse_eml_file()`**: Reads individual *.eml* files and extracts headers and body content
- **`parse_mbox_file()`**: Iterates through Unix mailbox archives containing multiple email threads
- **`parse_txt_file()`**: Handles plain-text conversation dumps

Each function accepts a `target` parameter to filter messages by sender or recipient, ensuring only relevant correspondence enters the training pipeline.

## Slack Export Parser

For Slack workspace data, [`tools/slack_auto_collector.py`](https://github.com/titanwings/distilly/blob/main/tools/slack_auto_collector.py) processes official JSON exports generated by Slack's native "Export" feature.

### Implementation and CLI Interface

The collector loads the JSON payload and walks the `messages` array to extract participant names, message content, and timestamps. It supports an optional `--name` command-line flag to limit output to specific participants. The parser normalizes Slack's proprietary data structure into the canonical Distilly message format, making it interchangeable with Feishu and DingTalk sources.

## DingTalk Export Parser

Mirroring the Slack implementation, [`tools/dingtalk_auto_collector.py`](https://github.com/titanwings/distilly/blob/main/tools/dingtalk_auto_collector.py) handles exports from the DingTalk admin console.

### JSON Processing Pipeline

This parser reads the top-level `messages` field from DingTalk JSON exports and applies the same normalization logic as the Slack collector. As implemented in `titanwings/distilly`, it yields a list of dictionaries identical to other parsers, ensuring the downstream persona-generation workflow treats DingTalk data uniformly regardless of source platform.

## Unified Message Schema

All Distilly built-in material parsers return a standardized Python list of dictionaries with the following schema:

```python
{
    "sender": str,      # Name or identifier of the author

    "content": str,     # The plain-text message body

    "timestamp": str    # Raw timestamp string (ISO, epoch, or source-specific format)

}

```

Because they share this canonical shape, the research, persona generation, and skill-writing modules can process Feishu, Slack, DingTalk, and email archives without format-specific logic. The `timestamp` field preserves the original string representation from each platform, allowing downstream components to handle date normalization according to their specific requirements.

## Practical Usage Examples

### Parsing Feishu JSON Exports

```python
from tools.feishu_parser import parse_feishu_json, extract_key_content, format_output

# Load JSON file and filter for specific sender

messages = parse_feishu_json("messages.json", target_name="张三")

# Classify into content buckets (long-form, decision, daily chat)

extracted = extract_key_content(messages)

# Generate formatted report for the pipeline

report = format_output("张三", extracted)
print(report)

```

### Processing Email Archives Programmatically

```python
from tools.email_parser import parse_eml_file, extract_key_content, format_output

# Parse single EML file with target filtering

emails = parse_eml_file("thread.eml", target="alice@example.com")

# Apply generic extraction and formatting

extracted = extract_key_content(emails)
print(format_output("alice@example.com", extracted))

```

### Command-Line Slack Collection

```bash
python tools/slack_auto_collector.py --file slack_export.json --name "bob" --output bob.txt

```

This command reads the Slack export, filters for participant "bob", and writes a plain-text summary ready for Distilly's downstream stages.

## Summary

- **Four core parsers**: Distilly ships with dedicated parsers for Feishu/Lark ([`feishu_parser.py`](https://github.com/titanwings/distilly/blob/main/feishu_parser.py)), Email ([`email_parser.py`](https://github.com/titanwings/distilly/blob/main/email_parser.py)), Slack ([`slack_auto_collector.py`](https://github.com/titanwings/distilly/blob/main/slack_auto_collector.py)), and DingTalk ([`dingtalk_auto_collector.py`](https://github.com/titanwings/distilly/blob/main/dingtalk_auto_collector.py)).
- **Multiple format support**: Handles official JSON exports, TXT files, EML/MBOX archives, and plain-text conversation dumps.
- **Unified output schema**: All parsers return identical dictionary structures with `sender`, `content`, and `timestamp` keys, enabling platform-agnostic processing.
- **Integrated classification**: The `extract_key_content()` function in [`feishu_parser.py`](https://github.com/titanwings/distilly/blob/main/feishu_parser.py) provides generic content bucketing for all material types.
- **Flexible access patterns**: Parsers support both programmatic Python imports and command-line execution with participant filtering.

## Frequently Asked Questions

### What file formats does Distilly support for Feishu exports?

Distilly accepts both official Feishu JSON exports (arrays of message objects) and handcrafted TXT files formatted as `<timestamp> <sender>: <content>`. The `parse_feishu_json()` and `parse_feishu_txt()` functions in [`tools/feishu_parser.py`](https://github.com/titanwings/distilly/blob/main/tools/feishu_parser.py) handle these formats respectively, filtering by `target_name` and normalizing them into the standard message schema.

### Can Distilly parse standard email archives like PST or MBOX files?

According to the Distilly source code, the [`tools/email_parser.py`](https://github.com/titanwings/distilly/blob/main/tools/email_parser.py) module supports *.eml* files, *.mbox* mailboxes, and plain-text email dumps via `parse_eml_file()`, `parse_mbox_file()`, and `parse_txt_file()`. PST files are not supported natively and require external conversion to one of these formats before processing.

### How does Distilly handle different timestamp formats across platforms?

Each built-in material parser preserves the raw timestamp string from its source platform—whether ISO format from Feishu, epoch timestamps from Slack, or custom strings from DingTalk. The unified schema stores these as strings in the `timestamp` field, allowing downstream modules to handle date normalization according to their specific requirements rather than enforcing a global format during ingestion.

### Is there a way to filter messages by specific participants before processing?

Yes. The Feishu parser accepts a `target_name` parameter in `parse_feishu_json()` and `parse_feishu_txt()`. The email parser uses a `target` parameter in `parse_eml_file()` and related functions. The Slack collector provides a `--name` CLI flag. All filters limit output to messages sent by or addressed to the specified participant before returning the normalized data structure, reducing noise in the downstream pipeline.