Distilly Built-In Material Parsers: Supported Formats and Implementation Guide

Distilly includes four built-in material parsers for Feishu/Lark, Email, Slack, and DingTalk that normalize diverse communication exports into a unified message schema for persona generation.

Distilly provides a dedicated ingestion layer that transforms raw enterprise communication logs into structured training data. Located in the repository's tools/ directory, these built-in material parsers handle JSON exports, email archives, and chat logs from major platforms. Each parser implements a normalization layer that converts disparate formats into a canonical dictionary structure suitable for the downstream research and skill-writing pipeline.

Feishu and Lark Message Parsers

The tools/feishu_parser.py module provides comprehensive support for Feishu (Lark) communication exports, handling both official and handcrafted formats.

Supported Formats and Functions

This parser offers two primary entry points for different export types:

  • parse_feishu_json(): Processes official JSON exports containing arrays of message objects
  • parse_feishu_txt(): Parses TXT exports formatted as <timestamp> <sender>: <content>

Both functions read the source file, normalize fields into sender, content, and timestamp keys, filter out system artifacts and non-target messages, and return a list of standardized dictionaries.

Generic Content Classification

The module exposes extract_key_content(), a versatile classifier that buckets parsed messages into long-form content, decision-type exchanges, and daily chat categories. Located in tools/feishu_parser.py, this function serves as a generic extractor used across multiple collectors to bridge raw parsing and Distilly's research pipeline.

Email Archive Parsers

Located in tools/email_parser.py, the email ingestion module supports three distinct archive formats common in enterprise environments.

EML, MBOX, and Plain Text Processing

The implementation provides three specialized functions:

  • parse_eml_file(): Reads individual .eml files and extracts headers and body content
  • parse_mbox_file(): Iterates through Unix mailbox archives containing multiple email threads
  • parse_txt_file(): Handles plain-text conversation dumps

Each function accepts a target parameter to filter messages by sender or recipient, ensuring only relevant correspondence enters the training pipeline.

Slack Export Parser

For Slack workspace data, tools/slack_auto_collector.py processes official JSON exports generated by Slack's native "Export" feature.

Implementation and CLI Interface

The collector loads the JSON payload and walks the messages array to extract participant names, message content, and timestamps. It supports an optional --name command-line flag to limit output to specific participants. The parser normalizes Slack's proprietary data structure into the canonical Distilly message format, making it interchangeable with Feishu and DingTalk sources.

DingTalk Export Parser

Mirroring the Slack implementation, tools/dingtalk_auto_collector.py handles exports from the DingTalk admin console.

JSON Processing Pipeline

This parser reads the top-level messages field from DingTalk JSON exports and applies the same normalization logic as the Slack collector. As implemented in titanwings/distilly, it yields a list of dictionaries identical to other parsers, ensuring the downstream persona-generation workflow treats DingTalk data uniformly regardless of source platform.

Unified Message Schema

All Distilly built-in material parsers return a standardized Python list of dictionaries with the following schema:

{
    "sender": str,      # Name or identifier of the author

    "content": str,     # The plain-text message body

    "timestamp": str    # Raw timestamp string (ISO, epoch, or source-specific format)

}

Because they share this canonical shape, the research, persona generation, and skill-writing modules can process Feishu, Slack, DingTalk, and email archives without format-specific logic. The timestamp field preserves the original string representation from each platform, allowing downstream components to handle date normalization according to their specific requirements.

Practical Usage Examples

Parsing Feishu JSON Exports

from tools.feishu_parser import parse_feishu_json, extract_key_content, format_output

# Load JSON file and filter for specific sender

messages = parse_feishu_json("messages.json", target_name="张三")

# Classify into content buckets (long-form, decision, daily chat)

extracted = extract_key_content(messages)

# Generate formatted report for the pipeline

report = format_output("张三", extracted)
print(report)

Processing Email Archives Programmatically

from tools.email_parser import parse_eml_file, extract_key_content, format_output

# Parse single EML file with target filtering

emails = parse_eml_file("thread.eml", target="alice@example.com")

# Apply generic extraction and formatting

extracted = extract_key_content(emails)
print(format_output("alice@example.com", extracted))

Command-Line Slack Collection

python tools/slack_auto_collector.py --file slack_export.json --name "bob" --output bob.txt

This command reads the Slack export, filters for participant "bob", and writes a plain-text summary ready for Distilly's downstream stages.

Summary

  • Four core parsers: Distilly ships with dedicated parsers for Feishu/Lark (feishu_parser.py), Email (email_parser.py), Slack (slack_auto_collector.py), and DingTalk (dingtalk_auto_collector.py).
  • Multiple format support: Handles official JSON exports, TXT files, EML/MBOX archives, and plain-text conversation dumps.
  • Unified output schema: All parsers return identical dictionary structures with sender, content, and timestamp keys, enabling platform-agnostic processing.
  • Integrated classification: The extract_key_content() function in feishu_parser.py provides generic content bucketing for all material types.
  • Flexible access patterns: Parsers support both programmatic Python imports and command-line execution with participant filtering.

Frequently Asked Questions

What file formats does Distilly support for Feishu exports?

Distilly accepts both official Feishu JSON exports (arrays of message objects) and handcrafted TXT files formatted as <timestamp> <sender>: <content>. The parse_feishu_json() and parse_feishu_txt() functions in tools/feishu_parser.py handle these formats respectively, filtering by target_name and normalizing them into the standard message schema.

Can Distilly parse standard email archives like PST or MBOX files?

According to the Distilly source code, the tools/email_parser.py module supports .eml files, .mbox mailboxes, and plain-text email dumps via parse_eml_file(), parse_mbox_file(), and parse_txt_file(). PST files are not supported natively and require external conversion to one of these formats before processing.

How does Distilly handle different timestamp formats across platforms?

Each built-in material parser preserves the raw timestamp string from its source platform—whether ISO format from Feishu, epoch timestamps from Slack, or custom strings from DingTalk. The unified schema stores these as strings in the timestamp field, allowing downstream modules to handle date normalization according to their specific requirements rather than enforcing a global format during ingestion.

Is there a way to filter messages by specific participants before processing?

Yes. The Feishu parser accepts a target_name parameter in parse_feishu_json() and parse_feishu_txt(). The email parser uses a target parameter in parse_eml_file() and related functions. The Slack collector provides a --name CLI flag. All filters limit output to messages sent by or addressed to the specified participant before returning the normalized data structure, reducing noise in the downstream pipeline.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →