# How the OpenHuman Document Synthesis Pipeline Works with tinydocs

> Discover how the OpenHuman document synthesis pipeline leverages tinydocs for efficient document processing including DOCX PPTX creation and PDF text extraction Streamline your workflow today

- Repository: [Tiny Humans/openhuman](https://github.com/tinyhumansai/openhuman)
- Tags: internals
- Published: 2026-08-30

---

**The OpenHuman document synthesis pipeline offloads all heavy document processing—building DOCX and PPTX files and extracting PDF text—to the external tinydocs TinyBus module, while the core maintains only the wire contract and orchestration logic.**

The tinyhumansai/openhuman repository implements a modular architecture where complex document generation runs in an isolated process. By routing operations through the TinyBus interface, the core system remains lightweight while supporting sophisticated document format handling through the specialized tinydocs module.

## Wire Contract and Module Registration

Integration begins with the wire contract defined in the `tinydocs-bus` crate. The core imports this contract via `pub(crate) use tinydocs_bus as format;` in [`src/openhuman/tools/impl/document/mod.rs`](https://github.com/tinyhumansai/openhuman/blob/main/src/openhuman/tools/impl/document/mod.rs)【src/openhuman/tools/impl/document/mod.rs†L42-L45】, ensuring the host and module share identical API definitions.

The module itself is registered in [`src/openhuman/modules/registry_part_01.rs`](https://github.com/tinyhumansai/openhuman/blob/main/src/openhuman/modules/registry_part_01.rs), where its ID, bus name (`ai.tinyhumans.tinydocs.Documents`), and object path are declared【src/openhuman/modules/registry_part_01.rs†L3-L13】. Before invoking operations, the core ensures the module is loaded using `ops::ensure_loaded` and establishes a TinyBus connection to the specified bus name【src/openhuman/modules/documents.rs†L34-L40】.

## Core Document Operations

The pipeline exposes three primary methods defined in `tinydocs_bus::names::methods`【src/openhuman/modules/documents.rs†L34-L40】:

- **`generate_docx`** – Constructs Word documents from structured JSON specifications
- **`generate_pptx`** – Builds PowerPoint presentations from template definitions
- **`extract_pdf`** – Performs raw text extraction from binary PDF files

Each operation translates potential failures into the core's `DocumentCallError` enum, covering variants such as `InvalidInput`, `GenerationFailed`, and `ExtractionFailed`【src/openhuman/modules/documents.rs†L334-L338】.

## Error Handling and Translation

When the tinydocs module returns errors, the documents layer maps them to strongly-typed variants in [`src/openhuman/modules/documents.rs`](https://github.com/tinyhumansai/openhuman/blob/main/src/openhuman/modules/documents.rs). The `DocumentCallError` enum distinguishes between input validation failures, generation timeouts, and extraction errors【src/openhuman/modules/documents.rs†L334-L338】. This translation layer ensures upstream components handle document failures consistently without depending on tinydocs-specific error codes.

## Tool Layer Abstraction

User-facing functionality surfaces through thin wrapper tools in [`src/openhuman/tools/impl/document/mod.rs`](https://github.com/tinyhumansai/openhuman/blob/main/src/openhuman/tools/impl/document/mod.rs). The `generate_document` and `generate_presentation` tools forward validated arguments to the documents module, mapping module-level errors back to tool-specific error types【src/openhuman/tools/impl/document/mod.rs†L71-L185】.

Input and output schemas reside in [`src/openhuman/tools/impl/document/types.rs`](https://github.com/tinyhumansai/openhuman/blob/main/src/openhuman/tools/impl/document/types.rs), defining the JSON contracts for document generation requests【src/openhuman/tools/impl/document/types.rs†L1-L2】.

## Multimodal Pipeline Integration

PDF extraction integrates with the multimodal agent through the `DocumentsTextExtractor` component in [`src/openhuman/agent/multimodal.rs`](https://github.com/tinyhumansai/openhuman/blob/main/src/openhuman/agent/multimodal.rs). This extractor routes PDF processing through the tinydocs module with explicit deadline parameters to prevent blocking the agent pipeline【src/openhuman/agent/multimodal.rs†L13-L53】.

```rust
// Extracting text from PDF with timeout protection
use openhuman::modules::documents::extract_pdf;

let pdf_path = "/path/to/file.pdf";
let text = extract_pdf(
    pdf_path,
    std::time::Duration::from_secs(30)
)
.await
.map_err(|e| match e {
    DocumentCallError::ExtractionFailed(msg) => format!("PDF parse error: {}", msg),
    _ => "Document processing failed".to_string()
})?;

println!("Extracted {} characters", text.len());

```

## Testing Strategy

Testing distinguishes between contract verification and integration testing. Unit tests in [`src/openhuman/modules/documents_tests.rs`](https://github.com/tinyhumansai/openhuman/blob/main/src/openhuman/modules/documents_tests.rs) validate bus constants and method name integrity without loading the external module【src/openhuman/modules/documents_tests.rs†L148-L161】.

Integration tests requiring the actual tinydocs binary are marked with `#[ignore]` in [`src/openhuman/tools/impl/document/document_tests.rs`](https://github.com/tinyhumansai/openhuman/blob/main/src/openhuman/tools/impl/document/document_tests.rs) and require the `OPENHUMAN_MODULE_PATH` environment variable to locate the compiled module【src/openhuman/tools/impl/document/document_tests.rs†L163-L167】.

```rust
// Generating a DOCX through the tool interface
use openhuman::tools::impl::document::{TOOL_NAME, DocumentInput};

let input = DocumentInput {
    title: "Report".to_string(),
    sections: vec![/* ... */],
    ..Default::default()
};

let result = core
    .run_tool(TOOL_NAME, serde_json::to_value(input).unwrap())
    .await?;

println!("Generated: {}", result["path"]);

```

## Summary

- The document synthesis pipeline delegates format-specific work to the external tinydocs module, communicating via the `ai.tinyhumans.tinydocs.Documents` TinyBus interface
- Three core operations—`generate_docx`, `generate_pptx`, and `extract_pdf`—handle Word, PowerPoint, and PDF processing respectively
- Error translation converts tinydocs failures into the `DocumentCallError` enum for consistent handling throughout the core
- Tool wrappers in `src/openhuman/tools/impl/document/` provide the public API surface while keeping implementation details encapsulated
- PDF extraction supports multimodal agents through `DocumentsTextExtractor` with configurable deadline timeouts
- Testing separates bus contract validation (unit tests) from full integration tests that require the built tinydocs binary

## Frequently Asked Questions

### How does OpenHuman handle errors from the tinydocs module?

The documents module translates all tinydocs errors into the `DocumentCallError` enum defined in [`src/openhuman/modules/documents.rs`](https://github.com/tinyhumansai/openhuman/blob/main/src/openhuman/modules/documents.rs)【src/openhuman/modules/documents.rs†L334-L338】. This includes variants like `InvalidInput` for schema violations, `GenerationFailed` for document construction failures, and `ExtractionFailed` for PDF parsing errors. The tool layer in [`src/openhuman/tools/impl/document/mod.rs`](https://github.com/tinyhumansai/openhuman/blob/main/src/openhuman/tools/impl/document/mod.rs) then maps these module errors to tool-specific error types for end-user consumption【src/openhuman/tools/impl/document/mod.rs†L71-L185】.

### What file formats does the document synthesis pipeline support?

According to the tinydocs bus interface defined in [`src/openhuman/modules/documents.rs`](https://github.com/tinyhumansai/openhuman/blob/main/src/openhuman/modules/documents.rs), the pipeline supports three primary formats: DOCX generation via `generate_docx`, PPTX generation via `generate_pptx`, and PDF text extraction via `extract_pdf`【src/openhuman/modules/documents.rs†L34-L40】. All format-specific logic resides in the external tinydocs module, with the core maintaining only the abstract operation contracts.

### How does the multimodal agent use PDF extraction?

The multimodal pipeline implements `DocumentsTextExtractor` in [`src/openhuman/agent/multimodal.rs`](https://github.com/tinyhumansai/openhuman/blob/main/src/openhuman/agent/multimodal.rs) to process PDF content. This component calls `extract_pdf` through the documents module with an explicit deadline parameter, ensuring PDF parsing does not block the agent indefinitely【src/openhuman/agent/multimodal.rs†L13-L53】. The extracted text feeds into subsequent multimodal processing stages.

### What is required to test the tinydocs integration?

Unit tests in [`src/openhuman/modules/documents_tests.rs`](https://github.com/tinyhumansai/openhuman/blob/main/src/openhuman/modules/documents_tests.rs) verify bus constants and method names without external dependencies【src/openhuman/modules/documents_tests.rs†L148-L161】. However, integration tests marked with `#[ignore]` in [`src/openhuman/tools/impl/document/document_tests.rs`](https://github.com/tinyhumansai/openhuman/blob/main/src/openhuman/tools/impl/document/document_tests.rs) require a compiled tinydocs module accessible via the `OPENHUMAN_MODULE_PATH` environment variable【src/openhuman/tools/impl/document/document_tests.rs†L163-L167】.