How the OpenHuman Document Synthesis Pipeline Works with tinydocs
The OpenHuman document synthesis pipeline offloads all heavy document processing—building DOCX and PPTX files and extracting PDF text—to the external tinydocs TinyBus module, while the core maintains only the wire contract and orchestration logic.
The tinyhumansai/openhuman repository implements a modular architecture where complex document generation runs in an isolated process. By routing operations through the TinyBus interface, the core system remains lightweight while supporting sophisticated document format handling through the specialized tinydocs module.
Wire Contract and Module Registration
Integration begins with the wire contract defined in the tinydocs-bus crate. The core imports this contract via pub(crate) use tinydocs_bus as format; in src/openhuman/tools/impl/document/mod.rs【src/openhuman/tools/impl/document/mod.rs†L42-L45】, ensuring the host and module share identical API definitions.
The module itself is registered in src/openhuman/modules/registry_part_01.rs, where its ID, bus name (ai.tinyhumans.tinydocs.Documents), and object path are declared【src/openhuman/modules/registry_part_01.rs†L3-L13】. Before invoking operations, the core ensures the module is loaded using ops::ensure_loaded and establishes a TinyBus connection to the specified bus name【src/openhuman/modules/documents.rs†L34-L40】.
Core Document Operations
The pipeline exposes three primary methods defined in tinydocs_bus::names::methods【src/openhuman/modules/documents.rs†L34-L40】:
generate_docx– Constructs Word documents from structured JSON specificationsgenerate_pptx– Builds PowerPoint presentations from template definitionsextract_pdf– Performs raw text extraction from binary PDF files
Each operation translates potential failures into the core's DocumentCallError enum, covering variants such as InvalidInput, GenerationFailed, and ExtractionFailed【src/openhuman/modules/documents.rs†L334-L338】.
Error Handling and Translation
When the tinydocs module returns errors, the documents layer maps them to strongly-typed variants in src/openhuman/modules/documents.rs. The DocumentCallError enum distinguishes between input validation failures, generation timeouts, and extraction errors【src/openhuman/modules/documents.rs†L334-L338】. This translation layer ensures upstream components handle document failures consistently without depending on tinydocs-specific error codes.
Tool Layer Abstraction
User-facing functionality surfaces through thin wrapper tools in src/openhuman/tools/impl/document/mod.rs. The generate_document and generate_presentation tools forward validated arguments to the documents module, mapping module-level errors back to tool-specific error types【src/openhuman/tools/impl/document/mod.rs†L71-L185】.
Input and output schemas reside in src/openhuman/tools/impl/document/types.rs, defining the JSON contracts for document generation requests【src/openhuman/tools/impl/document/types.rs†L1-L2】.
Multimodal Pipeline Integration
PDF extraction integrates with the multimodal agent through the DocumentsTextExtractor component in src/openhuman/agent/multimodal.rs. This extractor routes PDF processing through the tinydocs module with explicit deadline parameters to prevent blocking the agent pipeline【src/openhuman/agent/multimodal.rs†L13-L53】.
// Extracting text from PDF with timeout protection
use openhuman::modules::documents::extract_pdf;
let pdf_path = "/path/to/file.pdf";
let text = extract_pdf(
pdf_path,
std::time::Duration::from_secs(30)
)
.await
.map_err(|e| match e {
DocumentCallError::ExtractionFailed(msg) => format!("PDF parse error: {}", msg),
_ => "Document processing failed".to_string()
})?;
println!("Extracted {} characters", text.len());
Testing Strategy
Testing distinguishes between contract verification and integration testing. Unit tests in src/openhuman/modules/documents_tests.rs validate bus constants and method name integrity without loading the external module【src/openhuman/modules/documents_tests.rs†L148-L161】.
Integration tests requiring the actual tinydocs binary are marked with #[ignore] in src/openhuman/tools/impl/document/document_tests.rs and require the OPENHUMAN_MODULE_PATH environment variable to locate the compiled module【src/openhuman/tools/impl/document/document_tests.rs†L163-L167】.
// Generating a DOCX through the tool interface
use openhuman::tools::impl::document::{TOOL_NAME, DocumentInput};
let input = DocumentInput {
title: "Report".to_string(),
sections: vec![/* ... */],
..Default::default()
};
let result = core
.run_tool(TOOL_NAME, serde_json::to_value(input).unwrap())
.await?;
println!("Generated: {}", result["path"]);
Summary
- The document synthesis pipeline delegates format-specific work to the external tinydocs module, communicating via the
ai.tinyhumans.tinydocs.DocumentsTinyBus interface - Three core operations—
generate_docx,generate_pptx, andextract_pdf—handle Word, PowerPoint, and PDF processing respectively - Error translation converts tinydocs failures into the
DocumentCallErrorenum for consistent handling throughout the core - Tool wrappers in
src/openhuman/tools/impl/document/provide the public API surface while keeping implementation details encapsulated - PDF extraction supports multimodal agents through
DocumentsTextExtractorwith configurable deadline timeouts - Testing separates bus contract validation (unit tests) from full integration tests that require the built tinydocs binary
Frequently Asked Questions
How does OpenHuman handle errors from the tinydocs module?
The documents module translates all tinydocs errors into the DocumentCallError enum defined in src/openhuman/modules/documents.rs【src/openhuman/modules/documents.rs†L334-L338】. This includes variants like InvalidInput for schema violations, GenerationFailed for document construction failures, and ExtractionFailed for PDF parsing errors. The tool layer in src/openhuman/tools/impl/document/mod.rs then maps these module errors to tool-specific error types for end-user consumption【src/openhuman/tools/impl/document/mod.rs†L71-L185】.
What file formats does the document synthesis pipeline support?
According to the tinydocs bus interface defined in src/openhuman/modules/documents.rs, the pipeline supports three primary formats: DOCX generation via generate_docx, PPTX generation via generate_pptx, and PDF text extraction via extract_pdf【src/openhuman/modules/documents.rs†L34-L40】. All format-specific logic resides in the external tinydocs module, with the core maintaining only the abstract operation contracts.
How does the multimodal agent use PDF extraction?
The multimodal pipeline implements DocumentsTextExtractor in src/openhuman/agent/multimodal.rs to process PDF content. This component calls extract_pdf through the documents module with an explicit deadline parameter, ensuring PDF parsing does not block the agent indefinitely【src/openhuman/agent/multimodal.rs†L13-L53】. The extracted text feeds into subsequent multimodal processing stages.
What is required to test the tinydocs integration?
Unit tests in src/openhuman/modules/documents_tests.rs verify bus constants and method names without external dependencies【src/openhuman/modules/documents_tests.rs†L148-L161】. However, integration tests marked with #[ignore] in src/openhuman/tools/impl/document/document_tests.rs require a compiled tinydocs module accessible via the OPENHUMAN_MODULE_PATH environment variable【src/openhuman/tools/impl/document/document_tests.rs†L163-L167】.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →