How Meetily Generates Meeting Summaries with Local and Remote LLM Providers
Meetily generates meeting summaries through a multi-stage pipeline that supports seven LLM providers—both remote (OpenAI, Claude, Groq, OpenRouter) and local (Ollama, Built-in AI)—using intelligent chunking, hierarchical summarization, and template-driven report generation.
Meetily is an open-source meeting assistant that transforms raw transcripts into structured Markdown reports. The summary generation system, implemented in Rust, adapts to both cloud-based APIs and self-hosted models through a provider-agnostic architecture. This guide examines the complete pipeline from transcript to final report.
Chunking and Token Management for Large Transcripts
Long meetings exceed most models' context windows. Meetily addresses this through adaptive chunking controlled by a configurable token threshold.
In src/summary/processor.rs【54-L92】, the pipeline inspects rough_token_count of the raw transcript. When the count exceeds token_threshold (default ~4000 tokens) and the provider is local (Ollama or Built-in AI), the text splits into overlapping chunks via chunk_text. Overlap ensures continuity between segments.
// From processor.rs - chunking logic
if rough_token_count > token_threshold && is_local_provider {
let chunks = chunk_text(transcript, target_chunk_size, overlap_ratio);
// Each chunk processed independently, then combined
}
Remote providers with larger contexts may skip chunking, sending the full transcript in one request.
Per-Chunk Summarization via LLM Client
Each chunk flows through generate_summary in src/summary/llm_client.rs【94-L127】. The function assembles provider-specific HTTP requests using the LLMProvider enum:
pub enum LLMProvider {
OpenAI,
Claude,
Groq,
Ollama,
OpenRouter,
CustomOpenAI,
BuiltInAI, // Local sidecar engine
}
The client branches on match provider to construct proper authentication headers, request bodies, and endpoint URLs. For example, Claude requires anthropic-version headers while Groq uses OpenAI-compatible formatting.
Prompt construction happens in build_chunk_summary_user_prompt, combining:
- System instructions for meeting summarization
- The transcript chunk
- Optional user-provided context
Hierarchical Summary Combination
Multiple chunk summaries merge through a combine step. Individual summaries concatenate with --- separators, then feed back to the LLM via build_combine_summary_user_prompt. This produces a single intermediate summary preserving narrative flow across the entire meeting.
The approach mirrors MapReduce: parallel chunk processing followed by centralized reduction. For a 3-chunk meeting:
Chunk 1 summary →
└──→ Combined prompt → Intermediate summary
Chunk 2 summary → (full context)
│
Chunk 3 summary →
Template-Driven Final Report Generation
The intermediate summary (or original transcript for short meetings) renders through a Markdown template system defined in src/summary/templates.rs. Two key methods drive prompt construction:
template.to_markdown_structure()– Emits the target document skeletontemplate.to_section_instructions()– Provides per-section guidance
build_final_report_system_prompt assembles the final instruction set, directing the LLM to populate the template while enforcing strict output rules. The template system ensures consistent formatting across different meeting types (standups, retrospectives, 1:1s).
Multi-Language Output Handling
Meetily supports non-English summaries through a two-pass translation architecture:
English Target (Default)
Pipeline may apply normalize_markdown_to_english to strip stray non-English prose while preserving Markdown structure.
Non-English Target
- Generate English summary first (cached as canonical version)
- Execute
translate_markdownwith translation-specific system prompt - Return translated output as
final_markdown
// Pseudocode flow from processor.rs
let english_summary = generate_english_summary(&transcript).await?;
let final_output = if target_lang == "en" {
normalize_markdown_to_english(english_summary)?
} else {
translate_markdown(&english_summary, target_lang).await?
};
Cancellation and Error Resilience
All async operations accept CancellationToken from tokio_util. Signaled cancellation aborts immediately, returning an error string. LLM API failures propagate as human-readable messages rather than panics.
Complete Usage Examples
Local Ollama provider:
use meetily::summary::{
processor::generate_meeting_summary,
llm_client::LLMProvider,
};
use reqwest::Client;
let client = Client::new();
let provider = LLMProvider::Ollama;
let (final_md, english_md, chunks) = generate_meeting_summary(
&client,
&provider,
"llama-3.1-8b", // model name
"", // no API key for local
&transcript_text,
"", // optional custom prompt
"standard_meeting",
&template,
4000, // token threshold
None, // default Ollama endpoint
None, None, None, // generation params
None, // app-data dir
None, // cancellation token
Some("fr"), // target: French
None, None, None,
).await?;
Remote OpenAI provider:
let provider = LLMProvider::OpenAI;
let api_key = std::env::var("OPENAI_API_KEY")?;
let (final_md, _, _) = generate_meeting_summary(
&client,
&provider,
"gpt-4o-mini",
&api_key,
&transcript,
"",
"daily_standup",
&template,
4000,
None, None,
Some(1024), // max_tokens
Some(0.7), // temperature
Some(0.9), // top_p
None, None, None, None, None,
).await?;
Key Source Files
| Module | Responsibility |
|---|---|
processor.rs |
Orchestration: chunking, multi-level summarization, language handling, template rendering |
llm_client.rs |
Provider abstraction, HTTP construction, response parsing |
templates.rs |
Markdown template definitions and prompt helpers |
summary_engine.rs |
Built-in AI sidecar execution for local inference |
audio/transcription/engine.rs |
Transcript generation feeding the summary pipeline |
Summary
- Adaptive chunking protects local models from context overflow while allowing remote providers to process full transcripts
- Seven provider variants share a unified interface through
LLMProviderenum andgenerate_summaryfunction - Hierarchical summarization (chunk → combine → finalize) maintains coherence across long meetings
- Template system enforces consistent output structure across meeting types
- Translation architecture generates canonical English first, then localizes to target language
- Cancellation tokens ensure responsive UI behavior during lengthy operations
Frequently Asked Questions
How does Meetily handle transcripts longer than the model's context window?
Meetily splits long transcripts into overlapping chunks when rough_token_count exceeds token_threshold (default 4000 tokens) and the provider is local. Each chunk summarizes independently, then combines through a secondary LLM call. Remote providers with larger contexts may process the full transcript without chunking.
What is the Built-in AI provider and how does it differ from Ollama?
Built-in AI executes local inference through a sidecar engine (summary_engine.rs), embedding the LLM directly into Meetily's process. Ollama communicates via HTTP to an external Ollama server. Both run locally without API keys, but Built-in AI requires app_data_dir for model storage while Ollama needs an endpoint URL.
Can I customize the meeting report format?
Yes. Templates in templates.rs define Markdown structures with to_markdown_structure() and section instructions. Pass a template ID and Template struct to generate_meeting_summary. The LLM populates this skeleton, so custom templates yield custom output formats without code changes.
Why does non-English output go through English first?
This canonical intermediate approach ensures consistent quality. The English summary serves as cached ground truth for regeneration, comparison, or translation to other languages. It also isolates translation errors from core summarization quality.
How do I cancel an in-progress summary generation?
Pass a CancellationToken to generate_meeting_summary. Call .cancel() from your UI thread; the async operation detects this and aborts with an error message. This prevents wasted compute and improves perceived responsiveness.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →