# How Meetily Generates Meeting Summaries with Local and Remote LLM Providers

> Discover how Meetily generates meeting summaries using local and remote LLM providers like OpenAI Ollama and Groq for efficient AI-powered reports.

- Repository: [Zackriya Solutions/meetily](https://github.com/Zackriya-Solutions/meetily)
- Tags: how-to-guide
- Published: 2026-08-01

---

**Meetily generates meeting summaries through a multi-stage pipeline that supports seven LLM providers—both remote (OpenAI, Claude, Groq, OpenRouter) and local (Ollama, Built-in AI)—using intelligent chunking, hierarchical summarization, and template-driven report generation.**

Meetily is an open-source meeting assistant that transforms raw transcripts into structured Markdown reports. The summary generation system, implemented in Rust, adapts to both cloud-based APIs and self-hosted models through a provider-agnostic architecture. This guide examines the complete pipeline from transcript to final report.

## Chunking and Token Management for Large Transcripts

Long meetings exceed most models' context windows. Meetily addresses this through adaptive **chunking** controlled by a configurable token threshold.

In `src/summary/processor.rs【54-L92】`, the pipeline inspects `rough_token_count` of the raw transcript. When the count exceeds `token_threshold` (default ~4000 tokens) **and** the provider is local (Ollama or Built-in AI), the text splits into overlapping chunks via `chunk_text`. Overlap ensures continuity between segments.

```rust
// From processor.rs - chunking logic
if rough_token_count > token_threshold && is_local_provider {
    let chunks = chunk_text(transcript, target_chunk_size, overlap_ratio);
    // Each chunk processed independently, then combined
}

```

Remote providers with larger contexts may skip chunking, sending the full transcript in one request.

## Per-Chunk Summarization via LLM Client

Each chunk flows through `generate_summary` in `src/summary/llm_client.rs【94-L127】`. The function assembles provider-specific HTTP requests using the `LLMProvider` enum:

```rust
pub enum LLMProvider {
    OpenAI,
    Claude,
    Groq,
    Ollama,
    OpenRouter,
    CustomOpenAI,
    BuiltInAI,  // Local sidecar engine
}

```

The client branches on `match provider` to construct proper authentication headers, request bodies, and endpoint URLs. For example, Claude requires `anthropic-version` headers while Groq uses OpenAI-compatible formatting.

Prompt construction happens in `build_chunk_summary_user_prompt`, combining:
- System instructions for meeting summarization
- The transcript chunk
- Optional user-provided context

## Hierarchical Summary Combination

Multiple chunk summaries merge through a **combine step**. Individual summaries concatenate with `---` separators, then feed back to the LLM via `build_combine_summary_user_prompt`. This produces a single intermediate summary preserving narrative flow across the entire meeting.

The approach mirrors MapReduce: parallel chunk processing followed by centralized reduction. For a 3-chunk meeting:

```

Chunk 1 summary →
                 └──→ Combined prompt → Intermediate summary
Chunk 2 summary →     (full context)
                 │
Chunk 3 summary →

```

## Template-Driven Final Report Generation

The intermediate summary (or original transcript for short meetings) renders through a **Markdown template system** defined in [`src/summary/templates.rs`](https://github.com/Zackriya-Solutions/meetily/blob/main/src/summary/templates.rs). Two key methods drive prompt construction:

- `template.to_markdown_structure()` – Emits the target document skeleton
- `template.to_section_instructions()` – Provides per-section guidance

`build_final_report_system_prompt` assembles the final instruction set, directing the LLM to populate the template while enforcing strict output rules. The template system ensures consistent formatting across different meeting types (standups, retrospectives, 1:1s).

## Multi-Language Output Handling

Meetily supports non-English summaries through a **two-pass translation architecture**:

**English Target (Default)**
Pipeline may apply `normalize_markdown_to_english` to strip stray non-English prose while preserving Markdown structure.

**Non-English Target**
1. Generate English summary first (cached as canonical version)
2. Execute `translate_markdown` with translation-specific system prompt
3. Return translated output as `final_markdown`

```rust
// Pseudocode flow from processor.rs
let english_summary = generate_english_summary(&transcript).await?;
let final_output = if target_lang == "en" {
    normalize_markdown_to_english(english_summary)?
} else {
    translate_markdown(&english_summary, target_lang).await?
};

```

## Cancellation and Error Resilience

All async operations accept `CancellationToken` from `tokio_util`. Signaled cancellation aborts immediately, returning an error string. LLM API failures propagate as human-readable messages rather than panics.

## Complete Usage Examples

**Local Ollama provider:**

```rust
use meetily::summary::{
    processor::generate_meeting_summary,
    llm_client::LLMProvider,
};
use reqwest::Client;

let client = Client::new();
let provider = LLMProvider::Ollama;

let (final_md, english_md, chunks) = generate_meeting_summary(
    &client,
    &provider,
    "llama-3.1-8b",      // model name
    "",                   // no API key for local
    &transcript_text,
    "",                   // optional custom prompt
    "standard_meeting",
    &template,
    4000,                 // token threshold
    None,                 // default Ollama endpoint
    None, None, None,     // generation params
    None,                 // app-data dir
    None,                 // cancellation token
    Some("fr"),           // target: French
    None, None, None,
).await?;

```

**Remote OpenAI provider:**

```rust
let provider = LLMProvider::OpenAI;
let api_key = std::env::var("OPENAI_API_KEY")?;

let (final_md, _, _) = generate_meeting_summary(
    &client,
    &provider,
    "gpt-4o-mini",
    &api_key,
    &transcript,
    "",
    "daily_standup",
    &template,
    4000,
    None, None,
    Some(1024),      // max_tokens
    Some(0.7),       // temperature
    Some(0.9),       // top_p
    None, None, None, None, None,
).await?;

```

## Key Source Files

| Module | Responsibility |
|--------|----------------|
| [`processor.rs`](https://github.com/Zackriya-Solutions/meetily/blob/main/processor.rs) | Orchestration: chunking, multi-level summarization, language handling, template rendering |
| [`llm_client.rs`](https://github.com/Zackriya-Solutions/meetily/blob/main/llm_client.rs) | Provider abstraction, HTTP construction, response parsing |
| [`templates.rs`](https://github.com/Zackriya-Solutions/meetily/blob/main/templates.rs) | Markdown template definitions and prompt helpers |
| [`summary_engine.rs`](https://github.com/Zackriya-Solutions/meetily/blob/main/summary_engine.rs) | Built-in AI sidecar execution for local inference |
| [`audio/transcription/engine.rs`](https://github.com/Zackriya-Solutions/meetily/blob/main/audio/transcription/engine.rs) | Transcript generation feeding the summary pipeline |

## Summary

- **Adaptive chunking** protects local models from context overflow while allowing remote providers to process full transcripts
- **Seven provider variants** share a unified interface through `LLMProvider` enum and `generate_summary` function
- **Hierarchical summarization** (chunk → combine → finalize) maintains coherence across long meetings
- **Template system** enforces consistent output structure across meeting types
- **Translation architecture** generates canonical English first, then localizes to target language
- **Cancellation tokens** ensure responsive UI behavior during lengthy operations

## Frequently Asked Questions

### How does Meetily handle transcripts longer than the model's context window?

Meetily splits long transcripts into overlapping chunks when `rough_token_count` exceeds `token_threshold` (default 4000 tokens) and the provider is local. Each chunk summarizes independently, then combines through a secondary LLM call. Remote providers with larger contexts may process the full transcript without chunking.

### What is the Built-in AI provider and how does it differ from Ollama?

**Built-in AI** executes local inference through a sidecar engine ([`summary_engine.rs`](https://github.com/Zackriya-Solutions/meetily/blob/main/summary_engine.rs)), embedding the LLM directly into Meetily's process. **Ollama** communicates via HTTP to an external Ollama server. Both run locally without API keys, but Built-in AI requires `app_data_dir` for model storage while Ollama needs an endpoint URL.

### Can I customize the meeting report format?

Yes. Templates in [`templates.rs`](https://github.com/Zackriya-Solutions/meetily/blob/main/templates.rs) define Markdown structures with `to_markdown_structure()` and section instructions. Pass a template ID and `Template` struct to `generate_meeting_summary`. The LLM populates this skeleton, so custom templates yield custom output formats without code changes.

### Why does non-English output go through English first?

This **canonical intermediate** approach ensures consistent quality. The English summary serves as cached ground truth for regeneration, comparison, or translation to other languages. It also isolates translation errors from core summarization quality.

### How do I cancel an in-progress summary generation?

Pass a `CancellationToken` to `generate_meeting_summary`. Call `.cancel()` from your UI thread; the async operation detects this and aborts with an error message. This prevents wasted compute and improves perceived responsiveness.