# How the Document Parsing Trace Timeline Reconstructs Langfuse-Style Span Trees in WeKnora

> Learn how WeKnora reconstructs Langfuse-style span trees by persisting processing metadata to PostgreSQL and querying parent-child relationships for its Vue frontend timeline.

- Repository: [Tencent/WeKnora](https://github.com/tencent/WeKnora)
- Tags: internals
- Published: 2026-09-12

---

**WeKnora reconstructs Langfuse-style span trees by persisting hierarchical processing metadata to PostgreSQL and recursively querying parent-child relationships to render a nested timeline in the Vue frontend.**

In the Tencent/WeKnora project, every document ingestion pipeline is instrumented with distributed tracing to provide complete observability into multi-stage processing. The **document parsing trace timeline** visualizes this execution flow as a nested tree of spans, mirroring the Langfuse observability format. This implementation captures each pipeline stage—from document reading through chunking, embedding, and reranking—as hierarchical metadata that can be rendered natively in the UI or exported to external Langfuse instances.

## Root Trace Creation and Pipeline Instrumentation

When a request arrives at the `/api/v1/knowledge/create` HTTP endpoint, the system initializes a root trace that serves as the anchor for all subsequent operations.

### Initializing the Document Processing Trace

In [`internal/application/service/knowledge_process.go`](https://github.com/Tencent/WeKnora/blob/main/internal/application/service/knowledge_process.go), the Langfuse manager creates the top-level span using `StartSpan` with the context and span options. This establishes the root of the tree before any pipeline stages execute.

```go
ctx, rootSpan := langfuse.GetManager().StartSpan(ctx, langfuse.SpanOptions{
    Name: "document.process",
    Kind: "DocumentProcessing",
})

```

### Creating Child Spans for Each Stage

Each subsequent pipeline stage opens a child span under the current context. Throughout the processing service and agent engine files, the `StartSpan` method copies the active Langfuse trace context and records the span’s **ID**, **name**, **kind**, and **parent-span-ID**.

```go
// Inside the embedding stage
ctx, embedSpan := langfuse.GetManager().StartSpan(ctx, langfuse.SpanOptions{
    Name: "embedding",
    Kind: "Embedding",
})

// Later, when the stage completes
embedSpan.Finish(ctx, langfuse.SpanFinishOptions{
    Status:   "succeeded",
    Output:   json.RawMessage(`{"vectors":[…]}`),
    Duration: time.Since(startTime).Milliseconds(),
})

```

The `Finish` method implementation in [`internal/tracing/langfuse/tracer.go`](https://github.com/Tencent/WeKnora/blob/main/internal/tracing/langfuse/tracer.go) handles the persistence logic, ensuring every span captures its input/output payload and timing data.

## Cross-Process Context Propagation

For asynchronous operations such as background document processing, WeKnora serializes the tracing context to maintain span continuity across worker boundaries. The `types.TracingContext` struct defined in [`internal/types/tracing.go`](https://github.com/Tencent/WeKnora/blob/main/internal/types/tracing.go) carries the Langfuse trace ID, parent observation ID, and W3C `traceparent` header.

When workers receive `DocumentProcessPayload` tasks, they extract this context to ensure spans created in separate goroutines attach to the correct parent. This prevents orphaned observations and preserves the tree structure across process boundaries.

## Persistence Schema and Span Storage

### The knowledge_processing_spans Table Structure

Span metadata is stored in PostgreSQL via the `knowledge_processing_spans` table, defined in [`migrations/versioned/000055_knowledge_processing_spans.up.sql`](https://github.com/Tencent/WeKnora/blob/main/migrations/versioned/000055_knowledge_processing_spans.up.sql). The schema supports hierarchical relationships through foreign key constraints:

- `span_id` (Primary Key)
- `parent_span_id` (Foreign Key referencing another span in the same table)
- `name`, `kind`, `status`, `input`, `output` (JSONB), `duration_ms`, `metadata`

This relational structure enables recursive querying to reconstruct the full tree from flat database rows.

### Writing Span Metadata on Completion

When `span.Finish()` is invoked, the implementation writes the complete span record to the database using GORM. This includes serialized input/output JSON, error information, and millisecond-precision timing data, enabling the timeline to display exact durations and payload summaries for each stage.

## Reconstructing the Hierarchical Tree for the UI

### Recursive SQL Queries in the Repository Layer

The [`internal/application/repository/knowledge_span_repo.go`](https://github.com/Tencent/WeKnora/blob/main/internal/application/repository/knowledge_span_repo.go) file implements tree reconstruction through SQL queries that walk the hierarchy. The repository fetches spans for a specific `knowledge_id` and `attempt`, then assembles the tree by mapping `parent_span_id` relationships to their corresponding `span_id` values.

```go
type SpanNode struct {
    ID       string
    ParentID string
    Name     string
    Children []*SpanNode
}

// Repository queries all spans for the knowledge ID and attempt
var rows []KnowledgeProcessingSpan
db.Where("knowledge_id = ? AND attempt = ?", kid, attempt).Find(&rows)

// Build map[span_id]*SpanNode and link children to parents

```

### Vue Component Visualization

The frontend component [`frontend/src/components/knowledge-processing-timeline.vue`](https://github.com/Tencent/WeKnora/blob/main/frontend/src/components/knowledge-processing-timeline.vue) transforms the flat SQL result set into a nested tree structure. It calculates relative timing between parent and child spans, displays color-coded status indicators, and renders token usage metrics for each stage. This provides a Langfuse-compatible visual representation where users can expand and collapse pipeline stages.

## External Langfuse Export via OpenTelemetry

WeKnora can export the same span data to external Langfuse instances using the OpenTelemetry exporter in [`internal/tracing/langfuse/exporter.go`](https://github.com/Tencent/WeKnora/blob/main/internal/tracing/langfuse/exporter.go). This OTLP/HTTP implementation sends the identical span hierarchy to Langfuse endpoints, ensuring the UI view inside WeKnora remains semantically consistent with the native Langfuse interface. Developers can inspect traces in either tool without losing context or structure.

## Summary

- **Root traces** are created via `langfuse.GetManager().StartSpan` in [`knowledge_process.go`](https://github.com/Tencent/WeKnora/blob/main/knowledge_process.go) when document processing begins
- Each pipeline stage generates **child spans** with unique IDs linked via `parent_span_id` to form the tree
- **Async context propagation** uses `types.TracingContext` to maintain tree integrity across worker goroutines
- Span data persists to the **`knowledge_processing_spans`** table with full JSONB metadata and timing information
- The repository layer **recursively queries** parent-child relationships in [`knowledge_span_repo.go`](https://github.com/Tencent/WeKnora/blob/main/knowledge_span_repo.go) to reconstruct the hierarchy
- The **Vue timeline component** visualizes the nested structure with timing, status, and token consumption data
- Optional **OTel export** via [`exporter.go`](https://github.com/Tencent/WeKnora/blob/main/exporter.go) sends identical span data to external Langfuse instances

## Frequently Asked Questions

### How does WeKnora maintain span relationships across asynchronous workers?

WeKnora embeds a `types.TracingContext` (defined in [`internal/types/tracing.go`](https://github.com/Tencent/WeKnora/blob/main/internal/types/tracing.go)) into async task payloads such as `DocumentProcessPayload`. This struct carries the Langfuse trace ID, parent observation ID, and W3C `traceparent` header. When background workers process these tasks, they restore the context to ensure new spans attach to the correct parent, preventing orphaned branches in the tree.

### What database schema supports the document parsing trace timeline?

The timeline relies on the `knowledge_processing_spans` table defined in [`migrations/versioned/000055_knowledge_processing_spans.up.sql`](https://github.com/Tencent/WeKnora/blob/main/migrations/versioned/000055_knowledge_processing_spans.up.sql). Key columns include `span_id` as the primary key, `parent_span_id` as a self-referencing foreign key, and JSONB fields for `input`, `output`, and `metadata`. This schema enables recursive queries that reconstruct the full hierarchical tree from flat relational rows.

### Can the trace data be exported to external observability platforms?

Yes. The [`internal/tracing/langfuse/exporter.go`](https://github.com/Tencent/WeKnora/blob/main/internal/tracing/langfuse/exporter.go) file implements an OpenTelemetry exporter that transmits span data via OTLP/HTTP to external Langfuse instances. Because WeKnora uses native Langfuse span semantics internally, the exported data requires no transformation and renders identically in both the WeKnora UI and the Langfuse platform.

### Where is the root trace created when a document upload begins?

The root trace is created in [`internal/application/service/knowledge_process.go`](https://github.com/Tencent/WeKnora/blob/main/internal/application/service/knowledge_process.go) when the HTTP endpoint `/api/v1/knowledge/create` receives a request. The code calls `langfuse.GetManager().StartSpan(ctx, langfuse.SpanOptions{Name: "document.process"})` to establish the root span before the pipeline stages (DocReader, Chunking, Embedding) begin execution.