How the Document Parsing Trace Timeline Reconstructs Langfuse-Style Span Trees in WeKnora

WeKnora reconstructs Langfuse-style span trees by persisting hierarchical processing metadata to PostgreSQL and recursively querying parent-child relationships to render a nested timeline in the Vue frontend.

In the Tencent/WeKnora project, every document ingestion pipeline is instrumented with distributed tracing to provide complete observability into multi-stage processing. The document parsing trace timeline visualizes this execution flow as a nested tree of spans, mirroring the Langfuse observability format. This implementation captures each pipeline stage—from document reading through chunking, embedding, and reranking—as hierarchical metadata that can be rendered natively in the UI or exported to external Langfuse instances.

Root Trace Creation and Pipeline Instrumentation

When a request arrives at the /api/v1/knowledge/create HTTP endpoint, the system initializes a root trace that serves as the anchor for all subsequent operations.

Initializing the Document Processing Trace

In internal/application/service/knowledge_process.go, the Langfuse manager creates the top-level span using StartSpan with the context and span options. This establishes the root of the tree before any pipeline stages execute.

ctx, rootSpan := langfuse.GetManager().StartSpan(ctx, langfuse.SpanOptions{
    Name: "document.process",
    Kind: "DocumentProcessing",
})

Creating Child Spans for Each Stage

Each subsequent pipeline stage opens a child span under the current context. Throughout the processing service and agent engine files, the StartSpan method copies the active Langfuse trace context and records the span’s ID, name, kind, and parent-span-ID.

// Inside the embedding stage
ctx, embedSpan := langfuse.GetManager().StartSpan(ctx, langfuse.SpanOptions{
    Name: "embedding",
    Kind: "Embedding",
})

// Later, when the stage completes
embedSpan.Finish(ctx, langfuse.SpanFinishOptions{
    Status:   "succeeded",
    Output:   json.RawMessage(`{"vectors":[…]}`),
    Duration: time.Since(startTime).Milliseconds(),
})

The Finish method implementation in internal/tracing/langfuse/tracer.go handles the persistence logic, ensuring every span captures its input/output payload and timing data.

Cross-Process Context Propagation

For asynchronous operations such as background document processing, WeKnora serializes the tracing context to maintain span continuity across worker boundaries. The types.TracingContext struct defined in internal/types/tracing.go carries the Langfuse trace ID, parent observation ID, and W3C traceparent header.

When workers receive DocumentProcessPayload tasks, they extract this context to ensure spans created in separate goroutines attach to the correct parent. This prevents orphaned observations and preserves the tree structure across process boundaries.

Persistence Schema and Span Storage

The knowledge_processing_spans Table Structure

Span metadata is stored in PostgreSQL via the knowledge_processing_spans table, defined in migrations/versioned/000055_knowledge_processing_spans.up.sql. The schema supports hierarchical relationships through foreign key constraints:

  • span_id (Primary Key)
  • parent_span_id (Foreign Key referencing another span in the same table)
  • name, kind, status, input, output (JSONB), duration_ms, metadata

This relational structure enables recursive querying to reconstruct the full tree from flat database rows.

Writing Span Metadata on Completion

When span.Finish() is invoked, the implementation writes the complete span record to the database using GORM. This includes serialized input/output JSON, error information, and millisecond-precision timing data, enabling the timeline to display exact durations and payload summaries for each stage.

Reconstructing the Hierarchical Tree for the UI

Recursive SQL Queries in the Repository Layer

The internal/application/repository/knowledge_span_repo.go file implements tree reconstruction through SQL queries that walk the hierarchy. The repository fetches spans for a specific knowledge_id and attempt, then assembles the tree by mapping parent_span_id relationships to their corresponding span_id values.

type SpanNode struct {
    ID       string
    ParentID string
    Name     string
    Children []*SpanNode
}

// Repository queries all spans for the knowledge ID and attempt
var rows []KnowledgeProcessingSpan
db.Where("knowledge_id = ? AND attempt = ?", kid, attempt).Find(&rows)

// Build map[span_id]*SpanNode and link children to parents

Vue Component Visualization

The frontend component frontend/src/components/knowledge-processing-timeline.vue transforms the flat SQL result set into a nested tree structure. It calculates relative timing between parent and child spans, displays color-coded status indicators, and renders token usage metrics for each stage. This provides a Langfuse-compatible visual representation where users can expand and collapse pipeline stages.

External Langfuse Export via OpenTelemetry

WeKnora can export the same span data to external Langfuse instances using the OpenTelemetry exporter in internal/tracing/langfuse/exporter.go. This OTLP/HTTP implementation sends the identical span hierarchy to Langfuse endpoints, ensuring the UI view inside WeKnora remains semantically consistent with the native Langfuse interface. Developers can inspect traces in either tool without losing context or structure.

Summary

  • Root traces are created via langfuse.GetManager().StartSpan in knowledge_process.go when document processing begins
  • Each pipeline stage generates child spans with unique IDs linked via parent_span_id to form the tree
  • Async context propagation uses types.TracingContext to maintain tree integrity across worker goroutines
  • Span data persists to the knowledge_processing_spans table with full JSONB metadata and timing information
  • The repository layer recursively queries parent-child relationships in knowledge_span_repo.go to reconstruct the hierarchy
  • The Vue timeline component visualizes the nested structure with timing, status, and token consumption data
  • Optional OTel export via exporter.go sends identical span data to external Langfuse instances

Frequently Asked Questions

How does WeKnora maintain span relationships across asynchronous workers?

WeKnora embeds a types.TracingContext (defined in internal/types/tracing.go) into async task payloads such as DocumentProcessPayload. This struct carries the Langfuse trace ID, parent observation ID, and W3C traceparent header. When background workers process these tasks, they restore the context to ensure new spans attach to the correct parent, preventing orphaned branches in the tree.

What database schema supports the document parsing trace timeline?

The timeline relies on the knowledge_processing_spans table defined in migrations/versioned/000055_knowledge_processing_spans.up.sql. Key columns include span_id as the primary key, parent_span_id as a self-referencing foreign key, and JSONB fields for input, output, and metadata. This schema enables recursive queries that reconstruct the full hierarchical tree from flat relational rows.

Can the trace data be exported to external observability platforms?

Yes. The internal/tracing/langfuse/exporter.go file implements an OpenTelemetry exporter that transmits span data via OTLP/HTTP to external Langfuse instances. Because WeKnora uses native Langfuse span semantics internally, the exported data requires no transformation and renders identically in both the WeKnora UI and the Langfuse platform.

Where is the root trace created when a document upload begins?

The root trace is created in internal/application/service/knowledge_process.go when the HTTP endpoint /api/v1/knowledge/create receives a request. The code calls langfuse.GetManager().StartSpan(ctx, langfuse.SpanOptions{Name: "document.process"}) to establish the root span before the pipeline stages (DocReader, Chunking, Embedding) begin execution.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →