How PostHog Handles Temporal Workflow Activity Payloads and Limits: Avoiding the 2 MiB Constraint

PostHog avoids Temporal's ~2 MiB blob size limit by persisting large data outside activity boundaries and returning lightweight references, ensuring reliable execution of data-intensive background jobs.

PostHog relies on Temporal for orchestrating background workflows at scale. Understanding how the system handles Temporal workflow activity payloads and limits is critical for developers building tasks that process large datasets, such as insight exports or AI-generated reports.

Why Temporal Enforces a ~2 MiB Hard Limit on Activity Payloads

Temporal serializes activity inputs and outputs across a gRPC boundary that the platform caps at approximately 2 MiB per payload. When this limit is exceeded, the server rejects the request with a blobSizeLimitError (reference code TMPRL1103).

This constraint matters because the gRPC envelope is serialized once per activity execution. A single field containing a few megabytes of data will cause the entire workflow to fail before any downstream steps can run. For PostHog, where jobs like large Insight exports or AI-generated reports routinely produce multi-megabyte results, violating this limit would break critical background processes.

The Architectural Rule: Pass Large Data by Reference

The PostHog codebase enforces a strict design contract documented in AGENTS.md (lines 106-107):

Temporal activity payloads have a ~2 MiB hard limit — pass large data by reference, not by value. If a field could exceed ~256 KB once serialized, write it to Postgres / S3 from inside the activity and return only the reference (row ID, S3 key).

This 256 KB threshold serves as an early warning system. Rather than risking the hard 2 MiB boundary, developers proactively offload potentially large fields to persistent storage during activity execution.

Implementation Patterns in the PostHog Codebase

Data Type Definitions

In posthog/temporal/subscriptions/types.py (line 162), type definitions include explicit comments acknowledging the 2 MiB cap. These types deliberately retain only small metadata fields, omitting large blobs or marking them as size-bounded to prevent accidental serialization of oversized data.

Activity Implementation

The activity implementations in posthog/temporal/subscriptions/activities.py (line 258) demonstrate the recommended pattern with explicit comments warning about multi-MB query results. Before returning results, the activity persists heavy query output to Postgres as a JSONB snapshot, then returns only the primary key:


# posthog/temporal/subscriptions/activities.py

async def build_insight_delivery_snapshot(insight_id: int) -> dict:
    # Heavy query that may return millions of rows

    rows = run_hogql_query(insight_id)

    # Persist the raw rows in Postgres JSONB (no 2 MiB ceiling)

    snapshot = InsightDeliverySnapshot.objects.create(
        insight_id=insight_id,
        content_snapshot={"query_results": {"columns": [...], "results": rows}},
    )
    # Return only the snapshot primary key – < 100 bytes

    return {"snapshot_id": snapshot.id}

Error Handling Helpers

When payload violations do occur, posthog/temporal/common/errors.py defines custom exceptions that wrap Temporal’s blobSizeLimitError, surfacing clear diagnostic messages to developers and distinguishing TMPRL1103 failures from other runtime errors.

Practical Workflow Example

Workflows consume these references to reconstruct full datasets without ever touching the 2 MiB boundary. In posthog/temporal/subscriptions/workflows.py, the ProcessSubscriptionWorkflow calls the activity and subsequently loads the persisted data:


# posthog/temporal/subscriptions/workflows.py

@workflow.defn
class ProcessSubscriptionWorkflow:
    @workflow.run
    async def run(self, inputs: ProcessSubscriptionWorkflowInputs):
        # Call the activity – payload stays <2 MiB

        snapshot_ref = await workflow.execute_activity(
            build_insight_delivery_snapshot,
            inputs.insight_id,
            start_to_close_timeout=timedelta(minutes=10),
        )
        # Load the persisted data for further processing (e.g. email)

        snapshot = await sync_to_async(
            InsightDeliverySnapshot.objects.get
        )(id=snapshot_ref["snapshot_id"])
        await send_email_with_snapshot(snapshot.content_snapshot)

This pattern ensures the gRPC payload between the workflow and activity contains only the lightweight reference (under 100 bytes), while the actual megabyte-scale data travels through Postgres.

Enforcing Payload Limits at CI Time

PostHog protects against regressions through posthog/temporal/tests/test_subscriptions_workflows.py (lines 53-57). This regression test deliberately constructs a ~4 MiB query result and runs the workflow, asserting successful completion. If the activity ever attempts to return raw data instead of a reference, the workflow fails, blocking the CI pipeline.

Additionally, static analysis rules scan activity type definitions for fields whose default values could exceed the threshold, embedding the 2 MiB constraint into the repository's linting process.

Summary

  • Temporal enforces a ~2 MiB hard limit (TMPRL1103) on activity input and output payloads due to gRPC serialization constraints.
  • PostHog uses a 256 KB threshold as a safety margin; any field that might exceed this size is written to Postgres or S3 from within the activity.
  • Activities return references (row IDs or S3 keys) rather than large data objects, keeping payloads well under the limit.
  • Key files implementing this pattern include AGENTS.md (the rule definition), posthog/temporal/subscriptions/activities.py (implementation), and posthog/temporal/tests/test_subscriptions_workflows.py (regression testing).
  • Error handling in posthog/temporal/common/errors.py provides clear diagnostics when blob size errors occur.

Frequently Asked Questions

What is the exact payload size limit for Temporal activities?

Temporal caps activity input and output payloads at approximately 2 MiB (mebibytes) per serialization boundary. This limit is enforced at the gRPC level, and violations return the blobSizeLimitError (TMPRL1103) from the Temporal server.

What happens if an activity exceeds the 2 MiB payload limit?

The workflow will fail immediately upon attempting to serialize the oversized payload, before the activity result reaches any downstream steps. PostHog mitigates this by catching blobSizeLimitError in posthog/temporal/common/errors.py and by architectural rules that prohibit returning large values directly.

How should I structure activities that generate multi-megabyte reports?

Generate the report content inside the activity, persist it to Postgres or S3, create a reference record (e.g., InsightDeliverySnapshot), and return only the primary key or S3 key string. The calling workflow can then retrieve the full data using this reference in a subsequent step without violating the Temporal boundary.

Where does PostHog store large data instead of returning it in activity payloads?

Large data is typically stored in Postgres as JSONB snapshots (as seen in InsightDeliverySnapshot usage) or uploaded to S3 with reference keys returned to the workflow. This keeps the Temporal payload under the 2 MiB limit while maintaining data availability through standard database or object storage queries.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →