How to Create and Run Asynchronous Migrations in PostHog
To create and run asynchronous migrations in PostHog, subclass AsyncMigrationDefinition in the posthog/async_migrations/migrations/ directory, implement the operations property with AsyncMigrationOperationSQL instances or Python callables, and execute via the run_async_migrations management command or the Instance Settings UI.
PostHog ships with a dedicated async migration framework specifically designed for long-running data backfills and complex schema changes that cannot safely execute within standard Django migrations. Unlike synchronous migrations that block application startup, this system—implemented across posthog/async_migrations/definition.py and posthog/async_migrations/runner.py—orchestrates ClickHouse and PostgreSQL operations asynchronously through Celery, providing progress tracking, health monitoring, and automatic rollback capabilities.
Understanding the Async Migration Framework
The async migration system in PostHog centers on three core components:
AsyncMigrationDefinition– A declarative base class inposthog/async_migrations/definition.pythat describes migration metadata, version constraints, dependencies, and the sequence of operations.AsyncMigrationOperation/AsyncMigrationOperationSQL– Individual execution units that run either synchronous Python code or asynchronous SQL against ClickHouse or PostgreSQL.- Runner & CLI – The
AsyncMigrationRunnerclass inposthog/async_migrations/runner.pyorchestrates execution, persists state to the database, handles retries, and manages rollbacks, whileposthog/management/commands/run_async_migrations.pyprovides the command-line interface.
When triggered, the runner creates an AsyncMigration database row, executes each operation sequentially while storing current_operation_index, and updates status to Complete, Errored, or Rolled back based on outcome.
Step 1: Create the Migration File
All async migrations live in posthog/async_migrations/migrations/ and are auto-discovered at import time, requiring no manual registration.
- Create a new Python module under
posthog/async_migrations/migrations/using zero-padded numeric prefixes (e.g.,0010_my_backfill.py). - Define a class inheriting from
AsyncMigrationDefinitionwhere the class name matches the intended migration identifier.
# posthog/async_migrations/migrations/0010_my_backfill.py
from posthog.async_migrations.definition import AsyncMigrationDefinition, AsyncMigrationOperationSQL
class MyBackfillMigration(AsyncMigrationDefinition):
"""
Example async migration that backfills a new ClickHouse column.
"""
posthog_min_version = "1.45.0"
description = "Backfill `event.timestamp` for historic events."
depends_on = "person_distinct_id2"
parameters = {
"batch_size": (100_000, "Rows per batch", int),
}
The framework uses the class name as the migration name for dependency resolution and CLI targeting.
Step 2: Define Migration Logic and Constraints
Configure version gating, dependencies, and conditional execution to ensure safety across diverse PostHog deployments.
posthog_min_version/posthog_max_version– Tuple or string defining compatible PostHog versions to prevent execution on incompatible releases.depends_on– String name of another async migration that must complete before this one runs, establishing a directed acyclic graph.parameters– Dictionary mapping parameter names to tuples of(default_value, help_text, type_cast_function), exposing tunable settings in the UI.
Implement lifecycle hooks to control execution flow:
def is_required(self) -> bool:
# Skip for fresh installs where column already exists
return not self._column_already_populated()
def precheck(self):
# Verify ClickHouse version supports required functions
return (True, None)
def healthcheck(self):
# Abort if disk usage exceeds threshold
return (self._disk_usage_ok(), "Disk usage too high")
Step 3: Implement Operations
The operations property returns a list of AsyncMigrationOperation instances that execute sequentially.
SQL-Based Operations
Use AsyncMigrationOperationSQL for ClickHouse or PostgreSQL backfills. Set per_shard=True to execute simultaneously across all ClickHouse shards.
@property
def operations(self):
sql = """
INSERT INTO events_new SELECT *
FROM events_old
WHERE timestamp IS NULL
ORDER BY id
LIMIT %(batch_size)s
"""
return [
AsyncMigrationOperationSQL(
sql=sql,
sql_settings={"allow_experimental_object_type": 1},
rollback="ALTER TABLE events_new DROP COLUMN timestamp",
per_shard=True,
)
]
Each operation supports a rollback parameter containing SQL or a callable invoked if the migration fails or is manually reversed.
Python-Based Operations
For custom logic, subclass AsyncMigrationOperation and implement a synchronous fn method receiving the migration instance and parameters dictionary. Keep execution fast, as it runs in a Celery worker thread.
Step 4: Run the Migration
Command Line Execution
Use the Django management command defined in posthog/management/commands/run_async_migrations.py:
# List pending migrations and their status
python manage.py run_async_migrations --list
# Execute a specific migration by class name
python manage.py run_async_migrations MyBackfillMigration
# Rollback a failed migration
python manage.py run_async_migrations --rollback MyBackfillMigration
The runner respects MAX_CONCURRENT_ASYNC_MIGRATIONS=1 (default) to prevent resource contention.
UI and API Execution
Self-hosted administrators can trigger migrations via Instance Settings → Async Migrations. The frontend consumes posthog/api/async_migration.py endpoints to display status, progress percentages, and error messages returned by healthcheck().
Step 5: Monitor, Test, and Rollback
Health Checks and Progress
Implement progress(self, migration_instance) to return an integer 0-100 for UI rendering. The default implementation uses the operation index.
def progress(self, migration_instance):
# Custom progress calculation based on rows processed
return int(self._calculate_percentage())
Testing with AsyncMigrationTestCase
PostHog provides AsyncMigrationTestCase in posthog/async_migrations/test/ for isolated validation:
from posthog.async_migrations.test import AsyncMigrationTestCase
from posthog.async_migrations.migrations.0010_my_backfill import MyBackfillMigration
class MyBackfillTest(AsyncMigrationTestCase):
migration_class = MyBackfillMigration
def test_backfill(self):
self.assertFalse(self.migration_instance.is_complete())
self.run_migration()
self.assertTrue(self.migration_instance.is_complete())
Run tests with pytest posthog/async_migrations/test/.
Rollback Procedures
If an operation fails, the runner marks the migration as Errored. Invoke python manage.py run_async_migrations --rollback MigrationName to execute the rollback SQL or callable defined for each completed operation in reverse order.
Summary
- Define async migrations by subclassing
AsyncMigrationDefinitioninposthog/async_migrations/migrations/with version constraints (posthog_min_version), dependencies (depends_on), and UI parameters. - Implement operations using
AsyncMigrationOperationSQLfor database queries or customAsyncMigrationOperationsubclasses for Python logic, always providing rollback strategies. - Execute via
python manage.py run_async_migrations ClassNameor the Instance Settings UI, leveraging the runner inposthog/async_migrations/runner.pyfor state persistence and health monitoring. - Validate using
AsyncMigrationTestCaseto ensure idempotency and correctness before deployment.
Frequently Asked Questions
What is the difference between Django migrations and async migrations in PostHog?
Django migrations run synchronously during deployment and are suitable for schema changes and small data updates. Async migrations execute outside the deployment critical path via Celery workers, designed for long-running ClickHouse backfills or large PostgreSQL updates that would timeout or block standard migrations, as implemented in the posthog/async_migrations/ framework.
How do I ensure an async migration only runs when necessary?
Implement the is_required() method in your AsyncMigrationDefinition subclass to return True only when the backfill is actually needed. For fresh installations where the schema is already current, return False to skip execution and mark the migration as complete automatically.
Can I pass custom parameters to an async migration at runtime?
Yes. Define the parameters dictionary in your migration class with keys mapping to tuples of (default_value, description, type_function). These appear in the UI when triggering the migration manually, or can be passed via environment configuration. Access them in your SQL using %(parameter_name)s syntax or via the parameters dict in custom Python operations.
What happens if an async migration fails midway?
The runner stops execution, marks the migration status as Errored, and preserves the current_operation_index in the database. You can inspect logs via the UI or CLI (--list), fix the underlying issue, and resume. If rollback is required, run python manage.py run_async_migrations --rollback MigrationName to execute the rollback SQL defined for each completed operation in reverse order.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →