How to Use Celery with Apache Superset: Architecture, Setup, and Configuration
To use Celery with Apache Superset, configure CELERY_BROKER_URL and RESULT_BACKEND in superset_config.py, instantiate the shared Celery application in superset/tasks/celery_app.py, and launch workers using celery -A superset.tasks.celery_app worker.
Apache Superset relies on Celery to execute long-running or periodic background jobs—such as chart caching, email alerts, and asynchronous SQL queries—without blocking the main web application. This guide explains the architecture, configuration files, and operational steps required to integrate Celery into the superset-sh/superset codebase.
Understanding the Celery Architecture in Superset
The integration follows a standard distributed task queue pattern consisting of three core components that work together to offload work from the Superset web processes.
The Three Core Components
| Component | Role | Superset Integration |
|---|---|---|
| Broker (Redis, RabbitMQ, or Amazon SQS) | Queues tasks submitted by Superset. | Superset pushes tasks via celery.send_task() or apply_async() using the URL defined in CELERY_BROKER_URL. |
| Worker(s) | Pull tasks from the broker, execute the Python functions, and store results. | Workers import the shared app from superset/tasks/celery_app.py and run celery worker -A superset.tasks.celery_app. |
| Result Backend (Redis, Database, or S3) | Persists task results so Superset can poll for completion. | Superset checks task status via AsyncResult(task_id) against the backend configured in RESULT_BACKEND. |
Data Flow for Async Operations
- The Superset UI triggers an async operation (e.g., "Refresh chart cache" or "Run async query").
- Superset’s Python code creates a Celery task and sends it to the broker.
- A Celery worker picks up the task, executes the function defined in
superset/tasks/, and stores the output in the result backend. - Superset periodically polls the task state via
AsyncResultand updates the UI or caches the final result.
Configuring Celery in superset_config.py
All Celery-related settings reside in your custom configuration file, typically named superset_config.py. This file must be importable by both the Superset web processes and the Celery workers.
Essential Configuration Variables
# superset_config.py
from datetime import timedelta
# Required: Broker URL (Redis example)
CELERY_BROKER_URL = "redis://localhost:6379/0"
# Required: Result backend URL (can be same Redis instance, different DB)
RESULT_BACKEND = "redis://localhost:6379/1"
# Optional: Serialize results as JSON (recommended for Superset)
CELERY_RESULT_SERIALIZER = "json"
CELERY_TASK_SERIALIZER = "json"
# Optional: Task execution limits
CELERYD_TASK_TIME_LIMIT = 60 * 60 # 1 hour hard limit
CELERYD_TASK_SOFT_TIME_LIMIT = 55 * 60 # 55 minutes soft limit
Ensure this file is in your PYTHONPATH or explicitly set via the SUPERSET_CONFIG_PATH environment variable before starting Superset or Celery workers.
Creating the Celery Application
Superset centralizes the Celery app instance in superset/tasks/celery_app.py. This file creates the shared Celery object that both the web app and workers import, ensuring consistent configuration.
The celery_app.py Implementation
# superset/tasks/celery_app.py
from celery import Celery
from superset import config
celery_app = Celery(
"superset",
broker=config.CELERY_BROKER_URL,
backend=config.RESULT_BACKEND,
include=["superset.tasks"] # Explicitly include task modules
)
# Automatically discover tasks in superset.tasks and submodules
celery_app.autodiscover_tasks(["superset.tasks"])
This pattern allows Superset to use the same configuration object (superset.config) for both the web application and background workers, preventing configuration drift.
Defining and Running Celery Tasks
Tasks are defined in Python modules within superset/tasks/. Superset ships with built-in tasks for caching, reporting, and SQL execution, but you can extend functionality by adding custom tasks.
Built-in Task Categories
- Caching:
cache_chart,warm_cache(pre-computes visualization data) - Email Reports:
schedule_email_report(sends dashboards/charts via email) - SQL Lab:
execute_sql_query(runs long-running queries asynchronously)
Custom Task Example
# superset/tasks/custom_tasks.py
from superset.tasks.celery_app import celery_app
@celery_app.task(bind=True, name="superset.tasks.custom_tasks.process_data")
def process_data(self, datasource_id, query_params):
"""
Custom long-running data processing task.
"""
try:
# Business logic here
result = f"Processed datasource {datasource_id}"
return {"status": "success", "result": result}
except Exception as exc:
# Retry logic
raise self.retry(exc=exc, countdown=60, max_retries=3)
Starting Celery Workers
Workers must be started separately from the Superset web server. Run this command from the directory containing your superset Python package:
celery -A superset.tasks.celery_app worker \
--loglevel=INFO \
--concurrency=4 \
--hostname=worker1@%h
Key parameters:
-A superset.tasks.celery_app: Points to the Celery app instance--concurrency=4: Number of parallel worker processes (typically match CPU cores)--hostname: Unique worker identifier for monitoring
Running Periodic Tasks with Celery Beat
For scheduled jobs (e.g., nightly cache warming), run the Celery beat scheduler alongside your workers:
celery -A superset.tasks.celery_app beat --loglevel=INFO
Define schedules in superset_config.py:
from celery.schedules import crontab
CELERY_BEAT_SCHEDULE = {
"warm_cache_every_morning": {
"task": "superset.tasks.cache.warm_cache",
"schedule": crontab(hour=6, minute=0),
"args": ("dashboard_slug",),
},
}
Common Pitfalls and Troubleshooting
| Issue | Root Cause | Solution |
|---|---|---|
| Workers fail to start | CELERY_BROKER_URL is undefined or points to an unreachable host. |
Verify the broker URL format (redis://, amqp://) and network connectivity. |
| Tasks stuck in PENDING | No running workers or queue mismatch. | Ensure workers are running and autodiscover_tasks includes your task modules. |
| "Task revoked" errors | Workers running stale code after deployment. | Restart all Celery workers to reload task definitions from superset/tasks/. |
| Memory leaks in long tasks | Workers never release memory after heavy tasks. | Set CELERYD_WORKER_MAX_TASKS_PER_CHILD = 1000 to force worker restart after N tasks. |
| Result timeout issues | Default visibility timeout shorter than task duration. | Increase CELERY_BROKER_TRANSPORT_OPTIONS = {"visibility_timeout": 43200} (12 hours) for Redis. |
Summary
- Apache Superset uses Celery to handle asynchronous operations like chart caching, email reports, and long-running SQL queries.
- The integration requires three components: a message broker (Redis/RabbitMQ), Celery workers (running
superset/tasks/celery_app.py), and a result backend to persist task outcomes. - Configuration is centralized in
superset_config.pyviaCELERY_BROKER_URLandRESULT_BACKEND. - Workers are started with
celery -A superset.tasks.celery_app workerand must be restarted after code changes to pick up new task definitions. - For scheduled jobs, run Celery Beat alongside workers and define schedules in
CELERY_BEAT_SCHEDULE.
Frequently Asked Questions
What is the role of the Celery broker in Apache Superset?
The Celery broker acts as a message queue that holds tasks submitted by the Superset web application until Celery workers can process them. Superset pushes tasks to the broker via celery.send_task() or apply_async(), using the URL defined in CELERY_BROKER_URL. Common broker choices include Redis, RabbitMQ, and Amazon SQS.
How do I configure Redis as both the broker and result backend for Superset Celery?
In your superset_config.py, set both URLs to point to your Redis instance, using different database numbers to avoid key collisions:
CELERY_BROKER_URL = "redis://localhost:6379/0"
RESULT_BACKEND = "redis://localhost:6379/1"
Ensure the Redis server is running and accessible from both the Superset web processes and the Celery workers before starting services.
Why are my Superset Celery tasks stuck in the PENDING state?
Tasks remain in PENDING when no Celery worker is running to consume them, or when workers are connected to a different broker queue than the one Superset is publishing to. Verify that workers are started with the correct app instance (celery -A superset.tasks.celery_app worker) and that CELERY_BROKER_URL matches between superset_config.py and the worker environment.
How do I run scheduled periodic tasks with Celery Beat in Superset?
First, define your schedule in superset_config.py using CELERY_BEAT_SCHEDULE and crontab expressions. Then, run the Celery Beat scheduler process alongside your workers:
celery -A superset.tasks.celery_app beat --loglevel=INFO
For production deployments, consider running Beat as a singleton service or using a distributed scheduler like RedBeat to prevent duplicate task execution across multiple Beat instances.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →