How the Celery Worker Updates Scan Progress State During APK Analysis in MobileAudit

The Celery worker updates scan progress by repeatedly calling task.update_state(state='STARTED', meta={...}) inside the analysis.analyze_apk function, passing a metadata dictionary that tracks current percentage, total steps, and human-readable status messages.

In the mobileaudit repository, a Django-based Android security scanner, long-running APK analysis tasks run asynchronously via Celery. The system implements a real-time progress tracking mechanism that keeps the frontend synchronized with the backend analysis state without blocking the main application thread.

How the Progress State Architecture Works

The implementation follows a producer-consumer pattern where the Celery worker acts as the state producer and the Django view acts as the state consumer. When a user initiates a scan, the system creates a Scan model instance, enqueues a Celery task, and stores the returned task ID for later retrieval. The worker then publishes incremental progress updates to Celery's result backend, which the frontend polls through a dedicated endpoint.

The Complete Execution Flow

1. Task Initialization and State Seeding

The process begins in app/worker/tasks.py where the task_create_scan function receives the scan ID and immediately seeds the initial state. This establishes the metadata structure that subsequent updates will follow.


# app/worker/tasks.py

@shared_task
def task_create_scan(scan):
    current_task.update_state(
        state='STARTED',
        meta={'current': 1, 'total': 100, 'status': 'In Progress'}
    )
    analysis.analyze_apk(current_task, scan)

Source: [app/worker/tasks.py](https://github.com/mpast/mobileaudit/blob/main/app/worker/tasks.py) lines 12-14

2. Incremental Progress Updates During Analysis

The core logic resides in app/analysis.py within the analyze_apk function. After every major analysis phase—such as APK info extraction, certificate parsing, VirusTotal lookups, and decompilation—the function performs a dual update: it persists progress to the database and pushes the state to Celery.


# app/analysis.py

scan.status = 'Getting info of apk'
scan.progress = 5
scan.save()
task.update_state(
    state='STARTED',
    meta={'current': scan.progress, 'total': 100, 'status': scan.status}
)

Source: [app/analysis.py](https://github.com/mpast/mobileaudit/blob/main/app/analysis.py) lines 60-66

This pattern repeats throughout the analysis lifecycle at key milestones:

  • Certificate extraction (lines 70-71)
  • VirusTotal lookup (lines 74-79)
  • Decompilation phase (lines 86-90)
  • Vulnerability scanning (lines 95-99)
  • Finalization (lines 101-107)

Each call to task.update_state() overwrites the previous metadata, ensuring the latest progress is always available through Celery's result backend.

3. Storing the Celery Task Reference

When the task is first dispatched in app/views.py, the system captures the Celery task ID and associates it with the Scan model. This linkage enables the frontend to query the correct task state later.


# app/views.py – create_scan()

task_id = task_create_scan.delay(scan.id)   # returns AsyncResult

scan.task = task_id.id
scan.save()

Source: [app/views.py](https://github.com/mpast/mobileaudit/blob/main/app/views.py) lines 81-84

4. Retrieving Current State for the Frontend

The scan_state view handles AJAX polling requests by instantiating an AsyncResult object with the stored task ID. It extracts the metadata dictionary (accessed via job.info) and returns it as JSON for the progress bar rendering.


# app/views.py – scan_state()

scan = Scan.objects.get(pk=id)
job = AsyncResult(scan.task)
data = job.info if job.info else job.result
return HttpResponse(json.dumps(data), content_type='application/json')

Source: [app/views.py](https://github.com/mpast/mobileaudit/blob/main/app/views.py) lines 16-24

Reusable Implementation Patterns

Enqueuing a Scan with Progress Tracking

When implementing similar functionality in your own Django projects, ensure you capture the AsyncResult ID immediately after calling delay():

from app.worker.tasks import task_create_scan

def start_scan(scan):
    # scan is a saved Scan model instance

    async_res = task_create_scan.delay(scan.id)
    scan.task = async_res.id
    scan.save()

Updating Progress in Long-Running Functions

For any multi-step analysis process, structure your updates to maintain both database consistency and Celery state synchronization:

def long_job(task, scan):
    for step, (status, progress) in enumerate([
        ('Opening APK', 5),
        ('Extracting certificates', 15),
        ('Decompiling', 30),
        ('Scanning for patterns', 70),
        ('Finishing', 100),
    ], start=1):
        # Persist to database for durability

        scan.status = status
        scan.progress = progress
        scan.save()
        
        # Push to Celery result backend for real-time access

        task.update_state(
            state='STARTED',
            meta={'current': progress, 'total': 100, 'status': status}
        )

Summary

  • State updates occur via task.update_state(): The Celery worker calls this method after each analysis phase in app/analysis.py, passing a meta dictionary with current, total, and status keys.
  • Dual persistence strategy: The system updates both the Django Scan model (for database durability) and the Celery task state (for real-time polling).
  • Task ID storage is critical: The create_scan view stores the Celery task ID in scan.task to enable later retrieval via AsyncResult.
  • Frontend polling via scan_state: The view reconstructs the task result object and returns the metadata as JSON, enabling live progress bars without WebSocket complexity.

Frequently Asked Questions

How does the frontend receive progress updates without WebSockets?

The frontend polls the scan_state endpoint via standard AJAX requests. This endpoint instantiates AsyncResult(scan.task) to access the latest metadata from Celery's result backend and returns it as JSON. While not as immediate as WebSockets, this approach requires no additional infrastructure beyond the standard Celery result backend.

What happens if the Celery worker crashes during analysis?

Because the Scan model fields (status and progress) are saved to the database at each step (via scan.save()), the system maintains durable state even if the worker fails. The frontend can detect stalled tasks by checking if the Celery state remains STARTED without recent metadata changes, while the database preserves the last known progress.

Why does the code use state='STARTED' for all progress updates instead of custom states?

According to the mobileaudit source code, the implementation uses state='STARTED' consistently because Celery treats this as an active task state that prevents premature task cleanup. The actual progress granularity is encoded in the meta dictionary, allowing the frontend to distinguish between different analysis phases while Celery continues to recognize the task as running.

Can this pattern work with Redis or RabbitMQ as the broker?

Yes. The update_state() method stores metadata in the configured result backend, not the broker. As long as the Celery app is configured with a result backend (such as django-db or Redis) in app/worker/celery.py, the AsyncResult retrieval in scan_state will function correctly regardless of whether you use RabbitMQ or Redis as the message broker.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →