# How to Deploy and Run a Gunicorn Flask Application: A Complete Production Guide

> Deploy your Flask app in production with Gunicorn. Learn to configure workers, binding, and the module callable pattern for a robust application.

- Repository: [Pallets/flask](https://github.com/pallets/flask)
- Tags: how-to-guide
- Published: 2026-02-15

---

**Deploy a Flask application in production by installing Gunicorn and running it with the `module:callable` pattern, using multiple workers and proper binding settings.**

Flask is a WSGI-compatible web framework, which means you should never use the built-in development server for production workloads. To properly deploy and run a Gunicorn Flask application, you need a dedicated WSGI server that manages worker processes and handles concurrent requests efficiently. According to the official Flask documentation in `docs/deploying/gunicorn.rst`, Gunicorn is the recommended choice because it is pure Python, easy to install, and supports multiple worker types including synchronous and asynchronous gevent workers.

## Prerequisites and Installation

Before deploying your Flask application, you must set up a proper Python environment and install the required dependencies.

### Setting Up the Virtual Environment

Create an isolated environment to avoid conflicts with system packages. This ensures that Gunicorn and Flask are installed only for your project.

```bash
python -m venv venv
source venv/bin/activate  # On Windows: venv\Scripts\activate

```

### Installing Flask and Gunicorn

Install both packages using pip. The Flask source code in [`src/flask/app.py`](https://github.com/pallets/flask/blob/main/src/flask/app.py) defines the core `Flask` class that Gunicorn will serve, while Gunicorn provides the WSGI HTTP server.

```bash
pip install flask gunicorn

```

## Application Structure for Gunicorn Deployment

Gunicorn requires a Python module that exposes a callable application object. Flask supports two primary patterns for this: the simple app variable and the application factory pattern.

### Using an Application Factory (create_app)

The **application factory** pattern is the recommended approach for production deployments. As shown in [`examples/tutorial/flaskr/__init__.py`](https://github.com/pallets/flask/blob/main/examples/tutorial/flaskr/__init__.py), a factory function named `create_app()` returns a configured Flask instance. This pattern allows you to pass configuration parameters and works reliably with Gunicorn workers that spawn new processes.

```python
def create_app():
    from flask import Flask
    app = Flask(__name__)
    
    @app.route("/")
    def index():
        return "Factory-based Flask app"
    
    return app

```

### Simple App Variable Approach

For smaller applications, you can create the Flask instance directly in your module. Save this as [`hello.py`](https://github.com/pallets/flask/blob/main/hello.py):

```python
from flask import Flask

app = Flask(__name__)

@app.route("/")
def index():
    return "Hello, Flask with Gunicorn!"

```

## Running Gunicorn with Flask

Once your application structure is ready, you launch Gunicorn using the **module:callable** syntax. This tells Gunicorn which Python module to import and which callable object to use as the WSGI application.

### The Module:Callable Pattern

The command structure follows the pattern `module_name:callable`. For a simple app variable, use the module filename (without .py) followed by the variable name:

```bash
gunicorn -w 4 'hello:app'

```

For an application factory, append parentheses to indicate Gunicorn should call the function:

```bash
gunicorn -w 4 'hello:create_app()'

```

### Configuring Worker Processes

The `-w` flag specifies the number of **worker processes**. According to the Flask deployment documentation, a common heuristic is to use `CPU * 2` workers. The default of 1 worker is insufficient for production traffic because it cannot handle concurrent requests.

```bash
gunicorn -w 4 'myapp:create_app()'

```

### Binding and Network Configuration

By default, Gunicorn binds to `127.0.0.1:8000`. In containerized environments or virtual machines, you must bind to all interfaces using `0.0.0.0` to accept external traffic:

```bash
gunicorn -w 4 -b 0.0.0.0 'myapp:create_app()'

```

For production deployments exposed to the internet, place a reverse proxy (such as nginx or Apache) in front of Gunicorn to handle TLS termination, static file serving, and request buffering.

## Production-Ready Gunicorn Configuration

Beyond basic startup, configure logging and worker types to match your workload characteristics.

### Access Logging

To view HTTP request logs on stdout (useful for containerized environments), enable the access logfile option with a hyphen:

```bash
gunicorn -w 4 --access-logfile=- 'myapp:create_app()'

```

### Async Workers for High Concurrency

For I/O-bound applications with high concurrency requirements, switch from the default **sync** workers to **gevent** workers using the `-k` flag:

```bash
gunicorn -w 4 -k gevent 'myapp:create_app()'

```

This configuration allows each worker process to handle multiple concurrent connections without blocking, significantly improving throughput for database-heavy or API workloads.

## Summary

- **Flask requires a WSGI server** in production; the built-in development server is insufficient for handling concurrent traffic.
- **Gunicorn** is the recommended WSGI server for Flask applications, supporting both simple app variables and application factories via the `module:callable` syntax.
- Use the **application factory pattern** (`create_app()`) for better configuration management and compatibility with Gunicorn's worker processes.
- Configure **worker count** using the `CPU * 2` heuristic and bind to `0.0.0.0` in containerized environments.
- Enable **access logging** and consider **gevent workers** for high-concurrency production workloads.

## Frequently Asked Questions

### What is the difference between Flask's development server and Gunicorn?

Flask's built-in development server, defined in [`src/flask/cli.py`](https://github.com/pallets/flask/blob/main/src/flask/cli.py), is designed for local debugging and automatically reloads code changes. It is single-threaded and not optimized for security or performance. **Gunicorn** is a production-grade WSGI HTTP server that manages multiple worker processes, handles concurrent connections efficiently, and integrates with reverse proxies for TLS termination.

### How many Gunicorn workers should I use for my Flask app?

According to the Flask deployment documentation in `docs/deploying/gunicorn.rst`, a common heuristic is to use **CPU count multiplied by 2** (e.g., `-w 4` on a dual-core machine). The default of 1 worker is insufficient for production because it cannot handle multiple simultaneous requests, leading to request queuing and increased latency.

### Can I use Gunicorn with Flask application factories?

Yes, Gunicorn fully supports Flask's application factory pattern. Instead of passing a module-level app instance, use the syntax `module:create_app()` (including the parentheses). This tells Gunicorn to call the factory function and use the returned Flask instance as the WSGI application, as demonstrated in [`examples/tutorial/flaskr/__init__.py`](https://github.com/pallets/flask/blob/main/examples/tutorial/flaskr/__init__.py).

### How do I enable logging when running Flask with Gunicorn?

To view HTTP access logs on stdout, which is essential for containerized environments and log aggregation systems, add the `--access-logfile=-` flag to your Gunicorn command. This configuration streams request logs to standard output, allowing you to monitor traffic and debug issues in real-time without writing to disk.