DEV Community

Cover image for FastAPI, Celery, and Redis in Docker: A Complete Guide to Background Jobs Done Right
Akeem O. Salau
Akeem O. Salau

Posted on Originally published at zuqolab.com

FastAPI, Celery, and Redis in Docker: A Complete Guide to Background Jobs Done Right

If your API ever sends an email, generates a PDF, calls a slow third party service, or processes a file upload, you have almost certainly made your users wait for something they did not need to watch happen live. That wait is not a performance problem you fix with faster code. It is an architecture problem, and the fix is a background task queue.

This guide walks through building that queue properly with FastAPI, Celery, and Redis, all running in Docker. Not the toy version where a single task prints "hello" to the console, but the version with retries, idempotency, scheduled jobs, and the configuration details that separate a demo from something you can actually deploy.

Why Your API Should Never Do Slow Work In The Request Cycle

A request handler that blocks on a slow operation ties up a worker process for the entire duration of that operation. Do this often enough under load and your API becomes unresponsive even for requests that have nothing to do with the slow task. Users refresh, retry, and submit duplicate requests, which makes the problem worse, not better.

The fix is to separate "accept the work" from "do the work." Your API accepts the request, records what needs to happen, hands it off to something else, and responds immediately. That something else is a task queue. In the Python world, that almost always means Celery.

The Three Moving Parts

Three technologies do three distinct jobs here, and conflating them is where most confusion starts.

FastAPI is your web layer. It accepts HTTP requests, validates input, and returns responses quickly. It should never block on slow work.

Celery is your task execution engine. It runs Python functions outside the request response cycle, on separate worker processes that can live on separate machines entirely.

Redis plays two roles at once. As the broker, it is the queue that holds tasks waiting to be picked up by a worker. As the result backend, it stores the outcome of a task so something else can check on it later. You can use different technologies for each role (RabbitMQ as a broker, a database as a result backend) but Redis doing both is the simplest setup that still holds up in production.

Docker ties all of this together so that your API, your worker, your scheduler, and Redis itself run as separate, isolated, reproducible services that start with one command.

Structuring The Project The Right Way

A background task is still part of your application's business logic. It deserves the same separation of concerns as everything else. Here is a layout that keeps the Celery task thin and puts the actual logic where it belongs.

app/
  core/
    config.py
    celery_app.py
    database.py
  jobs/
    models.py
    schemas.py
    repository.py
    service.py
    tasks.py
    routes.py
  main.py
docker/
  Dockerfile
docker-compose.yml
.env

Enter fullscreen mode Exit fullscreen mode

The repository only queries and persists data, without committing. A narrow service applies business rules, also without committing. One orchestrating service per use case owns the transaction, commits it, and is the only place that triggers a Celery task. Routes just parse input, call a single service method, and translate exceptions into HTTP responses. This matters more than it looks like it does, because the most common Celery bug in production is dispatching a task before the database transaction that created its data has actually committed.

Configuring Celery For Production

Your Celery app needs more than a broker URL to behave well under real traffic.

# app/core/celery_app.py
from celery import Celery
from app.core.config import settings

celery_app = Celery(
    "worker",
    broker=settings.REDIS_URL,
    backend=settings.REDIS_URL,
    include=["app.jobs.tasks"],
)

celery_app.conf.update(
    task_serializer="json",
    result_serializer="json",
    accept_content=["json"],
    timezone="UTC",
    enable_utc=True,
    task_acks_late=True,
    worker_prefetch_multiplier=1,
    task_reject_on_worker_lost=True,
    result_expires=3600,
    task_routes={
        "app.jobs.tasks.generate_report": {"queue": "reports"},
        "app.jobs.tasks.send_notification_email": {"queue": "emails"},
    },
)

Enter fullscreen mode Exit fullscreen mode

A few of these settings carry real weight. task_serializerset to JSON instead of the default pickle avoids deserializing arbitrary objects on the worker side, which is both a security risk and a source of subtle bugs when your models change shape. task_acks_late means a task is only marked complete after it finishes successfully, not the moment a worker picks it up, so a crashed worker does not silently lose the task. worker_prefetch_multiplierset to 1 stops a single worker from grabbing a large batch of tasks upfront and starving other workers while it slowly works through them, which matters a great deal once your tasks vary in duration.

task_routessends different kinds of work to different queues. This is the detail most tutorials skip entirely, and it is the one that saves you the day a report generation task that takes two minutes stops a time sensitive notification email from going out for two minutes as well.

Writing Tasks That Do Not Fall Apart Under Load

A Celery task will run more than once eventually. A worker will crash mid task. A network call will time out and trigger a retry. Your task needs to survive running twice without creating duplicate side effects, a property called idempotency.

# app/jobs/tasks.py
from celery.exceptions import SoftTimeLimitExceeded
from app.core.celery_app import celery_app
from app.core.database import SessionLocal
from app.jobs.repository import JobRepository


@celery_app.task(
    bind=True,
    max_retries=5,
    soft_time_limit=50,
    time_limit=60,
)
def generate_report(self, job_id: int):
    db = SessionLocal()
    repo = JobRepository(db)

    try:
        job = repo.get(job_id)

        if job.status == "completed":
            return

        repo.update_status(job_id, "processing")
        db.commit()

        result_path = build_report_file(job_id)

        repo.update_status(job_id, "completed", result_path=result_path)
        db.commit()

    except SoftTimeLimitExceeded:
        repo.update_status(job_id, "failed")
        db.commit()

    except Exception as exc:
        db.rollback()
        countdown = 2 ** self.request.retries * 10
        raise self.retry(exc=exc, countdown=countdown)

    finally:
        db.close()
Enter fullscreen mode Exit fullscreen mode

Notice the task receives job_id, a plain integer, not a database model object. Passing ORM objects into a task is a common mistake because Celery serializes the arguments, and the data could be stale by the time the task actually runs minutes later. Fetching a fresh row inside the task guarantees you are working with current data.

The status check at the top is the idempotency guard. If this task runs twice because of a redelivered message, the second run sees the job already marked complete and exits without redoing the work. The retry uses exponential backoff, so a failing external dependency gets progressively more breathing room instead of being hammered with five retries in rapid succession. soft_time_limitand time_limitmake sure a task that hangs does not tie up a worker forever.

Wiring Celery Into Your FastAPI Routes

The route never calls the Celery task directly. The orchestrating service does, and only after its own transaction has committed.

# app/jobs/service.py
from sqlalchemy.orm import Session
from app.jobs.repository import JobRepository
from app.jobs.schemas import JobCreate
from app.jobs.tasks import generate_report


class CreateReportJobService:
    def __init__(self, db: Session):
        self.db = db
        self.repo = JobRepository(db)

    def execute(self, payload: JobCreate):
        job = self.repo.create(payload)
        self.db.commit()
        self.db.refresh(job)

        generate_report.delay(job.id)

        return job
Enter fullscreen mode Exit fullscreen mode
# app/jobs/routes.py
from fastapi import APIRouter, Depends, HTTPException
from sqlalchemy.orm import Session
from app.core.database import get_db
from app.jobs.schemas import JobCreate, JobRead
from app.jobs.service import CreateReportJobService

router = APIRouter(prefix="/jobs", tags=["jobs"])


@router.post("", response_model=JobRead, status_code=201)
def create_job(payload: JobCreate, db: Session = Depends(get_db)):
    service = CreateReportJobService(db)
    job = service.execute(payload)
    return job

Enter fullscreen mode Exit fullscreen mode

The ordering here is deliberate. Commit first, dispatch second. If you call generate_report.delay() before db.commit(), there is a real chance the worker picks up the task and queries the database for a row that is not there yet, because the transaction holding it has not been written to disk. This single ordering mistake accounts for a surprising share of "it works locally but fails randomly in production" Celery bugs.

Putting It All Together With Docker Compose

Each piece of this system runs as its own container, sharing the same codebase but running different commands.

# docker-compose.yml
services:
  web:
    build: .
    command: uvicorn app.main:app --host 0.0.0.0 --port 8000
    ports:
      - "8000:8000"
    env_file: .env
    depends_on:
      - redis
      - db

  worker:
    build: .
    command: celery -A app.core.celery_app.celery_app worker --loglevel=info -Q reports,emails --concurrency=4
    env_file: .env
    depends_on:
      - redis
      - db

  beat:
    build: .
    command: celery -A app.core.celery_app.celery_app beat --loglevel=info
    env_file: .env
    depends_on:
      - redis

  flower:
    image: mher/flower:2.0
    command: celery --broker=redis://redis:6379/0 flower --port=5555
    ports:
      - "5555:5555"
    depends_on:
      - redis

  redis:
    image: redis:7-alpine
    command: ["redis-server", "--appendonly", "yes"]
    volumes:
      - redis_data:/data

  db:
    image: postgres:16-alpine
    environment:
      POSTGRES_DB: app
      POSTGRES_USER: app
      POSTGRES_PASSWORD: app
    volumes:
      - pg_data:/var/lib/postgresql/data

volumes:
  redis_data:
  pg_data:

Enter fullscreen mode Exit fullscreen mode

The -Q reports,emails flag on the worker command tells it which queues to consume from, matching the routes defined in your Celery config earlier. In a larger deployment you would typically run separate worker containers for separate queues, so a flood of report generation jobs can never delay a time sensitive email. Redis is configured with append only file persistence enabled, so queued tasks survive a Redis restart instead of vanishing.

Scheduling Recurring Work With Celery Beat

Celery Beat handles work that needs to run on a schedule rather than in response to a request, things like nightly cleanup jobs or daily digest emails.

from celery.schedules import crontab
from app.core.celery_app import celery_app

celery_app.conf.beat_schedule = {
    "cleanup-expired-jobs-daily": {
        "task": "app.jobs.tasks.cleanup_expired_jobs",
        "schedule": crontab(hour=2, minute=0),
    },
}
Enter fullscreen mode Exit fullscreen mode

Run exactly one Beat process per deployment. Beat's entire job is to push scheduled tasks onto the queue at the right time, and running two instances means every scheduled task gets duplicated, which is a quiet, easy to miss bug that only shows up when someone notices they received the same digest email twice.

Watching Your Workers With Flower

Flower gives you a dashboard over your running workers, queued tasks, and task history, and it is already wired into the Docker Compose file above on port 5555. Once your stack is running, open it in a browser to see tasks move through pending, started, and succeeded states in real time, along with retry counts and failure traces. For anything beyond a side project, this visibility is not optional. The first time a queue silently backs up because a worker crashed and did not restart, you will want to have been watching.

The Production Checklist Most Tutorials Skip

A handful of habits separate a Celery setup that survives real traffic from one that quietly degrades.

Route different workloads to different queues so a slow task type cannot starve a fast one. Keep worker_prefetch_multiplierlow so tasks distribute evenly across workers instead of piling onto whichever worker grabbed them first. Set ignore_result=True on tasks where you never check the result, since storing results you never read wastes Redis memory for no benefit. Never call one task and wait synchronously for another task's result inside it, since that turns your asyncsystem into a slow, fragile synchronous one. Always pass plain identifiers into tasks rather than ORM objects, and always refetch fresh data inside the task itself. Set both a soft and a hard time limit on every task so a hung task cannot occupy a worker indefinitely. And restart your worker processes under a supervisor, so a crash is a blip instead of an outage.

None of these are exotic. They are the difference between a Celery setup that works in a demo and one that keeps working at two in the morning when nobody is watching it.

Wrapping Up

FastAPI, Celery, and Redis solve a problem that every growing API eventually runs into: work that takes too long to do while someone is waiting. The pattern is straightforward once it is laid out clearly. Accept the request fast, commit what you need to remember, hand the slow part to a worker, and build that worker to survive failure instead of assuming it will not happen. Docker makes the whole stack reproducible, Flower makes it observable, and a little discipline around idempotency and transaction ordering makes it something you can actually trust in production.

Top comments (0)