Learnings8 min read928 views

Scaling n8n: queue mode, workers and task runners

How main, workers, Redis and external task runners fit together in n8n queue mode, with routing, configuration responsibilities and validation checks.

Written byIdir Ouhab

On this page

What problem are we solving?

In a basic deployment, one n8n instance handles the editor, webhooks, workflow execution and triggers. As concurrency grows, those responsibilities can compete for resources. Queue mode separates production workflow execution from the main instance; Code nodes have their own task-runner requirements.

Cover image for Scaling n8n: queue mode, workers and task runners

The solution is to separate responsibilities:

ComponentWhat it does
MainServes the editor, manages the API and triggers (cron, polling)
WorkerExecutes the workflows (the heavy lifting)
Webhook ProcessorReceives incoming HTTP requests for webhooks
Task RunnerExecutes Code-node JavaScript or Python with isolation determined by its mode and configuration
RedisMessage queue that connects everything
PostgreSQLShared database
💡 Think of it like a restaurant: the Main is the maître d' who takes orders, Redis is the counter where order tickets are placed, the Worker is the chef, the Webhook Processor is the dedicated entrance for online delivery orders, and the Task Runners are the specialized prep cooks handling specific ingredients.

Visual architecture

text
Public traffic -> reverse proxy
  Editor, API and test webhooks -> Main
  Production and waiting webhooks -> Webhook processors

Main / webhook processors -> Redis job queue -> Workers
Main / webhook processors / workers <-> PostgreSQL

Worker's Code node <-> broker in worker <-> external runner

The worker retrieves the workflow from the database and stores execution results there. The runner executes Code-node tasks and returns their results through the broker. A main instance that executes manual workflows also needs a runner for those Code nodes. Queue architecture.

Before implementation

Scope of this guide: this is an architecture walkthrough. The former Compose example mixed n8n 1.71.3 with a runner-sidecar arrangement whose documented minimum is 1.111.0, and used incorrect connection settings. It has been removed. This article does not provide a deployment that has been tested end to end.

That minimum is a compatibility boundary, not a recommendation to deploy an old release. For implementation, choose a supported release, pin the n8n and runner images to the same version, and follow the official external-runner setup and queue-mode guide. Main instances, webhook processors and workers must also use matching n8n versions.

The deployment still needs its own secret management, persistent storage, network access rules, HTTPS proxy and recovery tests. The connections below explain what to configure; they are not a complete Compose file.

Understand each component

PostgreSQL

Stores workflows, credentials and execution data. Main instances, webhook processors and workers need database access. Use PostgreSQL for this distributed architecture; n8n does not support a distributed queue setup backed by SQLite.

Redis

Carries queued execution IDs and coordination messages. It is part of the execution path, so treating it as a disposable cache is a mistake. Use maxmemory-policy=noeviction for job storage: Redis then rejects writes that need more memory instead of evicting existing keys when the memory limit is reached. Plan persistence and memory headroom, and monitor memory usage. This policy does not prevent memory exhaustion or make failed jobs impossible. Redis eviction behavior.

Main and webhook processors

The main instance serves the editor and API and manages triggers. Queue mode sends production executions to workers. Manual executions run on main unless OFFLOAD_MANUAL_EXECUTIONS_TO_WORKERS=true sends them to workers too.

Optional webhook processors receive production webhook traffic and pass execution work to the queue. Separating them can reduce pressure on the editor; shared CPU, memory, Redis and database capacity can still become bottlenecks.

Workers

Workers consume queued executions and perform the workflow. The CLI flag n8n worker --concurrency=10 sets a ten-job limit unless N8N_CONCURRENCY_PRODUCTION_LIMIT overrides it. In n8n 2.0.0, an environment value other than -1 takes precedence over the flag. Ten is the flag's default, not a sizing recommendation. Runner task concurrency is a separate limit. Version 2.0.0 worker implementation.

External task runners

Each worker needs a companion runner container for Code-node tasks. Main needs one too when it performs manual executions; a webhook-only receiver does not need a runner merely because it receives HTTP requests.

Task runners are enabled by default from n8n 2.0. Native Python requires external mode. JavaScript can use internal mode, but n8n recommends external mode and additional hardening for production or sensitive data. Container separation is one layer: permissions, mounted files, exposed secrets and resource limits still matter. It does not guarantee that hostile code or an out-of-memory event cannot affect other services.

Route traffic to the right process

For a shared public domain and dedicated webhook processors, the documented default paths are:

RequestDestination
/webhook/*Webhook processor pool
/webhook-waiting/*Webhook processor pool
/webhook-test/*Main instance
Editor, API and other pathsMain instance

Update these rules if you customize endpoint names. Host ports depend on your network and proxy arrangement; 5680 is not a required webhook port. Match the public webhook URL and proxy configuration to the actual deployment. Official load-balancer guidance.

What to verify in a separate environment

A running container does not prove that a workflow works. Test editor access, a saved production workflow, production and test webhooks, a waiting webhook resumed by an external callback, manual execution, and the JavaScript or Python Code nodes you actually use. Verify the expected worker handles execution and that results persist.

Then test restarts and recovery with representative data and a checked backup. If a runner fails, inspect the broker address, network reachability, shared authentication token, matching image versions and logs. An unhealthy runner does not identify a single cause. These are validation steps for an implementation; they have not been executed for a deployment supplied by this article.

Configuration responsibilities

SettingWhere it belongsMeaning
EXECUTIONS_MODE=queueMain, workers and webhook processorsQueue-based production execution
DB_TYPE=postgresdb and DB_POSTGRESDB_*n8n processesConnection to the shared database
QUEUE_BULL_REDIS_HOST, QUEUE_BULL_REDIS_PORT, authentication settingsn8n processesConnection to the same Redis queue
N8N_ENCRYPTION_KEYMain, workers and webhook processorsSame securely stored encryption key for stored credentials
OFFLOAD_MANUAL_EXECUTIONS_TO_WORKERS=trueMainSend manual executions to workers
N8N_RUNNERS_MODE=externalEach n8n instance executing Code tasksUse external runners
N8N_RUNNERS_BROKER_LISTEN_ADDRESS=0.0.0.0n8n task brokerAccept runner connections over the private container network
N8N_RUNNERS_TASK_BROKER_URIRunner containerAddress of its n8n broker, for example http://n8n-worker:5679
N8N_RUNNERS_AUTH_TOKENBroker and its runnerSame securely generated authentication token on both sides
WEBHOOK_URLn8n configurationPublic webhook base URL

The runner connects to the broker in n8n. Keep broker port 5679 private; it is not the public editor port. The encryption key and runner authentication token serve different purposes.

For compatible 1.x releases, enable runners explicitly with N8N_RUNNERS_ENABLED=true. From 2.0, runners are enabled by default and that flag is no longer needed. Check the selected release's runner settings, including Python and dependency requirements, in the official guide.

When do you need this?

Use queue mode when you need to separate production execution from the editor or distribute work across workers. Add dedicated webhook processors when webhook intake needs to scale separately.

There is no useful universal threshold based only on executions per day. A short HTTP request and a workflow processing large files have very different costs. For sizing, test representative workflows and bursts, then measure queue wait time, execution duration, memory, CPU and database load before adding workers or raising concurrency. The old 4 GB / 2 vCPU minimum and daily-volume bands were not validated capacity limits.

Resources

About the author

Idir Ouhab

AI Deployment Engineer at OpenAI, trainer and host of Prompt&Play. I write about what I learn taking AI into production.

Your next step

Putting this into production?

Review the architecture, integrations, and risks of your AI system before the next step.

Share

Topics

  • n8n
  • queue mode
  • docker
  • workflow automation
  • scaling
  • automation tools
  • cloud
  • containers
  • performance optimization
Always active

Remembers your language and cookie choices. A separate session cookie keeps administrators signed in. These are not used for advertising.

Your choices are valid for 180 days in this browser. Optional purposes start switched off. If browser storage is unavailable, your choice lasts for this page only.

You can withdraw permission here at any time. If optional content has already loaded, the page reloads to stop it; unsent form changes may be lost.

How cookies and storage are used