This document describes the current stable version of Celery (5.7). For development docs, go here.
Periodic Tasks¶
Introduction¶
celery beat is a scheduler; It kicks off tasks at regular intervals, that are then executed by available worker nodes in the cluster.
By default the entries are taken from the beat_schedule setting,
but custom stores can also be used, like storing the entries in a SQL database.
You have to ensure only a single scheduler is running for a schedule at a time, otherwise you’d end up with duplicate tasks. Using a centralized approach means the schedule doesn’t have to be synchronized, and the service can operate without using locks.
Time Zones¶
The periodic task schedules uses the UTC time zone by default,
but you can change the time zone used using the timezone
setting.
An example time zone could be Europe/London:
timezone = 'Europe/London'
This setting must be added to your app, either by configuring it directly
using (app.conf.timezone = 'Europe/London'), or by adding
it to your configuration module if you have set one up using
app.config_from_object. See Configuration for
more information about configuration options.
The default scheduler (storing the schedule in the celerybeat-schedule
file) will automatically detect that the time zone has changed, and so will
reset the schedule itself, but other schedulers may not be so smart (e.g., the
Django database scheduler, see below) and in that case you’ll have to reset the
schedule manually.
Django Users
Celery recommends and is compatible with the USE_TZ setting introduced
in Django 1.4.
For Django users the time zone specified in the TIME_ZONE setting
will be used, or you can specify a custom time zone for Celery alone
by using the timezone setting.
The database scheduler won’t reset when timezone related settings change, so you must do this manually:
$ python manage.py shell
>>> from djcelery.models import PeriodicTask
>>> PeriodicTask.objects.update(last_run_at=None)
Django-Celery only supports Celery 4.0 and below, for Celery 4.0 and above, do as follow:
$ python manage.py shell
>>> from django_celery_beat.models import PeriodicTask
>>> PeriodicTask.objects.update(last_run_at=None)
Entries¶
To call a task periodically you have to add an entry to the beat schedule list.
from celery import Celery
from celery.schedules import crontab
app = Celery()
@app.on_after_configure.connect
def setup_periodic_tasks(sender: Celery, **kwargs):
# Calls test('hello') every 10 seconds.
sender.add_periodic_task(10.0, test.s('hello'), name='add every 10')
# Calls test('hello') every 30 seconds.
# It uses the same signature of previous task, an explicit name is
# defined to avoid this task replacing the previous one defined.
sender.add_periodic_task(30.0, test.s('hello'), name='add every 30')
# Calls test('world') every 30 seconds
sender.add_periodic_task(30.0, test.s('world'), expires=10)
# Executes every Monday morning at 7:30 a.m.
sender.add_periodic_task(
crontab(hour=7, minute=30, day_of_week=1),
test.s('Happy Mondays!'),
)
@app.task
def test(arg):
print(arg)
@app.task
def add(x, y):
z = x + y
print(z)
Setting these up from within the on_after_configure handler means
that we’ll not evaluate the app at module level when using test.s(). Note that
on_after_configure is sent after the app is set up, so tasks outside the
module where the app is declared (e.g. in a tasks.py file located by
celery.Celery.autodiscover_tasks()) must use a later signal, such as
on_after_finalize.
The add_periodic_task() function will add the entry to the
beat_schedule setting behind the scenes, and the same setting
can also be used to set up periodic tasks manually:
Example: Run the tasks.add task every 30 seconds.
app.conf.beat_schedule = {
'add-every-30-seconds': {
'task': 'tasks.add',
'schedule': 30.0,
'args': (16, 16)
},
}
app.conf.timezone = 'UTC'
Note
If you’re wondering where these settings should go then please see Configuration. You can either set these options on your app directly or you can keep a separate module for configuration.
If you want to use a single item tuple for args, don’t forget that the constructor is a comma, and not a pair of parentheses.
Using a timedelta for the schedule means the task will
be sent in 30 second intervals (the first task will be sent 30 seconds
after celery beat starts, and then every 30 seconds
after the last run).
A Crontab like schedule also exists, see the section on Crontab schedules.
Like with cron, the tasks may overlap if the first task doesn’t complete before the next. If that’s a concern you should use a locking strategy to ensure only one instance can run at a time (see for example Ensuring a task is only executed one at a time).
Scheduling groups and other workflows¶
beat schedules a single task by name, so you can’t pass a
group, chain, or chord signature directly to
add_periodic_task() (or a beat_schedule entry) – an entry only
stores a task name with its arguments, not a workflow.
To run a workflow periodically, wrap it in a regular task and schedule that task:
from celery import Celery, group
app = Celery()
@app.task
def add(x, y):
return x + y
@app.task
def run_add_group():
group(add.s(i, i) for i in range(10)).apply_async()
@app.on_after_configure.connect
def setup_periodic_tasks(sender: Celery, **kwargs):
sender.add_periodic_task(30.0, run_add_group.s(), name='add group every 30')
The wrapper only dispatches the workflow with apply_async() and returns. Don’t
call get() on the result inside the task to wait for it to finish: blocking on a
result from within a task ties up a worker process and is discouraged (see
Avoid launching synchronous subtasks). The group’s tasks run independently on the workers,
so the initiating task can return immediately.
Available Fields¶
task
The name of the task to execute.
Task names are described in the Names section of the User Guide. Note that this is not the import path of the task, even though the default naming pattern is built like it is.
schedule
args
Positional arguments (
listortuple).kwargs
Keyword arguments (
dict).options
Execution options (
dict).This can be any argument supported by
apply_async()– exchange, routing_key, expires, and so on.relative
If relative is true
timedeltaschedules are scheduled “by the clock.” This means the frequency is rounded to the nearest second, minute, hour or day depending on the period of thetimedelta.By default relative is false, the frequency isn’t rounded and will be relative to the time when celery beat was started.
Crontab schedules¶
If you want more control over when the task is executed, for
example, a particular time of day or day of the week, you can use
the crontab schedule type:
from celery.schedules import crontab
app.conf.beat_schedule = {
# Executes every Monday morning at 7:30 a.m.
'add-every-monday-morning': {
'task': 'tasks.add',
'schedule': crontab(hour=7, minute=30, day_of_week=1),
'args': (16, 16),
},
}
The syntax of these Crontab expressions are very flexible.
Some examples:
Example |
Meaning |
|
Execute every minute. |
|
Execute daily at midnight. |
|
Execute every three hours: midnight, 3am, 6am, 9am, noon, 3pm, 6pm, 9pm. |
|
Same as previous. |
|
Execute every 15 minutes. |
|
Execute every minute (!) at Sundays. |
|
Same as previous. |
|
Execute every ten minutes, but only between 3-4 am, 5-6 pm, and 10-11 pm on Thursdays or Fridays. |
|
Execute every even hour, and every hour divisible by three. This means: at every hour except: 1am, 5am, 7am, 11am, 1pm, 5pm, 7pm, 11pm |
|
Execute hour divisible by 5. This means that it is triggered at 3pm, not 5pm (since 3pm equals the 24-hour clock value of “15”, which is divisible by 5). |
|
Execute every hour divisible by 3, and every hour during office hours (8am-5pm). |
|
Execute on the second day of every month. |
|
Execute on every even numbered day. |
|
Execute on the first and third weeks of the month. |
|
Execute on the eleventh of May every year. |
|
Execute every day on the first month of every quarter. |
See celery.schedules.crontab for more documentation.
Solar schedules¶
If you have a task that should be executed according to sunrise,
sunset, dawn or dusk, you can use the
solar schedule type.
Solar schedules require the https://pypi.org/project/ephem/ library, so
to use them you must install Celery with the solar extra:
$ pip install celery[solar]
Example:
from celery.schedules import solar
app.conf.beat_schedule = {
# Executes at sunset in Melbourne
'add-at-melbourne-sunset': {
'task': 'tasks.add',
'schedule': solar('sunset', -37.81753, 144.96715),
'args': (16, 16),
},
}
The arguments are simply: solar(event, latitude, longitude)
Be sure to use the correct sign for latitude and longitude:
Sign |
Argument |
Meaning |
|
|
North |
|
|
South |
|
|
East |
|
|
West |
Possible event types are:
Event |
Meaning |
|
Execute at the moment after which the sky is no longer completely dark. This is when the sun is 18 degrees below the horizon. |
|
Execute when there’s enough sunlight for the horizon and some objects to be distinguishable; formally, when the sun is 12 degrees below the horizon. |
|
Execute when there’s enough light for objects to be distinguishable so that outdoor activities can commence; formally, when the Sun is 6 degrees below the horizon. |
|
Execute when the upper edge of the sun appears over the eastern horizon in the morning. |
|
Execute when the sun is highest above the horizon on that day. |
|
Execute when the trailing edge of the sun disappears over the western horizon in the evening. |
|
Execute at the end of civil twilight, when objects are still distinguishable and some stars and planets are visible. Formally, when the sun is 6 degrees below the horizon. |
|
Execute when the sun is 12 degrees below the horizon. Objects are no longer distinguishable, and the horizon is no longer visible to the naked eye. |
|
Execute at the moment after which the sky becomes completely dark; formally, when the sun is 18 degrees below the horizon. |
All solar events are calculated using UTC, and are therefore unaffected by your timezone setting.
In polar regions, the sun may not rise or set every day. The scheduler
is able to handle these cases (i.e., a sunrise event won’t run on a day
when the sun doesn’t rise). The one exception is solar_noon, which is
formally defined as the moment the sun transits the celestial meridian,
and will occur every day even if the sun is below the horizon.
Twilight is defined as the period between dawn and sunrise; and between sunset and dusk. You can schedule an event according to “twilight” depending on your definition of twilight (civil, nautical, or astronomical), and whether you want the event to take place at the beginning or end of twilight, using the appropriate event from the list above.
See celery.schedules.solar for more documentation.
Starting the Scheduler¶
To start the celery beat service:
$ celery -A proj beat
You can also embed beat inside the worker by enabling the
workers -B option, this is convenient if you’ll
never run more than one worker node, but it’s not commonly used and for that
reason isn’t recommended for production use:
$ celery -A proj worker -B
Beat needs to store the last run times of the tasks in a local database file (named celerybeat-schedule by default), so it needs access to write in the current directory, or alternatively you can specify a custom location for this file:
$ celery -A proj beat -s /home/celery/var/run/celerybeat-schedule
Note
To daemonize beat see Daemonization.
Health checks¶
Added in version 5.7.
By default there’s no way to check whether a running beat process is
still healthy. If you enable the
beat_enable_remote_control setting – or pass
--enable-remote-control
– beat joins the same remote-control exchange the workers use (as a
node named celerybeat@hostname) and answers
celery inspect ping:
$ celery -A proj inspect ping -t 5 -d celerybeat@$(hostname)
-> celerybeat@example.com: OK
pong
The command exits non-zero when nobody replies within the timeout.
Give --timeout room to spare: it
defaults to one second, and the probe has to establish a broker
connection of its own before it can ask anything.
Beat’s reply is the same {'ok': 'pong'} a worker sends, so nothing
downstream has to special-case it.
What a reply means¶
The control node runs in its own thread, so on its own a reply would
only prove that the process is alive and reaching the broker – not
that the schedule is advancing. To close that gap, beat stops answering
once the scheduler has not completed a pass for
beat_remote_control_max_tick_age seconds, which defaults to
twice the interval the scheduler settled on. A wedged scheduler
therefore fails the probe rather than passing it.
Silence is deliberate: celery inspect exits non-zero only when no node replies, so an error reply would leave the exit status at zero and the probe green. Beat logs a warning each time it declines, so the reason is visible in its own output.
Set beat_remote_control_max_tick_age to 0 to answer
regardless of tick age.
Choosing a probe¶
Only a liveness probe actually recovers a wedged beat: beat serves
no traffic and sits behind no Service, so marking a pod NotReady
removes nothing and starts no remediation. A readiness probe on beat
buys you visibility in kubectl get pods and gating for Deployment
rollouts – useful, but it will not restart anything.
The catch is that this check travels over the broker, so it fails whenever the broker is unreachable – during an ordinary broker restart, and for every beat pod at once. Restarting beat does not fix a broker outage, and a beat that restarts re-reads its schedule. Wired carelessly, a short blip becomes a cluster-wide restart storm.
So use a liveness probe, but give it a failureThreshold that rides
out a broker restart and an initialDelaySeconds long enough for the
control node to have connected, or the first probe kills a healthy pod:
livenessProbe:
exec:
command:
- /bin/sh
- -c
- celery -A proj inspect ping -t 10 -d celerybeat@$(hostname)
initialDelaySeconds: 60
periodSeconds: 60
failureThreshold: 5
With those numbers a wedged beat is restarted about fifteen minutes
after its last tick: ten waiting for
beat_remote_control_max_tick_age to elapse, which defaults
to twice the default scheduler’s five minute loop interval, then five
more for the probe to fail often enough. Lower the setting if that is
slower than you want to find out. A broker outage, by contrast, only
costs you a restart if it outlasts those last five minutes.
Add a readiness probe as well if you want the state surfaced in
kubectl and rollouts gated on beat coming up:
readinessProbe:
exec:
command:
- /bin/sh
- -c
- celery -A proj inspect ping -t 5 -d celerybeat@$(hostname)
initialDelaySeconds: 30
periodSeconds: 60
failureThreshold: 3
Comparison with a heartbeat file¶
A common alternative is to have the scheduler touch a file on every tick and let the probe check its age, usually through a custom scheduler subclass. That has no broker dependency at all, so a broker outage cannot take it down, and it proves the tick loop directly.
The trade-off is reach. A file is only visible from inside the
container, so it answers “is this beat alive” and nothing more.
Remote control answers the same question from anywhere that can talk to
the broker, which is what lets celery status and centralised
monitoring see beat alongside the workers. Pick the file if you only
need a local probe; pick remote control if you want beat visible in the
same place as everything else.
Interactions worth knowing¶
Beat shows up as a node in destination-less celery inspect ping and celery status output once this is enabled. If you count nodes in those replies, the count changes.
For the same reason, a destination-less broadcast with a reply limit –
app.control.ping(limit=1), say – may now come back with beat instead of a worker.Beat implements
pingand nothing else. Other remote-control commands are ignored silently,shutdownincluded:celery control shutdownstops your workers and leaves beat running.A beat scheduler embedded in a worker (
-B) never starts a control node, since the worker already answers for that process.Node names have to be unique. Two beats resolving to the same name – the same hostname, or containers on
network_mode: host– share one pidbox queue, so only one of them ever answers. Give each its own--hostnameif that can happen.Remote control needs fanout exchanges, so it is available on the RabbitMQ (AMQP) and Redis transports. On a transport without them beat logs one warning at startup and carries on without a control node.
If the control node gives up on the broker, it does not come back, so the probe fails until beat is restarted. That applies to the connection it makes at startup and to re-establishing one that dropped. See
beat_enable_remote_control.
Using custom scheduler classes¶
Custom scheduler classes can be specified on the command-line (the
--scheduler argument).
The default scheduler is the celery.beat.PersistentScheduler,
that simply keeps track of the last run times in a local shelve
database file.
There’s also the https://pypi.org/project/django-celery-beat/ extension that stores the schedule in the Django database, and presents a convenient admin interface to manage periodic tasks at runtime.
To install and use this extension:
Use pip to install the package:
$ pip install django-celery-beat
Add the
django_celery_beatmodule toINSTALLED_APPSin your Django project’settings.py:INSTALLED_APPS = ( ..., 'django_celery_beat', )
Note that there is no dash in the module name, only underscores.
Apply Django database migrations so that the necessary tables are created:
$ python manage.py migrate
Start the celery beat service using the
django_celery_beat.schedulers:DatabaseSchedulerscheduler:$ celery -A proj beat -l INFO --scheduler django_celery_beat.schedulers:DatabaseScheduler
Note: You may also add this as the
beat_schedulersetting directly.Visit the Django-Admin interface to set up some periodic tasks.