Quartz.NETQuartz.NET
Home
Features
Blog
Discussions
NuGet
GitHub
Home
Features
Blog
Discussions
NuGet
GitHub
  • Getting Started

    • Overview
    • Quartz 4 Quick Start
    • Tutorial
      • Using Quartz
      • Jobs And Triggers
      • More About Jobs & JobDetails
      • Job Data
      • More About Triggers
      • Querying Jobs and Triggers
      • Simple Triggers
      • Cron Triggers
      • RecurrenceTrigger
      • Time and TimeProvider
      • Trigger and Job Listeners
      • Scheduler Listeners
      • Job Execution Middleware
      • Job Stores
      • Configuration, Resource Usage and Building a Scheduler
      • Building a Scheduler Without a Host
      • Clustering
      • Execution Groups
      • Node Affinity (Preferred Node)
      • Testing
      • Compile-Time Checks
      • Declaring Jobs with Attributes
    • Configuration Reference
    • JSON Configuration
    • Cron Expression Reference
    • Multi-Tenancy
    • Comparison
    • Frequently Asked Questions
    • Best Practices
    • Before You Go Live
    • Operating a Cluster
    • Log Events
    • Tenancy Patterns
    • Database Schema
    • Database Schema Changes
    • Migration Guide
    • Troubleshooting
    • API Documentation
  • How To's
    • One-Off Job
    • Rescheduling Jobs
    • Retrying Failed Jobs
    • Job Continuations
    • Multiple Triggers
    • Job Template
    • Running Quartz under Aspire
    • Quartz.NET with Wolverine
    • Coming from Hangfire
    • Coming from TickerQ
    • Embedding Quartz in a Library
    • Running under an External Leader Election
    • Publishing Trimmed and Native AOT
    • Extending Quartz: what is open, what is closed, and how to ask
    • A Job Store of Your Own
    • A Driver Delegate for a New Database
    • Persisting a Custom Trigger Type
    • A Lock Handler of Your Own
  • Packages

    • Quartz Core Additions

      • Jobs
      • Serialization (System.Text.Json)
      • JSON Serialization
      • Plugins
    • Integrations

      • Aspire Integration
      • ASP.NET Core Integration
      • HTTP API
      • HTTP Client
      • Dashboard
      • Hosted Services Integration
      • Microsoft DI Integration
      • Multiple Schedulers with Microsoft DI
      • Observability
      • Redis Lock Handler
      • TimeZoneConverter Integration
    • 3rd Party Plugins for Quartz
  • Quartz 3.x

    • Getting Started

      • Quartz 3 Quick Start
      • Tutorial
        • Using Quartz
        • Library Overview
        • Jobs And Triggers
        • More About Jobs
        • More About Triggers
        • Execution Groups
        • Node Affinity (Preferred Node)
        • Simple Triggers
        • Cron Triggers
        • RecurrenceTrigger
        • Trigger and Job Listeners
        • Scheduler Listeners
        • Job Stores
        • Tuning the Scheduler
        • Configuration, Resource Usage and SchedulerFactory
        • Advanced (Enterprise) Features
      • Configuration Reference
      • JSON Configuration
      • Multi-Tenancy
      • Frequently Asked Questions
      • Best Practices
      • Tenancy Patterns
      • Troubleshooting
      • API Documentation
      • Database Schema
      • Database Schema Changes
      • Migration Guide
      • Miscellaneous Features
    • How To's

      • One-Off Job
      • Multiple Triggers
      • Job Template
      • Using the CronTrigger
      • Rescheduling Jobs
    • Packages

      • Quartz Core Additions

        • Dashboard
        • Jobs
        • Serialization (System.Text.Json)
        • Serialization (Newtonsoft Json.NET)
        • Plugins
      • Integrations

        • ASP.NET Core Integration
        • Hosted Services Integration
        • Microsoft DI Integration
        • Multiple Schedulers with Microsoft DI
        • OpenTelemetry Integration
        • OpenTracing Integration
        • Redis Lock Handler
        • TimeZoneConverter Integration
      • 3rd Party Plugins for Quartz
  • Old Releases

    • Quartz 2.x
      • Quartz 2 Quick Start
      • Tutorial
        • Lesson 1: Using Quartz
        • Lesson 2: Jobs And Triggers
        • Lesson 3: More About Jobs & JobDetails
        • Lesson 4: More About Triggers
        • Lesson 5: SimpleTrigger
        • Lesson 6: CronTrigger
        • Lesson 7: TriggerListeners and JobListeners
        • Lesson 8: SchedulerListeners
        • Lesson 9: JobStores
        • Lesson 10: Configuration, Resource Usage and SchedulerFactory
        • Lesson 11: Advanced (Enterprise) Features
        • Lesson 12: Miscellaneous Features of Quartz
        • CronTrigger Tutorial
      • Configuration Reference
      • Migration Guide
      • API Documentation
    • Quartz 1.x
      • Tutorial
        • Lesson 1: Using Quartz
        • Lesson 2: Jobs And Triggers
        • Lesson 3: More About Jobs & JobDetails
        • Lesson 4: More About Triggers
        • Lesson 5: SimpleTrigger
        • Lesson 6: CronTrigger
        • Lesson 7: TriggerListeners and JobListeners
        • Lesson 8: SchedulerListeners
        • Lesson 9: JobStores
        • Lesson 10: Configuration, Resource Usage and SchedulerFactory
        • Lesson 11: Advanced (Enterprise) Features
        • Lesson 12: Miscellaneous Features of Quartz
      • API Documentation
  • License

Clustering

A cluster is several scheduler instances sharing one database. Between them they load-balance the work — whichever node acquires a trigger runs it — and they fail over: when a node dies, the jobs it was running are recovered by the others, provided the job asked for it with RequestRecovery().

Clustering needs a persistent store; RAMJobStore has nothing to share. LocalTransactionJobStore — what UsePersistentStore registers — is the store to use.

Enabling it

builder.Services.AddQuartz(q =>
{
    q.ConfigureScheduler(options =>
    {
        // every node in the cluster shares this name: it is what makes them one cluster
        options.InstanceName = "orders";

        // ...and each needs its own id. Generating one is the easy way to be sure.
        options.GenerateInstanceId = true;
    });

    q.UsePersistentStore(store =>
    {
        store.UseSqlServer(connectionString);
        store.UseClustering();
    });
});

Three rules follow from what those settings mean:

  • InstanceName must be the same on every node. It is the SCHED_NAME column of every row, so two nodes with different names sharing a database are not a cluster — they are two schedulers that cannot see each other's work.
  • InstanceId must be different on every node. It is how a node recognises its own check-in row and its own fired triggers. GenerateInstanceId = true derives one at startup from the registered IInstanceIdGenerator — by default host name and timestamp — which is what the flat value AUTO used to mean. Set an explicit id instead when something outside Quartz needs to name the node, such as node affinity, which pins a trigger to an id and therefore needs one that survives a restart.
  • The rest of the configuration should match. Nodes may differ in thread pool size, and in anything else local to the process, but they must agree about the store: same tables, same table prefix, same serializer.

UseClustering() turns on database locking as well, because clustering has never worked without it.

Tuning the check-in

store.UseClustering(cluster =>
{
    cluster.CheckinInterval = TimeSpan.FromSeconds(10);
    cluster.CheckinMisfireThreshold = TimeSpan.FromSeconds(20);
});
OptionDefaultWhat it does
CheckinInterval00:00:07.5How often this node writes "still alive" to QRTZ_SCHEDULER_STATE.
CheckinMisfireThreshold00:00:07.5How long past a missed check-in another node waits before treating this one as dead and recovering its triggers.

Shorter intervals notice a dead node sooner, at the cost of more database traffic from every node all the time. The threshold is the one to raise if a node is being declared dead while it is merely busy or its database is slow — a false positive means its running jobs are recovered and run twice.

Caution

Never run clustering on separate machines, unless their clocks are synchronized using some form of time-sync service (daemon) that runs very regularly (the clocks must be within a second of each other). See https://www.nist.gov/pml/time-and-frequency-division/services/internet-time-service-its if you are unfamiliar with how to do this.

Caution

Never start (scheduler.Start()) a non-clustered instance against the same set of database tables that any other instance is running (Start()ed) against. You may get serious data corruption, and will definitely experience erratic behavior.

Caution

Monitor and ensure that your nodes have enough CPU resources to complete jobs. When some nodes are in 100% CPU, they may be unable to update the job store and other nodes can consider these jobs lost and recover them by re-running.

Batching trigger acquisition

Each node acquires the triggers it is about to fire in batches. By default a batch is one trigger, and that default is deliberate rather than conservative: at MaxBatchSize = 1 — with AcquireTriggersWithinLock left off, which is also the default — acquisition takes no cluster-wide lock at all. Raise it above 1 and every acquisition cycle takes the TRIGGER_ACCESS row lock, including the cycles that acquire nothing. On a lightly loaded cluster that is strictly more lock traffic for no batching.

Two settings decide the size of a batch, and neither does anything on its own:

OptionFlat keyDefault
Scheduler:MaxBatchSizequartz.scheduler.batchTriggerAcquisitionMaxCount1
Scheduler:BatchTriggerAcquisitionFireAheadTimeWindowquartz.scheduler.batchTriggerAcquisitionFireAheadTimeWindow00:00:00

MaxBatchSize is the upper bound on how many triggers one acquisition may take. The window is what decides how many it actually takes: after the first trigger, only triggers due within that window of it join the batch. With the window at zero — the default — a batch holds the triggers due at the same instant and nothing else, so raising MaxBatchSize alone changes nothing for a schedule whose fire times are spread out.

So: move the pair together, or leave them alone.

q.ConfigureScheduler(options =>
{
    options.MaxBatchSize = 10;
    options.BatchTriggerAcquisitionFireAheadTimeWindow = TimeSpan.FromSeconds(1);
});

That is worth doing when many triggers fire at once — a few hundred at the top of the hour — because one acquisition and one TRIGGERS_FIRED round trip replace one of each per trigger. It is not worth doing for a schedule of triggers a minute apart.

The price of the window is that triggers fire early, by up to the window. A one-second window on a schedule with second-level precision is a behaviour change, not a tuning knob.

MaxBatchSize may not exceed the thread pool's MaxConcurrency, and is rejected at startup if it does: triggers acquired beyond the number of threads available to run them are held by this node, unfireable by any other, until the pool drains.

Seeing the cluster

QRTZ_SCHEDULER_STATE is where the check-ins land, and IScheduler.QueryClusterNodes() is how to read it without writing SQL:

List<ClusterNode> nodes = await scheduler.QueryClusterNodes();

foreach (ClusterNode node in nodes)
{
    string marker = node.IsCurrentNode ? " (this node)" : "";
    Console.WriteLine($"{node.InstanceId}{marker}: {node.State}, last check-in {node.LastCheckInUtc:u}");
}

// The verdicts come from the same predicate the failover sweep applies, so a node reported
// Failed is one whose in-flight work the cluster is about to take over.
List<ClusterNode> failed = nodes.FindAll(node => node.State == ClusterNodeState.Failed);

Each ClusterNode carries the node's InstanceId, its LastCheckInUtc, the CheckInInterval that node was configured with, whether it IsCurrentNode, and a State of Alive, Overdue or Failed. The list is the node answering first, then the rest by instance id, and the node answering is always in it — even before its first check-in has written a row.

The verdict is decided by the same predicate the failover sweep uses, so the listing and the recovery it predicts cannot disagree: Failed means the next check-in pass will take this node's work over and delete its row, which is why a corpse is reported for a while and then disappears. Overdue is a missed check-in and nothing more; nothing is recovered from an overdue node. The verdicts are what this node believes, read off its own clock — which is another reason the clocks have to agree.

A scheduler that is not clustered answers with the one node it is, Alive, and both times null: there is no check-in history because there is nobody to keep one for. That is the honest answer rather than an empty list, so a caller need not branch on whether clustering is on.

To see what each node is doing, join the listing to QueryFireInstances on SchedulerInstanceId:

List<ClusterNode> nodes = await scheduler.QueryClusterNodes();
PagedResult<FireInstance> firings = await scheduler.QueryFireInstances(new FireInstanceQuery
{
    // both states: what a node is holding is as interesting as what it is running, and a
    // reservation left behind by a dead node is what recovery is about to clear
    State = null
});

foreach (ClusterNode node in nodes)
{
    int running = firings.Items.Count(firing =>
        firing.SchedulerInstanceId == node.InstanceId && firing.State == FireInstanceState.Executing);

    Console.WriteLine($"{node.InstanceId} ({node.State}) is running {running} job(s)");
}

The same listing is behind GET /schedulers/{name}/nodes in the HTTP API and the Cluster page of the dashboard.

When a node leaves

A node that is shut down rather than killed hands back what it can, and the two kinds of shutdown now differ less than they used to.

It stops firing before it stops running jobs. The scheduling loop is halted and waited for first, so every trigger it had reserved but not yet fired is released to WAITING for another node to pick up on its next pass, and a firing it had already committed — the fired-trigger row written, the trigger moved on to its next instant — is dispatched rather than dropped. This holds whether or not the shutdown waits for jobs; before 4.1 the unwaited kind could close the thread pool underneath its own loop and lose that occurrence outright, which nothing but RequestRecovery would ever get back.

It settles the firings it can, then leaves the rest. Shutdown(waitForJobsToComplete: true) waits for its executions, so nothing is left over at all. Shutdown(waitForJobsToComplete: false) — what AddQuartzHostedService does unless told otherwise — does not wait for the jobs, but it does give the executions already in flight a couple of seconds to report their completions before the job store is closed, because a completion issued after that is refused: the firing stays EXECUTING and, for a [DisallowConcurrentExecution] job, its trigger stays BLOCKED. A job still working when the window closes is abandoned, and what it leaves is what a crashed node leaves — a peer settles it once the check-in lapses, which is CheckinInterval + CheckinMisfireThreshold and then a grace period for an execution that may still be alive.

It does not delete its check-in row. Nothing removes a QRTZ_SCHEDULER_STATE row on the way down; it stays with the timestamp the node last wrote until a peer decides the node has failed and recovers it. So a node that leaves is declared failed on the same schedule as one that crashed — the argument for waiting for the jobs, and for the window above, is that a node with nothing left over has nothing to be recovered.

Asking for recovery

Failover recovers a dead node's executions, and only for jobs that asked. RequestRecovery() on the job is the whole of the asking:

q.AddJob<ChargeInvoicesJob>(j => j
    .WithIdentity("charge-invoices")
    .RequestRecovery());

It is off by default, and it means nothing without a persistent store. When a clustered scheduler decides a peer has stopped checking in, every fired-trigger row that peer left behind is examined: one whose job requested recovery becomes a new trigger in the RECOVERING_JOBS group, carrying the original trigger's job data; one whose job did not is deleted, and that occurrence is lost. The read side is IJobDetail.RequestsRecovery, and a firing that is a recovery says so through IJobExecutionContext.Recovering.

Which jobs should ask, exactly what recovery re-runs and what it deliberately does not, and how a recovered firing identifies itself, are in What RequestsRecovery re-runs, and when. Recovery is about an execution that was interrupted; a job that threw is a completed firing, and re-running that one is a retry policy on its trigger.

See also

  • Best Practices — what to ask for recovery, and how to make a job safe to re-run
  • Retrying Failed Jobs — a trigger's own answer to a job that throws
  • Operating a Cluster — rolling upgrades, instance ids in containers, sizing, backup
  • Running under an External Leader Election — when the election already exists
  • Quartz.Examples.Worker — a worker service with a persistent store, configured the way this lesson describes
Help us by improving this page!
Last Updated: 9/19/26, 9:30 PM
Contributors: Marko Lahma, Claude Opus 5 (1M context)
Prev
Building a Scheduler Without a Host
Next
Execution Groups