Troubleshooting
Common Quartz.NET problems, by symptom, with how to diagnose and fix each.
Scheduler Stops Executing Jobs
Symptoms: jobs stop firing after hours or days. No errors in the logs. The scheduler appears to be running, but no triggers fire.
Common causes:
- Thread pool exhaustion. Long-running jobs occupy every worker; other jobs wait and eventually misfire.
- Check the thread pool size (default 10):
ThreadPool:MaxConcurrencyin 4.x,quartz.threadPool.threadCountas a flat key on both versions. Raise it if you run many jobs at once. - Make sure jobs do not block threads indefinitely. Use cancellation tokens and timeouts.
- Consider
[DisallowConcurrentExecution]so one slow job cannot take every thread.
- Check the thread pool size (default 10):
- Database connectivity. Transient database errors during trigger acquisition can leave the scheduler unable to pick up new triggers.
- Check the connection string and the connection pool configuration.
- Make the connection pool at least the thread count + 3 (see Best Practices).
- Check the database server's logs for connection timeouts or deadlocks.
- Unhandled exceptions in listeners. An exception from an
IJobListener,ITriggerListenerorISchedulerListenercan disrupt the scheduling cycle.- Wrap listener code in try-catch (see Best Practices).
Diagnosis:
- Enable debug logging for the
Quartznamespace to see trigger acquisition. - Check
QRTZ_FIRED_TRIGGERSfor jobs that never completed. - Check
QRTZ_TRIGGERSfor triggers stuck in unexpected states (see the next section). - Check that the scheduler is still firing:
scheduler.StatusisSchedulerStatus.Runningin 4.x; on 3.x,scheduler.IsStartedistrueandscheduler.InStandbyModeisfalse.
Triggers Stuck in ACQUIRED State
Symptoms: triggers show TRIGGER_STATE = 'ACQUIRED' in the database but never fire. New triggers are not picked up.
Causes:
- The scheduler instance that acquired the trigger crashed or lost connectivity before firing it.
- A transient database error during the fire-and-complete cycle: the reservation was written, but the statement that would have fired or released it did not run.
Diagnosis:
-- Find stuck triggers
SELECT TRIGGER_NAME, TRIGGER_GROUP, TRIGGER_STATE, NEXT_FIRE_TIME
FROM QRTZ_TRIGGERS
WHERE TRIGGER_STATE = 'ACQUIRED';
-- Find fired triggers that never completed
SELECT * FROM QRTZ_FIRED_TRIGGERS
WHERE STATE = 'ACQUIRED';
Resolution: the store already does this. On both versions, clustered or not, RecoverStaleAcquiredTriggers runs on the persistent store's misfire loop, every MisfireHandlerFrequency (by default the misfire threshold, one minute).
- For each of this node's own fired-trigger rows still
ACQUIREDpast the stale threshold, it sets the trigger back toWAITING(fromACQUIREDorBLOCKED, since a[DisallowConcurrentExecution]job's trigger may have moved on) and deletes the row. - The stale threshold is twice the misfire threshold, with a floor of two minutes. The floor keeps it clear of normal acquisition, which takes at most one
IdleWaitTime(30 seconds by default) plus the time to fire. There is no separate setting; wideningMisfireThresholdwidens it.
So stuck rows disappear on their own, a minute or two after they stopped moving. The sweep does not touch:
- Rows with another instance id. Rows left by a node that is gone are cleaned up by cluster recovery once that node is declared failed; see Operating a Cluster (4.x). For this node's rows swept by a peer that decided it was gone, see Clock Skew Between Nodes.
- Rows in
EXECUTINGstate. They describe a job the node believes is running, and the node is the authority on that.
Wait one sweep. If nothing changes, find which instance id owns the rows. In 4.x, IScheduler.QueryFireInstances(new FireInstanceQuery { State = null }) lists them without SQL, and QueryClusterNodes() says which of those instance ids still exist.
Resolution, as a fallback:
- Restart the scheduler. A non-clustered scheduler frees every
ACQUIREDandBLOCKEDtrigger and deletes every fired-trigger row at startup. A clustered one does the same for its own rows on its first check-in. - Manual recovery. If a restart is not possible, put the stuck triggers back to
WAITING:
UPDATE QRTZ_TRIGGERS
SET TRIGGER_STATE = 'WAITING'
WHERE TRIGGER_STATE = 'ACQUIRED'
AND NEXT_FIRE_TIME < :currentTimeInMillis;
Warning
Update the database by hand only as a last resort, and never against a running cluster: the row you edit may be one a node is about to fire, and its paired fired-trigger row is left behind. Prefer the sweep or a restart.
Prevention:
- Size the database connection pool adequately.
- Run clustered if you run several scheduler instances; clustering recovers failed nodes automatically.
- Keep jobs short to narrow the window for failures.
A Lock Held by a Connection That Is Gone
Symptoms: every node of a cluster stops firing at the same moment, and the log is silent: no exception, no misfire, no SchedulerError, not even a retry. The processes are healthy and the scheduler reports itself running. In the database, a session is waiting on QRTZ_LOCKS, and the session blocking it belongs to a client that no longer exists.
Diagnosis: ask the database who is blocking whom, then whether the blocker's client still exists.
-- Oracle
SELECT sid, serial#, status, last_call_et, blocking_session, event
FROM v$session
WHERE blocking_session IS NOT NULL
OR sid IN (SELECT blocking_session FROM v$session WHERE blocking_session IS NOT NULL);
-- PostgreSQL: the blocker is the one sitting in 'idle in transaction'
SELECT pid, state, wait_event_type, wait_event, state_change, pg_blocking_pids(pid) AS blocked_by
FROM pg_stat_activity
WHERE backend_type = 'client backend';
-- SQL Server
SELECT session_id, blocking_session_id, wait_type, wait_time, command
FROM sys.dm_exec_requests
WHERE blocking_session_id <> 0;
-- MySQL
SELECT * FROM performance_schema.data_lock_waits;
It is this problem if the blocking session has been idle as long as the outage has lasted and belongs to a node whose process is gone.
Cause: a node held the TRIGGER_ACCESS row lock, and its connection died without telling the server.
- Killing the process sends a TCP reset, and the server ends the session at once. Killing a node therefore does not reproduce this.
- Cutting the network under a live process sends nothing: the socket is aborted on the client side and no FIN or RST reaches the server. The server keeps the session, its open transaction and the row lock until it notices the client is gone.
- When the node comes back, it connects on a fresh session and queues behind its own ghost.
Quartz cannot release that lock, and neither can any other client. A row lock belongs to the session that took it; only the server can end a session that is no longer there. A blocked lock statement returns nothing and throws nothing, so the handler's retry loop, the store's transient-failure handling and the scheduler's error listener all wait with it in silence. Since #3764, Quartz can make the wait finite and visible.
There are two fixes, for different halves of the problem. Apply both.
Make the wait finite
This turns a stall into a failure the scheduler reports and recovers from. It does not free the lock.
CommandTimeout:JobStore:CommandTimeoutin 4.x,quartz.jobStore.commandTimeoutfrom 3.22. It applies to every statement the store issues, the lock statement included, on every database. On Oracle, ODP.NET'sCommandTimeoutcancels the statement rather than ending the wait on the server. The cancel usually surfaces asORA-01013, and on some managed-driver versions asORA-03111. Either way Quartz treats it as a failed statement.- A wait timeout in the lock statement itself. Oracle has no session-level DML lock wait timeout, so on Oracle this is the cleaner option:
FOR UPDATE WAIT 20fails the statement withORA-30006after twenty seconds.
q.UsePersistentStore(s =>
{
s.UseSystemTextJsonSerializer();
s.UseOracle(connectionString);
s.ConfigureStore(options =>
{
// Bounds every statement the store issues, the lock statement included: what would
// have been a wait with no end becomes a failure the store retries and reports.
options.CommandTimeout = TimeSpan.FromSeconds(30);
// And Oracle's own wait timeout, written into the lock statement itself, which
// fails it with ORA-30006 after twenty seconds. {0} is the table prefix, and the
// @ parameter prefix is rewritten for the driver.
options.SelectWithLockSql =
"SELECT * FROM {0}LOCKS WHERE SCHED_NAME = @schedulerName AND LOCK_NAME = @lockName FOR UPDATE WAIT 20";
});
});
On 3.x the same statement is a flat key:
quartz.jobStore.selectWithLockSQL = SELECT * FROM {0}LOCKS WHERE SCHED_NAME = @schedulerName AND LOCK_NAME = @lockName FOR UPDATE WAIT 20
- PostgreSQL:
lock_timeoutaborts any statement that waits longer than it for a lock. Npgsql sets it per connection withOptions=-c lock_timeout=20000in the connection string. - MySQL:
innodb_lock_wait_timeoutalready bounds the waiter at 50 seconds by default.
What Quartz then does:
- The row-lock handler tries the statement
MaxRetrytimes (three by default, a second apart), then throwsLockException. - The scheduler thread reports the first failure through
ISchedulerListener.SchedulerErrorand backs offDbRetryIntervalbefore trying again. - The check-in and misfire loops log every
RetryableActionErrorLogThreshold-th consecutive failure as an error.
None of it is fatal. The scheduler picks up by itself the moment the lock is released; the cluster is now loud instead of silent.
On 4.x it is loud before the timeout expires, too. An acquisition waiting longer than JobStore:LockWaitWarningThreshold (30 seconds by default) logs warning 3716 once, naming the lock and how long it has waited. Every acquisition is measured on the quartz.jobstore.lock.wait.duration histogram, tagged with quartz.jobstore.lock. Alert on that warning: while a lock wait is in progress, nothing else in Quartz produces a signal.
Make the server drop the dead session
This half frees the lock. It is configured on the database server, not in Quartz.
| Database | Server setting | Default |
|---|---|---|
| Oracle | SQLNET.EXPIRE_TIME=n in the server's sqlnet.ora | 0 (off) |
| Oracle | MAX_IDLE_BLOCKER_TIME (19c and later, minutes, ALTER SYSTEM, per-PDB) | — |
| PostgreSQL | tcp_keepalives_idle / _interval / _count | 0: the operating system's values |
| PostgreSQL | idle_in_transaction_session_timeout | — |
| SQL Server | Keep Alive on the TCP/IP protocol, in SQL Server Configuration Manager | 30 seconds, one-second retransmission interval |
| MySQL | OS keepalive (net.ipv4.tcp_keepalive_time), and wait_timeout | Two hours on Linux; eight hours |
- Oracle
EXPIRE_TIMEis dead connection detection. It probes every n minutes and closes the connection when the client is gone, which ends the session and rolls its transaction back. Left at0, you depend on the operating system's TCP keepalive, typically two hours. A single-digit number of minutes is usual. - Oracle
MAX_IDLE_BLOCKER_TIMEends a session that has been idle while blocking another session for that long. A session running a long statement is never idle, so a long query or a slow job is not a candidate. The only Quartz session it can reach is one idle between the statements of a lock-holding transaction, which lasts milliseconds unless the process is paused; if one is caught there, the commit fails and the store retries. A few minutes is a good backstop next toEXPIRE_TIME. - PostgreSQL: setting the keepalives makes the server notice a gone client on its own schedule.
idle_in_transaction_session_timeoutends a session idle inside an open transaction, which is exactly the shape of a lock-holding ghost. - SQL Server enables keep-alive on every connection, unlike the operating system default, so an orphaned connection is usually dropped within a minute and this case rarely lasts long.
- MySQL: the dead session lives until the OS keepalive gives up or
wait_timeoutcloses the idle connection. Tune the keepalive:innodb_lock_wait_timeoutcovers the waiter, and nothing else covers the holder.
Warning
Client-side keepalive is not a substitute. ODP.NET's Keep Alive=true, (ENABLE=BROKEN) in a connect descriptor, and the client-side EXPIRE_TIME of newer clients help the client notice a dead server. None of them frees a lock held by a session on the server, which is the direction this failure runs in.
Rehearsing it
Reproduce it before you need to:
- In a clustered node, put a breakpoint after the lock statement, or suspend the process.
- Disable that machine's network adapter. Do not kill the process: that sends the reset that makes the server clean up.
- Every other node stops firing within one lock attempt. On 4.x, warning 3716 appears after
LockWaitWarningThresholdand the lock-wait histogram climbs. - With the server-side setting in place, the server drops the session, the lock is released and the cluster resumes on its own.
The Misfire Sweep Times Out
Symptoms: JobPersistenceException with an inner timeout from the misfire handler, repeating every minute; Handling the first N triggers of M misfired triggers in the log, never catching up; the scheduler otherwise alive but firing late.
Cause: the sweep does too much work per pass for the time it is allowed, or the query that finds misfired triggers is scanning. Three settings and one index decide it.
The sweep runs on every node. Each pass starts with a COUNT that takes no cluster-wide lock, to avoid paying for the lock when there is nothing to do: WHERE SCHED_NAME = ? AND MISFIRE_INSTR <> -1 AND NEXT_FIRE_TIME <= ? AND TRIGGER_STATE = ?.
- 4.x serves it from the acquisition index
IDX_QRTZ_T_NFT_STon(SCHED_NAME, TRIGGER_STATE, NEXT_FIRE_TIME ASC, PRIORITY DESC, MISFIRE_INSTR): two equalities, a range, andMISFIRE_INSTRin the index so the<> -1never leaves it. 4.0 drops the second index,IDX_QRTZ_T_NFT_ST_MISFIRE, that four dialects had (#3656): it led withMISFIRE_INSTR, compared with<>, so it could not seek, and no measured optimizer picked it (#3608). - 3.x still ships
IDX_QRTZ_T_NFT_ST_MISFIREand sweeps from it. - A schema older than the 3.20 index migration has a different index shape. On a large
QRTZ_TRIGGERS, this query is where a slow database first shows.
q.UsePersistentStore(s =>
{
s.UseSystemTextJsonSerializer();
s.UseSqlServer(connectionString);
s.ConfigureStore(options =>
{
// A pass handles at most this many triggers, then commits. Lower it when the
// sweep is timing out; the loop comes straight back for the rest.
options.MaxMisfiresToHandleAtATime = 20;
// How often the sweep runs. Defaults to MisfireThreshold.
options.MisfireHandlerFrequency = TimeSpan.FromMinutes(1);
// Applied to every statement the store issues, this one included.
options.CommandTimeout = TimeSpan.FromSeconds(30);
});
});
Resolution:
- Apply the current index set:
migrations/3.20on 3.x, or section 5 ofmigrations/4.0on 4.x. This helps most and costs least. - Lower
MaxMisfiresToHandleAtATime(default 20). It bounds one pass; the loop comes back for the rest after a 50 ms pause, so a smaller number means more, shorter transactions, not less progress. - Raise
CommandTimeout(JobStore:CommandTimeoutin 4.x) if the statements are slow rather than blocked. It applies to every statement the store issues, so a node waiting on the cluster-wide lock also waits this long before it can fail and retry. - Raise
MisfireThresholdif the schedule can tolerate more lateness. Fewer triggers cross the line, so there is less to sweep.
Warning
A non-clustered scheduler's startup sweep is unbounded on purpose, on both versions. It handles every misfired trigger in one pass, ignoring MaxMisfiresToHandleAtATime, so that a scheduler starting after a long outage catches up before it fires. It is the pass most likely to time out on a large schedule, and the batch size does not affect it; only the index and the timeout do. A clustered scheduler has no such pass: its startup work is the first cluster check-in, which recovers fired triggers rather than misfires, and the ordinary bounded sweep catches up afterwards.
Clock Skew Between Nodes
Symptoms: jobs run twice; a node logs This scheduler instance (…) is still active but was recovered by another instance in the cluster; nodes flip between Alive and Failed in the cluster listing with no matching outage.
Cause: clustered failure detection compares a timestamp one node wrote with another node's clock. A node whose clock runs ahead of a peer's by more than the slack writes off a healthy peer, releases its acquired triggers and re-runs its recovery-requesting jobs while the peer is still executing them.
The database's clock plays no part. LAST_CHECKIN_TIME holds the writing node's own clock reading, and no SQL in the store asks the server for the time (no GETDATE(), now() or SYSDATE). Setting the database server's clock does not fix this. Only the nodes' clocks agreeing with each other matters, within the failed node's stored check-in interval plus the deciding node's check-in misfire threshold.
Resolution: run a time-synchronisation service on every node; ordinary NTP is orders of magnitude inside the requirement. If you cannot guarantee that, or that the process gets CPU promptly (a starved process shows the same symptom with a perfect clock), widen the window with quartz.jobStore.clusterCheckinMisfireThreshold.
What the written-off node does, on 4.x. A node that finds its own QRTZ_SCHEDULER_STATE row gone on a check-in other than its first has been failed out by a peer. It:
- writes the row back; until then, the rest of the cluster does not see it;
- logs the warning above (event id
3501), plus one naming the peer that recovered it (3515), or saying the peer cannot be named because more than one node has a state row and no row records who recovered whom (3516); - counts the event on
quartz.cluster.recovery.triggerwithquartz.cluster.recovered.instance.idset to its own instance id. Alert on that series: it is a node reporting that it was written off while running. The count is 1, not a number of triggers; the peer's own measurement carries that number under the same attribute.
It does not recover its own fired triggers. The peer already released, rescheduled and deleted them, so a second pass would schedule a second recovery trigger for a firing already being replayed.
On 3.x the node logs the warning and otherwise carries on. Neither version makes this safe: the peer took over work from a process that is still running it, and only the clock fixes that. Rows left in ACQUIRED by the same event clean themselves up; see Triggers Stuck in ACQUIRED State.
Clocks in a cluster has the arithmetic, the default window, and why a pause matters more than an inaccurate clock. On 4.x, Operating a Cluster (4.x) states the exact predicate, and IScheduler.QueryClusterNodes() shows what each node believes about the others.
Misfire Handling
A misfire is a trigger whose scheduled fire time passed without the job running. Causes: the scheduler was shut down, no worker thread was free, or the system was under heavy load.
How It Works
- On startup, and periodically while running, Quartz finds triggers whose
NEXT_FIRE_TIMEis at or beforenow - misfireThreshold. - It applies each misfired trigger's misfire instruction.
The default misfire threshold is 60 seconds for a persistent store: JobStore:MisfireThreshold in 4.x, quartz.jobStore.misfireThreshold as a flat key on both versions.
The threshold instant itself counts as late. In 4.x the rule is the same everywhere: the in-memory store, the persistent store's periodic sweep, and the single-trigger path a resumed or unblocked trigger takes. On 3.x the persistent store's sweep uses strictly before, so a trigger due at exactly now - misfireThreshold is misfired in memory and, for one tick, not in the database.
Misfire Instructions by Trigger Type
Quartz 4.x names the instructions on one enum per family (SimpleTriggerMisfireInstruction, CronTriggerMisfireInstruction and so on). Quartz 3.x names the same values as constants under MisfireInstruction, sometimes spelled longer.
| Trigger Type | 4.x | 3.x | Behavior |
|---|---|---|---|
| SimpleTrigger | FireNow | FireNow | Fire immediately, remaining repeat count unchanged |
NowWithExistingCount | RescheduleNowWithExistingRepeatCount | Fire now, keep original repeat count | |
NowWithRemainingCount | RescheduleNowWithRemainingRepeatCount | Fire now, only remaining repeats | |
NextWithExistingCount | RescheduleNextWithExistingCount | Skip to next scheduled time, keep original count | |
NextWithRemainingCount | RescheduleNextWithRemainingCount | Skip to next scheduled time, remaining count | |
| CronTrigger | FireAndProceed | FireOnceNow | Fire immediately once, then resume schedule |
DoNothing | DoNothing | Skip missed firings, wait for next scheduled time | |
| RecurrenceTrigger | FireAndProceed (default) | — | Fire immediately once, then resume schedule |
DoNothing | — | Skip missed firings, wait for next scheduled time |
Every family also has:
IgnoreMisfires: fires every missed firing, as fast as it can.SmartPolicy: the default. ForCronTriggerandRecurrenceTriggerit fires once now and resumes; forSimpleTriggerit depends on the repeat count.
Choosing a misfire instruction covers which to pick.
Tuning
If triggers misfire often under normal load:
- Raise the thread pool size:
ThreadPool:MaxConcurrencyin 4.x,quartz.threadPool.threadCountas a flat key on both versions. - Raise the misfire threshold if small delays are acceptable:
JobStore:MisfireThresholdin 4.x,quartz.jobStore.misfireThresholdas a flat key on both. - Spread high-frequency triggers across several scheduler instances with clustering.
Job Deserialization Failures After Refactoring
Symptoms: after renaming a job class, changing its namespace or moving it to another assembly, the scheduler throws TypeLoadException or JobPersistenceException on startup.
Cause: QRTZ_JOB_DETAILS.JOB_CLASS_NAME stores the full type name, including namespace and assembly. When the type moves, the stored name no longer resolves. The job's trigger goes to ERROR rather than firing, because the failure is in building the job, not running it; see What the trigger states mean.
Resolution: rewrite the stored name.
UPDATE QRTZ_JOB_DETAILS
SET JOB_CLASS_NAME = 'NewNamespace.NewClassName, NewAssembly'
WHERE JOB_CLASS_NAME = 'OldNamespace.OldClassName, OldAssembly';
Run it during the deployment that renames the type. Afterwards, clear any triggers that reached ERROR with IScheduler.ResetTriggerFromErrorState.
Resolution: declare the rename (4.x). Map the old name to the new type, and every stored row carrying the old name keeps resolving. This suits a rolling deployment: nodes still on the old build write the old name while the new ones read it, so nothing needs rewriting while both run.
services.AddQuartz(q => q.UseTypeLoader(loader =>
{
// Old assembly-qualified name as stored, and the type it means now. Keep the entry until
// every row that could carry the old name has been rewritten or has aged out.
loader.Map("Acme.Jobs.NightlyReport, Acme.Jobs", typeof(NightlyRollupJob));
}));
The same map binds from configuration, so a rename can ship in appsettings.json with the deployment, without a rebuild:
{
"Quartz": {
"TypeLoader": {
"Aliases": {
"Acme.Jobs.NightlyReport, Acme.Jobs": "Acme.Jobs.NightlyRollupJob, Acme.Jobs"
}
}
}
}
How the map behaves:
- Where it applies: wherever Quartz turns a string into a type at run time: a stored
JOB_CLASS_NAME, a job named in XML or JSON scheduling data, aquartz.plugin.<name>.typekey. - Where it does not: the flat keys naming a scheduler's own components (job store, thread pool, serializer, lock handler, job factory, instance id generator, time provider, connection provider) are resolved while the service collection is still being built, before any options exist, and are not aliased. Each names a type in a file you can edit.
- Matching: a key matches the whole stored name, or the part before the comma that starts the assembly.
Acme.Jobs.NightlyReportcovers any assembly spelling after it;Acme.Jobs.NightlyReport, Acme.Jobscovers only that one. - Validation: an alias whose target names no loadable type fails at startup, naming both halves of the entry, rather than surfacing later as a
TypeLoadExceptionabout the dead name. - Scope: type loading is container-wide. A rename declared through any scheduler's builder applies to every scheduler in the container.
- Nothing is written back. A job read under an aliased name keeps the old spelling in
JOB_CLASS_NAME. That makes the alias safe during a rollout, and makes theUPDATEabove the way to retire it eventually. EnableDebuglogging forQuartz.Impl.SimpleTypeLoaderto see whether anything still hits an alias before you remove it.
Resolution: a type loader of your own (4.x). When the answer is not a table (a name resolved out of a plugin's AssemblyLoadContext, or a naming scheme rather than a list), implement ITypeLoader, a single method, and register it with UseTypeLoader<T>():
/// <summary>
/// Resolves the type names stored in JOB_CLASS_NAME, translating the ones that have since moved.
/// </summary>
public sealed class RenameAwareTypeLoader : ITypeLoader
{
// Old assembly-qualified name as stored, new type. Keep an entry until every row that could
// carry the old name has been rewritten or has aged out.
private static readonly Dictionary<string, Type> renamed = new(StringComparer.Ordinal)
{
["Acme.Jobs.NightlyReport, Acme.Jobs"] = typeof(NightlyRollupJob)
};
public Type? LoadType(string name)
{
if (string.IsNullOrEmpty(name))
{
return null;
}
if (renamed.TryGetValue(name, out Type? moved))
{
return moved;
}
// A name that cannot be resolved must throw rather than return null: Quartz only asks when
// it already knows a type is required.
return Type.GetType(name, throwOnError: true);
}
}
services.AddQuartz(q => q.UseTypeLoader<RenameAwareTypeLoader>());
- It replaces the loader for the whole container, and the declared map with it: the map is read by the loader Quartz ships.
- It must throw for a name it cannot resolve. Quartz asks only when it knows a type is required, so a
nullwould fail later with nothing to point at. Returnnullonly for a null or empty name.
The loader Quartz ships also maps Quartz's own 3.x → 4.0 renames, logging a warning each time so the configuration can be corrected:
Quartz.Spi.*asQuartz.Extensibility.*Quartz.Simpl.*asQuartz.Impl.*Quartz.Job.*asQuartz.Jobs.*Quartz.Plugin.*asQuartz.Plugins.*Quartz.Listener.*asQuartz.Listeners.*- the job stores' old names (
JobStoreTX,JobStoreCMT) and the assemblies merged into the core package
Prevention:
- Keep job class names and namespaces stable across releases.
- To rename, declare the alias in the deployment that renames the type (4.x), and update the database in a later one, once nothing hits the alias.
- Name the type in one place: a
public static readonly JobKeyon the job class, and registration throughAddJob<T>()rather than a type-name string.
Database Connection Issues
Symptoms: JobPersistenceException with an inner SqlException/NpgsqlException, intermittent "Couldn't obtain triggers" errors, or "Object cannot be cast from DBNull" errors.
Common causes:
- Connection pool too small. The pool runs out under load.
- Minimum: thread pool size + 3.
- Clustered setups need extra connections for cluster management.
- Connection timeouts. The database is slow or the network unreliable.
- Set the store's
CommandTimeout(JobStore:CommandTimeoutin 4.x,quartz.jobStore.commandTimeoutfrom 3.22), not a connection-string keyword. It bounds every statement the store issues and is the only setting that reaches the lock statement; not every driver has a connection-string equivalent, and ODP.NET has none. - Check network latency between the scheduler and the database server.
- If the statements are stuck rather than slow, and the whole cluster with them, see A Lock Held by a Connection That Is Gone.
- Set the store's
- Lock contention. Several scheduler instances compete for the same rows.
- Two schedulers share a name (
Scheduler:InstanceName, orquartz.scheduler.instanceName) only when they are meant to be one cluster, and then both must have clustering enabled. - Never point several non-clustered schedulers at the same tables (see Best Practices).
- Two schedulers share a name (
Datasource Configuration Example
services.AddQuartz(q =>
{
q.UsePersistentStore(s =>
{
s.UseSystemTextJsonSerializer();
s.UseSqlServer(connectionString);
// Ensure your connection string has an adequate pool size
// e.g., "...;Max Pool Size=25;"
});
});
Scheduler in Web Environments
IIS App Pool Recycling
By default IIS recycles and stops idle application pools, which stops the scheduler. Fixes:
- IIS 8+: configure the site as "Always Running" with preload enabled. See Application Initialization.
- Use the hosted service integration (recommended), so Quartz follows the ASP.NET Core application lifecycle:
services.AddQuartz(q =>
{
// configure jobs and triggers
});
services.AddQuartzHostedService(q => q.WaitForJobsToComplete = true);
- Run a separate process. For critical scheduling, run the scheduler as a Windows Service or a Linux systemd service instead of inside a web application.
Graceful Shutdown
To give jobs time to complete when the application shuts down:
services.AddQuartzHostedService(options =>
{
options.WaitForJobsToComplete = true;
});
Jobs should check IJobExecutionContext.CancellationToken so they respond to shutdown promptly. See Shutdown has a deadline for the time limits.
Common Error Messages
| Error | Likely cause | Resolution |
|---|---|---|
ObjectAlreadyExistsException | Scheduling a job or trigger whose key already exists | scheduler.RescheduleJob() to replace a trigger, or check first with scheduler.Exists() (3.x: scheduler.CheckExists()) |
JobPersistenceException | Database error in a job store operation | Check connectivity, pool size and query timeouts |
SchedulerException: Scheduler has been shutdown | Calling the scheduler after Shutdown() | Fix the application's scheduler lifecycle |
TypeLoadException on job execution | Job class renamed or moved | Declare the rename (4.x) or update JOB_CLASS_NAME in QRTZ_JOB_DETAILS; see Job Deserialization Failures |
JobExecutionException | Unhandled exception inside IJob.Execute() | Catch it in the job; see Best Practices |
