Monitoring and alerting

Varnish exposes its internal state through counters, readable with varnishstat. This page lists the counters that are worth alerting on, what a non-zero value means, and where the storage-related signals live on hosts running MSE4.

Counter names have the form PREFIX[.<ident>].<counter>, where <ident> is present when the group is scoped to a specific object: an MSE instance, an accounting key, a named observatory, and so on.

Counters prefixed c_ and most unprefixed counters are cumulative: alert on their rate of change, not their absolute value. Counters prefixed g_ are gauges and can be alerted on directly, and counters prefixed n_ count events. To scrape counters, use varnishstat -j for JSON output, or the stat VMOD for Prometheus format.

The suggested severities below are starting points:

  • page: user-visible impact is happening or imminent; act now.
  • warn: capacity or health problem building up; investigate soon.
  • watch: worth a dashboard panel and a look during business hours.

Where the Enterprise counters live

The counter groups below exist only in Varnish Enterprise. Groups shared with open source Varnish Cache, such as MAIN, MGT, VBE, SMA and SMF, are documented upstream; Enterprise adds counters to some of them. For per-counter descriptions, see the counter reference (man varnish-counters) shipped in the varnish-plus package, which is authoritative for the version you are running.

Group prefix Scope
GOV. Memory governor (global)
LICENSE. Global; only present when license enforcement is enabled
ACCG. Per <namespace>.<key> accounting key (vmod_accounting)
ACCG_DIAG. Global
BROTLI. Global
COND. Global
MSE. Per MSE instance
MSE_BOOK. Per MSE book
MSE_STORE. Per MSE store
MSE4. Per MSE4 instance
MSE4_MEM. Per MSE4 instance
MSE4_CAT. Per MSE4 category
MSE4_BOOK. Per MSE4 book
MSE4_STORE. Per <book>.<store>
MSE4_BANJRN. Per MSE4 book, identified by the book id. A stand-alone ban journal, configured with banfile_path, adds one unidentified counter set for that file
MSE4_LRU. Per leaf category, as (<category-path>). A category divided into several LRU lists gets one counter set per list, as (<category-path>,<n>)
SLICER. Global
OBSERVATORY_EVENT. / OBSERVATORY_AMOUNT. / OBSERVATORY_C_HISTO. / OBSERVATORY_G_HISTO. Per named observatory (vmod_stat)
RATELIMIT. / RLCOL. Global / per collection
VMOD_DATADOME. Global
VMOD_HTTP. Global
VMOD_IMAGE. Global
VMOD_KV. Per store
VMOD_NODES. / VMOD_NODES_CG. Global / per config group

MAIN counters to alert on

Thread and worker saturation

When Varnish runs out of worker threads, requests queue and are eventually dropped. See also the thread-related material in Troubleshooting.

Counter Meaning Severity
MAIN.threads_limited More threads were needed, but the thread_pool_max limit was reached. warn
MAIN.threads_failed Creating a thread failed (OS limits); see VSL Debug records for diagnostics. warn
MAIN.thread_queue_len Gauge: length of the session queue waiting for threads. Only updates once per second. Persistent non-zero values mean the thread pools cannot keep up. warn
MAIN.sess_queued A session had to wait for a thread. An early warning for the drop counters below. watch
MAIN.sess_dropped An HTTP/1 session was dropped because the queue was already too long (thread_queue_limit). Work is being refused. page
MAIN.req_dropped An HTTP/2 stream was refused because the queue was already too long. page
MAIN.sess_drop A session was silently dropped for lack of a worker thread. page

Session accept failures

Counter Meaning Severity
MAIN.sess_fail Failure to accept a TCP connection. The MAIN.sess_fail_* counters break down the cause; the most common operational one is sess_fail_emfile: no file descriptor was available – raise RLIMIT_NOFILE (see ulimit -n). page

Backend health

Counter Meaning Severity
MAIN.backend_fail Backend connection failures. warn
MAIN.backend_unhealthy Backend connections not attempted because the backend was marked unhealthy by its probe. warn
MAIN.backend_busy Backend connection not made because the backend’s max_connections was reached. warn
MAIN.backend_wait_fail A backend request waited in the backend wait queue and timed out without getting a connection. warn
MAIN.fetch_failed A backend fetch failed, all causes. warn

Workspace overflows

All of these mean a request ran out of its preallocated workspace memory; the fix is usually raising the corresponding workspace_* parameter. See Troubleshooting for the symptoms of each.

Counter Meaning Severity
MAIN.ws_client_overflow Out of space in workspace_client. warn
MAIN.ws_backend_overflow Out of space in workspace_backend. warn
MAIN.ws_session_overflow Out of space in workspace_session. warn
MAIN.ws_thread_overflow Out of space in workspace_thread. warn
MAIN.client_resp_500 A response failed with a 500 due to running out of workspace memory during delivery. Client-visible. page

Cache and VCL

Counter Meaning Severity
MAIN.n_lru_nuked Objects forcefully evicted from storage to make room for new objects. A steady rate means the cache is smaller than the working set. Stays at zero on MSE4 hosts – see the warning below. watch
MAIN.n_lru_limited More storage space was needed, but the nuke_limit was reached. warn
MAIN.vcl_fail A failure prevented VCL from completing, and the client received an error. page
MAIN.losthdr HTTP header overflows: a header did not fit and was dropped. warn
MAIN.sc_overload Sessions closed with error OVERLOAD (out of some resource). page

MSE4 counters

On hosts using MSE4, the storage health and pressure signals live in the MSE4* counter groups, not in MAIN.

Store and book health

Counter Meaning Severity
MSE4_STORE.<id>.online Gauge: non-zero when the store is in state ONLINE, zero in FAILING or OFFLINE. A store at 0 has been taken out of service, typically after an I/O error. page
MSE4_BOOK.<id>.online Gauge: same as above, for the book. A book going offline takes all of its stores with it. page
MSE4_STORE.<id>.c_checksum_verification_failed A chunk read from disk failed checksum verification: the object is removed from the cache and that delivery is aborted. A rising rate indicates disk corruption. warn
MSE4_BANJRN.<book>.g_overflow_bans Gauge: bans purged from the persisted ban journal for lack of space. Non-zero means some persisted objects may not be revivable on restart because their ban information was lost. Alert on every book; a stand-alone ban journal reports under MSE4_BANJRN.g_overflow_bans, without a book id. warn

Allocation pressure

Writes to MSE4 are write-behind: objects land in the memory cache and are flushed to the store asynchronously, so a full store does not directly delay clients as long as pruning keeps up. Pressure first becomes visible in the allocation queue and failure counters.

Counter Meaning Severity
MSE4_STORE.<id>.g_allocation_queue Gauge: threads currently queued waiting for store space. A persistently growing queue means pruning cannot keep up with the write rate – back-pressure on fetches. warn
MSE4_STORE.<id>.c_allocation_failure Store allocation failures, total. Breakdowns: c_allocation_failure_timeout (waited too long) and c_allocation_failure_noslot (no book slot). warn
MSE4_STORE.<id>.c_segment_prune / c_segment_pruned_objects Segments pruned for free space, and objects evicted by pruning. This is the MSE4 equivalent of LRU nuking; a steady rate is normal operation on a full store. watch
MSE4_BOOK.<id>.g_freeslot_queue Gauge: threads queued waiting for a free book slot. Persistent non-zero values mean the book is too small for the object count. warn

IO latency

Each store exposes IO latency histograms: MSE4_STORE.<id>.c_io_read_* and c_io_write_* with buckets at 10, 100, 1000, 10000, 100000, and 1000000 microseconds (plus an _up bucket above that). A shift of reads or writes into the slower buckets is an early signal of a struggling device.

Memory governor

With the memory governor enabled, sustained memory pressure has a distinct signature:

Counter Meaning Severity
GOV.g_overloaded Gauge: memory usage exceeds the target and the governor cannot release memory. Clears when usage falls back under the target. page
GOV.n_enter_overload Times the overloaded state was entered. warn
GOV.n_eviction_failed Times the governor failed to evict any object. Can grow if memory_target is set too low for the baseline (non-object) memory usage. warn
MSE4_MEM.g_space Gauge: memory headroom (target minus usage, floored at zero). Zero means the governor is evicting to hold the target. watch

®Varnish Software, Wallingatan 12, 111 60 Stockholm, Organization nr. 556805-6203