Configuration for clustering multiple Virtual Registry nodes together.
For more information on the internals of clustering: https://docs.varnish-software.com/varnish-enterprise/features/cluster
Peers can be listed statically under peers, fetched from a service with peer_group, or found in the environment the node runs in with discovery. Only one of the three can be configured, and the Supervisor rejects a configuration that sets more than one. To enable clustering for a specific virtual registry, set enable_cluster on the registry.
Example:
cluster:
token: secret
peers:
- url: http://orca-node-1:8080
- url: http://orca-node-2:8080
- url: http://orca-node-3:8080
tokencluster:
token: secret
Type: String
Shared secret used to authenticate cluster peers. All nodes in the cluster must be configured with the same token.
extra_tokenscluster:
token: new-secret
extra_tokens:
- secret
Type: List
Additional shared secrets accepted from peers, alongside token. Lets a cluster token be rotated with a rolling redeploy instead of a coordinated restart of every node at once: stage the new token here on every node first, promote it to token once every node accepts it, then remove the old value.
storage_replicascluster:
token: secret
storage_replicas: 1
Type: Integer
Default: unset, which replicates in full
How many nodes keep an on-disk copy of each cacheable object.
A clustered node caches every object it serves in memory and persists it to disk as well, so a cluster of N nodes holds N copies on disk and its usable persistent cache is no larger than one node’s. Setting storage_replicas: 1 persists an object only on the node that owns it, pooling the whole cluster’s disk into a single cache. A higher count keeps that many copies, on the nodes nearest in the hash ring, trading capacity back for redundancy. Left unset, an object is persisted on every node it passes through.
A node that does not persist an object still caches it in memory, so this decides what survives a restart rather than what a hit costs.
Requires persistent storage to shard, meaning configured varnish.storage.stores and the sup-persistence addon. Without them every object is memory-only already and the setting has no effect, which the Supervisor warns about at startup.
peer_groupDynamic cluster peer group configuration. Uses the same fields as remotes to point to a service that exposes the list of cluster peers.
cluster:
token: secret
peer_group:
url: http://peer-discovery-service:8080
peerscluster:
token: secret
peers:
- url: http://orca-node-1:8080
- url: http://orca-node-2:8080
- url: http://orca-node-3:8080
Type: List
List of individual peer nodes for a static cluster. Each entry uses the same fields as remotes, including name, which replaces the peer’s position in the generated <registry>_cluster_dir_<label> identifier.
All nodes in the cluster should be listed, including the node itself. Each node must be able to send HTTP requests to all other nodes, including itself. All inter-node communication happens over the regular HTTP(S) port.
discoverycluster:
token: secret
discovery:
enabled: true
aws:
tags:
- key: aws:autoscaling:groupName
values:
- orca-prod
Type: Object
Default: unset, which is off
Find the cluster’s peers in the environment the node runs in, instead of naming them. EC2 is the only environment supported, so turning discovery on means asking EC2.
The node polls for running instances, writes the ones it accepts to a file the cluster director reads, and repeats on an interval. An instance that comes up joins the cluster within one interval and one that goes away leaves it, so an Auto Scaling Group can grow and shrink without a configuration change on any node. The node discovers itself along with its peers, which is what a static peers list has to spell out.
A peer is reached on its private address, at the port of this node’s first varnish.http listener. Where there is no plain HTTP listener, the first varnish.https one is used instead and peer traffic goes over TLS. Every node in a cluster is deployed from the same configuration, so this node’s listeners describe its peers’ listeners too. The port is read once when the node starts, so moving it takes a restart rather than a reload.
A failed poll is logged and retried on the next interval, leaving the peers already discovered in place. When it is the first poll that fails, the node starts as a cluster of one and picks the others up once a poll succeeds.
Requires a license carrying the sup-clustering and sup-auth addons along with the vmod-nodes feature. The Supervisor names whichever of the three the license is missing and refuses to start.
enabledcluster:
discovery:
enabled: true
Type: Boolean
Default: false
Turn peer discovery on. EC2 is the only provider, so there is nothing further to select: the node reads its AWS credentials from its instance profile and the region from the instance it runs on, both through the standard AWS resolution chain, and needs the ec2:DescribeInstances permission. The region, the interval and the tag filter it ended up with are logged once at startup, since none of them are in the configuration file.
awsType: Object
Options for EC2 discovery. Omit it for the defaults, which poll every 15 seconds and accept every running instance in the region.
tagscluster:
discovery:
enabled: true
aws:
tags:
- key: aws:autoscaling:groupName
values:
- orca-prod
- key: Environment
values:
- staging
- dev
- key: orca-node
Type: List
Default: unset, which accepts every running instance in the region
Narrow what counts as a peer. An instance has to carry every key listed to be accepted, and where a key has values, any one of them matches. A key given without values matches on the tag being present at all, whatever it is set to.
An Auto Scaling Group deployment should filter on aws:autoscaling:groupName, the tag AWS sets on the instances it launches for the group. Left out entirely, every running instance in the region is a peer, including whatever else the account runs there.
intervalcluster:
discovery:
enabled: true
aws:
interval: 15
Type: Integer
Default: 15
How often to poll EC2 for the cluster’s peers, in seconds. This is the delay before an instance that has just come up starts taking cluster traffic, and before one that has gone away stops being sent any.