Recover blocked audit devices
Every operation with Vault is an API request or response. When you enable one or more audit devices, Vault logs those requests and responses in detail.
The following diagram shows Vault writing audit entries to a socket audit device:

As applications and users make requests to Vault, it writes those requests and responses to the audit device as described in the audit device documentation. Vault operates as expected and continues to respond to requests.
Vault must be able to write to its audit devices, and those devices must not block Vault from connecting. When a device does block, such as during a network outage, Vault can stop responding to requests until it can write to the audit device.
This tutorial provides examples of blocked audit device behavior, related Vault operational log messages, and suggestions for resolution.
Blocked audit device behavior
When an enabled audit device fails in a blocking manner, Vault requests do not complete until you recover the device.
The following diagram shows a blocked audit device condition. Vault has enabled a socket audit device on localhost at port 9090, but that device is not reachable.

Applications and users can still make requests, but Vault does not respond until it can write to the socket audit device again. Refer to the blocked audit devices section of the audit devices documentation for more details.
Enable redundant audit devices when possible to improve availability.
When you discover a blocked audit device, first restore Vault's ability to write to it.
Use your monitoring alerts and the Vault operational log output to identify the blocked device, then unblock it. Vault returns to service immediately.
The following sections describe how to discover and resolve each type of blocked audit device.
Blocked file audit device
A file audit device commonly blocks when the storage device holding its log file runs out of capacity. When capacity runs out, the audit device blocks and Vault stops servicing requests until you free or add storage.
The following diagram shows a file audit device whose storage has no capacity left:

Example file audit device log messages
To diagnose a blocked file audit device, check the Vault operational log output for ERROR level lines from the audit and core subsystems that resemble this example:
[ERROR] audit: backend failed to log response: backend=file/ error="write /mnt/log/vault-audit.log: file already closed"
[ERROR] core: failed to audit response: request_path=pki/issue/example-dot-com error="1 error occurred: * no audit backend succeeded in logging the response"
These log entries occur in pairs, and Vault repeats them for every request until you resolve the file audit device issue.
By analyzing the log lines, you can also learn the path to the device in question. In this example, the path is /mnt/log/vault-audit.log.
Examine the /mnt/log filesystem with the df command to check its current capacity.
$ df --human-readable --exclude-type=tmpfs --exclude-type=devtmpfs
Filesystem Size Used Avail Use% Mounted on
/dev/mapper/vagrant--vg-root 62G 2.0G 57G 4% /
/dev/sdb 390M 390M 0 100% /mnt/log
In this example output, the /mnt/log filesystem on the device /dev/sdb shows a Use% of 100%, so the filesystem has no capacity left for Vault to write to the audit device log. Increase the available capacity on that filesystem to restore Vault service to applications and users.
Restore the storage device at the path in the audit device configuration. You cannot resolve a blocked file audit device by adding another audit device or changing its path, because Vault cannot service those requests while the device blocks.
Blocked socket audit device
You can configure a socket audit device to use a TCP, UDP, or Unix socket type.
An extreme example of a blocked socket audit device is a closed TCP socket.
In the following example, a log aggregation agent listens locally on Vault servers through a TCP socket, and a Vault socket audit device communicates with this agent.
When the agent process stops, crashes, or otherwise stops listening on the TCP socket, Vault can no longer write to that device. Restore service to the listening process, or otherwise let Vault connect to it again.
When Vault writes to a socket audit device across a routed network, a filter or firewall can interfere with communication and block the device. Check the network path when you troubleshoot a socket audit device.
Example socket audit device log messages
To diagnose a blocked socket audit device, check the Vault operational log output for ERROR level lines from the audit subsystem that reference failures to log responses, as shown in this example:
[ERROR] audit: backend failed to log response: backend=socket/ error="2 errors occurred:
* write tcp 127.0.0.1:59660->127.0.0.1:9090: write: broken pipe
* dial tcp 127.0.0.1:9090: connect: connection refused
These log lines spell out important details about the issue.
- The socket audit device enabled at the path
socket/has a problem. - Vault reports two distinct errors:
- Vault cannot write to the socket and reports
broken pipe. - Vault cannot dial the socket and reports
connection refused.
- Vault cannot write to the socket and reports
These details are enough to help you decide that the service is not listening on the socket and is no longer accepting connections from Vault.
The log messages differ based on your environment, but follow this general pattern.
If you have a log aggregation and analysis stack ingesting Vault operational logging, consider configuring an alert for instances of no audit backend succeeded in logging the response to detect blocked device incidents.
Blocked syslog audit device
Process capabilities, user permissions, and the size of the data Vault writes can all block the syslog audit device.
To write to the system log, the Vault process user must hold the required capabilities, such as CAP_SYSLOG, along with the necessary permissions. Refer to capabilities(7) in the Linux Programmer's Manual on man7.org for the full list.
Cumulative Vault data, such as certificate revocation lists (CRLs) and LDAP groups, can also grow enough to exceed the UDP datagram size that the syslog protocol specification allows. This limit applies when Vault logs to remote syslog devices. Refer to the syslog protocol specification on ietf.org for the datagram size limit.
Example syslog audit device log messages
To diagnose a blocked syslog audit device, check the Vault operational log output for ERROR level lines from the audit and core subsystems. The following examples show lines that reference failures to log responses.
If syslog is not accessible on the system, you can observe errors like this when Vault tries to write to it, and when you first try to enable it.
[ERROR] enable audit mount failed: path=syslog/ error="Unix syslog delivery error"
[ERROR] core: failed to audit response: request_path=sys/audit/syslog error=1 error occurred:
* no audit backend succeeded in logging the response
Vault logs this example when you attempt to enable the audit device. The error Unix syslog delivery error can mean that the host does not run the syslog service, or that Vault cannot access it. SELinux configuration on the host often causes this restriction.
If Vault writes items to the syslog audit device that exceed the syslog host's configured maximum socket send buffer, Vault logs errors such as the following example.
[ERROR] audit: backend failed to log response: backend=syslog/ error=write unixgram ->/var/run/log: write: message too long
[ERROR] core: failed to audit response: request_path=pki/certs/ error=1 error occurred:
* no audit backend succeeded in logging the response
In this example, the audit device is available, but Vault cannot write entries that exceed the allowed size. The result is the same blocking behavior as an unreachable device.
The write: message too long error is the critical clue. Alert on this error. Vault writes to the syslog socket /var/run/log in this example, but that path can differ in your environment. The second error line holds another clue about the source of the issue in the request_path value.
Because the request_path value includes pki/certs, the usual cause is a list operation over many PKI certificates. That operation makes Vault write an excessively large audit device log entry.
Refer to socket(7) in the Linux Programmer's Manual on man7.org to learn about raising the /proc/sys/net/core/wmem_default kernel tunable, which increases the socket send buffer size. Switching to a TCP-based syslog listener also helps with larger log messages.
These write: message too long errors point to deeper problems, such as an unmaintained CRL in a PKI secrets engine or a large list of LDAP groups. Investigate those problems to find the use cases that generate oversized audit log entries.
Help and reference
- Audit devices
- File audit device
- Socket audit device
- syslog audit device
- Filtering concepts
- Audit filter API documentation
- Audit filter CLI documentation
- capabilities(7) in the Linux Programmer's Manual on man7.org