A lock-free ring buffer is a fixed-capacity queue that lets one thread publish data while another consumes it without a mutex. In C++, the most practical version is single-producer, single-consumer (SPSC): exactly one thread calls push, and exactly one thread calls pop. That constraint removes the need for compare-and-swap (CAS) in the hot path and makes correctness easier to reason about.
This pattern suits telemetry pipelines, market-data adapters, audio, network ingestion, logging, and AI inference systems where a producer must hand work to a consumer with predictable latency. It is not automatically faster than a mutex queue: a bounded queue can waste CPU while spinning, and a mutex may win when contention is low. Measure the complete pipeline, including batching, scheduling, cache misses, and backpressure.
Start with the ownership model
Before writing code, specify the queue contract:
- One producer, one consumer: the implementation below is safe only under this rule.
- Bounded storage:
pushreturnsfalsewhen full; it never allocates or blocks. - Non-blocking API: callers decide whether to retry, drop, batch, or apply backpressure.
- Object lifetime: the queue must outlive both worker threads, and destruction must happen only after they stop accessing it.
- Element requirements:
Tmust be safely assignable and readable by the relevant threads. For the simplest low-latency path, prefer trivially copyable or cheap-to-move types.
If several producers or consumers are required, do not simply add CAS around the indices. A producer can reserve a slot and then pause before writing its value, leaving a consumer unsure whether the slot is ready. MPMC queues need per-slot sequence state or another proven algorithm. For production infrastructure, review a tested library and validate its progress guarantees rather than adapting an SPSC queue.
How the indices publish ownership
Use two monotonically increasing counters: head identifies the next item the consumer may read, while tail identifies the next slot the producer may write. The counters can grow without wrapping; indexing uses a mask when capacity is a power of two.
The producer first checks the consumer’s published head with an acquire load. If space exists, it writes the element and then publishes the new tail with a release store. The consumer acquires tail before reading the element, then releases its updated head after it has finished reading. This creates a happens-before relationship around the payload, not just around the counters.
memory_order_relaxed is appropriate for a thread’s locally owned counter. Acquire and release are required at the hand-off points. Code that appears correct on x86 can fail on ARM or other weaker memory-ordering architectures if those operations are weakened incorrectly. This matters for Indian deployments using ARM cloud instances as well as mixed x86 and ARM fleets.
A practical C++20 SPSC implementation
The following implementation uses a power-of-two capacity, one unused slot to distinguish full from empty, and cache-line separation for the indices. It stores objects in a vector, so construct the queue before starting worker threads and do not resize it afterward.
#include <atomic>
#include <cstddef>
#include <memory>
#include <stdexcept>
#include <utility>
template <class T>
class SpscRingBuffer {
public:
explicit SpscRingBuffer(std::size_t requested_capacity)
: capacity_(next_power_of_two(requested_capacity + 1)),
mask_(capacity_ - 1),
storage_(std::make_unique<T[]>(capacity_)) {
if (requested_capacity == 0) {
throw std::invalid_argument("capacity must be greater than zero");
}
}
bool try_push(const T& value) {
const auto tail = tail_.load(std::memory_order_relaxed);
const auto head = head_.load(std::memory_order_acquire);
if (tail - head == capacity_ - 1) {
return false; // Full.
}
storage_[tail & mask_] = value;
tail_.store(tail + 1, std::memory_order_release);
return true;
}
bool try_push(T&& value) {
const auto tail = tail_.load(std::memory_order_relaxed);
const auto head = head_.load(std::memory_order_acquire);
if (tail - head == capacity_ - 1) {
return false;
}
storage_[tail & mask_] = std::move(value);
tail_.store(tail + 1, std::memory_order_release);
return true;
}
bool try_pop(T& result) {
const auto head = head_.load(std::memory_order_relaxed);
const auto tail = tail_.load(std::memory_order_acquire);
if (head == tail) {
return false; // Empty.
}
result = std::move(storage_[head & mask_]);
head_.store(head + 1, std::memory_order_release);
return true;
}
private:
static std::size_t next_power_of_two(std::size_t value) {
std::size_t result = 1;
while (result < value) {
result <<= 1;
}
return result;
}
const std::size_t capacity_;
const std::size_t mask_;
std::unique_ptr<T[]> storage_;
alignas(64) std::atomic<std::size_t> head_{0};
alignas(64) std::atomic<std::size_t> tail_{0};
};The queue has an effective capacity one less than its internal storage because an empty queue is represented by equal counters and a full queue is represented by a one-slot distance. The subtraction is safe while the counters do not overflow the representable range during the queue’s lifetime. For extremely long-running systems, use an unsigned counter type and document the intended wraparound assumptions, or adopt a sequence-number design.
Performance choices that actually matter
- Power-of-two capacity:
index & mask_avoids modulo division. Choose capacity from measured burst size, not habit. - False sharing: separate
headandtailonto different cache lines.64is common, butstd::hardware_destructive_interference_sizeis preferable where supported. - Allocation: allocate storage before entering the hot path. Avoid logging, exceptions, and dynamic allocation inside
try_pushandtry_pop. - Payload size: copying large objects can dominate synchronization costs. Store compact handles or pointers only when ownership and reclamation are explicit.
- Polling policy: a busy-spin can provide low latency but consumes a core. Use bounded spinning followed by
std::this_thread::yield, a notification mechanism, or a blocking queue when latency permits. - NUMA placement: pin producer and consumer deliberately on multi-socket servers. A queue crossing NUMA nodes may lose the benefit of a lock-free design.
For AI serving and data pipelines, the queue is one part of a larger latency budget. Pair measurements with LLM application performance monitoring in India practices: record enqueue delay, queue depth, drops, processing time, and end-to-end latency rather than reporting only operations per second.
Full, empty, and backpressure policy
Returning false is a policy boundary, not an error by itself. When full, the caller might drop the newest item, overwrite the oldest item, retry briefly, batch work, or slow the producer. The correct choice depends on semantics: dropping telemetry may be acceptable; dropping an order, payment event, or safety signal is not.
Expose counters for rejected pushes, empty polls, maximum observed depth, and processing age. These metrics make capacity decisions evidence-based. In a production service, connect them to alerts and load tests; the same disciplined approach used in full-stack AI engineering best practices applies to queue boundaries and failure handling.
Testing and verification
Test both functional behavior and the memory model:
- Run long producer-consumer stress tests with randomized pauses and different payload patterns.
- Compile with sanitizers, including
-fsanitize=threadwhere supported, and treat reports as defects. - Test on x86-64 and ARM64; do not assume x86 behaviour validates weaker architectures.
- Add assertions for the SPSC ownership rule in debug builds and document thread ownership in the API.
- Benchmark steady state, bursts, full-queue behaviour, empty polling, and realistic payload sizes.
- Use release compiler settings and inspect generated code only after correctness is established.
A mutex queue remains a valid baseline. Compare it against this ring buffer under the same thread affinity, batch size, workload, and scheduling conditions. If your application is an AI or industrial pipeline, queue tests should sit beside system-level tests such as those used for multi-agent AI manufacturing workflows, where uneven stage latency often creates bursts.
When not to build one yourself
Choose a standard concurrent queue, a well-reviewed library, or a blocking channel when you need MPMC semantics, dynamic capacity, safe object reclamation, condition-variable waiting, or a simpler maintenance burden. A lock-free queue is valuable when its bounded, non-blocking contract is a deliberate requirement—not because “lock-free” sounds faster.
For Indian engineering teams building low-latency infrastructure, the strongest implementation is the smallest one whose ownership rules, memory ordering, capacity policy, and measurements are explicit. Start with SPSC, verify it on every target architecture, and only move to MPMC after profiling proves that the simpler design is insufficient.