Introduction
The throughput of many modern computing systems is increasingly constrained not by arithmetic throughput alone, but by the cost of moving data through memory and I/O hierarchies. Zero-copy design addresses part of this problem by eliminating redundant payload replication wherever underlying ownership, protection, and lifetime semantics permit.
Zero-copy is not a single mechanism, but a family of copy-elision techniques. In some systems, the same physical pages are mapped into multiple address spaces. In others, the kernel forwards page references without bouncing the payload through user space. In still others, a DMA-capable device transfers data directly between memory domains. The common architectural principle is that software increasingly moves descriptors, mappings, offsets, and ownership state rather than repeatedly moving the bytes themselves.
Broadly speaking, four architecturally different mechanisms constitute this copy-elision family:
- Virtual memory remapping: mmap() and shared memory avoid replication by exposing existing pages through virtual-memory mappings.
- Kernel-mediated forwarding: sendfile() and splice() remove user-space bounce buffers from kernel-mediated I/O.
- Kernel bypass: RDMA, DPDK, and SPDK move portions of the data path into DMA-capable devices or user-space drivers.
- Device and interconnect: Coherent interconnects such as CXL and NVLink-C2C reduce the need for explicit transfers by making remote or heterogeneous memory directly addressable under defined coherence rules.
The Data Movement Bottleneck
In conventional I/O pipelines, a payload gets routed through several independently managed memory domains before reaching its destination:
$$\text{Storage} \rightarrow \text{Kernel Page Cache} \rightarrow \text{User Buffer} \rightarrow \text{Kernel Network Buffers} \rightarrow \text{NIC}$$
Not every transition in this path is a CPU-mediated copy. Storage controllers and NICs normally transfer data using Direct Memory Access (DMA), while conventional read(), write(), and send() paths may introduce full payload copies specifically between kernel-managed and user-space buffers. If a payload of size S undergoes n complete memory-to-memory copies, and each copy reads S bytes from one location and writes S bytes to another, the incremental memory traffic attributable to those copies is approximately 2*n*S.
This expression models redundant copy traffic, not total I/O traffic. DMA transfers, cache writebacks, protocol processing, coherence traffic, and interconnect transfers also consume bandwidth. The architectural objective of zero-copy design is to eliminate an unnecessary 2S term each time data would otherwise be replicated merely to cross a software boundary.
Redundant copies consume memory bandwidth without transforming the payload. They may displace useful cache lines and consume cache-fill, writeback, and coherence resources. The actual effect depends on payload size and the memcpy implementation. Optimized implementations may use non-temporal stores for sufficiently large streaming writes to reduce cache pollution and avoid unnecessary write allocation, but these stores do not universally bypass the cache hierarchy; their behavior depends on memory type, cache residency, and processor implementation. Conventional I/O may also require repeated transitions between user mode and kernel mode. At high operation rates, system-call processing, security mitigations, descriptor management, and kernel-stack traversal can therefore generate substantial aggregate costs.
Memory-Mapped and Shared-Memory Architectures
Virtual Memory Integration with mmap()
Instead of issuing a read() that copies file contents into an application-owned buffer, mmap() can establish a file-backed region in the application’s virtual address space. The kernel records the mapping, but typically populates individual page-table entries lazily as pages are accessed. If the corresponding file page is already resident in the page cache, the resulting page fault maps that page into the process directly; if it is not resident, the fault may first trigger storage I/O and populate the page cache.
Once mapped and resident, subsequent accesses are ordinary memory references rather than repeated read() calls, followed by page-cache-to-user-buffer copies. The cost comes from virtual-memory machinery: page faults, page-table creation, TLB pressure, page-cache management, and potentially storage latency on first access. For workloads with large sequential transfers, a conventional buffered read can sometimes outperform access patterns that generate many page faults.
Shared-Memory Inter-Process Communication
Conventional buffered IPC through pipes or sockets commonly introduces two application-visible payload copies: the sender copies data into a kernel-managed buffer, and the receiver copies it out again. Shared memory removes those payload transfers from the steady-state data path by mapping a common region into multiple processes.
Instead of exchanging the payload itself, the processes exchange offsets, lengths, sequence numbers, ownership flags, or buffer descriptors, often through a ring buffer or queue. The producer writes the payload once; consumers operate directly on the shared region. The operating system establishes mappings, enforces protection, schedules processes, and handles page faults. Instead of copying, the focus shifts to synchronization and ownership. Processes must agree on when a buffer is valid, when it can be reused, and which participant owns it at each point in the protocol. This is a recurring pattern in zero-copy design: eliminating data movement usually increases metadata and ownership-state movement.
Kernel Bypass and Network Data-Plane Optimization
In-Kernel Forwarding
Linux can transfer data between file descriptors without bouncing it through an application-owned user-space buffer. With splice(), kernel-managed page references can be moved through pipe buffers without repeatedly copying the underlying page contents between kernel and user address spaces. A file-to-network path, therefore, becomes:
$$\text{Storage} \rightarrow \text{Page Cache} \rightarrow \text{Kernel Networking} \rightarrow \text{NIC}$$
instead of
$$\text{Storage} \rightarrow \text{Page Cache} \rightarrow \text{User Buffer} \rightarrow \text{Kernel Networking} \rightarrow \text{NIC}$$
The networking stack itself is not removed. Socket state, TCP processing, checksumming, congestion control, and NIC descriptors remain. Depending on the protocol, hardware support, and offload configuration, the kernel may still copy or linearize data on some paths.
Remote Direct Memory Access (RDMA)
RDMA changes the abstraction more radically by allowing a network interface to operate directly on registered memory regions. A one-sided RDMA WRITE conceptually follows:
$$\text{Local Memory} \rightarrow \text{Local RNIC} \rightarrow \text{Network} \rightarrow \text{Remote RNIC} \rightarrow \text{Remote Memory}$$
without executing a request handler on the remote CPU for each transfer. One-sided READ, WRITE, and atomic operations let the initiator access registered remote memory directly. Two-sided SEND/RECV operations require the receiving side to post receive work requests and therefore retain a stronger message-passing model. In conventional sockets, receiving software decides where incoming data lands. In one-sided RDMA, the adapter places data directly into an authorized remote region.
User-Space Polling: DPDK and SPDK
DPDK and SPDK remove substantial portions of network and storage I/O processing from the conventional kernel data path. DPDK poll-mode drivers access NIC descriptor rings directly from user space and normally poll them rather than wait for per-packet interrupts. SPDK applies the same philosophy to NVMe. This trades one resource for another: interrupt handling and kernel-stack traversal decrease, but high-performance deployments often dedicate one or more CPU cores to polling. DPDK and SPDK are therefore better understood as kernel-bypass architectures than as zero-copy mechanisms in the narrow sense; copy elimination is one part of a larger effort to remove unpredictability from the I/O path.
Control-Plane Batching io_uring
io_uring can be treated as a compositional I/O framework, where batching, asynchronous completion, registered buffers, polling, and zero-copy operations are optimizations combined based on workload type. Its central contribution is reducing I/O submission and completion overhead: applications communicate with the kernel through shared submission and completion queues, and multiple asynchronous operations can be queued together so that system-call overhead amortizes across many requests. Registered buffers reduce repeated buffer-registration work, but an ordinary fixed-buffer read does not become payload-zero-copy on its own. Specific mechanisms such as zero-copy send eliminate a payload copy only when the networking path supports it.
Heterogeneous Computing and Device-to-Device I/O
In AI training, scientific computing, and GPU-accelerated systems, unnecessary movement occurs when data destined for an accelerator is staged through host memory instead of moving directly from storage or the NIC to GPU memory.
GPUDirect Storage, GPUDirect RDMA, and PCIe Peer-to-Peer
GPUDirect Storage provides a direct DMA path between supported storage stacks and GPU memory, avoiding a CPU-side bounce buffer when topology and software support permit it. GPUDirect RDMA provides the analogous capability for compatible network and PCIe peer devices. These capabilities are strongly topology-dependent: PCIe root-complex placement, switch configuration, filesystem support, IOMMU configuration, and device capability determine whether the direct path is available; when it is not, implementations fall back to staging through pinned system memory.
Hardware-Coherent Interconnects: CXL and NVLink-C2C
Compute Express Link defines several protocol classes over a common interconnect: CXL.io provides PCIe-like discovery and configuration; CXL.cache lets a capable device coherently access host memory; CXL.mem lets a host processor access memory attached to a CXL device using load/store semantics. Not every CXL device implements every protocol class, so CXL does not give every processor, accelerator, and memory device one universal coherent address space. The supported semantics depend on device class and system architecture. NVLink-C2C provides a related, platform-specific model. In architectures such as NVIDIA Grace Hopper, the CPU and GPU participate in hardware-coherent memory access across the interconnect, which differs from treating every form of GPU-to-GPU NVLink as having identical coherence semantics.
The broader shift is architectural: some data movement problems can increasingly be transformed into memory-access problems. Instead of copying a buffer from one physical domain into another, a device may access the original memory under appropriate coherence, protection, and topology constraints.
Architectural Trade-Offs: The Cost of Zero-Copy
Eliminating a copy introduces other costs: synchronization, ownership tracking, page mapping, memory registration, locality problems, pinned memory, and sometimes, more complicated failure modes. For small payloads, a straightforward memcpy() may be cheaper than the machinery required to avoid it.
Synchronization and Ownership Costs
Shared-buffer architectures require explicit coordination of ownership, visibility, and lifetime, using atomics, locks, memory barriers, sequence counters, or queue protocols. Under contention, cache-coherence traffic and cache-line migration can consume more cycles than the copy that the design was meant to eliminate. A useful first-order model compares a fixed buffer-handoff cost against the time to copy a payload of size S at effective throughput B. The break-even payload size satisfies the following:
$$\frac{S^{*}}{B_{\text{copy}}} = C_{\text{handoff}} \quad \Longrightarrow \quad S^{*} = B_{\text{copy}} \cdot C_{\text{handoff}}$$
For S > S*, avoiding the copy becomes attractive under this simplified model. The threshold is workload-specific and depends on cache state, NUMA placement, synchronization strategy, buffer reuse, consumer count, and actual copy bandwidth. The model above assumes a single consumer. With k independent consumers, naive replication may require approximately k payload copies, whereas a shared-buffer design avoids those full replicas. The coordination cost is not constant with k. Each consumer still requires metadata, synchronization, reference tracking, and potentially additional coherence traffic. These costs are generally much smaller than replicating a large payload for every consumer, which is why zero-copy architectures are particularly attractive in fan-out workloads.
NUMA and Locality
Zero-copy can eliminate replication while accidentally worsening locality. On a NUMA system, a thread accessing pages allocated on another node incurs remote-access penalties when cache misses, or coherence events require data to traverse the inter-socket fabric; subsequent accesses that hit in a local cache need not pay the full remote-memory penalty. In some workloads, copying a working set locally once is therefore cheaper than repeatedly fetching remote cache lines or servicing cross-socket coherence traffic. A zero-copy design can still generate substantial traffic across cache-coherence fabrics, memory controllers, PCIe links, or NUMA interconnects. The correct optimization target is total movement across the expensive boundaries of the actual machine, not the number of calls to memcpy.
TLB Pressure and Huge Pages
Large or sparsely accessed mappings can create substantial address-translation overhead because translations are cached at page granularity and the TLB can cover only a finite working set. With conventional 4 KiB pages, a 2 MiB huge page covers 512 base pages (2 MiB/4 KiB), while a 1 GiB huge page covers 262,144 base pages (1 GiB/4 KiB). Huge pages can therefore increase TLB reach substantially for large contiguous working sets. This is why they are common in DPDK, SPDK, RDMA, and database systems, though the trade-offs include memory fragmentation, harder allocation, and coarser mapping granularity.
Lifecycle and Memory Pinning
DMA-capable devices require that the memory they address remain valid and correctly mapped for the duration of an operation. Traditional RDMA registration and many high-performance user-space I/O designs therefore rely on pinned memory, and long-lived pins reduce the memory available for normal reclamation while they remain active. At sufficient scale, memory registration becomes a system resource-management problem. Modern RDMA implementations can use on-demand paging and MMU-notifier mechanisms on supported hardware, allowing the HCA to obtain or invalidate mappings dynamically rather than pinning the entire region for its lifetime. The general constraint nevertheless remains: memory mappings and their associated access permissions must remain valid for every hardware operation that can address them.
Ownership, Protection, and Failure Isolation
Copying creates an implicit ownership boundary: once a consumer has its own copy, the producer can typically modify or release the original without affecting that consumer. Zero-copy removes that boundary. A producer must not reuse a buffer while a NIC, GPU, or another process can still reference it. Consumers must obey the appropriate memory-ordering rules; DMA-capable devices must be restricted to authorized regions; completion semantics must determine when ownership can safely change; and buffer exhaustion must propagate through backpressure. At scale, buffer ownership, IOMMU mappings, and completion handling become correctness properties. This is why copying is sometimes the correct design choice. A copy can purchase isolation, simpler ownership, shorter lifetimes, improved locality, or easier failure recovery. The question is whether the complexity required to avoid copying is worth paying.
Closing Comments
Zero-copy should not be treated as a doctrine against copying. A copy is justified when it creates a clean ownership boundary, improves NUMA locality, decouples a producer’s lifetime from its consumers’, simplifies failure recovery, or reduces the coordination required around shared state. The architectural problem arises when data is copied only because successive software layers were designed around independent buffers. In such cases, the copy purchases no useful semantic property; it merely converts memory bandwidth and cache capacity into overhead.
This distinction becomes increasingly important as storage, networking, accelerators, and memory interconnects continue to increase aggregate data rates. Redundant software-mediated movement can consume a substantial fraction of the available memory and interconnect budget before useful computation begins. Page reuse, shared mappings, DMA, device-accessible memory, coherent interconnects, and descriptor-driven I/O therefore represent more than isolated performance optimizations. Collectively, they move system architecture toward a model in which processors increasingly coordinate data movement rather than acting as mandatory intermediaries through which every payload must pass. The cost does not disappear; it moves into synchronization, ownership, coherence, address translation, topology, and lifecycle management.
The correct engineering objective is therefore not to minimize the number of copies in isolation, but to minimize unnecessary data movement across the expensive boundaries of the machine. That requires examining the full path through caches, DRAM, NUMA links, PCIe fabrics, device memory, kernel buffers, and network interfaces, while accounting for the correctness properties purchased by each transfer. A successful zero-copy design is one in which removing a copy reduces total movement and latency without making ownership, isolation, or recovery disproportionately more complex. The governing principle is simple: keep a copy only when it provides more value than the data movement it costs.
References
- Pai, V. S., Druschel, P., & Zwaenepoel, W. (1999). IO-Lite: A Unified I/O Buffering and Caching System. 3rd USENIX Symposium on Operating Systems Design and Implementation (OSDI). https://www.usenix.org/conference/osdi-99/io-lite-unified-io-buffering-and-caching-system
- Recio, R., Culley, P., Garcia, D., Hilland, J., & Metzler, B. (2007). RFC 5040: A Remote Direct Memory Access Protocol Specification. IETF. https://www.rfc-editor.org/rfc/rfc5040
- DPDK Project. Data Plane Development Kit Documentation. https://doc.dpdk.org
- SPDK Project. Storage Performance Development Kit Documentation. https://spdk.io
- NVIDIA Corporation. GPUDirect Storage Overview Guide. https://docs.nvidia.com/gpudirect-storage/overview-guide/index.html
- NVIDIA Corporation. NVIDIA Grace Hopper Superchip Architecture Whitepaper. https://developer.nvidia.com/blog/nvidia-grace-hopper-superchip-architecture-in-depth/
- Basu, A., Gandhi, J., Chang, J., Hill, M. D., & Swift, M. M. (2013). Efficient Virtual Memory for Big Memory Servers. 40th International Symposium on Computer Architecture (ISCA). https://doi.org/10.1145/2485922.2485943
- Lameter, C. (2013). NUMA (Non-Uniform Memory Access): An Overview. ACM Queue, 11(7). https://queue.acm.org/detail.cfm?id=2513149
- Linux documentation, e.g., splice(), io_uring, pin_user_pages(), and others.
Disclosure: The author acknowledges the use of generative AI to refine portions of the paper. The author retains ownership of the overall content, including the technical material, core arguments, and analysis.