Fundamentals 25 min read

Inside read(): How Data Moves from Disk to Memory via Kernel & DMA

This article dissects the Linux read() system call, revealing how user-space applications, the kernel, and hardware controllers collaborate via DMA to move data from disk into memory, covering parameters, return values, blocking modes, file-type differences, and the copy_to_user overhead versus mmap zero-copy.

Deepin Linux
Deepin Linux
Deepin Linux
Inside read(): How Data Moves from Disk to Memory via Kernel & DMA

Misconception About read()

Many beginners assume read() is a self-contained operation where the application directly pulls data into memory. In reality, a user-space program cannot access hardware such as disks or network cards directly. A single read() system call involves collaboration among three actors: the application, the kernel, and the hardware controller. The hardware controller (via DMA) is the entity that actually writes data into memory; the kernel orchestrates scheduling and resource management; the application merely initiates the request and waits for the result.

What Is the read() System Call?

In the OS, read() retrieves data from an open file descriptor. Whether reading a file, network data, or terminal input, read() is the standard interface. Its core job is to copy data from kernel space into a user-space buffer.

Memory is divided into kernel space (high privilege, direct hardware access) and user space (lower privilege, where applications run). When an application needs data, it invokes read() to ask the kernel; the kernel interacts with the hardware, fetches the data, and places it into the user-space buffer. For example, a Python read()</call> ultimately relies on the <code>read() system call to move file data from kernel space to user space.

C Prototype and Parameters

#include <unistd.h>
ssize_t read(int fd, void *buf, size_t count);

fd (file descriptor) : Integer returned by open() identifying the file or device. It acts as an identity token so the kernel knows which object to read. Works for regular files, character devices, pipes, sockets, etc.

buf (user-space buffer pointer) : Pointer to a pre-allocated memory region in user space (e.g., char buffer[1024];). The kernel copies data here. Invalid or too-small buffers cause failures.

count (maximum bytes to read) : Upper limit on bytes to transfer. The call may return fewer bytes due to EOF, file type, or available data. Example: reading a 50-byte file with count=100 returns at most 50 bytes.

Return Values

> 0 : Success; value equals bytes actually read. May be less than count (e.g., socket or pipe reads).

= 0 : End of file (EOF). No more data.

< 0 : Error; check errno for specifics (e.g., EBADF for bad fd, EIO for I/O error).

Data Read Flow: From Disk to Memory

1. System Call Trigger

Calling read() in user code (glibc wrapper) executes a trap instruction ( syscall on x86_64, int 0x80 on 32-bit). CPU switches to kernel mode, saves user context, and jumps to the kernel's sys_read handler.

Analogy: An employee (user program) cannot grab company resources (hardware) directly; they submit a request form (system call) for the manager (kernel) to approve and process.
#include <unistd.h>
#include <fcntl.h>
#include <stdio.h>

int main(void) {
    int fd = open("test.txt", O_RDONLY);
    if (fd < 0) {
        perror("open failed");
        return 1;
    }
    char buf[64];
    ssize_t ret = read(fd, buf, sizeof(buf)-1);
    if (ret < 0) {
        perror("read failed");
        close(fd);
        return 1;
    }
    buf[ret] = '\0';
    printf("读到内容:%s
", buf);
    close(fd);
    return 0;
}
open

, read, close are glibc wrappers running in user space.

At read(fd, buf, sizeof(buf)-1), glibc loads syscall number and arguments into registers, then executes syscall.

CPU saves user context, enters kernel, finds sys_read via syscall number.

Key point: User programs cannot call kernel functions directly; they must use CPU trap instructions.

2. Kernel Work

Inside sys_read:

ssize_t sys_read(int fd, void __user *buf, size_t count)
{
    struct file *file = fget(fd);
    if (!file)
        return -EBADF;
    loff_t pos = file->f_pos;
    ssize_t ret = file->f_op->read(file, buf, count, &pos);
    file->f_pos = pos;
    fput(file);
    return ret;
}
fget(fd)

retrieves the struct file from the process's fd table. file->f_op->read is a function pointer; different file types (regular file, socket, pipe) provide different implementations. For disk files, it invokes block-device read logic.

Kernel translates file offset to disk sector numbers, builds an I/O request, and passes it to the block driver.

Driver commands the disk hardware; disk performs seek and rotation, reads sectors into the disk controller's hardware buffer.

No data is copied to user space yet. This step only initiates block I/O.

3. Data Transfer via DMA and Copy to User

Once the disk controller holds the data, DMA (Direct Memory Access) moves it from the controller buffer to the kernel's page cache (kernel buffer) without CPU involvement. The kernel then copies data from the page cache to the user buffer ( buf) using copy_to_user — a CPU copy.

Buffers involved:

Disk controller buffer: temporary staging to absorb speed mismatch.

Kernel page cache: caches hot data for future reads.

User buffer: application's destination for processing.

Example: video playback — controller buffer holds raw chunks, page cache retains recent frames, user buffer feeds the decoder.

Regular read path: Disk → DMA → kernel page cache → CPU copy ( copy_to_user) → user buffer.

4. Zero-Copy Alternative: mmap

#include <unistd.h>
#include <fcntl.h>
#include <sys/mman.h>
#include <stdio.h>

int main(void) {
    int fd = open("test.txt", O_RDONLY);
    if (fd < 0) {
        perror("open");
        return 1;
    }
    char *p = mmap(NULL, 64, PROT_READ, MAP_SHARED, fd, 0);
    if (p == MAP_FAILED) {
        perror("mmap");
        close(fd);
        return 1;
    }
    printf("mmap读到内容:%s
", p);
    munmap(p, 64);
    close(fd);
    return 0;
}
mmap

also triggers DMA to fill the page cache.

It maps the same physical pages into the process's virtual address space, eliminating the copy_to_user step (zero-copy).

Regular read always incurs one copy_to_user; kernel implementation roughly:

unsigned long copy_to_user(void __user *to, const void *from, unsigned long n);
to

: user buffer address (from read 's buf). from: kernel page cache address.

CPU performs this copy, which is the main overhead of traditional read.

Factors Influencing read() Behavior

1. File Type Differences

Regular files (text, images, binaries): predictable; reads up to count if data remains.

Device files (serial ports, disks): behavior depends on hardware. Serial read blocks until data arrives; disk read may wait for device readiness.

Pipes (inter-process communication): blocks when empty; waits for writer. Limited buffer size.

Sockets : TCP read returns partial stream data; may need multiple calls for a full message. Network latency/congestion causes delays. UDP read returns whole datagrams (or partial if buffer too small).

2. Blocking vs Non-Blocking Mode

Blocking (default): read sleeps until data arrives or error. Suitable for sequential, latency-sensitive tasks (e.g., video streaming).

Non-blocking ( O_NONBLOCK): read returns immediately with EWOULDBLOCK / EAGAIN if no data. Enables high-concurrency servers to poll many descriptors without stalling.

3. Kernel Scheduling and System Load

High-priority processes can preempt the reading process, delaying its CPU time. High system load (CPU, memory, disk I/O contention) increases latency — e.g., heavy disk writes starve read bandwidth.

Generalization and Clarification

The same core logic applies to network cards, pipes, etc.: kernel issues commands, hardware uses DMA to write memory, kernel forwards data to user. Legacy non-DMA devices required CPU-driven copying, but modern hardware and OSes universally use DMA.

Colloquially people say "the kernel reads data". The precise technical statement: the kernel schedules hardware to write data into memory, then delivers it to the user program.

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

system programmingDMAmmapzero-copyfile I/OLinux kernelblocking I/Oread system call
Deepin Linux
Written by

Deepin Linux

Research areas: Windows & Linux platforms, C/C++ backend development, embedded systems and Linux kernel, etc.

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.