CPU affinity, NUMA placement, and isolation

Note

System tuning tools are either about better understanding the system’s performance, or applying such knowledge to improve it. See our common system tuning thoughts for some generally applicable considerations about that.

Controlling where a workload runs cuts contention for central processing unit (CPU) resources and reduces memory latency. Without care, manual restrictions can also leave cores idle or overload a single memory node. This guide covers CPU affinity, non-uniform memory access (NUMA) policies, CPU isolation, and how they interact.

Manual placement works best when you already know how a workload uses threads, memory, and devices together. Always measure the default scheduler against a representative workload first, since tuning for lower latency can reduce overall throughput or starve background services.

CPU and memory topology

A NUMA system groups CPUs and memory into nodes. A CPU can access memory on another node, but that access generally has higher latency than access to local memory. Nodes do not necessarily correspond one-to-one with processor sockets: a socket can contain several NUMA nodes.

lscpu(1) maps logical CPU numbers to nodes, sockets, and cores. numactl(8) reports the CPUs, memory capacity, free memory, and relative distances for each NUMA node. Ubuntu provides lscpu and taskset(1) in the util-linux package; numactl and numastat(8) are in the numactl package.

To inspect the topology visible to the kernel:

$
lscpu -e=CPU,NODE,SOCKET,CORE,ONLINE
CPU NODE SOCKET CORE ONLINE
  0    0      0    0    yes
  1    0      0    1    yes
  2    1      0    2    yes
  3    1      0    3    yes
$
numactl --hardware
available: 2 nodes (0-1)
node 0 cpus: 0 1
node 0 size: 1930 MB
node 0 free: 1614 MB
node 1 cpus: 2 3
node 1 size: 1965 MB
node 1 free: 1710 MB
node distances:
node     0    1
   0:   10   20
   1:   20   10

CPU lists in taskset and numactl use logical CPU numbers, not core or socket numbers. With simultaneous multithreading (SMT), several logical CPUs share a physical core. Rows with the same socket and core identify these siblings. Consecutive CPU numbers are not necessarily separate cores or members of the same node.

In numactl --hardware, lower distance numbers indicate closer nodes. In this output, the relative access latency from a core to memory on the same node is 10, and the relative access latency to memory on the other node is 20, which means that it would take twice as long. These numbers reflect relative distance rather than measured latency. If the command reports only node 0, there are no remote nodes to balance against, but CPU affinity still applies. Keep in mind that virtual machines (VMs) report the virtual topology presented by their hypervisor, which does not always mirror the host hardware (and the host would need to pin virtual CPUs to a topology that the guests can rely on).

For cache sharing and the relationship between CPUs and attached devices, see hardware topology with hwloc.

CPU affinity

CPU affinity is the set of logical CPUs on which a thread may run. The scheduler can move the thread within that set. Pinning a thread to a single CPU removes that choice; allowing several CPUs retains some scheduling flexibility.

The Linux scheduler already balances tasks for CPU locality. Setting explicit affinity helps when you need deterministic placement or want cooperating threads to share cache levels. However, it hurts performance when the chosen CPUs saturate while other cores sit idle. Affinity does not reserve CPU cycles, prevent other tasks from using those cores, or migrate existing memory.

The following examples assume that CPUs 2 and 3 are online and available to your workload. In the topology shown above, those two CPUs belong to node 1.

In the examples below, replace ./workload with your executable and its arguments, and 1864 with the process ID (PID) of a running workload.

To start a program pinned to CPUs 2 and 3:

$
taskset -c 2-3 ./workload

Spawned threads and child processes inherit this affinity mask, though applications can alter their own masks later.

When targeting a running process ID, taskset affects only the main thread unless you pass --all-tasks (-a) to include every thread. To query every thread in a process:

$
taskset --all-tasks --pid --cpu-list 1864
pid 1864's current affinity list: 0,1
pid 1866's current affinity list: 0,1
pid 1867's current affinity list: 0,1
pid 1868's current affinity list: 0,1

To move all threads to CPUs 2 and 3:

$
taskset --all-tasks --pid --cpu-list 2-3 1864
pid 1864's current affinity list: 0,1
pid 1864's new affinity list: 2,3
pid 1866's current affinity list: 0,1
pid 1866's new affinity list: 2,3
pid 1867's current affinity list: 0,1
pid 1867's new affinity list: 2,3
pid 1868's current affinity list: 0,1
pid 1868's new affinity list: 2,3

Take note of original affinity masks before altering them. Applying a single mask across an entire process overwrites any internal thread placement configured by the application. You can adjust your own processes, but changing tasks owned by another user requires root privileges. A process also cannot use CPUs disallowed by its control group (cgroup), regardless of requested affinity.

Affinity dictates where code runs, not how the scheduler shares time among tasks. Tools like nice(1) and renice(1) adjust relative scheduling priorities instead. Lower nice values grant higher priority, while higher values ensure background jobs yield to time-sensitive work.

NUMA memory policies

CPU placement and memory placement are separate decisions. Under the default memory policy, the kernel generally allocates a new private memory page on the node of the CPU that first triggers a page fault for it. This is often called first-touch allocation. If one thread initializes memory before worker threads start on other nodes, those workers can end up accessing remote memory.

Automatic NUMA balancing, when enabled, can move eligible memory pages and tasks to improve locality. Explicit CPU and memory restrictions limit the placements it can choose. Changing affinity with taskset does not itself migrate pages that a process has already allocated.

numactl launches a program with a CPU binding, a memory policy, or both. --cpunodebind=0 selects the available CPUs belonging to node 0, whereas --physcpubind=2-3 selects logical CPUs 2 and 3. Despite its name, --physcpubind does not select whole physical cores. Neither option sets a memory policy.

For memory placement, numactl provides options such as --preferred and --membind. While --preferred expresses a soft preference that falls back to other nodes under memory pressure, --membind strictly confines allocations to the specified nodes.

The main memory-policy choices are:

Option

Allocation behavior

Trade-off

--localalloc

Prefer the node of the CPU making the allocation, with fallback to other nodes.

Locality follows the allocating thread, which may move between nodes.

--preferred=0

Prefer node 0, with fallback to other nodes.

Keeps a preferred location without making it a strict requirement.

--membind=0

Allocate only from node 0.

Allocation can fail when that node lacks memory, even if other nodes have free memory.

--interleave=all

Distribute allocations round-robin across the allowed nodes, with fallback under memory pressure.

Can spread memory-bandwidth demand but increases remote accesses for threads running on one node.

To run a workload on node 0 and prefer memory allocations from that same node:

$
numactl --cpunodebind=0 --preferred=0 -- ./workload

To pin strictly to CPUs 2 and 3 and require memory allocation exclusively from node 1:

$
numactl --physcpubind=2-3 --membind=1 -- ./workload

This policy applies to the launched process and its children. It governs future allocations rather than migrating existing shared or file-backed pages already resident in memory. NUMA-aware software can also configure policies for specific memory ranges programmatically using numa(3) from the libnuma-dev package. As with CPU affinity, cgroup memory limits constrain the nodes available to the program.

numactl --show reports the policy of numactl itself, not an external running process. To inspect where a running workload holds resident memory across nodes, run numastat -p:

$
numastat -p 1864
Per-node process memory usage (in MBs) for PID 1864 (workload)
                           Node 0          Node 1           Total
                  --------------- --------------- ---------------
Huge                         0.00            0.00            0.00
Heap                         0.00            0.00            0.00
Stack                        0.02            0.00            0.02
Private                    578.89          190.78          769.67
----------------  --------------- --------------- ---------------
Total                      578.91          190.78          769.69

This workload started on node 0 and was moved to the CPUs of node 1 by the earlier taskset command. Most of its memory remains on node 0, which is what makes those accesses remote.

The columns show memory resident on each node, in megabytes. They do not measure how often the CPUs access remote memory. Similarly, the numa_hit and numa_miss counters from numastat without arguments describe allocation decisions, not processor cache hits and misses. Memory on several nodes is not by itself evidence of a problem.

Placement for systemd services

Changes made with taskset disappear when a service restarts. For persistent placement across restarts, systemd provides CPUAffinity=, NUMAPolicy=, and NUMAMask= in systemd.exec(5).

[Service]
CPUAffinity=2-3
NUMAPolicy=preferred
NUMAMask=1
$
systemctl show -p CPUAffinity -p NUMAPolicy -p NUMAMask workload.service
CPUAffinity=2-3
NUMAPolicy=preferred
NUMAMask=1

These directives set initial process policies rather than dedicated CPU reservations. To enforce boundaries through cgroup cpuset controls, use AllowedCPUs= from systemd.resource-control(5). Constraining one service does not keep unconstrained units off the same CPUs.

CPU isolation

Isolation shields designated CPUs from general system work and scheduler balancing. This prevents context switches from interrupting latency-sensitive tasks, but leaves fewer CPUs for the rest of the system. Pinning alone does not isolate cores: it only dictates where a workload runs, without keeping other tasks away.

The isolcpus kernel command-line parameter removes CPUs from normal scheduler load balancing. For example:

isolcpus=domain,2-3

This removes logical CPUs 2 and 3 from standard scheduling domains. Tasks will not run on these cores unless assigned there explicitly using affinity masks or cpusets. Because the scheduler will not balance load across isolated cores automatically, you must position individual worker threads deliberately. Note that any other process with matching affinity can still run on these CPUs.

Domain isolation configured via isolcpus is locked at boot. Updating or clearing the isolated list requires a reboot. If you need dynamic isolation at runtime without rebooting, use cgroup v2 cpuset partitions instead. Limiting cpuset.cpus alone is not equivalent to an isolated scheduling partition; consult the kernel’s CPU isolation documentation for partition setup requirements.

To inspect the active kernel parameters and isolated CPUs:

$
cat /proc/cmdline
BOOT_IMAGE=/vmlinuz-7.0.0-30-generic root=UUID=90d99162-58b5-439f-b459-a3d205abce11 ro console=tty1 console=ttyS0 isolcpus=domain,2-3
$
cat /sys/devices/system/cpu/isolated
2-3

An empty isolated file means no CPUs have this form of boot-time domain isolation. It does not report every form of isolation, such as runtime cpuset partitions.

isolcpus=domain only affects thread balancing. It does not disable device interrupts, suppress kernel background threads, or stop the scheduler clock tick. Hardware IRQ steering (irqbalance(1)) must be configured separately. You can combine isolation with nohz_full to disable the timer tick when a single task runs on a core, and isolcpus=managed_irq to divert managed device interrupts, though neither offers a complete guarantee against system jitter.

Warning

Reserve enough housekeeping CPUs for operating-system tasks, device drivers, and SSH access. Remember SMT siblings: an unconstrained task running on a sibling thread will still contend for execution pipelines on the same core. Always verify boot isolation on a test host with remote console access before modifying production bootloaders.

Placement for virtual machines and packet processing

For virtual machines managed by Libvirt, pinning and memory policies belong in the host’s domain XML. A guest’s virtual CPU (vCPU) index does not map directly to host CPU numbers: setting affinity inside the VM only constrains guest processes to guest vCPUs, leaving the hypervisor’s underlying execution threads floating across the host.

Libvirt separates these controls:

Domain setting

Purpose

cputune/vcpupin

Map each guest vCPU to an allowed set of host logical CPUs.

cputune/emulatorpin

Place emulator threads separately from vCPU threads.

cputune/iothreadpin

Place configured input/output (I/O) threads.

numatune/memory

Set the host NUMA memory policy and node set for the guest.

cpu/numa

Describe the NUMA topology visible inside the guest.

Guest-visible topology alone does not establish host placement. A practical configuration pins a guest’s vCPU threads and memory to matching host nodes, while leaving enough host capacity for emulator and I/O work. Strict memory binding can prevent a guest from starting if the chosen nodes run out of memory. Host CPU and node numbers also need adjustment when migrating guests between physical hosts. See Libvirt for guest-management instructions and the upstream references below for domain configuration details.

Checking the effect

Verify affinity masks and node-resident memory with taskset --all-tasks --pid --cpu-list and numastat -p. To monitor CPU utilization and voluntary or involuntary context switches per thread, use pidstat(1) from the sysstat package:

$
pidstat -t -u -w -p 1864 1 1
Linux 7.0.0-30-generic (numa2) 	09/15/26 	_x86_64_	(4 CPU)

18:49:51      UID      TGID       TID    %usr %system  %guest   %wait    %CPU   CPU  Command
18:49:52        0      1864         -  200.00    0.00    0.00    0.00  200.00     1  workload
18:49:52        0         -      1864    0.00    0.00    0.00    0.00    0.00     1  |__workload
18:49:52        0         -      1866   69.00    0.00    0.00   31.00   69.00     2  |__workload
18:49:52        0         -      1867   65.00    0.00    0.00   36.00   65.00     2  |__workload
18:49:52        0         -      1868   68.00    0.00    0.00   33.00   68.00     3  |__workload

18:49:51      UID      TGID       TID   cswch/s nvcswch/s  Command
18:49:52        0      1864         -      0.00      0.00  workload
18:49:52        0         -      1864      0.00      0.00  |__workload
18:49:52        0         -      1866      0.00    115.00  |__workload
18:49:52        0         -      1867      0.00    129.00  |__workload
18:49:52        0         -      1868      0.00    115.00  |__workload

This command takes a 1-second sample across all threads of PID 1864 (omitting final summary averages). Note that the CPU column shows which core ran the task during that snapshot, not its allowed mask, and context-switch counts do not measure cross-node migrations. Always measure application throughput and latency under identical load before and after tuning: verifying an affinity mask only confirms that placement took effect, not that it improved performance.

Further reading