Python DISCUSSION

Why don't Python threads make my CPU-bound loop faster on a multi-core machine?

Started by atultorane global interpreter lockPython threadingmultiprocessingCPU-bound tasksconcurrent.futures
5 replies 248 views 6 participants
Latest activity · 30 Sep 2026

Why don't Python threads make my CPU-bound loop faster on a multi-core machine?

atultorane Python Forum
#1

I have a pure-Python function that runs an FIR filter over a block of samples and takes about 8 seconds for one file. To process four files I started four threading.Thread objects on a quad-core machine and expected about 8 seconds in total. It still takes around 32 seconds, and the system monitor shows only one core busy.

Yet the same threading approach clearly speeds up another script of mine that collects data from several instruments over TCP. Why do threads help in one case and not in the other, and what should I use for the CPU-heavy case?

Community replies 5

Re: Why don't Python threads make my CPU-bound loop faster on a multi-core machine?

#2

Standard CPython has a global interpreter lock (GIL): only the thread that holds it may execute Python bytecode. The lock protects the interpreter's internal state, such as reference counts. Your four filter threads therefore take turns on one core. The interpreter asks the running thread to give up the lock every 5 ms by default (sys.getswitchinterval() returns 0.005), so you get concurrency without parallelism.

Four jobs of 8 s each on what is effectively one core is the 32 s you measured.

Re: Why don't Python threads make my CPU-bound loop faster on a multi-core machine?

#3

The network script speeds up because a thread releases the GIL while it is blocked on I/O: waiting on a socket, reading a file, sleeping, waiting for a serial port. While one thread waits for an instrument, the others run. If each of four instruments takes 2 s to answer, the sequential version needs 8 s and the threaded version about 2 s, because the time is spent waiting, not computing.

The rule of thumb is threads or asyncio for I/O-bound work and processes for CPU-bound work. concurrent.futures.ThreadPoolExecutor is the convenient way to run the I/O case and collect the results.

Re: Why don't Python threads make my CPU-bound loop faster on a multi-core machine?

#4

For the CPU-bound case use separate processes, each with its own interpreter and its own GIL: with concurrent.futures.ProcessPoolExecutor() as ex: results = list(ex.map(filter_file, paths)). Four files on four cores then take roughly 8 s plus start-up.

The costs are different from threads. Arguments and return values are pickled and copied between processes, so pass file names rather than huge arrays where you can and return compact results. The worker function must be importable, which means defined at module top level and not a lambda, and the pool must be created under if __name__ == "__main__": on platforms that spawn. Starting a process is far more expensive than starting a thread, so batch small tasks.

Re: Why don't Python threads make my CPU-bound loop faster on a multi-core machine?

#5

Before parallelising, check whether the inner loop should be in Python at all. NumPy runs such loops in compiled code: numpy.convolve(x, taps) or scipy.signal.lfilter(taps, 1.0, x) replaces the per-sample Python loop and is commonly tens of times faster or more, which may remove the need for parallelism altogether.

Many of those library routines also release the GIL while they run, so threads that call them can use several cores. Numba or Cython are the next step for loops that cannot be expressed as array operations. Profile first with cProfile or time.perf_counter() so that you optimise the part that is actually slow.

Re: Why don't Python threads make my CPU-bound loop faster on a multi-core machine?

#6

Two qualifications. The GIL does not make your own code thread-safe: counter += 1 is several bytecodes, and a thread switch between the load and the store can lose an update. Shared state still needs a threading.Lock or a queue.Queue.

And the picture is changing. Python 3.13 introduced an optional free-threaded build without the GIL, separate from the default build. Extension modules have to support it and single-threaded code runs somewhat slower in it, so check your dependencies before relying on it. The default build behaves as described above.

TEP COMMUNITY