I have a pure-Python function that runs an FIR filter over a block of samples and takes about 8 seconds for one file. To process four files I started four threading.Thread objects on a quad-core machine and expected about 8 seconds in total. It still takes around 32 seconds, and the system monitor shows only one core busy.
Yet the same threading approach clearly speeds up another script of mine that collects data from several instruments over TCP. Why do threads help in one case and not in the other, and what should I use for the CPU-heavy case?




