Minisforum Says Disable Four of the MS-R1’s 12 CPU Cores – Is Linux Scheduling Them Properly?

What happens under background load?

So far the CP8180 has presented Linux with an unusual scheduling problem and the kernel has handled it extremely well. A single CPU-heavy task stays on the fastest cores, eight demanding tasks occupy the eight Cortex-A720 cores without touching the Cortex-A520s, and a ninth task is the point at which one of the smaller cores is brought into service.

Those tests give every process the same priority, though, whereas a real desktop or workstation workload is more complicated. Some work matters more than other work, so I wanted to see what happens when a normal-priority CPU-intensive process appears while the machine is already busy with low-priority background jobs.

I started eight continuously running Python processes at nice level 19. Linux initially placed all eight on the eight Cortex-A720 cores. I then launched a normal-priority single-threaded sysbench CPU test without setting any CPU affinity, and this time the results were noticeably less impressive.

Test Events/sec
No background load, average 1063.08
8 low-priority jobs, unpinned run 1 898.03
8 low-priority jobs, unpinned run 2 1008.77
8 low-priority jobs, unpinned run 3 854.37
8 low-priority jobs, unpinned average 920.39
8 low-priority jobs, forced to CPU 0 1108.29 average

The automatically scheduled foreground workload averages 920.39 events per second, around 13.4% below the unloaded baseline. By contrast, explicitly forcing the same foreground workload onto CPU 0 produces 1107.36, 1108.69 and 1108.81 events per second. The fastest core is therefore still capable of essentially full performance even while all eight background jobs continue to run.

The reason for the poorer automatic results becomes clearer when we look at where Linux places the foreground process. In all three runs it spends most of its sampled time on CPU 10, one of the 2.5 GHz Cortex-A720 cores, but it also spends some time on the much slower Cortex-A520 cores.

Run Foreground placement Result
1 80.4% CPU 10, 19.6% CPU 2 898.03 events/sec
2 98.2% CPU 10, 1.8% CPU 3 1008.77 events/sec
3 73.2% CPU 10, 26.8% CPU 2 854.37 events/sec

The relationship is hard to miss. The more time the normal-priority workload spends on an A520, the lower its performance.

What happens to the background work?

I repeated the test while sampling both the foreground process and all eight low-priority jobs. Before sysbench starts, the eight background jobs occupy the eight A720 cores:

CPU 0
CPU 1
CPU 6
CPU 7
CPU 8
CPU 9
CPU 10
CPU 11

Once the normal-priority process appears, Linux does recognise that some rearrangement is needed. One of the low-priority jobs spends part of its time on CPU 2, an A520. Across all background samples, 94.5% remain on A720 cores and 5.5% are on an A520. That is sensible as far as it goes because the scheduler is willing to move low-priority work away from a fast core to make room for something more important.

The problem is that it does not do so consistently enough. The normal-priority sysbench process itself still spends 12% of its sampled time on CPU 5, another A520:

CPU 5: 12.0%
CPU 6: 88.0%

That run produces only 884.81 events per second, around 16.8% below the unloaded baseline and roughly 20% below the performance available when I explicitly force the same foreground workload onto CPU 0.

This is the first test where I think Linux’s scheduling behaviour is clearly less than ideal. The scheduler understands the topology, knows that the A520 cores have much lower capacity and even begins shifting low-priority work to an A520 when a normal-priority process arrives. What it does not do consistently is keep the more important workload on the faster cores while relegating the nice-19 work to the slower ones.

Should the four small cores be disabled?

After all of these tests, my answer is no, at least not on the grounds that Linux is unable to schedule this processor properly. For equal-priority CPU-heavy workloads, its behaviour on the MS-R1 is excellent.

The kernel sees five distinct CPU performance levels, and its scheduler capacity figures match measured performance almost uncannily well across the eight Cortex-A720 cores. A single demanding process runs on the fastest pair. Eight heavy processes occupy all eight A720s without touching an A520. Add a ninth and Linux starts using exactly one small core. Give it twelve CPU-heavy processes and all twelve cores are put to work.

The Cortex-A520 cores also provide useful extra throughput when enough parallel work exists. With eight threads, sysbench produces around 8,180 events per second. With twelve threads and all cores available, the best runs reach around 9,868 events per second, roughly 20.6% more. Disabling the four small cores would throw that additional processing capacity away.

There is an important qualification, however. Under mixed-priority load, Linux does not always make the decision I would prefer. A normal-priority CPU-heavy task can spend measurable time on an A520 while nice-19 background work remains on faster A720 cores, and explicitly placing the important task on a fast core recovers the lost performance. I therefore wouldn’t describe the scheduling as perfect.

Nor do these tests support the idea that the four Cortex-A520 cores should routinely be disabled. For most of the workloads I tested, Linux makes very good use of the CP8180’s unusual topology, while the smaller cores provide useful additional performance when sufficient parallel work is available.

Minisforum may have good reasons for recommending eight-core operation for particular workloads, especially software that is sensitive to synchronisation, inter-core communication or memory behaviour. Those are separate issues from the scheduler-placement tests I’ve carried out here, and I would not generalise these results to every possible workload.

For general Linux use, I’d leave all 12 cores enabled. The more interesting result is that the Linux scheduler already understands this unusual processor remarkably well. The weakness I found is not an inability to distinguish fast cores from slow ones. It appears in the more specific case where process priorities and heterogeneous cores interact under load.

That is a much narrower problem than simply saying that four of the MS-R1’s twelve cores should be switched off.

Pages in this article:
Page 1: Introduction and CPU Layout
Page 2: Per-Core Performance
Page 3: One CPU-Heavy Task
Page 4: The Eight-Task Test
Page 5: Explicit CPU Affinity
Page 6: Background Load and Conclusions


Complete list of articles in this series:

Minisforum MS-R1 ARM Mini Workstation
IntroductionIntroduction to the series and interrogation of the Mini Workstation
BenchmarksBenchmarking the Minisforum MS-R1 ARM Mini Workstation
PowerTesting and comparing the power consumption
BIOSExploring the BIOS
UbuntuTesting the Ubuntu 26.04 LTS image
SchedulingMinisforum Says Disable Four of the MS-R1’s 12 CPU Cores
Subscribe

Please read our Comment Policy before commenting.

Notify of
guest
0 Comments
Oldest
Newest Most Voted