Hi Gregory,
Thanks for sharing the working branch.
Some of us at Samsung are testing the series (both v4 & v5) on a
compression capable CXL expander and are observing encouraging results.
Good things first! We feel that the isolation using private nodes and
the capability selection using NODE_PRIVATE_CAP_* work functionally well
for the compressed memory use cases.
On improvement points, we observed much higher 'page_faults',
'allocation_stalls' etc. which contributed to the higher tail latencies
in some tests. I hope these are already part of the optimization plans.
The higher latencies might have resulted from the write_protection
applied to the private node, I believe.
On the tools side, we used MLC and Taobench for the tests.
Note that our intention for this phase of experiments was to test only
the 'private node' layer based isolation, and not the cram and below
layers for the compressed memory. Therefore we did not enable/utilize
the compression capability in the hardware for these set of experiments
(enabling compression would require the cram level ballooning/memory
shrinking and cxl level interfacing driver to the compression device).
That would be different set of experiments where we would be testing the
cram balloon shrinkers and our alternate algorithms (upstream targeted).
So cram and below layers were used only for enumeration of private node
regions for these tests and not for the run-time memory shrinkers.
======= Test Methodology ==============
We explored these 3 cases:
1) (Baseline – existing CXL infra): DRAM 16GB + CXL 32GB. Here we
used the existing CXL driver infra to enable the device. No private node
is involved. And compression is disabled on device.
2) (private node in default write protected path): DRAM 16GB + CXL
Private Node 32GB. Here private node infra is used to enable the device.
And compression is disabled on device.
3) (private node with no write protection): DRAM 16GB + CXL Private
Node 32GB. Here private node infra is used to enable the device without
the write protection/fencing enabled. Compression is disabled on device.
We could not complete these runs as they resulted in kernel panics.
Guess the code path is not stable yet for this (We had hoped that case 1
and case 3 results would be similar). We also observed some unmovable
page warnings logs in 'dmesg' before crash (might be related to the
panic). Adding the log snippets at the end.
========= Results summary ===================
- TaoBench: ‘private node’ showed much higher latencies. We believe this
could be due to the migration back into the DRAM for overwrites (due to
write-protection) based on the kernel VM stats.
- MLC: For lower inject delay, the private node shows better latency;
but when inject delay goes high, the baseline is better.
Overall kernel VM stats suggested much higher ‘page faults, ‘allocation
stalls’ etc. in ‘private node’ case which is corroborated with the results.
========== Setup ======================
CPU: 144 Cores (single socket), DRAM: 16GB, CXL: 32GB,
Kernel: 7.2.0-rc2+ (Private Node), 6.18.32 (Baseline)
Private node opt CAPs:
#define CRAM_NP_CAPS (NODE_PRIVATE_CAP_RECLAIM |
NODE_PRIVATE_CAP_HOTUNPLUG | \
NODE_PRIVATE_CAP_DEMOTION |
NODE_PRIVATE_CAP_USER_NUMA | \
NODE_PRIVATE_CAP_NUMA_BALANCING |
NODE_PRIVATE_POLICY_WRITE_FENCE)
Enabled 'numa balancing' and 'demotion'.
$ echo 1 > /proc/sys/kernel/numa_balancing
$ echo true > /sys/kernel/mm/numa/demotion_enabled
========== TaoBench Experiments ====================
CMD:
$ python3 benchpress_cli.py run tao_bench_standalone -i '{"bind_mem": 0,
"bind_cpu": 0, "memsize": 32, "test_time": 300, "warmup_time": 0,
"clients_per_thread": 100, "set_get_ratio": "1:9"}'
1. (baseline) DRAM 16GB + CXL 32GB ==>
$ grep -A8 "^ALL STATS" benchmark_metrics_27cfff0e/client_0.log
===================================================
Type Avg. Latency p50 p95 p99
----------------------------------------------------
Sets 2.75215 2.73500 4.57500 5.79100
Gets 2.84236 2.81500 4.76700 5.95100
2. (private node in default write protected path)
DRAM 16GB + CXL Private Node(uncompressed) 32GB ===>
$ grep -A8 "^ALL STATS" benchmark_metrics_5181fda5/client_0.log
===================================================
Type Avg. Latency p50 p95 p99
----------------------------------------------------
Sets 5.58395 4.73500 12.35100 21.75900
Gets 4.83344 4.19100 10.43100 17.91900
Vmstat metrics comparison of both cases for Taobench
($cat /proc/vmstat):
-------------------------------------------------------------------
Metric | Baseline | CXL Private Node | Ratio
-------------------------------------------------------------------
pgfault 37,121,808 171,512,616 4.6x
pgactivate 5,360,928 161,947,594 30.2x
pgreuse 4,346,396 407,397 10.7x less
pgrefill 88,775,854 436,486,518 4.9x
pgdemote_kswapd 5,933,439 129,586,750 21.8x
pgdemote_direct 24,837 35,505,114 1429x
allocstall_normal 469 153,600 327x
kswapd_low_wmark_hit_quickly 671 27,821 41.5x
pageoutrun 1,333 29,217 21.9x
numa_pte_updates 24,508,077 0 na
numa_hint_faults 21,152,264 0 na
numa_pages_migrated 3,330,897 0 na
numa_hit 46,667,202 353,812,703 7.6x
---------- Takeaways -----
Private (CRAM) node performance is lower compared to a normal CXL memory
allocation path. VM stats shows more page_faults and allocation_stalls
on the private node case. Could it be the allocator waits for migration
path to demote pages to private node? Or the actual hot pages in dram
got demoted to private node to make space during overwrites? We will try
further analysis on this.
========== MLC Experiments ====================
CMD:
$./mlc --loaded_latency -j0 -c0 -b1g -k1-15 -W5 -r
Experiment Results:
1. (Baseline) DRAM 16GB + CXL 32GB
Inject Latency Bandwidth
Delay (ns) MB/sec
================
00000 686.18 40137.5
00002 686.69 40124.0
00008 709.74 40362.7
00015 711.75 40350.0
00050 710.72 40357.4
00100 712.22 40316.0
00200 695.29 40026.5
00300 602.91 38845.3
00400 482.28 36240.4
00500 410.03 33606.6
00700 275.13 26352.5
01000 241.82 18851.7
01300 230.91 14695.8
01700 215.27 11403.2
02500 217.76 7900.9
03500 194.42 5787.0
05000 192.17 4166.4
09000 190.36 2475.0
20000 188.09 1312.1
2. (private node in default write protected path)
DRAM 16GB + CXL Private Node(Uncompressed) 32GB
Inject Latency Bandwidth
Delay (ns) MB/sec
================
00000 320.09 12186.2
00002 543.42 22726.1
00008 482.11 28385.2
00015 488.26 26304.6
00050 485.19 26898.4
00100 477.65 24668.9
00200 483.68 21273.8
00300 486.44 17655.9
00400 480.41 16470.0
00500 481.30 13589.9
00700 493.84 9960.8
01000 488.66 8267.7
01300 491.31 6433.6
01700 476.97 6211.2
02500 472.97 5037.0
03500 466.05 3533.8
05000 446.44 2783.1
09000 417.03 1877.5
20000 387.22 1035.3
Vmstat metrics comparison of both cases for MLC
($cat /proc/vmstat):
-------------------------------------------------------------------
Metric | Baseline | CXL Private Node | Ratio
-------------------------------------------------------------------
pgfault 10,287,329 46540,782 4.6x
pgactivate 1141468 65,098,272 57x
pgrefill 260,699 47,592,751 182.5x
pgdemote_kswapd 1,085,783 26,063,754 24x
pgdemote_direct 4,562,539 15,809,485 3.5x
pgmigrate_fail 41,929 2,763,725 65.9x
allocstall_normal 161 425 2.6x
kswapd_low_wmark_hit_quickly 47 763 16.2x
pageoutrun 55 936 17x
Takeaway: Our MLC tests show a trade-off between the two configurations.
During low inject delay, the private Node is faster. But when inject
delay goes high, the baseline shows better latency. Could it be the TLB
cache effects?
--------------------------------------------------------------------
Case 3 dmesg log snippet:
[ 82.525741] page: refcount:1 mapcount:0 mapping:0000000000000000
index:0x0 pfn:0x5f5800
[ 82.525753] flags:
0x17ffffc0002000(reserved|node=0|zone=2|lastcpupid=0x1fffff)
[ 82.525762] raw: 0017ffffc0002000 ffefc85717d60008 ffefc85717d60008
0000000000000000
[ 82.525764] raw: 0000000000000000 0000000000000000 00000001ffffffff
0000000000000000
[ 82.525766] page dumped because: unmovable page
..............
We will update on the compression enabled experiments further.
~ Arun
On 23-07-2026 10:23 pm, Gregory Price wrote:
> On Thu, Jul 23, 2026 at 02:08:31PM +0530, Arun George/Arun George wrote:
>> On 21-07-2026 01:03 am, Gregory Price wrote:
>> We intend to test this series on a compression capable CXL expander
>> hardware. Since compressed ram example (cram) is not part of this
>> series, how do you suggest to do that? Do you have a version of cram
>> module compatible with this series?
>>
>> ~Arun
>
> Hi Arun,
> > I plan to RFC the new setup for compressed ram this a bit later, but I
> will share my working branch with you for testing.
>
> ~Gregory