2026-09-20 02:03:20
Integrated GPUs have become a crucial component in recent laptop chips, thanks to a push for better graphics performance in ultraportable devices. AMD and Intel have both launched impressive iGPU implementations in recent years. Qualcomm’s laptop push needs a powerful iGPU in this landscape, and that’s where the Snapdragon X2 Elite’s Adreno X2-90 enters the picture.
Adreno X2 expands on the foundation set by the Snapdragon X Elite’s Adreno X1, but brings an enhanced architecture that aims for higher throughput. Shader Processors (SPs) act as the basic building block for Qualcomm’s Adreno GPUs. Each SP contains a pair of Micro Shader Processor Texture Processor (uSPTPs), which have execution units, register files, and private texture caches. At a higher level, Qualcomm groups pairs of SPs into slices. A fully enabled X2-90 part has eight SPs grouped into four slices, compared to six SPs in three slices on Adreno X1. Qualcomm also scaled up clock speeds from 1.5 to 1.85 GHz. A larger, higher clocked GPU will put more pressure on the memory subsystem, so Qualcomm scaled up cache sizes and increased bandwidth throughout the memory hierarchy.
With all these changes, Adreno X2 has nearly twice the theoretical compute throughput of its predecessor. It’s dwarfed by AMD’s highest end iGPU offering, which brings in more of everything. AMD’s more standard “Strix Point” chip may be a better comparison, but I don’t have that on hand, so I’ll be doing performance comparisons across a smattering of interesting devices. A special thanks goes out to ASUS for supplying the Zenbook A16 for testing.
Qualcomm’s Adreno line traditionally emphasizes high throughput for basic FP32 and FP16 operations. Adreno X2’s uSPTPs continue that emphasis, with 128 FP32 lanes that appear to be arranged in two execution unit partitions. FP16 benefits from double rate execution like on many modern GPUs.
Other instruction categories don’t necessarily enjoy the same high throughput, but Adreno X2 improves over its predecessor in several areas. INT32 adds appear to execute at full rate, though it’s hard to approach theoretical throughput on them compared to FP32 operations. 32-bit integer multiplies execute at half rate, which compares well to Intel and AMD’s GPU architectures, and aligns with Nvidia. Integer performance overall is a huge improvement over the rather poor situation on Adreno X1, where even basic integer operations executed at less than half rate. Better INT32 units help 64-bit integer performance too, since GPUs carry those out using multiple INT32 operations. Like previous Adreno GPUs, Adreno X2 lacks FP64 support.
Testing with Nemes’s Vulkan-based suite shows substantial compute improvements compared to Adreno X1. Adreno X2 also compares well to older iGPUs. AMD’s Radeon 780M can achieve higher FP32 throughput if AMD can use VOPD dual-operation instructions or wave64 mode, but Adreno X2 isn’t far off even in that case. However, Adreno X2 seems to need a lot of work in flight to approach theoretical compute throughput, even compared to larger GPUs like AMD’s Strix Halo.
OpenCL experimentation with varying dispatch sizes shows that Adreno X2 can reach beyond TFLOPS with FP32 adds, or 6 TFLOPS with multiply-adds. Adreno X2 just needs roughly 8x more parallelism to reach its maximum throughput. Further testing with OpenCL with large dispatches shows Adreno X2 getting close to theoretical throughput in floating point operations. However, INT32 throughput barely improves over Vulkan.
Adreno X2 continues to lack hardware fused multiply-add (FMA) capability like its predecessor. FMA avoids intermediate rounding between the multiply and add, reducing numerical error. Asking for that using OpenCL’s fma() built-in function results in very poor throughput. AMD, Intel, and Nvidia’s GPUs all have native FMA hardware and can provide FMA behavior without throughput loss.
Each Adreno X2 uSPTP has 128 KB of register file capacity, down from 192 KB on Adreno X1. Register file design tends to be difficult because they have to provide immense bandwidth to feed a GPU’s wide vector units. Smaller register files might have let Qualcomm increase clock speeds while keeping power and area under control. Adreno X2 also ditches Adreno X1’s wave128 mode, and only supports wave64. Smaller wave sizes can let register file capacity go further because each thread can allocate more registers before impacting theoretical occupancy.
With the register file capacity change, Adreno X2’s occupancy should behave like Adreno 730’s.
Qualcomm didn’t disclose Adreno X2’s theoretical occupancy, but I tried to guess by running a DRAM latency test with different dispatch sizes, and the same access pattern used across all workitems. If the GPU can keep all workitems in flight, latency shouldn’t be much higher than with a single workitem. However, running out of wave slots would force some workitems to wait until another finishes and its wave slot frees up, increasing observed runtime.
I suspect each Adreno X2 uSPTP partition has somewhere between 6 to 8 wave slots, with 8 being a more likely figure. Latency starts increasing with a 12288 workitem dispatch (6x wave64 per uSPTP partition), and more than doubles when going from 16384 to 18432 workitems.
Qualcomm also did not disclose instruction cache capacity, but basic testing with increasingly large loops shows that throughput from a single thread stays high with up to 8192 FP32 add statements in the loop. Absolute throughput figures suggest Adreno X2 can’t dual-issue FP32 adds from a single thread, possibly because each wave64 thread is locked to a single uSPTP partition. Throughput remains reasonable at larger instruction footprints.
Adreno X2 gets a beefed up memory hierarchy compared to its predecessor, but continues to inherit Adreno-style memory subsystem features that make it unusual next to other modern GPUs. uSPTPs continue to have texture caches that don’t serve compute global memory accesses. These texture caches are now 4 KB in size, but that’s still very small. The smallest texture cache I’ve seen outside of Qualcomm’s Adreno GPUs is the 8 KB one on AMD’s Terascale.

Global memory accesses go straight to Adreno X2’s 128 KB cluster caches. Cluster cache accesses have slightly higher latency than texture cache hits, but remains reasonable in an absolute sense. However, that means Qualcomm has higher minimum latency for global memory accesses than AMD. The gap is especially wide if AMD can carry out scalar optimizations. Qualcomm does not have a separate scalar cache hierarchy. Even if I use the constant memory type, that appears to be backed by the same cluster cache. The cluster cache does fine in bandwidth terms, and can deliver 64B/cycle to each uSPTP.
A 2 MB L2 cache sits after the cluster caches, and matches AMD’s 2 MB L2 on Strix Halo and Strix Point. L2 latency is reasonably good, and better than AMD’s. L2 capacity is small next to what Intel has been doing, but Qualcomm does have a system level cache to catch L2 misses, making them less reliant on L2 capacity. AMD’s Strix Halo uses a similar strategy.
L2 bandwidth appears to be a bit above 1 TB/s, or just over 32B/cycle per uSPTP. Like upper level caches on many GPUs, Adreno X2’s cluster cache is either write-through or read-only. Writes go to L2, which services them at about half the read rate.
Atomic operations are often handled at L2 because it’s the first cache level shared across the GPU. Adreno X2 provides decent throughput for INT32 atomic adds. Counting each add as two 4B accesses (read-modify-write) gives 632.64 GB/s of effective bandwidth.
Testing cross-thread latency using atomic_cmpxchg on global memory shows excellent performance from Adreno X2, and a big improvement over Adreno X1. “Core to core” latency on Adreno X2 is on par with AMD’s latest GPU architectures, and better than on Intel Meteor Lake’s iGPU.
The Snapdragon X2 Elite’s 8 MB system level cache (SLC) has rather high latency at just over 200 ns, or roughly 150 ns above L2 hit latency. SLC bandwidth appears to be around 380 GB/s, which is comfortably higher than DRAM bandwidth. AMD’s high end Radeon 8060S has both higher bandwidth and lower latency to its 32 MB SLC.
For DRAM, Qualcomm has a rather aggressive setup with a 192-bit LPDDR5X-9523 configuration. I could get just over 150 GB/s of DRAM bandwidth from the GPU. That’s better than what a typical 128-bit LPDDR5X setup can achieve, but sits under 70% of theoretical. A potential positive there is that GPU bandwidth demands can’t pressure CPU-side accesses to the same degree as other iGPU implementations, where the iGPU can gobble up a larger fraction of theoretical DRAM bandwidth.
Adreno X2’s DRAM latency situation is a bit strange. Latency plateaus at around 274 ns with moderate test sizes, but heads for the hills when test coverage exceeds 64 MB. Testing with a 128B stride doubles apparent cluster cache capacity, suggesting Adreno X2’s cluster cache has 64B lines. Apparent L2 capacity remains unchanged until stride length exceeds 4 KB, after which larger strides push out both the L2 and 64 MB inflection points. My interpretation is that Adreno X2 has a virtually addressed L2, and uses 4 KB pages for address translation with a 16K entry TLB placed after the L2. High latencies at large test sizes imply very high L2 TLB miss penalties. Raw DRAM access latency, excluding address translation penalties, is likely around that reasonable 274 ns figure. However, applications making accesses with poor locality over a large memory footprint may experience much higher latency.
Qualcomm has a long history of embracing tiled rendering, and Adreno X2 doubles down on that with a massive 21 MB block of on-chip storage called “Adreno High Performance Memory”. This HPM fills the same function as GMEM (Graphics Memory) in prior Adreno GPUs, and holds render targets to contain intermediate tile state while a tile is being rasterized. For scale, Adreno X1 had 3 MB of GMEM, while the Snapdragon 8+ Gen 1’s Adreno 730 had 2 MB of GMEM. Qualcomm says they sized HPM to handle QHD+ frames, and the math checks out. A 1440P frame with 32 bits per pixel (RGBA, 8 bits per channel) would occupy 14.7 MB. ASUS equipped the Zenbook A16 with a 2880x1800 OLED screen. Doing the same math would give 20.7 MB per frame, which should barely fit within HPM.
Rather than doing tiled rendering in the traditional sense with perhaps a dozen or so tiles per frame, Qualcomm seems to be throwing a curveball by trying to make the whole screen one tile. Tiled rendering comes with potential inefficiencies. For example, a triangle that overlaps two tiles will have to go through the rasterizer twice, creating extra rasterizer work. As one tile finishes, wave slots might start freeing up across the shader array. The GPU might have extra pixel work available, but has to hold it back because the whole point of tiled rendering is to serialize the rasterization process at the tile level. Those potential inefficiencies suddenly go away if the entire screen is one tile. At the same time, Qualcomm would still get the advantage of being able to contain intermediate tile state in on-chip memory. A large cache could do this naturally of course, but a cache requires extra power and area for tag and state arrays.

HPM is an excellent move if applications stick to the conventional rasterization API, but I’m worried that HPM’s applicability might be limited for general purpose compute. While rasterization remains the bedrock of PC gaming and likely will for the foreseeable future, the modern rasterization process sometimes includes a compute component. Newer rendering engines often have features that leverage compute, like Unreal Engine 5’s Nanite. HPM isn’t useless in these scenarios, because Qualcomm can allocate some HPM for use as a software managed scratchpad. Then, HPM can take on the same role as AMD’s Local Data Share or Nvidia’s Shared Memory.

I tried to see how many workgroups the GPU could run simultaneously with different amounts of OpenCL local memory allocated. From those experiments, Adreno X2 only appears capable of allocating 1 MB of HPM as local memory. Local memory allocations that aren’t a power of two bring that down further. 1 MB of local memory isn’t bad for a GPU of Adreno X2’s size. Adreno X1 could only allocate 384 KB of local memory across the GPU. AMD’s Radeon 8060S has 2.5 MB of Local Data Share (LDS) across the GPU, which makes sense considering its larger size. Still, having 21 MB of storage and only being able to use a small fraction of it for compute doesn’t feel great.
Local memory latency improves over Adreno X1, which is impressive because HPM is seven times larger than Adreno X1’s GMEM. However, AMD, Intel, and Nvidia’s latest GPUs still provide better latency to local memory. It’s easier to do that when accesses don’t have to traverse a cross-GPU interconnect.
Thread-to-thread latency through local memory also improves on Adreno X2. But again, AMD’s core-private LDS on their latest GPUs is faster.
Atomic adds on local memory usually enjoy higher throughput than ones through global memory, because most GPUs have dedicated atomic ALUs implemented at per-core local memory instances. That’s not the case for Adreno X2, which is able to sustain 32 INT32 atomic adds per cycle across the entire GPU, or 4 per cycle on a SP basis. For comparison, AMD’s RDNA3.5 can do 32 INT32 atomic adds per cycle at each WGP’s Local Data Share. Intel’s Xe-LPG in Meteor Lake also provides high throughput, with 16 INT32 atomic adds per Xe Core cycle.
Local memory bandwidth testing shows that HPM can deliver as much bandwidth as the GPU’s texture or cluster caches.
Software conventionally uses copy APIs to get data to and from the GPU. With OpenCL’s clEnqueueReadBuffer and clEnqueueWriteBuffer, copying data to the GPU is much faster than going the other way around. Prior Adreno GPUs acted like this too. Curiously, pinning the workload to different cores changes measured bandwidth, but only affects movement from CPU to GPU memory. Qualcomm might be using a CPU core to do the transfer, rather than offloading data movement to DMA engines like many other GPUs.
Newer APIs can enable zero-copy behavior. OpenCL’s Shared Virtual Memory (SVM) API is one example, and also maintains the same virtual addresses to let CPU and GPU code use the same pointers. Adreno X2 advertises atomics support with SVM, meaning that the CPU and GPU can theoretically exchange data while a GPU kernel is running. However, I wasn’t able to get the GPU and CPU to see each other’s writes with atomic operations.
I was able to test with fine grained sharing, which lets the CPU see the GPU’s writes after the GPU kernel finishes. Making writes visible between the CPU and GPU is fast enough to show that Qualcomm isn’t copying the entire 256 MB test buffer under the hood. However, kernel launch and synchronization overhead is higher than on other platforms especially if the test is run from an E-Core.
FluidX3D simulates fluid behavior using the lattice Boltzmann method (LBM), and has historically been a challenging workload for Adreno GPUs. FluidX3D uses FMA operations in its default configuration, resulting in poor performance on Adreno compared to just about anything else. FluidX3D can be modified to use a “legacy_fma” setting, which uses multiply-add operations instead of fused multiply-add. But even with that modification, Adreno X2 still struggles. It regresses compared to Adreno X1, and is worlds away from last generation iGPUs.

Memory capacity and memory bandwidth tend to be persistent constraints for FluidX3D. Therefore, FluidX3D can be built to use FP16S and FP16C modes, which use 16-bit floating point formats for storage. These 16-bit storage formats are converted back to FP32 for computation to minimize precision loss, effectively increasing compute pressure to reduce pressure on the memory subsystem. FP16S uses the standard FP16 format, letting format conversions take advantage of fast-path hardware on most GPUs. FP16C is a custom format and needs software format conversion, but provides better precision for the typical values that FluidX3D sees.
Neither of these modes help Adreno X2, indicating that compute is the limitation rather than bandwidth. FP16C puts more pressure on compute because of the extra instructions required for format conversion, and actually regresses compared to the full FP32 mode. I’ve included the GTX 1050 3 GB as a comparison because it’s the kind of older, lower-midrange discrete GPU that newer iGPUs should easily beat. Adreno X2 has more than twice as much compute throughput, memory bandwidth, and cache capacity than the GTX 1050 3 GB, so losing to it is disappointing.
Folding at Home runs protein folding simulations. The project has a rather dated benchmark called FAHBench, which I’m running under binary translation. I briefly tried to build an arm64 version, but it seems like the time investment to do so isn’t worth it. Adreno X2 turns in a passable result on the default dhfr (dihydrofolate reductase) workitem, though losing to the GTX 1050 3 GB is again not a good look.
FAHBench also includes a larger nav (voltage gated sodium channel) workload that tends to place more pressure on the memory subsystem, compared to the more compute bound dhfr workload. I had several GPUs fail the accuracy check in this workload. I’m showing their performance anyway because their calculation results only differ in the 5th or 6th significant digits (for the GTX 1050 and Adreno X2 respectively). They weren’t off by orders of magnitude, indicating that they were making a good faith effort at the work involved.
Nvidia’s GTX 1050 3 GB is 36% faster than the Adreno X2 with the nav workitem, while it was just 13.7% faster with dhfr. Adreno X2 seems to struggle with memory bound workloads. I wonder if its relatively low TLB coverage (for a GPU) combined with high TLB miss costs hold it back in less cache-friendly compute workloads.
3DMark’s Wild Life Extreme should be a familiar workload for Qualcomm, because it’s a Vulkan-based test that targets mobile devices. Adreno X2 performs very well in Wild Life Extreme, pulling comfortably ahead of Meteor Lake’s iGPU and taking up a reasonable position relative to AMD’s largest iGPU. Fire Strike Extreme is a heavy DirectX11 workload, which I had to run under binary translation. I’m only presenting the graphics score for that test. Fire Strike Extreme is a tougher challenge for Adreno X2, but it still manages a comfortable lead over Intel’s older iGPU.
Adreno X2 also does well in Cyberpunk 2077’s built-in benchmark, where it provides twice the performance of Intel’s Meteor Lake iGPU and Nvidia’s old GTX 1050 3 GB.
Cyberpunk 2077 performance is vastly better on Adreno X2 compared to the prior Adreno X1, which achieved 24.3 FPS at 1080P with the low preset. Adreno X2 easily beats this at the medium preset, which I’m using because Cyberpunk 2077’s CPU-side load isn’t trivial like in 3DMark’s graphics test. High framerates could bring binary translation overhead into the picture, and I wanted to avoid that. Overall Qualcomm seems to be doing well in rasterization, in contrast to the poor performance in GPGPU workloads.
3DMark’s Solar Bay test showcases raytraced reflections in a simple scene. It’s a lighter workload meant to fit within the capabilities of typical cell phones while still using raytraced effects. Adreno X2 turns in an exceptional performance here. It outperforms Meteor Lake’s iGPU and AMD’s Radeon 780M by more than a factor of two, and isn’t too far off AMD’s Radeon 8060S. It’s also a huge improvement over Adreno X1.
Adreno X2 brings DirectX Raytracing (DXR) 1.1 to Qualcomm’s GPU line, letting a wide range of PC games tap into Adreno’s hardware raytracing accelerators. 3DMark has a more demanding Port Royal test that uses DXR, which targeted high end gaming PCs from a few generations ago. Compared to Solar Bay, Port Royal makes heavier use of compute shaders and places less emphasis on the conventional rasterization pipeline.

Curiously, Adreno X2 achieves the same score as Meteor Lake’s iGPU in Port Royal. Qualcomm also loses ground relative to Strix Halo’s massive iGPU, and only achieves 50.8% of the Radeon 8060S’s score compared to 74.5% in Solar Bay. It’s also fun to note that a score of 1678 corresponds to 7.77 average FPS, which isn’t ideal for most games.
Cyberpunk 2077 is something of a raytracing showcase title. I’m testing that by taking the medium preset above, enabling raytraced reflections, and setting the raytraced lighting slider to medium. Bringing those raytraced effects into the picture destroys performance on all three iGPUs. None of them deliver a comfortably playable framerate, but there is a concerning shift in placement. Adreno X2 no longer ties or beats Meteor Lake’s iGPU, and now loses by a large margin.
There’s certainly potential in Qualcomm’s raytracing implementation, as 3DMark’s tests show. Qualcomm’s slides indicate that each uSPTP has a raytracing unit (RTU) with vaguely similar capabilities to Intel’s RTAs. Both implementations accelerate BVH traversal and intersection testing, and both embed a L0 cache into the raytracing unit.
I wonder what holds Qualcomm back in Cyberpunk 2077 with raytracing enabled. I doubt it’s a CPU-side bottleneck, because the GPU stays at or near 100% throughout the benchmark run. Maybe it’s because Cyberpunk 2077 has a huge BVH that covers the entire city. When profiling with AMD’s tools on RDNA2, Cyberpunk 2077’s BVH had over 59K nodes, compared with 11.6K nodes in 3DMark Port Royal. Or, maybe Cyberpunk 2077’s raytraced effects put a higher burden on the compute pipelines. I haven’t done the profiling necessary to figure out why Adreno X2 behaves the way it does.
Adreno X2 is a step forward for Qualcomm’s GPU line in many respects. General performance benefits from higher clocks, more SPs, and a beefier memory subsystem. Specific classes of workloads can stand to benefit more thanks to Adreno X2’s huge HPM block and improved raytracing units. Adreno X2 subjectively feels good enough to not hold Qualcomm’s laptop push back. I tried a few games that I’ve been playing recently, and performance felt appropriate for a mainstream iGPU.
However, Qualcomm’s latest iGPU still contains a lot of glass jaw performance cases for a modern GPU. Adreno X2 inherits Adreno X1’s tendency to fall over in GPGPU compute. Architectural improvements are certainly evident in compute microbenchmarks, but those observations don’t translate into good performance in the compute workloads I tried. It feels like Qualcomm took a specific set of graphics workloads and optimized for them, while AMD, Intel, and Nvidia set out to build more general purpose designs with robust performance characteristics across a broader range of use cases.
For now, I think Qualcomm has bigger fish to fry. Binary translation overhead was the biggest issue I faced when trying to game on the device, because it’s easy to get bound by CPU performance when running x86-64 binaries. When playing back an Age of Empires II game through CaptureAge for example, the Snapdragon X2 Elite Extreme couldn’t provide enough CPU performance to maintain realtime playback, while Strix Halo had enough spare CPU power to fast forward. But in the years to come, Qualcomm’s GPU will have a tough matchup on their hands. Today’s iGPU landscape is characterized by brutal competition. AMD’s Radeon 8060S and Intel’s Arc B390 can approach the performance of lower-midrange mobile discrete GPUs in both graphics and compute. The thin and light laptop trend will likely continue, and drive AMD and Intel to create ever more ambitious designs. I look forward to seeing what Qualcomm has in store for their future designs, and hope they’ll challenge the best that AMD and Intel have to offer.
2026-09-10 17:46:18
The PC market is one of the toughest areas for a CPU designer to compete in. Consumers in the PC segment expect high performance across a wide range of applications, and famously expect their devices to reach beyond a constrained, curated set of use cases. Microsoft’s own Surface RT prominently failed more than a decade ago because it could not run the programs that PC users expect to just work on Windows. Recent 64-bit Arm cores from Qualcomm and others are fast enough to satisfy the performance component of the PC equation, but software compatibility presents a tougher roadblock. PC software is traditionally built for x86-64. Getting developers to offer 64-bit Arm (aarch64) versions of their programs is a slow and gradual process. Some programs may never get aarch64 ports because they’re no longer under active development and were only distributed in binary form. Arm’s PC market chances therefore ride on Microsoft’s efforts to ensure x86-64 binaries can run seamlessly on aarch64 hosts.
Windows 11’s newest binary translator, dubbed Prism, enables this by translating x86-64 instructions to aarch64 ones. Binary translation is challenging because x86-64 instructions sometimes don’t map to aarch64 ones in a straightforward manner. Translation also has to be fast to minimize program launch delays and avoid consuming excessive CPU time for doing the translation. All of this means binary translation comes with a performance penalty compared to running equivalent native code.
Here, I’m checking out what that penalty looks like using Geekbench 7. Geekbench 7 comes in both aarch64 and x86-64 versions, providing an opportunity to compare performance with binary translation against native aarch64 performance. John Poole (Founder of Primate Labs, Creator of Geekbench) has kindly provided a pro key, which lets me profile individual workloads. Of course, one benchmark suite can’t reflect behavior across a wide range of applications, but I consider this a good starting point. As for hardware, I’m testing with the Snapdragon X2 Elite Extreme X2E-96-100 in an Asus Zenbook A16 laptop that was sampled by Asus as well as Arm instances available on Microsoft Azure.
Like Geekbench 6, Geekbench 7 has a number of workloads that leverage vector extensions when available. The suite is slanted towards high IPC, compute-bound workloads. Running workloads through Intel’s Software Development Emulator (SDE) set to expose Haswell’s feature set shows more than half the workloads using AVX and 256-bit vector width. I’m running each workload with 200 iterations to ensure the vast majority of counted instructions come from the workload rather than Geekbench’s test harness.
With 200 iterations, each workload executes roughly a few trillion instructions. Performance counter data from various aarch64 cores show roughly similar instruction counts when executing Geekbench 7’s native aarch64 binary, with occasional exceptions. Qualcomm’s cores curiously report higher retired instruction counts than Neoverse N1 and N2 in many tests, which suggests there may be some inaccuracy in hardware performance monitoring.
Executing Geekbench 7’s x86-64 version under binary translation results in dramatically inflated instruction counts across the board, except in PDF Viewer. A Geekbench 7 workload will execute roughly twice as many aarch64 instructions as x86-64 ones when going through binary translation, as a rule of thumb.
Microsoft hints that binary translation doesn’t work the same way across all aarch64 CPUs. Different CPUs support different instruction set extensions, some of which may let Prism more closely map x86-64 instructions to aarch64 ones. An aarch64 CPU can also hypothetically provide guarantees beyond what the ISA requires, like maintaining store ordering even though aarch64 lets one core observe stores from another out of program order.
Prism is optimized and tuned specifically for Qualcomm Snapdragon processors. Some performance features within Prism require hardware features only available in the Snapdragon X series, but Prism is available for all supported Windows 11 on Arm devices with Windows 11 24H2.
If Prism is optimized specifically for Qualcomm’s cores, it doesn’t make a difference from the instruction count side. Ampere Altra’s Neoverse N1 is the oldest core I’ve tested on, and performance counters show very similar executed instruction counts.
Similar instruction counts can of course hide large differences. Performance is ultimately what matters in the end, and that’s affected both by the core architecture and the nature of the generated instructions.
Geekbench 7’s score shows harsh penalties from binary translation. Every core, including Qualcomm’s, suffers several generations worth of performance loss. However, the cores in the Snapdragon X2 Elite have much higher baseline performance than Neoverse N1 and N2. The Snapdragon X2 Elite has a 6-core cluster of “performance” cores, and two 6-core cluster of “prime” cores. “Performance” cores are 6-wide and run at 3.6 GHz. Since they fill the same role as efficiency cores in other chips, I’ll call them E-Cores for simplicity. “Prime” cores are 9-wide and run at 5 GHz. They aim for maximum performance, so I’ll call them P-Cores.
Qualcomm’s E-Cores manage to outperform Neoverse N1 even when taking a binary translation penalty. The same applies to the P-Cores against Neoverse N2. There’s really no arguing with much wider cores running at very high clock speeds, even with binary translation penalties in play.
Qualcomm’s cores also take a slightly lower penalty from binary translation. It’s hard to tell whether this is down to Qualcomm-specific optimizations in Prism, or whether it’s down to Qualcomm’s cores having much higher throughput. Instruction count increases from binary translation remind me of playing with Claude’s C Compiler (CCC). CCC’s case involved creating a lot of extra instructions off the critical path, which an out-of-order CPU can often absorb. Recent high performance cores usually leave most of their core width unused, so extra core width can mitigate higher instruction counts to some extent. But I’m not sure that’s the case here, because Qualcomm’s 4-wide E-Core comes off with a lighter penalty than 5-wide Neoverse N2.
Individual workloads suffer to varying degrees when run under binary translation. Navigation is a low IPC workload that’s heavily bound by branch prediction and backend memory latency, and takes a comparatively low penalty when run under binary translation. At the other end of the spectrum, Video Player is a well vectorized workload that takes advantage of AVX instructions. Score differences in that test are absolutely massive. Photo Editor and Photo Library are in a similar situation.
Workload characteristics under binary translation stay mostly the same on Arm’s Neoverse N1 and N2.
Performance counters show Qualcomm’s cores putting a dent in binary translation overhead by dipping into unused core width. However, the IPC increase is minor compared to the instruction count overhead, which explains the large binary translation penalty. Snapdragon X2 Elite Extreme’s P-Core manages a huge IPC increase in the Video Player workload, but gains elsewhere tend to be muted.
Qualcomm’s E-Core is in a similar situation, though as a narrower core it’s able to make much better use of its core width. Text Processing, HDR, and Video Player basically have the core running up against its 4-wide width limitation when running under binary translation. However, the meager IPC increases seen on Qualcomm’s P-Core in Text Processing and HDR suggest that widening the E-Core might do little, as the workload might soon run into other limitations.
Neoverse N2 starts off at lower IPC and curiously fails to gain much IPC when running x86-64 under binary translation. The core performs fine in an absolute sense when looking at IPC for native workloads, but it doesn’t look great next to Qualcomm’s latest efficiency optimized core. No workload on Neoverse N2 can average above 3 IPC, with or without binary translation in play.
Neoverse N1 is the oldest core in this set, and unsurprisingly has the hardest time. N1 in Ampere Altra spearheaded Arm’s push to get a foothold in the server market, and succeeded because the core could deliver adequate performance while being more density optimized than its x86-64 contemporaries. But that was years ago. Qualcomm’s E-Core is also 4-wide, and leaves it in the dust.
Comparatively good performance from Qualcomm’s E-Core serves as a reminder that core width means relatively little compared to other architecture characteristics. Qualcomm gives their modern E-Core a much larger out-of-order engine than Neoverse N1. Doing so lets the core keep more instructions in flight, reducing the impact of short duration delays on individual instructions.
Accounting for dispatch-stage pipeline throughput shows a largely similar picture for workloads running natively and under binary translation, which isn’t a surprise because they should be doing the same high level work. On Qualcomm’s P-Core, the biggest difference is more utilized pipeline slots, along with some movement towards being less frontend bound.
Qualcomm’s narrower E-Core uses a huge portion of its core width in many tests, especially the x86-64 versions running under binary translation. A narrower core is much easier to feed, and pipeline slots lost to frontend reasons are far fewer compared to on the wider P-Core.
Detailed performance monitoring is harder on Azure because the platform doesn’t pass through top-down counters that account at the pipeline slot granularity. Roughly sketching things out with cycle-level accounting (STALL_FRONTEND and STALL_BACKEND) however shows that Neoverse N2 loses more pipeline slots to both backend and frontend reasons than Qualcomm’s E-Core. On one hand, feeding a 5-wide pipeline is harder than a 4-wide one. Also, Qualcomm has larger core-private caches and a sizeable 12 MB L2 for its E-Core cluster with 21 cycle latency. Missing the 1 MB L2 on Neoverse N2 exposes the core to ~100 cycle L3 latency, which will be difficult for any core to cope with. Binary translation doesn’t change much on that front.
Backend stalls are even more of a problem on Neoverse N1, which also has to contend with >100 cycle L2 miss latency. N1 is more severely affected because of its smaller out-of-order engine.
A comparison with Qualcomm’s E-Core shows just how rough the situation is for Neoverse N1. Qualcomm’s core is designed to keep more than twice as many instructions in flight, even though it has the same core width and similar execution resources. Extra instructions from binary translation increase pressure on out-of-order resources, and Qualcomm’s beefier core has a better chance of absorbing that overhead.
Prism stores translated code in C:\Windows\XtaCache. Writing generated code to disk serves two functions. First, it lets Prism avoid the overhead of re-translating code every time a binary gets launched. Second, it lets Windows users examine translated code to understand how x86-64 binaries are able to run on aarch64 hosts. This is important because it lets them empathize with the hardware. XtaCache’s efficacy on the latter point has been limited because Microsoft forgot to document their .jc cache format, but thankfully others have stepped in to do so. Unfortunately, digging through translated code is still a very time consuming exercise. So here, I’ll stick to giving a couple of examples to give an idea of what Prism’s translation looks like.
Geekbench 7’s Video Player workload stands out because it runs at much higher IPC when under binary translation, and also suffers from a huge executed instruction count increase. Profiling Video Player on Skylake with VTune shows that the hottest basic block is a 17 instruction loop. This loop traverses an array and performs fused multiply-add operations using AVX FMA3 instructions.
Prism translates this to a 69 instruction sequence on both Neoverse N1 and Snapdragon X2 Elite. AVX instructions get mapped onto NEON ones, which only have 128-bit vector width. Qualcomm’s latest cores support SVE, but using it likely wouldn’t make much difference because Qualcomm’s cores implement SVE with 128-bit vector length. Also, fixed width x86-64 instructions wouldn’t map well to SVE’s variable vector length anyway.
Instruction overhead comes from performing 256-bit vector operations using multiple NEON instructions, replicating x86-64’s more flexible addressing modes, and having to perform memory accesses with separate instructions rather than specifying both a load and a math operation with a single instruction. For example, vbroadcastss ymm8, [r8+r10*4] loads a single precision scalar (32-bit) value and broadcasts it across all vector lanes. Prism uses three aarch64 instructions. One loads the value, another replicates it across a 128-bit NEON register’s lanes, and a final one copies the value to another NEON register that represents the upper 128 bits of ymm8.
Fused multiply-add instructions get even more expansion. vfmadd231ps ymm7, ymm8, [rax + r9*4 - 0x13c] uses x86-64’s most complex addressing mode, with a base, scaled index, and immediate offset. Prism does the addressing calculation with two instructions, and is able to mash both the index scaling and addition to the array base address into a single add-with-scale. Then, a load-pair (LDP) instruction specifies a 256-bit load into a pair of 128-bit NEON registers. Prism carries out the FMA operation with a pair of FMLA instructions, but curiously spills NEON register holding the high half of ymm7 to the stack. It does this again for all eight of the loop’s FMA instructions, which doubles cache bandwidth demands. None of the spilled registers are modified before the next loop iteration, so spilling and reloading them is completely unnecessary, and potentially harmful in a loop that already has high bandwidth demands. Toward the end of the loop body, Prism faithfully translates software prefetch instructions into aarch64 equivalents, including extra ALU instructions to generate addresses. Then, loop counter increment and control flow instructions translate directly into aarch64 equivalents, though of course with a different jump target address.
A human translator would avoid Prism’s unnecessary register spills. They could also omit prefetch instructions, because linearly traversing an array creates a predictable access pattern that a hardware prefetcher would likely pick up on. They could also eliminate a few instructions here and there, like by using a single register to hold both halves of ymm8 because the broadcast load’s destination is never modified. Prism does deserve credit though for simplifying address generation for six out of the eight FMA operations by emitting just a subtraction, recognizing that rax + r9*4 has already been generated.
I also took a peek at Geekbench 7’s HDR workload because its binary translated version barely gains any IPC. The hottest block from profiling on Skylake is a 36 instruction loop with scalar AVX operations. That should make it easier to translate because Prism won’t have to map AVX’s wider vector width onto NEON.
Prism translates this into somewhere over 60 instructions. The complete translation didn’t seem to be available in the translation cache when I grabbed it from my Neoverse N1 VM.
Prism introduces significant instruction overhead by trying to maintain a register mapping where x86 SSE registers map to the equivalently numbered aarch64 NEON register. For example, within the instruction sequence Geekbench 7 does a FP add (vaddss xmm5, xxm4, [rdi+rbx*4]) and stores the result to memory (vmovss [rdi+rbx*4], xmm5). Prism’s aarch64 translation puts the FP add result in s18, and then moves it to v5 before storing it in order to use v5 to represent xmm5. Additional overhead comes from inexplicable stores to the stack, as well as extra instructions that seemingly generates the carry flag for a right shift instruction. Many x86-64 instructions set flags, and not all have flag-setting equivalents in aarch64. Prism seems to prefer preserving the flag generating behavior rather than expend compute power seeing if the flags are used.
Prism generates mostly the same sequence for Qualcomm’s Snapdragon X2 Elite, but with differences around loads and stores. Qualcomm’s newer cores support FEAT_LPCPC2, which lets it use ldapurb (load-acquire with an immediate offset) instead of ldaprb (load-acquire). Being able to take an immediate offset simplifies address generation, cutting out some extra instructions. In both cases, Prism translates loads to load-acquire instructions rather than aarch64’s basic LDR to ensure memory ordering.
Binary translation is key to making Arm on Windows viable in the PC market. Prism’s translations are mostly straightforward. x86-64 instructions get their behavior faithfully replicated by equivalent aarch64 sequences, so much that it’s possible to look through translated code and easily see which sequences map to x86-64 instructions. I suspect Prism does this to avoid consuming excessive compute power, because what it’s doing looks achievable with a linear pass through each block with minimal state kept in memory. However, doing the translation in such a straightforward manner misses potential optimizations that a human could carry out. Replicating x86-64 addressing in aarch64 often requires extra instructions. x86-64 compilers also take advantage of instructions that combine a load and math operation. Those have to be broken into two instructions in aarch64, and may require using an extra temporary register to hold the loaded value.
On the hardware side, high performance cores can often absorb dead code with unused core width. That sometimes happens when Prism generates dead code, but Prism doesn’t generate excessive amounts of dead code the way Claude’s C Compiler does. Instruction overhead from replicating address generation or copying values between registers to maintain Prism’s register mapping can easily lengthen critical dependency chains. There’s not much an out-of-order engine can do about that.
The ultimate answer to binary translation’s consequences is to have a high performance architecture with few weaknesses in the first place. Qualcomm seems to be going for that with their latest Snapdragon X2 Elite. A 9-wide core running at 5 GHz is nothing to sneeze at. Large instruction caches (128 KB on the E-Core, 192 KB on the P-Core) and ample out-of-order resources are good for performance in general, and crucially help mitigate the instruction overhead from binary translation. I’m still trying different use cases on the Asus Zenbook A16 UX3607OA laptop, but it feels surprisingly normal to use outside of a few corner case scenarios like gaming. I assume the Windows on Arm situation will only get better over time as more applications get aarch64 ports, and higher performance cores mean performance gets more acceptable even when applications have to run through binary translation.
2026-09-08 10:00:00
Editor’s Note (9/8/2026): Arm has reached out to clarify that the maximum IPC increase of the C2-Ultra core over the C1-Ultra core using the parameters shown in the endnotes is 7% along with the memory bandwidth improvement not factoring into the IPC improvement. The article has been edited accordingly.
Hello you fine Internet folks,
While Arm’s last announcement was about their new datacenter focused CPU, the Arm AGI CPU. Their latest set of announcements bring a much wider focus. The incoming C2-Ultra CPU and G2-Ultra NX GPU IP will see broad use across flagship phones, with latter G2 Ultra already shipping in Xiaomi’s XRING O3 chip that launched on August 24th. While the new Neoverse CSS N4 IP will see usage in the datacenter space.
Hope y’all enjoy!
Starting off with a comparison to the Cortex X925, the last ARM core we have good data on, with the limited information provided to us by Arm on C2 Ultra. We see limited changes when comparing to the two generation old Cortex X925.

Across both cores you have the same overall layout featuring 10 wide decode, 8 simple ALUs, 6 lanes of FP, 3 branch ports, and same 4 load/2 store config. So for C2 Ultra to get it’s performance boost, we’re looking at more iterative changes targeting a few critical structures like the branch predictor and OoO buffers.
Arm was unfortunately very vague with how it accomplished these branch predictor improvements. We don’t know if they modified BTB sizes, return stacks, or the branch algorithm as a whole.
Likewise they mention C2-Ultra has a larger execution window compared to C1-Ultra along with improved speculation, but go into no explicit details here.
All this means that C2-Ultra spends less time waiting for data compared to C1-Ultra.
Now, looking at the claimed uplift from C1-Ultra to C2-Ultra we see that Arm is claiming a 15% peak performance uplift with an average 12% uplift in traditional benchmarks and workloads.
However, the endnotes provide some important context for these claims.
Firstly, these numbers are from FPGA simulations so actual hardware may see different uplifts. Secondly, the C2-Ultra platform’s results are estimated with an 8.5% higher clock compared to the C1-Ultra along with a larger 3 MB L2 cache on the C2-Ultra platform that C1-Ultra does support and does ship with in the 3 MB L2 configuration in Xiaomi’s XRING O3 and Samsung’s Exynos 2600. According to Arm, while the memory system is providing nearly twice the bandwidth for the CPU benchmarks, the improvement to the memory subsystem is not factoring into the performance uplift.
When factoring the increased clock speed, the average uplift of C2-Ultra over C1-Ultra is 3.2% and this increase does include the extra 1 MB of L2 that the C2-Ultra has which may be adding a boost in some workloads.
Moving to the peak increase of 7% IPC over C1-Ultra that Arm reached out with, I assume that this number also does factor in the L2 cache increase as well. Geekbench 6 does appear to benefit from a larger L2 cache and a 3MB L2 seems to catch most L1 misses where a smaller L2 cache may not.
From my perspective, it does appear as if in the tested workloads, C2-Ultra hasn’t improved the performance per clock much compared to C1-Ultra on average with some workloads benefiting from a larger L2 cache that some designers may opt for.
Arm says that C2-Ultra uses 38% less power compared to C1-Ultra. However, this number does factor in node and implementation improvements so how much of this 38% decrease comes from the microarchitectural improvements is up in the air.
Arm also announced C2-Nano and C2-Pro as well, however these use the same underlying microarchitecture as the C1-Nano and C1-Pro cores.
Moving to the GPU IP, Arm says that Mali G2-Ultra NX is “The largest re-architecting of the GPU IP in 7 generations."
The amount of math that a G2-Ultra NX shader core can do hasn’t changed. It still has 128 FMA units which equates to 256 FP32 FLOPs per clock or 512 FP16 FLOPs per clock. The maximum number of shader cores allowed in a G2-Ultra NX GPU, 24 G2-Ultra NX shader cores, also hasn’t changed.
G2-Ultra NX has increased the number of registers that a warp can access from 64 to 128 along with improving the granularity of register access so that G2-Ultra can now allocate registers at a granularity of 16 registers. And along with these improvements, the register file has also increased by 25%.
Arm has also improved the RT units in G2-Ultra NX by making the triangle structure that the RT unit works on more compact which according to Arm reduces the DRAM traffic by 13% due to an elimination in redundant data which allows more data to fit into cache.
Arm has also added Opacity Micromaps to G2-Ultra NX which brings a desktop GPU feature down into the mobile sector.
Another desktop GPU feature that Arm is integrating into G2-Ultra NX is a matrix accelerator into the Shader Core. This matrix accelerator can do up to 1,024 INT8 MACs per clock or 512 INT16 MACs per clock and it can clock up to twice the clock that the Execution Engines can clock to. However, a glaring omission is that the matrix accelerator only supports INT8 and INT16 but not formats such as lower precision like FP8 or the larger yet very popular format BF16 which also isn’t supported on the standard ALUs either.
Arm also doesn’t require every Shader Core in a G2-Ultra NX GPU to have a matrix accelerator but does require a minimum of 6 “NX cores” for the Ultra branding.

Xiaomi decided for their G2-Ultra NX implementation on their brand new XRING O3 that only half of the Shader Cores would implement the matrix accelerator.
Comparing the structure sizes of the cores with and without the matrix accelerator, a G2-Ultra NX shader core with a NX unit ends up being about 1.88 mm^2 and without the matrix unit a G2-Ultra shader core is about 1.55 mm^2 which means that the matrix unit adds about 21% to the area of a shader core.
The added matrix units have allowed Arm to introduce their own version of ML-powered upscaling called Neural Super Sampling (NSS).
Arm is also introducing their Frame Rate upscaling technology with the G2 generation.
With all of changes to the GPU IP, Arm is claiming up to 14% improvement in games and up to 24% in Ray Tracing benchmarks compared to last generation.
However, the G2-Ultra NX GPU is clocking about 11% higher compared to the G1-Ultra GPU in these comparisons which imply that these changes may not improve non-RT games much.
On the server side, Arm has also announced that Neoverse N4 CSS will become available for their partners.
Sadly, this is the one slide that Arm had about Neoverse CSS N4. When asked, Arm did give out more information:
CPU: 8–128 Neoverse N4 cores per die, the widest core-count range offered in a Neoverse CSS, with frequencies up to 3.8 GHz.
Scaling: Supports multi-chiplet and multi-socket designs for scaling beyond a single die.
L1 cache: 64KB instruction and 64KB data cache per core.
L2 cache: Up to 2MB private L2 cache per core.
System-level cache: Up to 256MB shared cache per die.
Memory: Supports DDR5 or LPDDR6, providing flexibility across capacity, bandwidth, power and system design.
I/O: Up to 128 lanes of PCIe Gen 6/7 and CXL 4.0.
Chiplet connectivity: Arm chip-to-chip interconnect with support for UCIe or partner-specific PHYs.
There are still a number of questions about what core is N4 using, what is the width of the memory bus on a single die, what is the number of PCIe lanes on a single die, among others.
Something that I should mention is that Arm’s technical disclosure this time around is disappointing. To Arm’s credit, the company did answer questions when asked and that responsiveness is appreciated, however it is frustrating that those questions were necessary to fill gaps that should have been addressed in the presentation. If we had been given the quality of information comparable to prior Arm announcements, we would have spent longer on analyzing the microarchitecture in our more typical fashion.
C2-Ultra is a good example of this problem, with claims about improved branch prediction and speculation giving little information of what actually changed, while the performance figures combine those core changes with higher clocks, a larger L2 cache, and more memory bandwidth. Moving to the power figures which include process and implementation improvements, we are left with the question about how much is the new microarchitecture contributing to the power decrease. These numbers are particularly frustrating for an announcement about CPU IP where comparisons at equivalent clocks and memory configurations would have been useful.
G2-Ultra NX gives more details, particularly around register allocation, the addition of matrix accelerators, and the improvements to the ray tracing units, however describing it as “The largest GPU rearchitecting in seven generations” creates an expectation of technical depth that wasn’t delivered. The matrix accelerator’s limited precision support also raises questions about its usefulness beyond the workloads Arm has chosen to target and I would have liked Arm to spend time explaining in the presentation why. And CSS N4 takes the lack of detail to an extreme with one slide with basic information about the product having to be asked.
Arm has IP and products worth talking about, but the presentation does a poor job of communicating them. The technical substance should be in the presentation from the beginning rather than something we have to assemble through follow-up questions and emails.
2026-08-31 01:31:05
Hello you fine Internet folks,
Today we're at Hot Chips 2026! We join IBM to talk about their just revealed Dual-ISA z/ architecture + ARM design with Christian Zoellin, the Core Design Team Leader and Christian Jacobi, the CTO of Systems Development!
Hope y’all enjoy!
The transcript below has been edited for conciseness and readability.
George Cozma: Hello, you fine internet folks. We’re here at Stanford during Hot Chips 2026. And the very first Hot Chips presentation this year was from IBM, talking about their next-generation z/Architecture, along with the really, really cool thing that they did there, which we’ll get to in a second. But I have two folks from IBM here. If you’d like to introduce yourselves.
Christian Jacobi: Hey, I’m Christian Jacobi. I’m an IBM Fellow and the CTO for Systems Development.
Christian Zoellin: My name is Christian Zoellin. I’m a Distinguished Engineer and I was the leader of the core design team for this generation.
George Cozma: Awesome. So what is so cool about z—I’ll call it z18 for this conversation, but the next-generation z?
Christian Zoellin: That it supports both the z/Architecture instruction set and the Arm instruction set.
George Cozma: So I think the last time we’ve seen something like this, I believe, was the original Itanium. What you guys did was slightly different. You guys actually integrated the decoders into the core. Tell me a bit about how are the decoders a single unit with multiple modes, or is it two separate decoders?
Christian Zoellin:It is a decode pipeline. Let me start there. There’s a decode pipeline, and a lot of details of the decode pipeline are shared between the two instruction set architectures. But then the individual decoders that do the actual decoding from the 16 to 48 bits in the z/Architecture or the 32-bit instructions in the Arm architecture, those are separate decoders that we construct.
George Cozma: Okay, awesome. Now, what is really interesting to me is that z is a big-endian ISA, which means that the most significant byte is first and least significant byte is last, excuse me. But Arm is little-endian, or really bi-endian, but has a little-endian mode. How did you guys deal with that swap of endians?
Christian Zoellin: In the load-store unit, basically the data cache is organized in words anyway, but we support unaligned accesses on any byte boundary. So there is a structure there that formats these word accesses out of the cache into the actual word that was requested by the load instruction. And in there, we just add all of the swapping to support big-endian versus little-endian.
George Cozma: Okay. So there’s no software that has to be implemented; the hardware will do it automatically?
Christian Zoellin: Correct.
George Cozma: Awesome. And speaking of sort of the load and store unit, something that z has is what’s known as strong ordering, whereas Arm is weak ordering. So that means it’s easy to implement, but does that mean that all load and store instructions are strong when they leave the core?
Christian Zoellin: Correct. They’re always strongly ordered. We have a lot of infrastructure to speculate across these ordering boundaries, and so for us, from a performance point of view, we have all the hardware to do this efficiently and quickly. And so, we simply reuse that hardware and stay strongly ordered. As a matter of fact, there’s a feature in the Arm architecture to enforce strong ordering, and for us, that bit basically does nothing.
George Cozma: I believe that is TSO, correct?
Christian Zoellin: Correct.
George Cozma: Yes, so you guys have TSO automatically then. So, moving on to how much sort of area/extra transistors did this use? Was it much, or was it not a ton?
Christian Zoellin: It definitely wasn’t a ton. But we already talked about the decoders, how those are separate decoders; there’s certainly transistors in that. There’s also certain features that we added that the z/Architecture did not have, but we had to implement. One example is the BF16 and FP16 floating-point formats, and for those, we also added new data flows, and those are extra transistors. But if you look at the overall floor plan, these things are small, tiny specks compared to our huge BTB branch prediction structures or to our large instruction and data caches.
George Cozma: And so speaking of what can be shared and reused, how many structures are being reused? I assume it’s the vast majority, correct?
Christian Zoellin: All of those big ones especially. That’s key. The translation lookaside buffers, the caches, the physical register files for the GPRs and vector registers, all of those are exactly the same and used as-is.
George Cozma: Cool. So moving sort of more into a business case use, why did you guys add Arm to z?
Christian Jacobi: So, the Arm software ecosystem has grown really rapidly over the last loosely decade, driven a lot by the hyperscalers deploying Arm in their data centers, right? So, for us, it’s a huge opportunity to bring all of that software closer to the mission-critical data and transactions that our clients are running. So it’s a huge expansion in terms of flexibility of where clients place that workload. In many cases, it does make sense to have that workload close latency-wise to the data and transactions, but also in the same sort of operational environment, so that from a security, from an availability perspective, it kind of all ties together. That, I would say, is one important use case.
The other important use case is we have clients who do massive workload consolidation projects, in particular on LinuxONE, sometimes running thousands of, for example, MongoDB databases on a single mainframe footprint, LinuxONE footprint. But if you look at these projects, they don’t consist just of like the main database, right? They have endpoint security software, they have backup software, monitoring, observability—a lot of software in a total solution stack.
And we have a great ecosystem team that would work with ISVs to port software from wherever it was initially coded onto Linux on z, but of course, there’s so much software out there, so many ISVs out there, we really can’t win them all to come to the platform. So, it’s been complex sometimes to get the ISVs to support our platform whenever a client wanted to do such a big consolidation project. Having the Arm ecosystem natively available on our platform solves a lot of that, so we believe we can see many more of these large-scale consolidation scenarios just because we have so much more of a sort of complete software ecosystem through the adoption of Arm technology into our platform.
George Cozma: And speaking of that, have you considered adding a third ISA? I know a lot of people have been talking about RISC-V, that was a big thing yesterday during the tutorials. But what about POWER? What about implementing POWER as a third ISA?
Christian Jacobi: Yeah, well, I mean right now we’re working on this project to integrate Arm. We’ve learned a lot in this experience. But if you also look at the use cases that we have for different enterprise systems, IBM does have our Power Systems, IBM has our Mainframe Systems; they are sort of in separate swim lanes going after different kinds of workloads. There’s really no good business case for us to bring the Power architecture into the mainframes and then try to compete with ourselves on the Power Systems. That wouldn’t really make any sense for us.
George Cozma: Absolutely. And I believe during the presentation, you made a comment about the amount of instructions in Arm. Really sort of going back into the hardware, what’s the real difference between CISC and RISC because, as you said, they’re kind of misnomers?
Christian Zoellin: So let me get to the most extreme CISC instructions right away. We have CISC instructions that as part of the instruction have a parameter block in memory that has 15 different parameters that get consumed by the instruction to produce the right result. That’s how we implement compression, that’s how we implement crypto on the mainframe. We have instructions that have sub-instructions and function codes, and they’re library routines in some sense.
RISC does not have that. RISC has to do all the operations on the register set, and as such, they just have many more instructions for all these substeps, and things that for us are one large instruction in a RISC instruction set architecture are usually programs that execute hundreds of instructions to do the same task. And so that’s how you get to this instruction inflation to a certain degree.
George Cozma: Well, I guess the other thing that was talked about beyond just next-generation Telum, is your next-generation Spyre—so not Spyre 2, but next-generation Spyre. You guys have added HBM to that. I know that there was a Meta paper, I believe about two or three years ago, talking about how many errors HBM would sort of not create, but would have. IBM is all about ultra-reliability. How are you guys sort of dealing with that new memory that may have higher errors?
Christian Zoellin: I’m actually not familiar with the details of that. Are you?
Christian Jacobi: Yeah, to a degree. I’m not the super deep expert, but first of all, it’s important to note we’re not on the bleeding edge of HBM technology, so we’re consuming a technology that has matured and has a lot lower error rate than at the bleeding edge. We’re also not running it at the highest speed levels, right? That’s another component in designing for reliability.
And then we do a lot of testing. We’re really focusing on shipped product quality in terms of really finding those production failures early. We do things like very deep tests for the processors. For example, we do a lot of burn-in to actually run the processors in an oven at a very high temperature to find the early failures before we even ship them. If the processor breaks on our manufacturing floor, that’s much better than if it breaks in a client’s data center, right? So that’s the kind of things we do for reliability. It’s not only about error correction and error checking; it’s also how you weed out the errors in your manufacturing processes.
George Cozma: Okay. And I guess the move over to HBM has sort of increased the amount of business case that you can now target. What are the broader cases that you’re looking at now?
Christian Jacobi: Yeah, it’s actually the other way around. It’s how the business cases, the use cases evolve that informs how we need to adapt the design and the architecture, basically, right? So when we introduced Telum 1 and Telum 2 with the on-chip AI acceleration, it was really a lot about relatively small models doing fraud detection inside transactions with very low latency so that you wouldn’t hold up a credit card swipe, for example.
Then when we introduced Spyre, it was a lot about how we do enhanced protection in such use cases where you use slightly more complex models, but then you could also run, I’d say, small language models on a group of cards to get some generative AI going. As this has evolved over the last few years, we’re really seeing use cases come up for things like document processing when you’re adjudicating insurance claims, for example. So that’s just one example of how a business process can benefit from large language model utilization.
The same is true when we’re thinking about AI ops where you use sort of an agentic workflow to actually have the system monitor itself, self-heal, self-optimize, those kinds of things. So, we’ve seen this shift from relatively small models for risk and fraud detection to more complex models for risk and fraud detection, small language models, and now we’re seeing the shift even in the enterprise space towards more agentic loops. And so that has determined that we need to invest more heavily in memory bandwidth to be able to run these models.
We haven’t talked about the details, I won’t talk about the details today, but of course that’s not only about a single chip; it’s really about designing a system consisting of many of those chips. We’ve talked about the 4 terabytes per second of memory bandwidth on a single chip; that system will of course then use multiple chips to get to significantly more than that 4 terabytes so that we can run many models in a heterogeneous environment, and large models, and very good output tokens per second performance to support these agentic workloads that we see on the horizon.
George Cozma: Awesome. You’ve been asked this question, so you’re going to be asked a different question. But CZ, what’s your favorite type of cheese to end this interview with?
Christian Zoellin: Roquefort. That’s a blue cheese, a French blue cheese made out of sheep’s milk. Goes very well with red wine.
George Cozma: Ooh. And Christian, what new cheeses have you tried?
Christian Jacobi: Right now I’m on a Gruyère kick.
George Cozma: Ooh! Just Gruyère or sort of the Gruyère family?
Christian Jacobi: Um, I really like Gruyère right now.
George Cozma: Okay. Okay. Well, thank you so much for sitting down with me. Good luck with your new chips, as always. Would love to test them one day, but thank you so much for sitting down.
Christian Jacobi: Our pleasure. Thank you.
Christian Zoellin: Thank you.
2026-08-30 15:25:37
Memory expansion has been attractive for years, and is more relevant than ever because ML models have an insatiable appetite for memory capacity. In response, XCENA has worked with Samsung to create a CXL memory expansion device that can also host SSDs and do compute. The device is called MX1, where MX stands for “Memory Xcelerator”.
On the memory expansion front, the MX1 can host up to 2 TB of DDR5 memory and connects to the host via a PCIe 6/CXL 3.2 x8 interface. The MX1 therefore has 128 GB/s of bandwidth to the host, or 64 GB/s in each direction. XCENA exposes another eight downstream PCIe 6 lanes that can be used to connect SSDs. SSD storage can be exposed as memory, with the MX1’s attached DRAM acting as a cache. The combination of DDR5 slots and downstream PCIe lanes lets the MX1 make a massive amount of memory capacity visible to the host. However, MX1’s most exciting feature is arguably a substantial amount of onboard compute.
MX1’s chip hosts 3072 RISC-V cores, divided into clusters of 32 cores that share L2 caches and a data TLB. Sets of four clusters form a “subsystem”, which acts as the smallest job allocation unit. MX1 has 24 subsystems, letting it run 24 independent jobs at a time. An in-house NoC connects the subsystems to L3 cache and memory. Two Arm Cortex A53 cores handle control functions. The MX1 is fabricated using Samsung’s 4nm process and consumes 40 W, implying that each each RISC-V core draws a bit under 13 mW. The board consumes 90W when accounting for power consumption from 4 DIMMs.
XCENA uses a lot of small cores because they’re targeting data-parallel workloads, where single threaded performance is less important than utilizing high memory bandwidth while maximizing power efficiency. It’s a strategy with parallels to Intel’s Xeon Phi, which similarly used a large number of low-clocked, relatively weak cores to take on highly parallel tasks.
Each RISC-V core uses in-order execution and runs at a pedestrian 1.1 GHz. XCENA’s cache hierarchy is almost GPU-like because it tries to avoid address translation overhead and varies cache sharing wildly at the upper levels. Each core has a 4 KB virtually addressed L1 data cache. Data-side memory accesses don’t go through address translation unless they miss L1D. 128 KB L2 data caches are shared across a cluster, as are TLBs for accelerating address translation. The L2 data cache is virtually indexed and physically tagged (VIPT), much like L1D caches in many conventional CPUs. From the instruction side, sets of four RISC-V cores share a 8 KB instruction cache. XCENA aims to contain kernel hot loops within this instruction cache, while a cluster-level 128 KB L2 instruction cache handles larger instruction footprints.
Instruction fetches operate directly on physical addresses and do not use virtual memory. Instruction accesses therefore don’t need address translation or TLBs. XCENA sets aside predefined device physical addresses for code, and restricts the program counter to those regions. Doing so prevents the RISC-V cores from accidentally jumping to data. XCENA handles process level isolation by isolating jobs at subsystem (128-core) boundaries. Presumably they also divide up the code region to give each subsystem its own code segment, which would prevent one process from accidentally executing another’s code.

MX1’s programming model has parallels to OpenCL or CUDA. One kernel gets invoked many times, and each invocation uses an index to figure out what data it should process. Specifically, mu::getTaskIdx() is analogous to OpenCL’s get_global_id(). This model encourages code sharing across many cores, so sharing L1 instruction caches makes sense. If a loop is small enough, the four cores that share an instruction cache may fetch the same address at the same time, likely letting the instruction cache satisfy multiple fetches with a broadcast read.
On the data side, MX1 uses virtual memory and operates on the same virtual addresses as host code. Host and MX1 code can therefore share pointers, much like with OpenCL’s SVM. XCENA’s software sets up page tables to maintain the same mappings as the host’s. Each RISC-V core has a virtually addressed 4 KB L1 data cache, letting the core avoid address translation on a L1D hit. The cluster-level L2 cache is virtually addressed and physically tagged (VIPT), much like L1D caches on many CPUs. L2 indexing proceeds in parallel with a lookup in the cluster-shared TLB. The TLB has 1024 entries for 64 KB pages, and 8 entries for 1 GB pages. 64 KB pages allow more TLB coverage than typical 4 KB pages, but operating systems tend to use smaller page sizes to reduce overhead when paging to disk, copying pages, or clearing pages. However, XCENA expects the OS to give CXL memory special treatment, and larger page sizes may make sense for a giant block of expansion memory. 64 KB pages also align with typical SSD block sizes. As an aside, using different sets of TLB entries for different page sizes makes sense because it lets the TLB use different indexing schemes for each page size.
XCENA takes advantage of RISC-V’s extensibility to implement a custom Vector Processing Engine (VPE) at the subsystem level. Each RISC-V core gets a VPE command queue, and can ask the VPE to accelerate a variety of vector operations. Likely, XCENA uses special instructions to enqueue messages into VPE command queues, and expects code to treat it as a giant shared coprocessor. The VPE supports FP32 and FP16, and provides ~3 TFLOPS of dot product throughput across the chip. If the VPEs run at 1.1 GHz like the cores, then each VPE can sustain 128 FLOPS per cycle. Curiously, the VPEs don’t appear to accelerate integer operations. Perhaps XCENA expects code to run integer operations directly on the RISC-V cores. 3072 RISC-V cores at 1.1 GHz would be get roughly 3T integer operations per second if each core completes one operation per cycle. Another curiosity is that XCENA’s API exposes the VPU using built-in functions that return an error code. Code has to explicitly check for error conditions like overflow and invalid accesses, suggesting that VPU instructions don’t raise exceptions.
MX1 can host SSDs, which are presented to the host as CXL memory. XCENA calls this “Infinite Memory”. MX1 can have SSDs attached in RAID and has matched bandwidth on both its downstream PCIe and upstream PCIe/CXL links, meaning that it can theoretically saturate its bandwidth to the host using SSDs alone. However, SSDs have high latency compared to DRAM. MX1 can mitigate that latency by using its attached DDR5 to cache SSD contents. Caching works with 64 KB pages, with an on-chip 1024 entry map cache. The map cache acts like a TLB, and tracks DRAM pages mapped to SSD-backed addresses. If an access misses in the map cache, it causes a page fault that’s handled by firmware running on the MX1’s RISC-V cores. Firmware handles the cache miss by fetching data from the SSD and updating the mapping. I’m confused because a 1024 entry map cache only covers 64 MB with 64 KB pages, but XCENA’s documentation suggests the cache defaults to 16 GB and has adjustable capacity in 16 MB steps. I’m not sure how that works given the map cache structure.
To further exploit attached DDR5 when using SSD as memory, users can configure a “pinned prefix” where SSD-backed addresses are pinned to DRAM. The pinned prefix region size can be configured in 16 MB steps. Prefixing implies that pinned memory can only cover a contiguous address space, and doesn’t have the flexibility of page-level caching. Therefore, pinned prefix memory is best used to keep a frequently accessed buffer in DRAM. XCENA’s site gives an example with 115.5 GB of pinned memory, out of 231 GB of attached DRAM.
To mitigate SSD latency, the MX1 can also be set up to prefetch from the SSD. Prefetch helps keep the IO path busy, and helps overlap data fetches with compute execution. XCENA hopes to use pinning and prefetch to mitigate a SSD’s low performance compared to DRAM. The MX1 can also run SSDs in RAID to increase bandwidth.
MX1 and LPDDR5X-PIM both place compute close to memory, where they can exploit high internal memory bandwidth that's not accessible over the host interface. MX1 presents a more convincing case than LPDDR5X-PIM because it can act like an accelerator with its own onboard memory. Software doesn’t have to face the tradeoffs and complexity associated with PIM mode switching, and using MX1’s compute won’t block memory accesses from other threads. Host threads and MX1 RISC-V cores could theoretically work on the same buffers, using standard multithreading techniques like locks to ensure ordering.
Working with caches should also be more straightforward than with LPDDR5X-PIM. Type 3 CXL devices (CXL.mem) can use snoops to back-invalidate host cache lines, letting the device make its results visible without resorting to making memory regions uncacheable. Similarly, the CXL.mem device can track read-for-ownership requests from the host and understand when its own compute cores can safely modify data. I’m not sure whether MX1 fully leverages this capability for its RISC-V cores, but Samsung indicates that the device can use CXL memory semantics to share its memory across CPUs and GPUs. Hopefully, that means the device can use snoops to integrate with processor memory subsystems.
As with LPDDR5X-PIM, MX1 doesn’t offer a lot of throughput. For perspective, an Nvidia GeForce GTX 1080 with similar onboard memory bandwidth has 8.8 TFLOPS of FP32 compute, compared to the 3 TFLOPS available from the MX1. But raw throughput isn’t the point. Rather, MX1 offers a way to mitigate some memory expansion downsides. MX1-side compute can avoid the constrained host interface, and avoids the power overhead of traversing the CXL link. A hypothetical accelerator with 8.8 TFLOPS of compute would be limited to just 64 GFLOPS if it had to load one byte per FLOP over CXL. MX1 could exceed 200 GFLOPS in the same scenario, even if it also missed cache.
MX1’s approach to near-memory compute shows promise. CXL memory expanders will often have higher bandwidth to their DRAM pool than what the CXL link can handle. Regular host accesses will leave a large chunk of that bandwidth on the table, which could lead to host-side compute getting memory bandwidth bound if the host can’t serve enough accesses out of its caches. Latency is also a problem because expansion memory will have higher access latency than directly attached memory. A near-memory compute option can mitigate those memory expansion problems without the downsides of attempting the same thing underneath a normal memory controller, as LPDDR5X-PIM does.
I hope to see companies continue to explore memory expanders and near-memory compute in the future. Perhaps in the distant future, this technology can make its way down to consumers. A lot of us have spare DRAM lying around from previous builds, and would be very interested in using it to mitigate the cost of buying new memory. A side of extra compute wouldn’t hurt either.
2026-08-29 13:36:33
In-memory compute has been an attractive proposition for many years because compute within a memory chip can exploit its higher internal bandwidth. Additionally, in-memory compute avoids the long latency path between DRAM and traditional compute cores. At Hot Chips 2026, Samsung discusses their continued pursuit of in-memory compute with their PIM (Processing-in-Memory) push. They’re implementing MAC units within LPDDR5X chips, while preserving the chip’s ability to interface with a standard memory controller.
DRAM chips are internally divided into banks, each with their own read and write logic. During a normal DRAM access, the memory controller selects a bank, activates a row within it, and then accesses data via column access strobe (CAS) commands. Bandwidth is limited by the chip’s external DRAM interface. Even if the memory controller could activate all of the banks simultaneously, it wouldn’t be able to get its hands the full bandwidth available across all the banks.
Samsung’s LPDDR5X-PIM is like a normal LPDDR5X-9600 chip with 16 banks, but places a PIM (Processing-in-Memory) block at each bank. These PIM blocks access their attached DRAM bank without being constrained by the chip’s external bus. Together, they can utilize the chip’s internal bandwidth across all 16 banks, which comes out to 614 GB/s. For comparison, regular DRAM accesses can hit two banks in parallel and max out at 76.8 GB/s.
PIM blocks internally consist of a MAC tree with surrounding register files and control logic. A 1024-bit instruction register file holds up to 64 16-bit instructions. A 4 kbit source register file is meant for activation vectors, and supplies one source operand for the MAC array. Samsung expects software to load model weights into DRAM, so the attached DRAM block supplies the second operand. Model weights can be scaled before the MAC computation, with scale factors coming from a 2 kbit scale register.
The PIM block’s MAC array supports a variety of low precision formats. Numbers from Samsung’s presentation suggest each PIM block’s MAC array can sustain four INT8 or FP8 MAC operations per data clock, or eight per cycle when not counting the double data rate. Throughput doubles for 4-bit input weights, bringing package-wide compute throughput to 2.4 TOPS.
This isn’t a very high figure, but an implementation with many LPDDR5X chips will have higher aggregate throughput. For example, eight LPDDR5X chips together would have 9.6 INT8 TOPS, which just about matches the NPU in Intel’s Meteor Lake. That would also be an expensive setup, because eight 16 GB LPDDR5X chips would correspond to 128 GB of system memory.
One highlight of LPDDR5X-PIM is that it stays within the standard LPDDR5X protocol while exposing compute capabilities that aren’t part of the memory standard. Samsung achieves this by setting aside special row addresses, which act like MMIO addresses of sorts. Each channel has a pair of predefined rows for mode control. Activating one of those rows sets the chip to single-bank mode, while the other sets the chip to multi-bank mode. Single-bank is the regular mode, while multi-bank applies commands across all 16 banks to exploit the chip’s internal bandwidth.
Special per-bank rows change how read and write commands behave. Activating one of these special rows makes read and write commands access PIM registers instead of regular DRAM bank contents (PIM Registers Activated mode). Samsung envisions a ML use case where software loads model weights into DRAM while the chip is in normal single-bank mode. Then, software switches into multi-bank mode and enters PIM Registers Activated mode. This lets code write activation values into PIM source registers, set scale factors in PIM scale registers, and specify an operation that’s filled into PIM instruction registers.
Because the chip is in multi-bank mode, each PIM register write gets broadcast across all 16 banks. PIM compute therefore works like a very constrained SIMD processor, where the operation, scale factor, and one source operand are the same across all banks. Samsung does allow writing PIM registers in single-bank mode, but that functionality is meant for debugging purposes. Each DRAM packet is 256 bits (BL=16) Filling each source register takes 16 write commands. Doing that one bank at a time across each of the 16 banks would mean 256 write commands, turning host to PIM register write bandwidth into the limiting factor.
After priming PIM registers, software switches back into multi-bank mode and issues read commands. Instead of reading DRAM contents, these read commands initiate computations and get results accumulated into PIM vector register files. Then, write commands tell PIM blocks to write VRF contents back into the DRAM banks.
PIM has to handle reordering that a normal memory controller might carry out. When code sets up PIM by activating the bank, PIM conventionally sets up its instruction register files so that instructions sequentially access each source register element. For instance, the first instruction would reference the first source register element, the second instruction would reference the second source register element, and so on. However, that falls apart if the memory controller reorders accesses. Samsung gets around this with an Address Align Mode (AAM), which makes each instruction infer its source register index from the column address being accessed.
When the host finishes using in-memory compute and wants to read results, it switches the DRAM chip back into single-bank mode. Then, regular DRAM reads and writes will start accessing DRAM contents as normal.
Samsung internally achieved huge performance gains when taking advantage of LPDDR5X-PIM, compared to using standard LPDDR5X. The chip’s ability to operate with a standard memory controller is impressive, and Samsung has been very creative in how they approached the problem.
Repurposing standard DRAM commands should simplify hardware, but software challenges look steep. Because PIM modes change the meaning of DRAM access commands, software can’t use PIM and carry out regular memory accesses at the same time. That applies even across threads, because memory controllers and DRAM chips are oblivious to what thread an access is for. If a non-PIM thread reads from memory while another is using PIM, the first thread could cause an unintended computation and get incorrect results into the PIM VRFs. A write from the non-PIM thread could cause PIM blocks to write VRF data back to the wrong address.
Samsung deals with this by having the host isolate a PIM region in memory. I can’t think of an easy way to do this in a typical system without compromising memory bandwidth and PIM performance. Hardware normally interleaves addresses across channels, which lets common access patterns naturally utilize bandwidth across those channels. PIM uses per-channel rows to control single/multi-bank mode changes, so dropping interleaving and designating memory channels as PIM-only would be the only reasonable way to create a PIM region. Then, non-PIM applications wouldn’t be able to take advantage of bandwidth from channels reserved for PIM. PIM code would miss out on bandwidth and compute from non-PIM channels. The latter could be a significant issue because per-chip compute throughput isn’t that high.
Multitasking issues could persist even after isolating a PIM region. If an application wants to use PIM and take advantage of multithreading, it would have to guard PIM region accesses with locks to prevent cases where one thread tries to do PIM compute while another attempts regular memory accesses. Things get even worse with a modern multitasking operating system, where multiple processes could try to use PIM without being aware of each other. I’m not sure there’s a good way to handle that besides making the operating system run PIM compute code segments with all other threads blocked and interrupts disabled. Handling interrupts or context switches with PIM feels like a nightmare for the OS in any case. Preempting a PIM thread would mean bringing the memory channel out of PIM mode and saving PIM state. The OS would have to read out instruction, source, scale, and vector register file across each bank and save it somewhere. Only allowing a single running thread with no task switching would leave multithreaded performance on the table, and could lead to system responsiveness issues if code spends too long in PIM compute sections.
PIM compute breaks a memory subsystem’s expectations about DRAM behavior because DRAM can generate memory values that the cache hierarchy never knows about. Caches can also break PIM behavior by absorbing accesses meant to trigger PIM operations. Samsung therefore recommends mapping PIM memory as uncacheable. That’s problematic because modern CPUs and GPUs rely heavily on caching to mitigate DRAM latency. Performance on uncacheable memory will be extremely slow because the CPU or GPU cores will spend far more time stalled waiting on memory.
Skipping caches isn’t the only problem. PIM reads act like MMIO accesses because they cause computations that affect PIM VRF values, rather than just retrieving data. CPUs also mitigate memory latency by initiating loads before they know that load data will actually be needed. Branch prediction lets CPUs issue instructions before the core knows for certain that those instructions will be executed. Prefetchers observe memory access patterns and attempt to load data into cache before instructions request that data. If the CPU loads data that turns out to unneeded later on, that’s fine because loads normally won’t cause incorrect program behavior. Unfortunately that’s not true with PIM, where reads trigger computations that modify PIM VRF contents.
Working with a PIM region will likely mean making memory accesses non-speculative as well as non-cacheable. Running a CPU without caching, prefetching, or out-of-order execution will cripple performance.
Setting aside PIM mode difficulties, in-memory compute poses high level challenges for software. Each PIM block only has fast access to its locally attached DRAM bank. All other input data has to be brought in through the DRAM chip’s comparatively constrained external interface. PIM blocks can’t directly exchange data with each other, so the host has to move data using regular DRAM reads and writes if one PIM block needs to use results generated by another.
Samsung’s LPDDR5X-PIM can theoretically go into any server, desktop, laptop, or even mobile device thanks to its ability to work with standard memory controllers. However, that doesn’t mean it’ll be easy to use with typical hardware and software paradigms. PIM mode switching throws a wrench into the works for multitasking operating systems. Modifying DRAM contents under the hood and attaching side effects to read commands breaks CPU caching, prefetching, and out-of-order execution.
I don’t think there’s an easy way to use in-memory compute without changes throughout the memory subsystem. For example, something like should make software adoption easier:
Expand the DRAM interface to add a set of compute commands, avoiding mode switch complexity
Have the memory controller act like a peer CPU core from a cache coherency perspective. Before using in-memory compute commands, the memory controller issues read-for-ownership (RFO) requests for all affected cache lines. That lets the memory controller obtain any modified data and write it back to DRAM before starting in-memory compute, ensuring that in-memory compute results reflect the latest CPU-side writes. Then, the memory controller holds ownership of affected cache lines until in-memory compute operations complete, letting CPU cores observe in-memory compute results without needing to invalidate or bypass caches
Add a new set of CPU instructions like “rep macb” that perform multiply-accumulate operations over a block of memory with fixed multiplicand/scale factors and undefined numerical characteristics. The CPU can choose whether to use in-memory compute (if supported by DRAM) or generate a sequence of internal ops (if operating over a small set of data that’s already in cache).
With those hardware changes, software would be able to use in-memory compute from a multitasking operating system without reserving memory or losing thread-level parallelism to PIM-related locks and synchronization. A transparent CPU instruction avoids the problem of shipping hardware specific binaries, and allows forward-compatible code that automatically takes advantage of new hardware capabilities including different in-memory compute implementations. It also lets hardware use implementation-specific knowledge and real-time data (like a no-fill-on-miss cache lookup) to make the best decision about where to carry out compute. I don’t like the software alternative of reserving memory regions, marking them uncacheable, and blocking threads. There’s just too many tradeoffs around performance, memory capacity, and responsiveness.