Skip to content

][GSD-13279] Silent wrong results on Arc A770 (DG2): reused per-dispatch private surface never made resident again — evictUnusedAllocations() unbinds it permanently (regression since ee21f7c717) #973

Description

@AIVirtuoso

Pre-submission Checklist

  • I am using the latest GPU driver version (releases)
  • I have searched for similar issues and found none

GPU Hardware

Intel Arc A770 16 GB (DG2-512 / ACM-G10), SPARKLE A770 TITAN OC Edition (SA770T-16GOC), single discrete GPU, no integrated GPU in use.

DRI Devices Information

$ ls -ls /dev/dri/*
0 crw-rw----+ 1 root video  226,   0 ago 10 13:16 /dev/dri/card0
0 crw-rw-rw-  1 root render 226, 128 ago 10 13:15 /dev/dri/renderD128

$ ls -la /dev/dri/by-path/
lrwxrwxrwx 1 root root  8 ago 10 13:15 pci-0000:09:00.0-card -> ../card0
lrwxrwxrwx 1 root root  8 ago 10 13:14 pci-0000:09:00.0-platform-simple-framebuffer.0-card -> ../card0
lrwxrwxrwx 1 root root 13 ago 10 13:15 pci-0000:09:00.0-render -> ../renderD128

GPU Detailed Information (lspci output)

$ sudo lspci -vvv -k -s 0000:09:00.0
09:00.0 VGA compatible controller: Intel Corporation DG2 [Arc A770] (rev 08) (prog-if 00 [VGA controller])
	Subsystem: Sparkle Computer Co., Ltd. Device 3937
	Control: I/O+ Mem+ BusMaster+ SpecCycle- MemWINV- VGASnoop- ParErr- Stepping- SERR- FastB2B- DisINTx-
	Status: Cap+ 66MHz- UDF- FastB2B- ParErr- DEVSEL=fast >TAbort- <TAbort- <MAbort- >SERR- <PERR- INTx-
	Latency: 0, Cache Line Size: 64 bytes
	Interrupts: unknown pin routed to IRQ 77, MSI(X) routed to IRQ 77
	IOMMU group: 19
	Region 0: Memory at fb000000 (64-bit, non-prefetchable) [size=16M]
	Region 2: Memory at 7800000000 (64-bit, prefetchable) [size=16G]
	Expansion ROM at fc000000 [disabled] [size=2M]
	Capabilities: [40] Vendor Specific Information: Intel Capabilities v1
		CapA: Peg60Dis- Peg12Dis- Peg11Dis- Peg10Dis- PeLWUDis- DmiWidth=x4
		      EccDis- ForceEccEn- VTdDis- DmiG2Dis- PegG2Dis- DDRMaxSize=Unlimited
		      1NDis- CDDis- DDPCDis- X2APICEn- PDCDis- IGDis- CDID=0 CRID=0
		      DDROCCAP- OCEn- DDRWrtVrefEn+ DDR3LEn+
		CapB: ImguDis- OCbySSKUCap- OCbySSKUEn- SMTCap- CacheSzCap 0x0
		      SoftBinCap- DDR3MaxFreqWithRef100=Disabled PegG3Dis-
		      PkgTyp- AddGfxEn- AddGfxCap- PegX16Dis- DmiG3Dis- GmmDis-
		      DDR3MaxFreq=2932MHz LPDDR3En-
	Capabilities: [70] Express (v2) Endpoint, IntMsgNum 0
		DevCap:	MaxPayload 128 bytes, PhantFunc 0, Latency L0s <64ns, L1 <1us
			ExtTag+ AttnBtn- AttnInd- PwrInd- RBE+ FLReset+ SlotPowerLimit 0W TEE-IO-
		DevCtl:	CorrErr- NonFatalErr- FatalErr- UnsupReq-
			RlxdOrd+ ExtTag+ PhantFunc- AuxPwr- NoSnoop+ FLReset-
			MaxPayload 128 bytes, MaxReadReq 128 bytes
		DevSta:	CorrErr- NonFatalErr- FatalErr- UnsupReq- AuxPwr- TransPend-
		LnkCap:	Port #0, Speed 2.5GT/s, Width x1, ASPM L0s L1, Exit Latency L0s <64ns, L1 <1us
			ClockPM- Surprise- LLActRep- BwNot- ASPMOptComp+
		LnkCtl:	ASPM L1 Enabled; RCB 64 bytes, LnkDisable- CommClk-
			ExtSynch- ClockPM- AutWidDis- BWInt- AutBWInt- FltModeDis-
		LnkSta:	Speed 2.5GT/s, Width x1
			TrErr- Train- SlotClk- DLActive- BWMgmt- ABWMgmt-
		DevCap2: Completion Timeout: Range B, TimeoutDis+ NROPrPrP- LTR+
			 10BitTagComp+ 10BitTagReq+ OBFF Not Supported, ExtFmt+ EETLPPrefix-
			 EmergencyPowerReduction Not Supported, EmergencyPowerReductionInit-
			 FRS- TPHComp- ExtTPHComp-
			 AtomicOpsCap: 32bit- 64bit- 128bitCAS-
		DevCtl2: Completion Timeout: 50us to 50ms, TimeoutDis-
			 AtomicOpsCtl: ReqEn-
			 IDOReq- IDOCompl- LTR+ EmergencyPowerReductionReq-
			 10BitTagReq- OBFF Disabled, EETLPPrefixBlk-
		LnkCap2: Supported Link Speeds: 2.5GT/s, Crosslink- Retimer- 2Retimers- DRS-
		LnkCtl2: Target Link Speed: 2.5GT/s, EnterCompliance- SpeedDis-
			 Transmit Margin: Normal Operating Range, EnterModifiedCompliance- ComplianceSOS-
			 Compliance Preset/De-emphasis: -6dB de-emphasis, 0dB preshoot
		LnkSta2: Current De-emphasis Level: -6dB, EqualizationComplete- EqualizationPhase1-
			 EqualizationPhase2- EqualizationPhase3- LinkEqualizationRequest-
			 Retimer- 2Retimers- CrosslinkRes: unsupported, FltMode-
	Capabilities: [ac] MSI: Enable+ Count=1/1 Maskable+ 64bit+
		Address: 00000000fee00000  Data: 0000
		Masking: 00000000  Pending: 00000000
	Capabilities: [d0] Power Management version 3
		Flags: PMEClk- DSI- D1- D2- AuxCurrent=0mA PME(D0+,D1-,D2-,D3hot+,D3cold-)
		Status: D0 NoSoftRst+ PME-Enable- DSel=0 DScale=0 PME-
	Capabilities: [100 v1] Alternative Routing-ID Interpretation (ARI)
		ARICap:	MFVC- ACS-, Next Function: 0
		ARICtl:	MFVC- ACS-, Function Group: 0
	Capabilities: [420 v1] Physical Resizable BAR
		BAR 2: current size: 16GB, supported: 256MB 512MB 1GB 2GB 4GB 8GB 16GB
	Capabilities: [400 v1] Latency Tolerance Reporting
		Max snoop latency: 1048576ns
		Max no snoop latency: 1048576ns
	Kernel driver in use: xe
	Kernel modules: i915, xe

Resizable BAR is enabled (BAR 2 at the full 16 GB). Note the link reads Speed 2.5GT/s, Width x1 — that is what lspci reports with the GPU idle in ASPM L1; under load the reproducer sustains 1.3-2.5 GB/s host-to-device, an order of magnitude more than a real 2.5GT/s x1 link could carry, so the link is not actually running at x1 when it matters.

Driver Version

26.18.38308.1

Installed GPU Driver Packages

NixOS, so there is no dpkg/rpm database to query. The relevant store paths:

intel-compute-runtime-26.18.38308.1
level-zero-1.28.5
intel-graphics-compiler-2.34.4
intel-gmmlib-22.10.0

The Level Zero device itself reports driver_version 1.15.38308, i.e. the same 26.18.38308.1 build named in the Driver Version field.

The Level Zero loader in use is the system libze_loader.so.1 (1.28.5); the SYCL runtime and UR adapters come from the pip intel-sycl-rt / intel-cmplr-lib-ur packages inside a Python venv, versions given under oneAPI Version below.

Driver Installation Details

  • Installation method: NixOS 26.05 declarative configuration (hardware.graphics.extraPackages pulling intel-compute-runtime from nixpkgs), not the Intel apt repository.
  • Kernel driver: xe, forced on this device via boot parameters i915.force_probe=!56a0 xe.force_probe=56a0 (both modules are present; i915 is loaded but bound to nothing). i915 is untested — see item 7 of "What was open" in Additional Notes.
  • The application runs inside an FHS-compatible shell so that a normal Python venv with the pip oneAPI runtimes works on NixOS.

Linux Distribution

Other (please specify below)

Other Linux Distribution

NixOS 26.05 (Yarara), BUILD_ID=26.05.20260814.02e0898

Kernel Version & Boot Parameters

$ uname -r
7.1.7

$ cat /proc/cmdline
initrd=\EFI\nixos\...-initrd-linux-7.1.7-initrd.efi init=/nix/store/...-nixos-system.../init
i915.force_probe=!56a0 xe.force_probe=56a0 root=fstab loglevel=4
video=HDMI-A-4:e drm.edid_firmware=HDMI-A-4:edid/SAM_1080p.bin
lsm=landlock,yama,bpf

$ lsmod | grep -E 'i915|xe'
xe                   4530176  3
i915                 5103616  0

Actual Behavior

A sequence of host-to-device copies whose source is a file-backed mmap silently breaks device modules that are already loaded at the time of the copies — but only those modules the driver has put on the per-dispatch private-memory path. Nothing in the module is altered; what is lost is the GPU mapping of the private surface its kernels spill to, so from then on those kernels read their spilled private state back as zeros for the remaining life of the process. Nothing reports an error.

The single-file SYCL reproducer (linked under Source Code / Reproducer) does it in under ten seconds with a 1 GiB file, one device allocation and a warm page cache: 942 queue::memcpy calls totalling ~10 GiB, all reading the same region of the mapping.

0. Root cause: the reused private surface is never made resident again

CommandListCoreFamily::allocateOrReuseKernelPrivateMemory() (level_zero/core/source/cmdlist/cmdlist_hw.inl) declares the private surface resident only on the dispatch that allocates it:

if (!allocToReuseFound) {
    privateAlloc = kernelImp->allocatePrivateMemoryGraphicsAllocation();
    privateAllocsToReuse.push_back({sizePerHwThread, privateAlloc});
    this->commandContainer.addToResidencyContainer(privateAlloc);   // <-- only here
}
kernel->patchCrossthreadDataWithPrivateAllocation(privateAlloc);

Every later dispatch takes the cached allocation and only re-patches the pointer. So after the first dispatch, nothing keeps the surface's binding alive — and Drm::bindBufferObject() (shared/source/os_interface/linux/drm_neo.cpp) calls evictUnusedAllocations() whenever a bind ioctl fails:

auto ret = changeBufferObjectBinding(this, osContext, vmHandleId, bo, true, forcePagingFence);
if (ret != 0) {
    ...->evictUnusedAllocations(false, isAsyncFence);
    ret = changeBufferObjectBinding(...);          // retry
}

A file-backed mmap source is what makes those binds fail — the 512 EPERM DRM_IOCTL_XE_VM_BIND userptr binds in section 3 below. The sweep unbinds everything not always-resident, not locked, and past its completion tag (evictUnusedAllocationsImplevictImplmakeBOsResident(..., bind = false)), which includes that private surface. It is never re-bound, its VA stays unmapped, and because DG2 creates the VM with DRM_XE_VM_CREATE_FLAG_SCRATCH_PAGE (see "Why zeros, not stale data" in Additional Notes) the kernel's private reads return zeros and its private writes are dropped, silently, forever.

PrintBOBindingResult=1 shows exactly this, grepping the three arms for the 89 MiB (22784 × 4096) private surface:

arm private-surface bind history result
174 kernels (per-dispatch), file source bind; unbind during the copies; never re-bound wrong 131072
173 kernels (not per-dispatch), same copies bind; unbind at the same point; bind again before the second run clean
174 kernels, anonymous source bind once; unbind only at teardown clean

Thirteen other BOs are unbound by that same sweep and every one of them is re-bound by the next operation that lists it as resident. The per-dispatch private surface is the only casualty.

Proposed fix — restores the behaviour from before ee21f7c717:

     if (!allocToReuseFound) {
         privateAlloc = kernelImp->allocatePrivateMemoryGraphicsAllocation();
         privateAllocsToReuse.push_back({sizePerHwThread, privateAlloc});
-        this->commandContainer.addToResidencyContainer(privateAlloc);
     }
+    this->commandContainer.addToResidencyContainer(privateAlloc);
     kernel->patchCrossthreadDataWithPrivateAllocation(privateAlloc);

ee21f7c717 ("fix: Use cmdlist residency container for reused private allocs", 2023-09-18) replaced patchAndMoveToResidencyContainerPrivateSurface() — which did the patch and residencyContainer.push_back(alloc) on every dispatch — with a residency add in the allocation branch only. First release containing it: 23.39.27427.19. The code is still exactly this at the tip of master today, and git log -S 'addToResidencyContainer(privateAlloc)' on that file returns exactly one commitee21f7c717 — so nothing has touched it since.

Duplicates are not a concern: removeDuplicatesFromResidencyContainer() already runs in close() for regular command lists and at the top of executeCommandListImmediateWithFlushTaskImpl() for immediate ones, so the extra push per dispatch never reaches submission.

That patch is built and tested here. 26.18.38308.1 rebuilt with those two lines moved, loaded via ZE_ENABLE_ALT_DRIVERS, interleaved same-session A/B with each run's strace proving which libze_intel_gpu.so it opened:

arm SYCL reproducer (174 kernels, file source) PyTorch 2.13.0+xpu torch.cat reproducer
stock 26.18.38308.1 corrupt 3/3 (wrong 131072, untouched 65536) corrupt 3/3 (cat=24576 wrong every rep)
same build + the patch above clean 3/3 clean 3/3

And with the patch the bind trace shows the missing event appear: bind / unbind (during the copies) / bind again before the next dispatch — the same shape a module not on the per-dispatch path already had.

The SYCL column is the deterministic one (that build is corrupt 10/10 without the patch); the PyTorch reproducer is stochastic, so its 3/3 is corroboration rather than proof on its own. The application session below is the third, independent line of evidence.

The same thing in my real workload, which is where I hit this. One long-lived ComfyUI server, 30 image generations at a fixed seed, with PrintBOBindingResult=1 counting events for the 89 MiB private surface. I score each image by correlation against a known-good image from the same seed that I checked by eye — correct images land at 0.43-0.85 depending on kernel path, the failure mode at 0.13-0.20, and the two have never overlapped:

driver good images private surface shape of the session
stock 1 / 30 1 bind, 1 unbind correct at generation 1; wrong from generation 2 to 30
patched 30 / 30 8 binds, 8 unbinds every unbind followed by a re-bind; correct throughout

The flip is visible as it happens: generation 1 runs with unbinds=0 and is correct, generation 2 runs with unbinds=1 and is wrong, and it never recovers — which is exactly what I had been seeing long before I knew the mechanism: it worked, then at some point images started coming out corrupted, and only restarting the server helped. On the patched driver the eviction sweeps still occur seven more times and cost nothing, because the surface is made resident again.

Two independent confirmations of the mechanism, same reproducer, same file:

  • MakeEachAllocationResident=2 ("bind all created allocations in flush"): clean 3/3 where the default is corrupt 10/10.
  • OverrideNumComputeUnitsForScratch=8192 puts the 128-kernel build (clean by default) over the threshold → Per Dispatch 1corrupt, at unchanged kernel count and image size. The threshold, not the size, selects the broken path.

1. The discriminator is your own debug output

NEOReadDebugKeys=1 PrintDebugMessages=1, two builds of the identical source differing only in how many private-memory kernels share the module. The whole stderr diff is one line (the other is a PCI barrier address):

128 kernels:  Private Memory Per Dispatch 0 for modulePrivateMemorySize 11945377792 \
                  subDevices 1 globalMemorySize 16225245593
192 kernels:  Private Memory Per Dispatch 1 for modulePrivateMemorySize 17918066688 \
                  subDevices 1 globalMemorySize 16225245593

modulePrivateMemorySize is exactly Σ perHwThreadPrivateMemorySize × 4096 over the module's kernels (computeUnitsUsedForScratch: 4096 on this device, also from your debug output). When that sum exceeds globalMemorySize, the module is switched to per-dispatch private memory — and that is exactly the set of modules that get corrupted.

The corruption boundary and the flag flip coincide to the individual kernel. Three different privateMemSize values, adjacent kernel counts, 2/2 runs each:

privateMemSize kernels modulePrivateMemorySize vs 16,225,245,593 Per Dispatch after the copies
22,784 B 173 16,144,924,672 under 0 clean 2/2
22,784 B 174 16,238,247,936 over 1 corrupt 2/2
12,032 B 329 16,214,130,688 under 0 clean 2/2
12,032 B 330 16,263,413,760 over 1 corrupt 2/2
6,656 B 595 16,221,470,720 under 0 clean 2/2
6,656 B 596 16,248,733,696 over 1 corrupt 2/2

One kernel either side of the switch, at three per-kernel sizes that differ by 3.4x, is the difference between correct and wholly wrong output.

Consequences worth stating explicitly, each measured:

  • It is not image size and not kernel count. 595 kernels are clean at 6,656 B while 174 are corrupt at 22,784 B. Separately: a 5,673,584 B image holding 16 private-memory kernels is clean, and an image with 255 ordinary kernels plus one private-memory kernel is clean.
  • The budget is per module. Rebuilding the same 174 kernels with -fsycl-device-code-split=per_kernel — same total private-memory demand, same process, split into one module per kernel — gives 0 modules on the per-dispatch path and is clean 2/2.
  • Free device memory is irrelevant; the comparison is against the total. Holding 4 GiB or 8 GiB of device USM for the whole run changes nothing: 173 stays clean, 174 stays corrupt.

2. What the copies have to look like

All rows below are one build (174 kernels, the minimal corrupting one), each differing from the reference row in exactly one way:

variation result
942 copies from an mmap'd file, warm cache (reference) corrupt
the same, cold page cache corrupt
MAP_SHARED instead of MAP_PRIVATE corrupt
the file on tmpfs (/dev/shm) instead of disk corrupt
942 separate device allocations instead of one corrupt
source is an anonymous mmap clean
source is ordinary malloc'd host memory clean
11,304 copies, 120 GiB from anonymous memory (12x the work) clean
942 device allocations, no copies clean
nothing at all between the two kernel runs clean

So the copies are what matter, the allocations are not, the page cache is not, and the volume is not — the source must be a file-backed mapping. tmpfs behaving like disk says this is about the pages belonging to a file, not about I/O or writeback.

3. Both conditions are needed — and the staging churn alone is harmless

A file-backed source makes xe refuse to bind the source pages, so the runtime falls back to staging. strace -c -e trace=ioctl, everything else identical:

over-threshold, file (CORRUPT) under-threshold, file (clean) over-threshold, anon (clean)
ioctl calls 8,342 8,347 1,666
of which EPERM 512 512 0
DRM_IOCTL_XE_VM_BIND 3,653 3,655 571
DRM_IOCTL_XE_GEM_CREATE / GEM_CLOSE 326 / 326 326 / 326 70 / 70
DRM command 0x0a (DRM_XE_WAIT_USER_FENCE 3,141 3,143 571

¹ strace decodes command 0x0a under etnaviv's table as DRM_IOCTL_ETNAVIV_PM_QUERY_DOM; on xe that number is DRM_XE_WAIT_USER_FENCE (include/drm/xe_drm.h:105).

Every one of the 512 errors is the same call:

ioctl(3, DRM_IOCTL_XE_VM_BIND, ...) = -1 EPERM (Operation not permitted)

The middle column is the important one. A module under the threshold issues the very same 512 failed binds, the same staging traffic and the same GEM churn — and comes through perfectly correct. So the staging path is not by itself the defect; it is only lethal to a module on the per-dispatch private-memory path.

4. Within a corrupted module, only privateMemSize > 0 kernels are wrong

The reproducer builds an in-image control: the same copy written with scalar arguments, privateMemSize = 0, in the same module, run on the same queue at the same moment.

kernel, same module, same queue wrong elements never written
privateMemSize = 22784 131072 / 131072 65536
privateMemSize = 0 0 0

This is why the fault looks op-specific in PyTorch: of 20 ops checked against CPU in a corrupted process, only torch.cat and torch.stack are wrong (index_select, gather, scatter, index_put, take, embedding, contiguous, clone, repeat, roll, flip, narrow+copy_, add, where, sum, argmax, matmul, _foreach_add are all correct). Those report privateMemSize=0; cat's reports 22,784, and cat's module is the only one in the process that prints Per Dispatch 1.

5. What the kernel-driver trace does and does not show

Traced as root with xe tracepoints (xe_bo_create, xe_bo_move, xe_bo_validate, xe_vma_bind, xe_vma_unbind, xe_vma_evict, xe_vma_invalidate, xe_vm_rebind_worker_exit) over the same 2x2, with the corrupting arm as an in-run positive control and buffer integrity checked (entries-in-buffer == entries-written in all four arms):

over-file (CORRUPT) over-anon under-file under-anon
xe_bo_create of the 89 MiB private surface 1 1 1 1
xe_vma_evict 0 0 0 0
xe_vma_invalidate 0 0 0 0
xe_bo_move, vram0 → system 0 0 0 0
  • Nothing is migrated in any arm — no TTM move, no xe_vma_evict, no xe_vma_invalidate. Every xe_bo_move in every arm is system → vram0 and they are all over before the copies start. So the surface is never moved out of VRAM; its mapping is what disappears.

  • xe_vma_unbind in the same traces confirms the root cause from the KMD side. Counting events for the private surface's exact range (start=0xd556aaa00000, i.e. the 89 MiB surface):

    arm xe_vma_bind xe_vma_unbind
    over-threshold, file (CORRUPT) 1 1
    over-threshold, anon (clean) 1 1 (at teardown)
    under-threshold, file (clean) 2 2
    under-threshold, anon (clean) 1 1 (at teardown)

    The under-threshold file arm is unbound by the same sweep and re-bound before its next dispatch; the corrupting arm is unbound and never re-bound. Two structurally different instruments — your PrintBOBindingResult in the UMD and xe tracepoints in the KMD — agree line for line.

  • Address layout is identical across the four arms, and the file arm's staging binds are at host VAs nowhere near the private surface.

  • The copies do not write outside their destination. With 8 MiB sentinel-filled guard allocations either side of the destination slab: 0 of 2,097,152 words disturbed on each side, in a run that simultaneously reported REPRODUCED: 131072 of 131072.

6. The kernel arguments reaching the GPU are provably correct

Interposing zeKernelSetArgumentValue shows the 1376-byte by-value argument is byte-correct at the call in a corrupted process — right pointers, right offsets, right element counts, identical to a healthy run. Re-issuing that same block immediately before zeCommandListAppendLaunchKernel (14 times, verified) changes nothing. The kernel behaves as though the fields it reads after its first branch — the ones the compiler spilled to private memory — were zero, while a field read before the branch and kept in a register survives.

7. It is silent, and permanent for the process

Every zeKernelSetArgumentValue, zeCommandListAppendLaunchKernel, zeCommandQueueExecuteCommandLists, zeEventHostSynchronize and zeFenceHostSynchronize returns ZE_RESULT_SUCCESS (118,650 launches / 39,735 copies / 31,855 syncs checked in one run). Nothing recovers it: freeing the whole allocation, a fresh queue, 200 unrelated kernels and 50 further launches all leave it broken. Only a new process.

Ordering decides it and it is not a one-time event: a module created after the copies is healthy (3/3), and the same module is then broken by a second round of copies (2/2, inside one process). So it is not "the first big allocation poisons the device" — any such copy sequence damages whatever qualifying modules are resident at that moment.

Expected Behavior

A dispatch must keep its own private surface resident. Concretely: a private allocation taken from the per-dispatch reuse cache should be added to the residency container on every dispatch that uses it, as it was before ee21f7c717, so that a later evictUnusedAllocations() sweep cannot leave a kernel's private memory unmapped.

More generally: copying data to the device must not change the behaviour of unrelated, already-loaded device modules, and a kernel must read the private state it wrote and the arguments the host set. Silently returning zeros is the worst case for a compute workload: in my case it produced plausible-looking but entirely wrong output with no diagnostic anywhere. If a private surface genuinely cannot be kept mapped, failing the launch would at least be visible.

Reproduction Rate

Always reproduces - 100%

Steps to Reproduce

# any file >= 512 MiB; its contents are never used
dd if=/dev/urandom of=/tmp/src.bin bs=1M count=1024

icpx -fsycl -O2 -DBLOAT_N=174 xpu_image_corruption_repro.cpp -o repro
./repro --file /tmp/src.bin --same-src --one-alloc --no-drop-cache   # exit 1

# the same source, one kernel fewer — under the threshold
icpx -fsycl -O2 -DBLOAT_N=173 xpu_image_corruption_repro.cpp -o repro173
./repro173 --file /tmp/src.bin --same-src --one-alloc --no-drop-cache  # exit 0

# the same 174 kernels, split per kernel
icpx -fsycl -O2 -DBLOAT_N=174 -fsycl-device-code-split=per_kernel \
     xpu_image_corruption_repro.cpp -o repro_pk
./repro_pk --file /tmp/src.bin --same-src --one-alloc --no-drop-cache  # exit 0

# and which side of the switch any build is on:
NEOReadDebugKeys=1 PrintDebugMessages=1 ./repro --file /tmp/src.bin \
    --same-src --one-alloc --no-drop-cache 2>&1 | grep 'Per Dispatch'

# the root cause, straight out of your own logging: the 89 MiB private surface
# is bound once, unbound during the copies, and never bound again
NEOReadDebugKeys=1 PrintBOBindingResult=1 ./repro --file /tmp/src.bin \
    --same-src --one-alloc --no-drop-cache 2>&1 | grep 93323264
# repro:    bind ... size: 93323264      <- one bind
#           unbind ... size: 93323264    <- during the trigger, never re-bound
# repro173: bind / unbind / bind / unbind   <- re-bound before the 2nd run

# keeping every allocation resident avoids it entirely (clean 3/3)
NEOReadDebugKeys=1 MakeEachAllocationResident=2 ./repro --file /tmp/src.bin \
    --same-src --one-alloc --no-drop-cache   # exit 0

What the reproducer does, and why each step is there:

  1. Builds a module holding BLOAT_N kernels that require private memory (privateMemSize=22784 each on this device; -DBATCH_SIZE varies that).
  2. Runs one of them once, so the module is created, and checks it — correct.
  3. Issues 942 queue::memcpy calls, ~10 GiB total, from an mmaped file into a single device allocation, all reading the same region.
  4. Runs the same kernel again into an output buffer pre-filled with a sentinel, so "never written" and "written with zeros" are distinguishable. This matters: a freshly allocated output buffer is recycled memory and will mislead you about which of the two happened.
  5. Prints the result and exits 1 if corrupted.

Expected output at BLOAT_N=174:

before upload : wrong 0, untouched 0   (private=0 control: wrong 0)
  10.04 GiB, 942 copies, 1 device allocations (4 s, 2481 MB/s)  [warm cache]
after upload  : wrong 131072, untouched 65536   (private=0 control: wrong 0)

The wrong/untouched counts are identical on every corrupting run; only the copy line's time and throughput vary (4-8 s warm on this machine).

Observed rate (the evidence behind the Reproduction Rate dropdown): 10 / 10 consecutive runs in that minimal form, identical wrong-element counts every time, plus 2/2 at each of the six boundary points in the table under Actual Behavior. Once it fires it is permanent for that process.

At the PyTorch level the same day, interleaved A/B on the same 12.24 GiB file, each arm printing its own build version: stock 2.14.0.dev20260811+xpu reproduced 2/2, the same commit rebuilt with per_kernel was clean 2/2. (The PyTorch-level reproduction is stochastic because its caching allocator batches the copies differently run to run; the direct SYCL form above is not.)

Is this a regression?

  • Yes, this is a regression - functionality that previously worked is now broken

Last Known Working Driver Version

23.35.27191.42 - newest tag not containing ee21f7c (git tag --contains); not run here: pre-xe-uAPI, will not zeInit on this machine

First Known Failing Driver Version

23.39.27427.19 - first tag containing ee21f7c (git tag --contains); reproduced here on 26.18.38308.1, which still carries the same code

API Call Logs

Captured with a small dlsym interposer rather than unitrace (the UR Level Zero adapter dlopens libze_loader and resolves through dlsym, so plain LD_PRELOAD symbol interposition does not see the calls). Key results, all from corrupted processes:

[ze_count] launch=118650 err=0 | exec=... err=0 | copy=39735 err=0
           | evsync=31855 err=0 | fence=... err=0 | last_rc=0x0
[ze_count] setarg=... err=0 | indirect=... err=0 | allocdev=... err=0

The 1376-byte by-value argument decoded at zeKernelSetArgumentValue in a corrupted process, byte-identical to the healthy case:

[ze_argdump] arg1 sz=1376 zerobytes=715/1376 hash=3b65d4f7e8310d30
             output=ffffffffff870000
             input[0..1]=ffffffffff860000 ffffffffff861400
             offset[0..2]=0 5 12  dimSize[0..2]=5 7 11
             nElements[0..2]=1280 1792 2816
             -> launch groups 3 x 3 x 1

and the kernel properties that separate affected from unaffected kernels (this is PyTorch's cat, whose module is the one printing Per Dispatch 1; its five by-value arguments total 1424 B, of which the metadata struct above is 1376 B):

[ze_argdump] props args=5 local=0 private=22784 spill=0 maxSG=16  ...CatArrayBatchedCopy_alignedK_contig...
[ze_argdump] props args=8 local=0 private=0     spill=0 maxSG=32  ...VectorizedGatherKernel...
[ze_argdump] props args=3 local=0 private=0     spill=0 maxSG=32  ...VectorizedElementwiseKernel...AddFunctor...

Full logs available on request; I can re-capture with unitrace if you prefer that format.

strace Logs

The ioctl counts are in "Actual Behavior" section 3 — they are a result, not a crash log. Full traces available on request. There is no failing syscall other than the 512 EPERM from DRM_IOCTL_XE_VM_BIND, which occur identically in a run that stays correct.

System Logs / dmesg Output

The corrupting runs are silent, and that is part of the report. Marking the journal, running three corrupting reproductions and reading back:

$ MARK=$(date '+%F %T'); ./repro ... ; ./repro ... ; ./repro ...   # 3x REPRODUCED
$ journalctl -k --since "$MARK"
-- No entries --

Zero kernel messages for the reported flow. The machine's log is not empty over the week, though, and rather than let that look like something I hid, here is what dmesg | grep -i -E 'i915|xe|drm|gpu' | tail -n 100 contains after 7.1 days of uptime spent entirely on this investigation — a week of deliberate stress tests, not a normal machine:

[ 22521.463365] xe 0000:09:00.0: [drm] VM worker error: -12
[ 22531.553533] xe 0000:09:00.0: [drm] exec queue reset detected
...
[ 31669.172346]  handle_mm_fault+0xee/0x2f0
[ 31669.172346]  drm_gpusvm_get_pages+0x203/0x910 [drm_gpusvm_helper]
[ 31669.172359]  xe_vma_userptr_pin_pages+0xc2/0xd0 [xe]
[ 31669.172514]  vm_bind_ioctl_ops_parse+0x336/0x970 [xe]
[ 31669.172577]  xe_vm_bind_ioctl+0xd15/0x1bb0 [xe]
...
[112739.457009] xe ...: [drm] Tile0: GT0: Timedout job: seqno=4294967169, ... in benchdnn [295165]
[112739.511099] xe ...: [drm] Xe device coredump has been created
[211691.581057] python[545244]: segfault at 776f792a4000 ip 0000776eeb388469 \
                  error 4 in libze_intel_gpu.so.1.15.38308[788469,776eeac00000+ab2000]
[528778.547894] repro256[1538897]: segfault at 10 ip 000071448080c610 \
                  error 4 in libze_intel_gpu.so.1.15.38308[80c610,714480000000+ab2000]
[592707.436707] xe ...: [drm] Tile0: GT0: Engine memory CAT error: class=bcs, guc_id=6

What each one is:

  • VM worker error: -12 and exec queue reset detected — recurring through the week. The -ENOMEM is consistent with the VRAM-oversubscription tests I ran deliberately (holding 15 GiB on a 16 GB card), and the 2026-08-16 cluster is definitely those; I have not traced every earlier cluster to a specific experiment, so I will not claim more than that. What I can state positively is the marked-window result above: a corrupting reproducer run emits nothing.
  • The xe_vm_bind_ioctlxe_vma_userptr_pin_pagesdrm_gpusvm_get_pageshandle_mm_fault stack (2026-08-10) is a host page-allocation failure during a userptr pin, from the same oversubscription work. I mention it only because it is in the same bind path this report's trigger exercises: per the mechanism, any failed bind runs evictUnusedAllocations(), so a machine that produces real -ENOMEM bind failures has a second way to reach the same sweep. I have not tried to reproduce corruption that way.
  • Timedout job ... in benchdnn + coredump (2026-08-11) — unrelated oneDNN microbenchmark work.
  • Two userspace segfaults inside libze_intel_gpu.so.1.15.38308, at different sites, both unreproduced — see the last item of "What was open" in Additional Notes.
  • Engine memory CAT error: class=bcs (2026-08-17 09:53) — a single occurrence that I cannot attribute to a particular experiment. It is a copy engine, not the compute engine the reproducer uses, and it postdates every measurement in this report.

Happy to attach the complete dmesg as a file if you want the surrounding context of any of these.

Backtrace (if crash or hang occurred)

Not applicable — no crash and no hang in the reported flow. The application runs to completion and produces wrong numbers.

Source Code / Reproducer

One file, ~680 lines, SYCL onlyxpu_image_corruption_repro.cpp, in a gist at https://gist.github.com/AIVirtuoso/0995fd491bd822f91cd6008e5fdf0ac3

No PyTorch, no model weights, no other dependency; the only external input is any file of at least 512 MiB, whose contents are never used.

It contains the kernel, the trigger and the check, and every control is a command-line flag, so each row of the tables above is one invocation:

--same-src     read the same region every copy (so a 1 GiB file suffices)
--one-alloc    one device allocation instead of 942
--no-copy      allocate but never copy
--host-src     copy from anonymous host memory instead of a mapping
--anon-mmap    copy from an anonymous mmap
--shared       MAP_SHARED instead of MAP_PRIVATE
--skip-upload  do nothing between the two kernel runs
--no-drop-cache  leave the page cache warm
--repeat N     N times the copy volume
--guard N      N MiB sentinel-filled guard allocations either side of the slab
--ballast G    hold G GiB of device USM for the whole run
--extra-lib    dlopen another module and check its kernel too
-DBLOAT_N=k    k private-memory kernels in the module (173 clean / 174 corrupt)
-DBATCH_SIZE=k sets privateMemSize (64→22784 B, 32→12032 B, 16→6656 B)

The kernel is a transcription of the concatenation kernel that first showed this in a real workload; the only thing that matters about it here is that IGC gives it privateMemSize=22784. The privateMemSize=0 in-module control is built in and reported on every run.

The same gist also contains xpu_cat_corruption_repro.py, the PyTorch-level form of the same fault (torch + safetensors only, generates its own 12.24 GiB file, ~2.5 min, stochastic). Happy to attach either file to this issue directly if you would rather have them here than in a gist.

Command Line / Application Details

./repro     --file /tmp/src.bin --same-src --one-alloc --no-drop-cache  # exit 1, corrupt
./repro173  --file /tmp/src.bin --same-src --one-alloc --no-drop-cache  # exit 0, correct
./repro_pk  --file /tmp/src.bin --same-src --one-alloc --no-drop-cache  # exit 0, correct

Runtime is under ten seconds each. Exit status is 0 when the kernel is still correct and 1 when the corruption reproduced, so it drops straight into a CI job.

oneAPI Version (if applicable)

Reproduces on both:

Intel(R) oneAPI DPC++/C++ Compiler 2026.0.0 (2026.0.0.20260331)
Intel(R) oneAPI DPC++/C++ Compiler 2026.1.0 (2026.1.0.20260617)

with matching intel-sycl-rt / intel-cmplr-lib-ur 2026.0.0 and 2026.1.0 respectively. Level Zero loader 1.28.5, IGC 2.34.4, gmmlib 22.10.0.

Screenshots / Video

No response

Additional Notes

On the checklist. The driver is 26.18.38308.1, the latest packaged in nixpkgs. Two older releases were downloaded and attempted (23.35.27191.9 and 23.39.27427.23) but neither initialises on this machine — see On regression — so 26.18.38308.1 is the only release with a result, in either direction.

On regression. The box is ticked on code archaeology plus a controlled A/B in current code — not on running the old releases, which I tried and could not do here. Precisely what is established:

  1. The commit made an unconditional residency declaration conditional. ee21f7c717 (2023-09-18) replaced an unconditional kernelImp->patchAndMoveToResidencyContainerPrivateSurface(privateAlloc) — which did patchCrossthreadDataWithPrivateAllocation() and residencyContainer.push_back(alloc) — with a bare patchCrossthreadDataWithPrivateAllocation(), moving the residency add into the if (!allocToReuseFound) branch.
  2. Before it, reused surfaces were resident on every dispatch. At ee21f7c717^, appendLaunchKernelWithParams copied the whole kernel residency container into the command container on every append (cmdlist_hw_xehp_and_later.inl:357-363), so the per-dispatch push reached the submission list each time.
  3. All three neighbouring commits were correct. c06ddfc7b8 (2022-11-08) introduced the per-dispatch path allocating and declaring residency every dispatch; 5807d512b3 (2023-08-31) introduced the reuse cache with the residency call still unconditional; 3b3e17e738 (2023-09-04) introduced the if (!allocToReuseFound) block but left residency outside it. Nothing has changed it since: git log -S 'addToResidencyContainer(privateAlloc)' on that file returns exactly ee21f7c717, and the tip of master today still has the add inside the branch.
  4. Restoring the pre-commit semantics on today's driver fixes it — the two-line-move A/B above, corrupt 3/3 → clean 3/3, plus the 30-generation application session. This isolates the semantic the commit changed, on current code, on the affected hardware.

The one thing I could not do is run the 23.35 and 23.39 builds themselves. The exact boundary tags have no published binaries, but the two release lines do — intel-level-zero-gpu_1.3.27191.9 (23.35.27191.9, which git tag --contains confirms is without the commit) and 1.3.27427.23 (23.39.27427.23, with it) — so I downloaded both with their matching libigdgmm12 and loaded each via ZE_ENABLE_ALT_DRIVERS: the loader dlopens them fine and then zeInit returns ZE_RESULT_ERROR_UNINITIALIZED (and ZE_RESULT_ERROR_UNSUPPORTED_FEATURE on one of the loader's init paths), leaving no GPU for the SYCL runtime to select. I did not isolate why beyond the obvious candidate: this A770 is bound to the xe KMD, and those builds date from before the xe uAPI settled — their linux/xe/ still issues DRM_IOCTL_XE_MMIO, which current NEO no longer uses at all. Note for anyone repeating this: without ONEAPI_DEVICE_SELECTOR=level_zero:gpu the SYCL runtime silently falls back to the OpenCL backend, which does not have the bug, and the reproducer then prints a clean result under a driver that never loaded — the version banner is the tell. On an i915-bound machine the 23.35-vs-23.39 bisect should be straightforward, and I would expect it to be decisive; I am happy to boot this card on i915 and run it myself if that would help.

It reproduces on everything I have tested, which spans two oneAPI runtimes and two compilers on the one driver:

reproduces
oneAPI 2026.0.0 / DPC++ 2026.0.0 / torch 2.13.0+xpu yes
oneAPI 2026.1.0 / DPC++ 2026.1.0 / torch 2.14.0.dev20260811+xpu yes

Workaround for application code. -fsycl-device-code-split=per_kernel, which keeps every module under the threshold and off the per-dispatch path. The documentation says this option "is not expected to produce incorrect code", i.e. it should be semantically neutral — that changing it changes correctness is itself the argument that this is a driver defect rather than an application one. On PyTorch's real Shape.cpp it takes one 542,952-byte image to 160 images of at most 6,588 bytes, and the shipped module from 3,444,848 B to 33,248 B.

Debug keys tried, none of which changes the outcome (with PrintDebugSettings=1 confirming the runtime echoes each key it accepts as "Non-default value of debug variable"): EnableCopyWithStagingBuffers=0 on the file arm and =1 on the anonymous arm, TreatNonUsmForTransfersAsSharedSystem =0/=1, EnableBOMmapCreate=0 (forcing GEM_USERPTR), EnableDeviceUsmAllocationPool=0, EnableUsmAllocationPoolManager=0.

No key found that controls the per-dispatch switch itself. ForcePerDispatchPrivateMemorySize, AllocatePrivateMemoryPerDispatch, ForceAllocatePrivateMemoryPerDispatch, DisablePrivateMemoryPerDispatch, EnablePrivateMemoryPerDispatch, ForcePrivateMemoryPerDispatch and PrivateMemoryPerDispatch are all unknown to this driver — PrintDebugSettings=1 echoes only real variables and echoed none of them.

Keys that do change the outcome, all of them evidence for the residency diagnosis rather than workarounds:

  • MakeEachAllocationResident=2clean 3/3 on the arm that is otherwise corrupt 10/10.
  • OverrideNumComputeUnitsForScratch — moves the threshold in both directions. =8192 makes the clean 128-kernel build corrupt; =2048 puts the 174-kernel build under the threshold but undersizes the surface, so its kernel is wrong before any copies (wrong 65536, untouched 32768), which incidentally shows that 4096 (= 512 EUs × 8 threads) is the real thread count and not a conservative guess.
  • MakeEachAllocationResident=1 hangs this reproducer before its first print. Not investigated; mentioning it in case it is unexpected.
  • DisableScratchPages=1 does not make the failure loud: PrintXeLogs=1 confirms getFlagsForVmCreate 1,0,1 and gemVmCreate f=0x2 (no scratch-page bit), and the run still ends wrong 131072, untouched 65536 with an empty journalctl -k.

Also ruled out, each by measurement: the by-value kernel argument (a hand-written kernel taking a byte-identical 1424-byte argument set, indexed the same way, is correct in the same corrupted process — as long as its module is small); argument delivery; module size in bytes; kernel count; the allocations; the page cache; H2D volume; the device USM pool; memory pressure and occupancy; and the hardware (30 minutes of saturating dense GEMM with retention checks is bit-exact, and a second process on the same GPU at the same moment is correct while the first is corrupt).

How I ran into this. My image-generation workload (ComfyUI on an A770) started producing pure noise after loading a 12.5 GB checkpoint, with no error anywhere. The shape of it is the session table under Actual Behavior: a fresh process is almost always correct for its first generation and wrong for every one after, which from the outside just looks like "it worked, then it started producing garbage, and only a restart helps". It took me a long time to get from that symptom to this report; the silence is the expensive part.

What was open, and what the answers turned out to be

This report originally carried nine open questions. Seven now have answers, from your source plus new measurements on this hardware; two are still open (items 7 and 9), as is one sub-part of item 4 — how to make this class of fault loud on DG2.

  1. What does Private Memory Per Dispatch 1 change? Ownership of the private surface and, decisively, who keeps it resident. Off the per-dispatch path, KernelImp::initialize() allocates it per kernel and puts it in the kernel's internalResidencyContainer, so every submission makes it resident again. On the per-dispatch path, allocateOrReuseKernelPrivateMemory() declares residency only on the dispatch that allocates it — see the "Root cause" section at the top. The path itself is sound; the bookkeeping is not.

  2. Does it collide with the staging allocator? No — they never touch the same memory. The private surface's range is left unmapped, and the staging BOs bound afterwards are at unrelated host VAs. The link is the error path: Drm::bindBufferObject() calls evictUnusedAllocations() on any failed bind, and a file-backed source is what makes binds fail.

  3. A supported way off the path, or a loud failure? Not needed as a knob if the residency add is restored. For the record: MakeEachAllocationResident=2 avoids the fault (3/3), OverrideNumComputeUnitsForScratch shifts the threshold but is unsafe downwards, and application-side the only clean answers are -fsycl-device-code-split=per_kernel or not sourcing H2D copies from a file-backed mmap.

  4. Why zeros, not stale data? The pages are not mapped at all: writes to an unbound VA are dropped and reads return zeros, which is exactly the observed "half never written, everything written is zero". DG2 absorbs it silently because isDisableScratchPagesSupported() is false before Xe2, so IoctlHelperXe::getFlagsForVmCreate() sets DRM_XE_VM_CREATE_FLAG_SCRATCH_PAGE. Xe2 and later default to disabling scratch pages, which may be why this has only shown up on Arc A-series. Still open: forcing DisableScratchPages=1 on DG2 does not make it loud (details under "Debug keys"), so I have no way to turn this class of fault into an error on this platform. If you know one, I will use it.

  5. Which globalMemorySize is authoritative? They are the same number: Device::getGlobalMemorySize() returns 16,225,245,593 and deviceInfo.globalMemSize is that value alignDowned to a page (16,225,243,136) — the 2,457 B is the alignment remainder. The per-dispatch decision reads the unaligned value, which is why it predicts the boundary points; KernelHelper::checkIfThereIsSpaceForScratchOrPrivate() reads the aligned one. Worth making consistent, but not the bug.

  6. Which versions to bisect? Answered from the history instead — see On regression above. ee21f7c717 (2023-09-18), first contained in 23.39.27427.19 by git tag --contains; the newest tag without it is 23.35.27191.42. I could not run either era on this machine (xe KMD), so if you want the bisect confirmed on silicon it needs an i915-bound card.

  7. Does it depend on xe vs i915? — STILL OPEN. The defect is above the KMD (Level Zero residency bookkeeping) and the sweep is generic DRM code, so I expect it to be KMD-independent; but the trigger here is xe returning EPERM for a userptr bind of file-backed pages, and I have not tested whether i915 refuses the same import or whether the upstream i915 path (DrmMemoryOperationsHandlerDefault, residency expressed per-execbuf) is exposed at all. Switching this machine needs a reboot with i915.force_probe=56a0; say the word and I will run the 2x2 again there.

  8. Is computeUnitsUsedForScratch = 4096 intended for DG2? It is exact, not conservative: maxSubSlice × MaxEuPerSubSlice × (ThreadCount / EUCount) = 32 × 16 × 8 = every hardware thread on an A770, and lowering it to 2048 makes the kernel wrong before any copies, so the slots above 2048 really are used. Mild remaining question: the module sum in checkIfPrivateMemoryPerDispatchIsNeeded() adds every kernel's whole-device surface as though all of them were resident simultaneously, which is what makes a 174-kernel module of a 22 KB/thread kernel exceed a 16 GB card. With the residency fix that is a tuning question, not a correctness one — but if the sum is meant to be an upper bound on concurrent private memory, it is very pessimistic.

  9. Two libze_intel_gpu segfaults — STILL OPEN, and probably separate. Both are in the dmesg excerpt above, at different sites, and I symbolized both against the shipped library:

    • repro256, 2026-08-16: ip 0x80c610 is the first instruction of NEO::MultiGraphicsAllocation::getGraphicsAllocation(uint32_t) const (mov 0x10(%rdi),%rax) with %rdi == 0 — a null-object dereference (faulting address 0x10). Consistent with an unchecked null from a lookup for a pointer the driver does not own.
    • python, 2026-08-13: ip 0x788469 is inside NEO::CommandStreamReceiver::baseWaitFunction(unsigned long volatile*, NEO::WaitParams const&, unsigned long) at mov (%r15),%rax — a read of the tag pointer it is polling, faulting on a mapped-looking address (0x776f792a4000) rather than on null.

    Neither is reproducible: I could not trigger the first in four attempts (--extra-lib; memcpy to null / freed / non-USM pointers all behave correctly), and the second happened once. I am not claiming either is this bug — I list them because they are visible in the logs above and both are inside your library. The related artefact that is reproducible is oversubscription: holding 15 GiB of device USM on a 16 GB card gives UR_RESULT_ERROR_DEVICE_LOST plus xe 0000:09:00.0: [drm] VM worker error: -12. I will file separately if I can reduce any of them.

This issue was created with the help of an LLM.

Metadata

Metadata

Assignees

No one assigned

    Labels

    OS: LinuxIssue specific to Linux distributions (Ubuntu, Fedora, RHEL, etc.)Type: BugGeneral bug report, unexpected behavior or crash

    Type

    No type

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions