Skip to content

JIT: Safe overlapping struct assignment corrupts data on current main and .NET 10 #135183

Description

@benaadams

A fully safe C# struct assignment corrupts a partially overlapping destination on current main and .NET 10.0.12 Windows x64. No unsafe code, pointer casts, GC references, UInt256 dependency or arithmetic is needed.

Reproduction

Create a console project targeting net10.0, build Release, and use this complete program:

using System;
using System.Runtime.CompilerServices;
using System.Runtime.InteropServices;

[StructLayout(LayoutKind.Sequential)]
struct Words { public ulong A, B, C, D; }

[StructLayout(LayoutKind.Explicit, Size = 40)]
struct Overlap
{
    [FieldOffset(0)] public Words Input;
    [FieldOffset(8)] public Words Output;
}

class Program
{
    [MethodImpl(MethodImplOptions.NoInlining)]
    static void Copy(in Words input, out Words output) => output = input;

    static int Main()
    {
        Overlap value = new() { Input = new Words { B = 1 } };
        Copy(in value.Input, out value.Output);
        Console.WriteLine($"[{value.Output.A}, {value.Output.B}, {value.Output.C}, {value.Output.D}]");
        return value.Output.A == 0 && value.Output.B == 1
            && value.Output.C == 0 && value.Output.D == 0 ? 0 : 1;
    }
}

PowerShell:

dotnet build -c Release
$env:DOTNET_EnableHWIntrinsic = '0'
$env:DOTNET_TieredCompilation = '0'
$env:DOTNET_JitDisasm = 'Program:Copy'
dotnet --fx-version 10.0.12 bin/Release/net10.0/YourProject.dll

Expected: [0, 1, 0, 0], exit 0.
Actual: [0, 1, 1, 0], exit 1.

The same copy also fails with default tiered execution. On the tested AVX-capable machine, enabling hardware intrinsics produces a single 32-byte load followed by a store and the sample passes.

CIL and native instructions

Copy emits:

ldarg.1
ldarg.0
ldobj Words
stobj Words
ret

Both .NET 10.0.12 and the freshly built main JIT emit this in FullOpts with intrinsics disabled:

movups xmm0, xmmword ptr [rcx]
movups xmmword ptr [rdx], xmm0
movups xmm0, xmmword ptr [rcx+0x10]
movups xmmword ptr [rdx+0x10], xmm0
ret

Here rdx = rcx + 8. The first store overwrites bytes needed by the second load. ldobj followed by stobj should copy the loaded value; the IL does not use cpblk. The union and its typed in/out references are expressed entirely in safe C#.

Confirmed scope

  • .NET 10.0.12 Windows x64, SDK 10.0.401. Installed runtime .version: 95017c711e6afc1085133d440e42b4bd78155701.
  • Current upstream main at d90dbf43be153cc7ba7f49bb271e0b7e56a81891, reporting .NET 12.0.0-dev. Release JIT, VM, corerun and CoreLib freshly built in an isolated clean checkout; other managed framework libraries retained from an existing local testhost. JIT-enabled native build (FEATURE_DYNAMIC_CODE_COMPILED=1), ReadyToRun disabled. The captured instructions above confirm the method used the fresh JIT.
  • Also reproduced the FullOpts copy on the installed .NET 11.0.0-preview.7.26381.103 Windows x64 runtime.

There is also a concrete library manifestation: multiplying one by 256-bit limbs [0, 1, 0, 0], with output eight bytes into that operand, takes a struct-copy shortcut. On .NET 10.0.12, default tiered execution gives [0, 1, 1, 0], and FullOpts gives [0, 0, 0, 0]. On tested main, default tiered execution still gives [0, 1, 1, 0], while FullOpts gives the correct [0, 1, 0, 0]. The same library DLL was used for both runtimes. The standalone reproduction above isolates the remaining block-copy failure from multiplication and inlining.

Related history

Related to #7539, whose 2017 reproduction uses unsafe pointer casts. This issue supplies a safe typed-reference reproduction and confirmed current-main/released-runtime scope. Also related to #133858 / #133877 and #134411: the promotion fix does not eliminate the general block-copy failure, and #134411 already notes the interaction with #7539.

Because this remains reproducible on current main, the scope includes a current JIT fix as well as consideration of supported release backports.

Suggested fix direction (not yet implemented or benchmarked)

Preserve the loaded-value snapshot when lowering struct assignment to a block copy. In Lowering::LowerCopyBlockStore, copies with potentially overlapping source and destination should not use an interleaved load/store expansion unless non-overlap has been established.

For this 32-byte, GC-reference-free x64 case, the existing BlkOpKindUnrollMemmove / CodeGen::genCodeForMemmove path already provides the required load-all-before-store ordering, with register allocation support in lsraxarch.cpp. With intrinsics disabled, the desired sequence is:

movups xmm0, xmmword ptr [rcx]
movups xmm1, xmmword ptr [rcx+0x10]
movups xmmword ptr [rdx], xmm0
movups xmmword ptr [rdx+0x10], xmm1
ret

A lowering change would need to use the memmove unroll threshold and register budget, and prepare source/destination addresses as that path expects; simply switching gtBlkOpKind on an already contained memcpy node would not be sufficient. Larger copies need an overlap-safe fallback, such as the applicable memmove helper or an explicit snapshot temporary. Proven-disjoint copies can retain the existing expansion. GC-containing layouts, volatile accesses and other architectures require their own existing safety constraints to be preserved; this reproduction is the reference-free x64 case.

Suggested regression coverage: the safe explicit-layout reproduction with intrinsics enabled and disabled; default tiered and FullOpts execution; destination before and after source; exact alias and disjoint storage; and sizes crossing scalar/SIMD chunk boundaries. The standalone no-inline copy should remain in coverage independently of the physical-promotion regression.

This is a source-based implementation suggestion, not a tested patch. It addresses the remaining block-copy failure; the distinct .NET 10 FullOpts promotion manifestation still needs separate backport assessment.

Activity

  1. added
    area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI
    on Oct 4, 2026
  2. dotnet-policy-service commented on Oct 4, 2026

    @dotnet-policy-service
    Contributor

    Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
    See info in area-owners.md if you want to be subscribed.

  3. EgorBo commented on Oct 4, 2026

    @EgorBo
    Member

    Yes, it's exactly the same issue as #7539
    We discussed the fix here #134411 (comment)

    Not saying we shouldn't fix it, but it likely will be a regression for many existing (not related to this edge-case) cases, because I assume in most cases JIT won't be able to tell that two pointers don't overlap.
    This bug effectively requires unsafe code (under unsafe-v2, although the developer can mark this one as explicit safe) to reproduce.

  4. benaadams commented on Oct 4, 2026

    @benaadams
    MemberAuthor

    Not saying we shouldn't fix it, but it the fix likely will be a regression for many existing (not related to this edge-case) cases, because I assume in most cases JIT won't be able to tell that two pointers (of the same type) don't overlap. This bug effectively requires unsafe code (under unsafe-v2, although the developer can mark this one as explicit safe) to reproduce.

    It occurs with AVX off though in this example (guess it depends on size of struct and vector size?); but also uses AVX instructions creating the bug 🤷‍♂️

    I assume it doesn't occur when instrinsics are on because it uses AVX2 for the copy

    So it is a bit weird

  5. benaadams commented on Oct 4, 2026

    @benaadams
    MemberAuthor

    This bug effectively requires unsafe code (under unsafe-v2, although the developer can mark this one as explicit safe) to reproduce.

    You used to be fun 😉

    The context here is I am trying to formally verify some CIL however this JIT behaviour means the result is indeterminate/undefined which is very problematic

    What is unsafe in v2? Refs, field offsets, out?

  6. benaadams commented on Oct 4, 2026

    @benaadams
    MemberAuthor

    Ah overlapping fields/struct unions 🤔

  7. EgorBo commented on Oct 4, 2026

    @EgorBo
    Member

    It occurs with AVX off though in this example (guess it depends on size of struct and vector size?); but also uses AVX instructions creating the bug 🤷‍♂️

    It doesn't need SIMD to run into this behavior, can be reproduced even with all SIMD disabled

    What is unsafe in v2? Refs, field offsets, out?
    Ah overlapping fields/struct unions 🤔

    All fields in ExplicitLayout structs now require explicit safe or unsafe keyword, see your example under the new rules: link
    It's done because you can do many very evil/unsafe things with ExplicitLayout (read undefined values, create misaligned pointers, introduce AVE/TypeSafety bugs for refs).

    In this case it can be marked as safe, although, don't you agree that it's a very nieche case to do overlap like this? 🙂

  8. hamarb123 commented on Oct 4, 2026

    @hamarb123
    Contributor

    What sort of cases would regress btw? Doing load, store, load, store vs load, load, store, store (and similar for longer sequences, up to whatever is supported being unrolled) doesn't seem like it'd be the end of the world to me, excepting increased register usage which might hurt a very small number of functions if they are using all the registers or something?

  9. EgorBo commented on Oct 4, 2026

    @EgorBo
    Member

    What sort of cases would regress btw? Doing load, store, load, store vs load, load, store, store (and similar for longer sequences, up to whatever is supported being unrolled) doesn't seem like it'd be the end of the world to me, excepting increased register usage which might hurt a very small number of functions?

    The memmove semantic is currently limitted by 4 regs (max possible is 5 allowed by our LSRA) - it's less than what memcpy may use. And, well, requesting 4(5) GPRs likely will introduce a lot of spills on x64 (around half of all available volatile regs).

    Also, memmove expansion doesn't support addressing mode yet (can be implemented).

    Again, not saying it can't be done, just need to inspect the impact on all other code by this scenario.

  10. benaadams commented on Oct 4, 2026

    @benaadams
    MemberAuthor

    although, don't you agree that it's a very nieche case to do overlap like this?

    Is niche and hope people wouldn't do it

    I'm trying create a formal verification to prove what CIL does; most of the Unsafe is ok and can verify that you aren't allowed to do things like GC holes, oob, uninit memory etc

    Issue here is its simple assignment (var ref) and does different things on different platforms; so doesn't have a verifiable behaviour 😭

  11. benaadams commented on Oct 4, 2026

    @benaadams
    MemberAuthor

    Maybe I can add a iff caller is offset aliasing they are bad and should feel bad rule 🤔

  12. tannergooding commented on Oct 4, 2026

    @tannergooding
    Member

    I'm trying create a formal verification to prove what CIL does

    A general issue you're going to run into is that what the JIT currently does is not strictly what any JIT is allowed to do. ECMA-335 leaves a lot of things as undefined or unverifiable behavior, strictly because they are unsafe and there's a lot of nuance even with regards to what hardware may do in various scenarios.

    Most unsafe code, like explicitly overlapping memory, is just such a case.

  13. benaadams commented on Oct 4, 2026

    @benaadams
    MemberAuthor

    Can you put overlapping fields into a double unsafe level that I can just reject the assembly if you have it enabled? 😅

    Though C/C++ unions like overlapping fields

  14. tannergooding commented on Oct 4, 2026

    @tannergooding
    Member

    Though C/C++ unions like overlapping fields

    And both notably have their own nuances as well.

    C effectively (not literally) defines unions as working like memcpy, not like memmove. But then explicitly leaves this case as UB (this is from C23, which is a refinement of the older versions where they didn't previously give the allowance for exact overlap):

    If the value being stored in an object is read from another object that overlaps in any way the storage
    of the first object, then the two objects shall occupy exactly the same storage and shall have qualified
    or unqualified versions of a compatible type; otherwise, the behavior is undefined.

    So you may observe the same kind of "tear" here:

    struct S1 { uint64_t x, y; };
    struct S2 { uint64_t a; S1 b; };
    union U { S1 m; S2 n; };
    
    U u;
    u.m = { 0, 1 };
    u.n.b = u.m;

    C++ then has similar:

    If the left operand and the right operand identify overlapping objects, the behavior is undefined unless the overlap is exact and the type is the same.

    However, they are much stricter with unions in general and do not define it working like memcpy. They rather require strict object lifecycles and so even union u { float f; int32_t i; } is UB, expecting you to use bit_cast instead. The only special allowance they really have is for the "Common Initial Sequence", so like if you have struct S1 { int32_t x, y, z; } and struct S2 { int32_t x, y, z, w; } then given union U { S s1; S2 s2; } it is legal to access s1.x/y/z and s2.x/y/z regardless of which is "live", while you are restricted from accessing s2.w unless s2 is live.

    -- Noting this is a high level overview, it is not exact and is leaving some nuance on the floor.

  15. 39 remaining items

  16. tannergooding commented on Oct 5, 2026

    @tannergooding
    Member

    Us regressing things to explicitly support unsafe code working seems like a very bad direction to take. We should be encouraging users to not do these things and to move away from such anti-patterns instead.

    It makes us less competitive all to support something that is already bad/dangerous and which devs should not be doing; which means less reasons for people to target C#/.NET

  17. EgorBo commented on Oct 5, 2026

    @EgorBo
    Member

    Just to re-iterate on the fix: it will likely affect all block copies everywhere (my quick fix had massive diffs) since in many cases JIT won't have any hints on whether src overlap with dst or not. I don't see any complains about the current behavior besides these new Ben's issues and Jan's 9yo issue. Just pointing out that the proper fix will likely affect a lot of code 🤷.

    Memmove may request up to 4(5?) regs from LSRA (for normal copies it is always just 1 reg) - it's quite a lot for win-x64/SysV with only 7/9 caller-saved registers available - very likely will lead to a lot of spilling.
    Memmove for arm64 can only handle 64 bytes today (4*16 simd vectors), LSRA's limitation is 5 regs, so can be increased to 80. Memcpy can unroll up to 128b today.

    Not saying we shouldn't do it, just to understand the impact.

  18. tannergooding commented on Oct 5, 2026

    @tannergooding
    Member

    Edit: Egor said what I was going to say at basically the same time


    I'd also note that this isn't limited to just Read/WriteUnaligned, the same general implication exists for all struct reads.

    Given a ref Byte3 s1 and ref Byte3 s2 you can observe the same issues. It is effectively requiring memmove semantics (as if a temporary buffer were used) for all ref T or T* because it could partially overlap another T and still be correctly aligned.

    This is massively impacting to a lot of codegen and optimizations being done. I expect it will be several hundred thousand bytes of disasm churn, if not more.

  19. jkotas commented on Oct 5, 2026

    @jkotas
    Member
    • I agree we do not want to be pessimizing ordinary code to handle overlaps that can only happen with unsafe code. static void Copy(in Words input, out Words output) => output = input; not handling overlaps correctly is fine.

    • We need to have a way for writing unsafe code that is guaranteed to be correct in presence of overlaps. I think making Unsafe.ReadUnaligned/WriteUnaligned compatible with overlaps is the most straightforward way to do that. We need numbers for impact on real world code. @EgorBo Is that something you can prototype and collect? I expect that we will find that it is not a problem.

  20. tannergooding commented on Oct 5, 2026

    @tannergooding
    Member

    Would some different ReadOverlapped/WriteOverlapped be sufficient?

    One of the issues, IMO, with reusing ReadUnaligned/WriteUnaligned is that this bug has very little to do with unalignment. The Byte3 struct case, for example, is always aligned; and that it is already heavily utilized throughout a lot of code.

    I believe the distinction is an important one and that users should have the flexibility of both, not be pessmized into unaligned. ins functioning differently than a regular struct copy would, aside from working if the platform doesn't have built-in unaligned memory support

  21. jkotas commented on Oct 5, 2026

    @jkotas
    Member

    Would some different ReadOverlapped/WriteOverlapped be sufficient?

    It is non-intuitive for an API that takes pointer to the buffer as void* or as ref byte to impose assumptions about the bytes in the buffer to be something else. I think we would want to fix the existing APIs even if invent something more specialized based on data.

    users should have the flexibility of both

    I am not convinced. I can be convinced by real world examples that show the flexibility is actually needed.

  22. tannergooding commented on Oct 5, 2026

    @tannergooding
    Member

    I am not convinced. I can be convinced by real world examples that show the flexibility is actually needed.

    I'm not convinced we need both either, just that I expect forcing this restriction on Unsafe.Read/WriteUnaligned is going to have perf impact and unintended consequences we don't want. One where an explicit compiler barrier API or alternative APIs for overlapped data would solve the issue without requiring changing the existing APIs

    It is non-intuitive for an API that takes pointer to the buffer as void* or as ref byte to impose assumptions about the bytes in the buffer to be something else.

    Not quite sure what you're saying here? The consideration was not about imposing assumptions, it was about having an API where the user can state "this may be overlapped data, so handle it in a way that avoids issues". It is the same categorization as the current ones which are "this may be unaligned data, so handle in a way that avoids issues". Some compiler barrier API lets you solve the same issues as well (for both aligned and unaligned data).

    I think we would want to fix the existing APIs even if invent something more specialized based on data.

    I'm still not convinced these are broken or problematic. This is how it's worked for 10+ years (likely much farther given the unaligned. ins prefix, even if it was rarely used). We've had 1 real report (this issue) and the issue you found when reviewing the Buffer.MemoryCopy code

    My expectation is that this type of overlapping is essentially non-existent in practice and it is why we don't get reports. I expect rather that we will get more reports by trying to go and change the behavior, especially perf wise.

    But I guess we can wait and see SPMI diffs from such an experiment.

  23. jkotas commented on Oct 5, 2026

    @jkotas
    Member

    Not quite sure what you're saying here? The consideration was not about imposing assumptions

    The current implementation of Unsafe.Read/WriteUnaligned is imposing assumption about the buffers being non-overlapping, and that such behavior is non-intuitive given the API signature.

    Some compiler barrier API lets you solve the same issues as well

    I think this solution is very bug prone. You have to make sure to add the barrier around any place where you are receiving a buffer from (safe) user code to prevent the read/write from being combined with another read/write in a bad way (like in the BitConverter example above). I am perfectly fine with leaving some perf on the table to get less bug prone APIs.

  24. tannergooding commented on Oct 5, 2026

    @tannergooding
    Member

    The current implementation of Unsafe.Read/WriteUnaligned is imposing assumption about the buffers being non-overlapping

    You could inversely state that changing it is then imposing an assumption that they are overlapping. I would personally think that changing it is less intuitive because it isn't the default and it isn't the assumption given to regular reads/writes. I'd further says its unintuitive given how these APIs are documented today and it being fairly well understood its just giving access to the unaligned. prefix.

    I think this solution is a very bug prone

    No more bug prone than any other unsafe code, including multi-threading where you have to remember to insert barriers, fences, volatile, etc.

    It is a feature other languages, including Rust, provide. GCC/Clang provide it as asm volatile("" ::: "memory");, MSVC via _ReadWriteBarrier() or via std::atomic_signal_fence(std::memory_order_seq_cst). Swift via atomicMemoryFence(ordering:), etc

    If you are dealing with overlapping memory, which you should know either via an explicit Overlaps check or similar (as you need to guard and handle that for many other reasons anyways). Then you need to insert some explicit handling depending on if it is element aliasing or partial element overlapping

    I am perfectly fine with leaving some perf on the table to get less bug prone APIs.

    This is where I don't think its bug prone. I think this is not even really an issue for code today and so we're not helping anyone and are hurting them instead.


    It's entirely possible we're not going to see eye to eye here either.

    I am very much on the side that I think this is likely to show up as problematic for perf and usability if we change the behavior, to benefit something that we've had a single real report around and where that report was with very unsafe code that doesn't even match the types of layout you can construct in most other languages.

    Happy to be proven wrong, but it seems like its going in the opposite direction of what we want given all the factors that exist here.

  25. jkotas commented on Oct 5, 2026

    @jkotas
    Member

    It is a feature other languages, including Rust, provide. GCC/Clang provide it

    This does not mean that they got it right and we should copy their designs. The strict aliasing optimizations in C/C++ are a mess. This repo, Linux kernel and number of other projects disable it to avoid undefined behavior introduced by optimizations.

    we change the behavior

    I see it as making the behavior deterministic. Deterministic behavior = good. I expect that we will find that the perf impact is non-existent.

  26. EgorBo commented on Oct 5, 2026

    @EgorBo
    Member
    • I agree we do not want to be pessimizing ordinary code to handle overlaps that can only happen with unsafe code. static void Copy(in Words input, out Words output) => output = input; not handling overlaps correctly is fine.
    • We need to have a way for writing unsafe code that is guaranteed to be correct in presence of overlaps. I think making Unsafe.ReadUnaligned/WriteUnaligned compatible with overlaps is the most straightforward way to do that. We need numbers for impact on real world code. @EgorBo Is that something you can prototype and collect? I expect that we will find that it is not a problem.

    @jkotas suprisingly the impact on diffs seem to be quite low: +213 bytes accross all SPMI diffs if we only these:

    case NI_SRCS_UNSAFE_Read:
    case NI_SRCS_UNSAFE_ReadUnaligned:
    case NI_SRCS_UNSAFE_Write:
    case NI_SRCS_UNSAFE_WriteUnaligned:
    

    I had to apply a few tricks for alias-analysis:

    • TYP_REF never overlaps with locals/implicit-byref/return-buffers
    • locals don't overlap with each other (or any addresses JIT sees the actual offset between them)
    • Partially overlapping structs with GC pointers are UB -- not sure about this one yet.

    Change: https://github.com/dotnet/runtime/compare/main...EgorBo:runtime-1:jit-unsafe-read-write-overlap?expand=1

    For https://godbolt.org/z/z686W1Mr6 issue the codegen is now:

           movzx    rax, word  ptr [rcx]
           movzx    r8, word  ptr [rcx+0x01]
           mov      word  ptr [rdx], ax
           mov      word  ptr [rdx+0x01], r8w

    PS: but if we do it for all block copies, the impact is still huge.

  27. hamarb123 commented on Oct 5, 2026

    @hamarb123
    Contributor

    @EgorBo @jkotas my 2 cents are:

    • NI_SRCS_UNSAFE_ReadUnaligned and NI_SRCS_UNSAFE_WriteUnaligned are meant to be identical to unaligned. 1 ldind/stind, and we shouldn't just be changing the managed API (we should keep them in sync and also change the applicable set of IL instructions in presence of unaligned prefix imo)
    • I don't think we should do for NI_SRCS_UNSAFE_Read or NI_SRCS_UNSAFE_Write for similar reasons
    • We shouldn't have partially overlapping structs with GC pointers as UB in general - it seems reasonable to me that InlineArray2<object> could overlap with another one for example - there is an easy workaround if people get bad codegen due to this (just check RuntimeHelpers.IsReferenceOrContainsReferences and don't use unaligned read/write in that case, as they can never actually be misaligned) - however, if you have struct { object; nint; } then it can't overlap with itself at any offset other than 0 seems reasonable in this case
    • Also, @EgorBo it should work if just 1 is unaligned read/write I think (idk if you implemented that, haven't checked) - consider if I have Span<int> and I access/cast as {ReadOnly}Span<byte> and use BinaryPrimitives.Read/WriteInt32 on it for one access, and access normally for the other access - this seems like it should work fine too just like if both are via BinaryPrimitives (and similarly for Byte3 for example)
  28. benaadams commented on Oct 5, 2026

    @benaadams
    MemberAuthor

    but if we do it for all block copies, the impact is still huge.

    Having an escape hatch is probably enough rather than applying gobally?

  29. benaadams commented on Oct 5, 2026

    @benaadams
    MemberAuthor

    Diffs look quite good actually; though maybe they will be worse in methods that use more registers

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMIuntriagedNew issue has not been triaged by the area owner

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions