user0/rust - Forgejo: Beyond coding. We Forge.

user0/rust

Author	SHA1	Message	Date
Peter Jaszkowiak	cc8b95cc54	add `overflow_checks` intrinsic	2025-11-08 10:57:35 -07:00
Augie Fackler	e3e342a90b	rustc_codegen_llvm: adapt for LLVM 22 change to pass masked intrinsic alignment as an attribute This was a bit more invasive than I had kind of hoped. An alternate approach would be to add an extra call_intrinsic_with_attrs() that would have the new-in-this-change signature for call_intrinsic, but this felt about equivalent and made it a little easier to audit the relevant callsites of call_intrinsic().	2025-10-23 17:23:01 -04:00
Camille Gillot	5dfbf67f94	Replace NullOp::SizeOf and NullOp::AlignOf by lang items.	2025-10-23 00:38:28 +00:00
bors	cf8346dd4c	Auto merge of #147476 - ehuss:cold-attribute-test, r=chenyukang Add a test for the cold attribute This adds a test for the cold attribute to verify that it actually does something, and that it applies correctly in all the positions it is expected to work.	2025-10-21 07:41:32 +00:00
bors	fd847d4d5d	Auto merge of #142696 - ZuseZ4:offload-device1, r=oli-obk Offload host2 r? `@oli-obk` A follow-up to my previous gpu host PR. With this, I can (in theory) run a sufficiently simple Rust function on GPUs. I tested it on AMD, where the amdgcn tartget of rustc causes issues due to Addressspace castings, which might not be valid. If I (manually) fix them, I can run the generated IR on an AMD GPU. This should conceptually also work on NVIDIA or Intel. I updated the dev-guide acordingly: https://rustc-dev-guide.rust-lang.org/offload/usage.html I am unhappy with the amount of standalone functions in my offload code, so in my second commit I bundled some of the code around two structs which are Rust versions of the LLVM/Offload structs which they represent. The structs themselves only have doc comments. Since I directly lower everything to llvm-ir I didn't saw a big value in modelling the struct member variables.	2025-10-20 10:17:29 +00:00
Manuel Drehwald	b56d555a36	fix host code	2025-10-19 09:28:39 -07:00
Camille Gillot	97f88f5603	Generalize the non-freeze and needs_drop handling.	2025-10-17 16:28:37 +00:00
Camille Gillot	031929c369	deduced_param_attrs: check Freeze on monomorphic types.	2025-10-15 21:21:05 +00:00
Matthias Krüger	041ecb124a	Rollup merge of #146949 - pmur:murp/improve-ppc-inline-asm, r=Amanieu Add vsx register support for ppc inline asm, and implement preserves_flag option This should address the last(?) missing pieces of inline asm for ppc: * Explicit VSX register support. ISA 2.06 (POWER7) added a 64x128b register overlay extending the fpr's to 128b, and unifies them with the vmx (altivec) registers. Implementations details within gcc/llvm percolate up, and require using the `x` template modifier. I have updated the inline asm to implicitly include this for vsx arguments which do not specify it. ~~Support for the gcc codegen backend is still a todo.~~ * Implement the `preserves_flags` option. All ABI's, and all ISAs store their flags in `cr`, and the carry bit lives inside `xer`. The other status registers hold sticky bits or control bits which do not affect branch instructions. There is some interest in the e500 (powerpcspe) port. Architecturally, it has a very different FP ISA, and includes a simd extension called SPR (which is not IBM's cell SPE). Notably, it does not have altivec/fpr/vsx registers. It also has an SPE accumulator register which its ABI marks as volatile, but I am not sure if the compiler uses it.	2025-10-15 07:09:54 +02:00
Matthias Krüger	10e535a163	Rollup merge of #147638 - alessandrod:indirect-res, r=wesleywiser bpf: return results larger than one register indirectly Fixes triggering the "only small returns supported" error in the BPF target.	2025-10-14 19:47:30 +02:00
Paul Murphy	4945d21ed9	Implement ppc/ppc64 preserves_flags option for inline asm Implemented preserves_flags on powerpc by making it do nothing. This prevents having two different ways to mark `cr0` as clobbered. clang and gcc alias `cr0` to `cc`. The gcc inline documentation does not state what this does on powerpc* targets, but inspection of the source shows it is equivalent to condition register field `cr0`, so it should not be added.	2025-10-14 10:05:07 -05:00
Paul Murphy	3c09d4a582	Allow vector-scalar (vs) registers in ppc inline assembly Where supported, VSX is a 64x128b register set which encompasses both the floating point and vector registers. In the type tests, xvsqrtdp is used as it is the only two-argument vsx opcode supported by all targets on llvm. If you need to copy a vsx register, the preferred way is "xxlor xt, xa, xa".	2025-10-14 09:52:56 -05:00
Alessandro Decina	056c2da339	bpf: return results larger than one register indirectly Fixes triggering the "only small returns supported" error in the BPF target.	2025-10-13 16:52:12 +00:00
Ben Kimock	029579d177	Change int-to-ptr transmute lowering back to inttoptr	2025-10-10 20:14:23 -04:00
Eric Huss	50e1884aa1	Add a test for the cold attribute This adds a test for the cold attribute to verify that it actually does something, and that it applies correctly in all the positions it is expected to work.	2025-10-09 20:24:45 -07:00
Stuart Cook	4e3e7ce078	Rollup merge of #147457 - the8472:slice_fill_memset2, r=RalfJung,joboet specialize slice::fill to use memset when possible It helps const eval performance https://github.com/rust-lang/miri/issues/4616, debug builds and the gcc backend. Previously attempted in https://github.com/rust-lang/rust/pull/83245 but reverted due to unsoundness https://github.com/rust-lang/rust/issues/87891 around potentially-uninitialized types. This PR only handles primitives where the problem does not arise. split off from https://github.com/rust-lang/rust/pull/147294	2025-10-09 18:43:22 +11:00
The 8472	99ab27f90c	specialize slice::fill to use memset when possible LLVM generally can do this on its own, but it helps miri and other backends.	2025-10-08 20:14:24 +02:00
Matthias Krüger	ffba05ee29	Rollup merge of #146865 - folkertdev:kcfi-only-reify-dyn-compatible, r=rcvalle kcfi: only reify trait methods when dyn-compatible fixes https://github.com/rust-lang/rust/issues/146853 Only generate a `ReifyShim` for trait method calls if the trait is dyn-compatible. Until now kcfi would generate a `ReifyShim` whenever a trait method was cast to a function pointer. But technically the shim is only needed for dyn-compatible traits (where the method might end up in a vtable). Up to this point that was only slightly inefficient, but in combination with c-variadic trait methods it is wrong. For c-variadic trait methods the generated shim is incorrect, and that is why c-variadic methods make a trait no longer dyn-compatible: we should simply never generate a `ReifyShim` that is c-variadic. With this change the documentation on `ReifyReason` is now actually correct: > If KCFI is enabled, creating a function pointer from a method on a dyn-compatible trait. This includes the case of converting `::call`-like methods on closure-likes to function pointers. cc ```@maurer``` ```@workingjubilee``` r? ```@rcvalle```	2025-10-07 19:39:06 +02:00
Manuel Drehwald	7c8fe29fd6	solve autodiffv2.rs FIXME and make identical_fnc test more robust	2025-10-05 03:07:51 -04:00
Matthias Krüger	a84359c5a7	Rollup merge of #147315 - ZuseZ4:fix-ad-batching-test, r=jieyouxu bless autodiff batching test This pr blesses a broken test and unblocks running rust in the Enzyme CI: https://github.com/EnzymeAD/Enzyme/pull/2430 Enzyme is the plugin used by our std::autodiff and (future) std::batching modules, both of which are not build by default. In the near future we also hope to enable std::autodiff in the Rust CI. This test is the only one to combine two features, automatic differentiation and batching/vectorization. This combination is even more experimental than either feature on its own. I have a wip branch in which I enable more vectorization/batching and as part of that I'll think more about how to write those tests in a robust way (and likely change the interface). Until that lands, I don't care too much about what specific IR we generate here; it's just nice to track changes. r? compiler	2025-10-04 17:11:13 +02:00
Manuel Drehwald	12cfad9a8b	update autodiff batching test	2025-10-03 18:12:22 -04:00
dianqk	c2a03cefd8	debuginfo: Use `LocalRef` to simplify reference debuginfos If the `LocalRef` is `LocalRef::Place`, we can refer to it directly, because the local of place is an indirect pointer. Such a statement is `_1 = &(_2.1)`. If the `LocalRef` is `LocalRef::Operand`, the `OperandRef` should provide the pointer of the reference. Such a statement is `_1 = &((*_2).1)`. But there is a special case that hasn't been handled, scalar pairs like `(&[i32; 16], i32)`.	2025-10-03 08:08:22 +08:00
dianqk	8da04285cf	mir-opt: Eliminate dead statements even if they are used by debuginfos	2025-10-02 14:58:59 +08:00
dianqk	1bd89bd42e	codegen: Generate `dbg_value` for the ref statement	2025-10-02 14:55:51 +08:00
dianqk	571412f819	mir-opt: Eliminate dead ref statements	2025-10-02 14:55:50 +08:00
Stuart Cook	7b0236fbd8	Rollup merge of #147200 - ZuseZ4:fix-autodiff-emptry-ret, r=Zalathar Fix autodiff empty ret regression closes https://github.com/rust-lang/rust/issues/147144 The two gsoc summer projects caused a bit of churn, which was to be expected, especially since we don't run autodiff in CI yet. This adds a void return testcase that we should have had anyway, and fixes the regression. r? `@Zalathar` (Just guessing since I've seen you in a few LLVM PRs and Oli is probably still busy. Feel free to reroll!)	2025-10-01 22:15:01 +10:00
Manuel Drehwald	de189fa982	updating tests to not break from new typetree metadata	2025-09-30 22:47:43 -04:00
Manuel Drehwald	28ffbab353	add empty struct ret testcase	2025-09-30 22:47:43 -04:00
Jacob Pratt	b310eb91ab	Rollup merge of #146457 - alexcrichton:wasm-no-exn-instructions, r=bjorn3 Skip cleanups on unsupported targets This commit is an update to the `AbortUnwindingCalls` MIR pass in the compiler. Specifically a new boolean is added for "can this target possibly unwind" and if that's `false` then terminators are all adjusted to be unreachable/not present. The end result is that this fixes rust-lang/rust#140293 for wasm targets. The motivation for this PR is that currently on WebAssembly targets the usage of the `C-unwind` ABI can lead LLVM to either (a) emit exception-handling instructions or (b) hit a LLVM-ICE-style codegen error. WebAssembly as a base instruction set does not support unwinding at all, and a later proposal to WebAssembly, the exception-handling proposal, was what enabled this. This means that the current intent of WebAssembly targets is that they maintain the baseline of "don't emit exception-handling instructions unless enabled". The commit here is intended to restore this behavior by skipping these instructions even when `C-unwind` is present. Exception-handling is a relatively tricky and also murky topic in WebAssembly, however. There are two sets of instructions LLVM can emit for WebAssembly exceptions, Rust's Emscripten target supports exceptions, WASI targets do not, the LLVM flags to enable this are not always obvious, and additionally this all touches on "changing exception-handling behavior should be a target-level concern, not a feature". Effectively WebAssembly's exception-handling integration into Rust is not finalized at this time. The best idea at this time is that a parallel set of targets will eventually be added which support exceptions, but it's not clear if/when to do this. In the meantime the goal is to keep existing targets working while still enabling experimentation with exception-handling with `-Zbuild-std` and various permutations of LLVM flags. To that extent this commit does not blanket disable these landing pads and cleanup routines for WebAssembly but instead checks to see if panic=unwind is enabled or if `+exception-handling` is enabled. Tests are updated here as well to account for this where, by default, using a `C-unwind` ABI won't affect Rust codegen at all. If `+exception-handling` is enabled, however, then Rust codegen will look like native platforms where exceptions are caught and the program aborts. More-or-less I've done my best to keep exceptions working on wasm where it's possible to have them work, but turned them off where they're not supposed to be emitted. Closes rust-lang/rust#140293	2025-09-29 21:37:50 -04:00
Matthias Krüger	c29fb2e57e	Rollup merge of #144197 - KMJ-007:type-tree, r=ZuseZ4 TypeTree support in autodiff # TypeTrees for Autodiff ## What are TypeTrees? Memory layout descriptors for Enzyme. Tell Enzyme exactly how types are structured in memory so it can compute derivatives efficiently. ## Structure ```rust TypeTree(Vec<Type>) Type { offset: isize, // byte offset (-1 = everywhere) size: usize, // size in bytes kind: Kind, // Float, Integer, Pointer, etc. child: TypeTree // nested structure } ``` ## Example: `fn compute(x: &f32, data: &[f32]) -> f32` Input 0: `x: &f32` ```rust TypeTree(vec![Type { offset: -1, size: 8, kind: Pointer, child: TypeTree(vec![Type { offset: -1, size: 4, kind: Float, child: TypeTree::new() }]) }]) ``` Input 1: `data: &[f32]` ```rust TypeTree(vec![Type { offset: -1, size: 8, kind: Pointer, child: TypeTree(vec![Type { offset: -1, size: 4, kind: Float, // -1 = all elements child: TypeTree::new() }]) }]) ``` Output: `f32` ```rust TypeTree(vec![Type { offset: -1, size: 4, kind: Float, child: TypeTree::new() }]) ``` ## Why Needed? - Enzyme can't deduce complex type layouts from LLVM IR - Prevents slow memory pattern analysis - Enables correct derivative computation for nested structures - Tells Enzyme which bytes are differentiable vs metadata ## What Enzyme Does With This Information: Without TypeTrees (current state): ```llvm ; Enzyme sees generic LLVM IR: define float ``@distance(ptr`` %p1, ptr %p2) { ; Has to guess what these pointers point to ; Slow analysis of all memory operations ; May miss optimization opportunities } ``` With TypeTrees (our implementation): ```llvm define "enzyme_type"="{[]:Float@float}" float ``@distance(`` ptr "enzyme_type"="{[]:Pointer}" %p1, ptr "enzyme_type"="{[]:Pointer}" %p2 ) { ; Enzyme knows exact type layout ; Can generate efficient derivative code directly } ``` # TypeTrees - Offset and -1 Explained ## Type Structure ```rust Type { offset: isize, // WHERE this type starts size: usize, // HOW BIG this type is kind: Kind, // WHAT KIND of data (Float, Int, Pointer) child: TypeTree // WHAT'S INSIDE (for pointers/containers) } ``` ## Offset Values ### Regular Offset (0, 4, 8, etc.) Specific byte position within a structure ```rust struct Point { x: f32, // offset 0, size 4 y: f32, // offset 4, size 4 id: i32, // offset 8, size 4 } ``` TypeTree for `&Point` (internal representation): ```rust TypeTree(vec![ Type { offset: 0, size: 4, kind: Float }, // x at byte 0 Type { offset: 4, size: 4, kind: Float }, // y at byte 4 Type { offset: 8, size: 4, kind: Integer } // id at byte 8 ]) ``` Generates LLVM: ```llvm "enzyme_type"="{[]:Float@float}" ``` ### Offset -1 (Special: "Everywhere") Means "this pattern repeats for ALL elements" #### Example 1: Array `[f32; 100]` ```rust TypeTree(vec![Type { offset: -1, // ALL positions size: 4, // each f32 is 4 bytes kind: Float, // every element is float }]) ``` Instead of listing 100 separate Types with offsets `0,4,8,12...396` #### Example 2: Slice `&[i32]` ```rust // Pointer to slice data TypeTree(vec![Type { offset: -1, size: 8, kind: Pointer, child: TypeTree(vec![Type { offset: -1, // ALL slice elements size: 4, // each i32 is 4 bytes kind: Integer }]) }]) ``` #### Example 3: Mixed Structure ```rust struct Container { header: i64, // offset 0 data: [f32; 1000], // offset 8, but elements use -1 } ``` ```rust TypeTree(vec![ Type { offset: 0, size: 8, kind: Integer }, // header Type { offset: 8, size: 4000, kind: Pointer, child: TypeTree(vec![Type { offset: -1, size: 4, kind: Float // ALL array elements }]) } ]) ```	2025-09-28 18:13:11 +02:00
Matthias Krüger	c772af78e9	Rollup merge of #146732 - durin42:llvm-22-less-assumes, r=nikic tests: relax expectations after llvm change 902ddda120a5 LLVM 22 is able to drop assumes that seem to not help further optimizations, which actually seems to dramatically _help_ further optimizations in some of our small test cases. I'm a little unclear how to fix the last failure, in `tests/codegen-llvm/issues/issue-122600-ptr-discriminant-update.rs`: ``` -; Function Attrs: mustprogress nofree norecurse nosync nounwind willreturn memory(argmem: readwrite, inaccessiblemem: write) uwtable +; Function Attrs: mustprogress nofree norecurse nosync nounwind nonlazybind willreturn memory(argmem: readwrite, inaccessiblemem: write) uwtable define void ``@update(ptr`` noundef captures(none) %s) unnamed_addr #0 { start: - %_3.sroa.0.0.copyload = load i8, ptr %s, align 1 - %0 = trunc nuw i8 %_3.sroa.0.0.copyload to i1 - %1 = xor i1 %0, true - tail call void ``@llvm.assume(i1`` %1) store i8 1, ptr %s, align 1 ret void } ``` I'm just not conversant enough in LLVM IR to follow the changes here. ``@rustbot`` label llvm-main r? nikic	2025-09-27 21:25:57 +02:00
Augie Fackler	99456cc015	tests: use max-llvm-major-version instead of ignore-llvm-version	2025-09-26 13:32:03 -04:00
bors	7ac0330c6d	Auto merge of #147037 - matthiaskrgr:rollup-xtgqzuu, r=matthiaskrgr Rollup of 8 pull requests Successful merges: - rust-lang/rust#116882 (rustdoc: hide `#[repr]` if it isn't part of the public ABI) - rust-lang/rust#135771 ([rustdoc] Add support for associated items in "jump to def" feature) - rust-lang/rust#141032 (avoid violating `slice::from_raw_parts` safety contract in `Vec::extract_if`) - rust-lang/rust#142401 (Add proper name mangling for pattern types) - rust-lang/rust#146293 (feat: non-panicking `Vec::try_remove`) - rust-lang/rust#146859 (BTreeMap: Don't leak allocators when initializing nodes) - rust-lang/rust#146924 (Add doc for `NonZero*` const creation) - rust-lang/rust#146933 (Make `render_example_with_highlighting` return an `impl fmt::Display`) r? `@ghost` `@rustbot` modify labels: rollup	2025-09-25 20:35:49 +00:00
Matthias Krüger	958d1438b6	Rollup merge of #142401 - oli-obk:pattern-mango, r=petrochenkov Add proper name mangling for pattern types requires adding demangler support first https://github.com/rust-lang/rustc-demangle/pull/81 needed for https://github.com/rust-lang/rust/pull/136006#discussion_r2139792593 as otherwise we will have symbol collisions	2025-09-25 18:15:08 +02:00
Stuart Cook	46e25aa7a3	Rollup merge of #146766 - nikic:global-alloc-attr, r=nnethercote Add attributes for #[global_allocator] functions Emit `#[rustc_allocator]` etc. attributes on the functions generated by the `#[global_allocator]` macro, which will emit LLVM attributes like `"alloc-family"`. If the module with the global allocator participates in LTO, this ensures that the attributes typically emitted on the allocator declarations are not lost if the definition is imported. There is a similar issue when the allocator shim is used, but I've opted not to fix that case in this PR, because doing that cleanly is somewhat gnarly. Related to https://github.com/rust-lang/rust/issues/145995.	2025-09-25 20:31:56 +10:00
bors	15283f6fe9	Auto merge of #146338 - CrooseGit:dev/reucru01/AArch64-enable-GCS, r=Urgau,davidtwco Extends AArch64 branch protection support to include GCS Extends existing support for AArch64 branch protection to include support for [Guarded Control Stacks](https://community.arm.com/arm-community-blogs/b/architectures-and-processors-blog/posts/arm-a-profile-architecture-2022#guarded-control-stack-gcs:~:text=Extraction%20or%20tracking.-,Guarded%20Control%20Stack%20(GCS),-With%20the%202022).	2025-09-24 13:04:19 +00:00
Oli Scherer	739e89980f	Add proper name mangling for pattern types	2025-09-23 10:59:29 +00:00
Nikita Popov	bc7986ec79	Add attributes for #[global_allocator] functions Emit `#[rustc_allocator]` etc. attributes on the functions generated by the `#[global_allocator]` macro, which will emit LLVM attributes like `"alloc-family"`. If the module with the global allocator participates in LTO, this ensures that the attributes typically emitted on the allocator declarations are not lost if the definition is imported.	2025-09-23 10:21:17 +02:00
bors	ce4beebecb	Auto merge of #146683 - clarfonthey:safe-intrinsics, r=RalfJung,Amanieu Mark float intrinsics with no preconditions as safe Note: for ease of reviewing, the list of safe intrinsics is sorted in the first commit, and then safe intrinsics are added in the second commit. All recently added float intrinsics have been correctly marked as safe to call due to the fact that they have no preconditions. This adds the remaining float intrinsics which are safe to call to the safe intrinsic list, and removes the unsafe blocks around their calls. --- Side note: this may want a try run before being added to the queue, since I'm not sure if there's any tier-2 code that uses these intrinsics that might not be tested on the usual PR flow. We've already uncovered a few places in subtrees that do this, and it's worth double-checking before clogging up the queue.	2025-09-22 14:35:46 +00:00
Folkert de Vries	f51fb9178e	kcfi: only reify trait methods when dyn-compatible	2025-09-22 12:30:47 +02:00
Reuben Cruise	06819d95c0	Extends branch protection tests to include GCS	2025-09-22 11:29:54 +01:00
Stuart Cook	46be365a60	Rollup merge of #146831 - taiki-e:powerpc-clobber, r=Amanieu Support ctr and lr as clobber-only registers in PowerPC inline assembly Follow-up to rust-lang/rust#131341. CTR and LR are marked as volatile in all ABIs, but I skipped them in rust-lang/rust#131341 due to they are currently marked as reserved. `dd7fda5700/compiler/rustc_target/src/asm/powerpc.rs (L209-L212)` However, they are actually only unusable as input/output of inline assembly, and should be fine to support as clobber-only registers as discussed in [#t-compiler > ppc/ppc64 inline asm support](https://rust-lang.zulipchat.com/#narrow/channel/131828-t-compiler/topic/ppc.2Fppc64.20inline.20asm.20support/with/540413845). r? ````@Amanieu```` or ````@workingjubilee```` cc ````@programmerjake```` ````@rustbot```` label +O-PowerPC +A-inline-assembly	2025-09-22 20:25:14 +10:00
Stuart Cook	40db498a0f	Rollup merge of #146791 - folkertdev:readonly-not-pure, r=nikic,joshtriplett emit attribute for readonly non-pure inline assembly fixes https://github.com/rust-lang/rust/issues/146761 Provide a better `MemoryEffects` to LLVM when an inline assembly block specifies `readonly` but not `pure`. That means that the assembly block may not perform any writes, but that there still may be side effects from its instructions. I haven't been able to find a case yet where this actually matters, though. So the test checks that the right attribute is applied, but the generated assembly is equivalent to not specifying `readonly` at all. r? ````@nikic```` cc ````@Amanieu````	2025-09-22 20:25:13 +10:00
ltdk	055e05a338	Mark float intrinsics with no preconditions as safe	2025-09-21 20:37:51 -04:00
Folkert de Vries	3565b0699d	emit attribute for readonly non-pure inline assembly	2025-09-21 21:16:06 +02:00
The 8472	5f0a68eb22	regression test for https://github.com/rust-lang/rust/issues/117763	2025-09-21 19:54:43 +02:00
Taiki Endo	f4b876867d	Support ctr and lr as clobber-only registers in PowerPC inline assembly	2025-09-21 13:48:22 +09:00
Augie Fackler	e32b975e84	tests: relax expectations after llvm change 902ddda120a5 LLVM 22 is able to drop assumes that seem to not help further optimizations, which actually seems to dramatically _help_ further optimizations in some of our small test cases.	2025-09-19 13:56:19 -04:00
Karan Janthe	664e83b3e7	added typetree support for memcpy	2025-09-19 04:02:20 +00:00
Karan Janthe	e1258e79d6	autodiff: Add basic TypeTree with NoTT flag Signed-off-by: Karan Janthe <karanjanthe@gmail.com>	2025-09-19 04:02:19 +00:00

1 2 3 4

151 commits