Fix i16x8.relaxed_dot_i8x16_i7x16_s pairwise add - #2252
brendandahl wants to merge 1 commit into
Conversation
Allow i16x8.relaxed_dot_i8x16_i7x16_s to use either wrapping or saturating pairwise addition via $R_idot instead of hardcoding saturating addition. ARM NEON uses smull/smull2/addp (signed multiply with wrapping i16 add), whereas x86-64 uses vpmaddubsw (signed/unsigned multiply with saturating i16 add). Add a spec test covering both lowerings.
|
The spec should now match the original pseudo code and matches what is actually implemented by v8, jsc, and spidermonkey. There are also some issues with |
|
FWIW I believe that the specification could be implemented as is (with saturating addition) at a similar cost to what the Wasm engines are doing currently using instructions introduced by the second version of the Scalable Vector Extension (SVE2) to the Arm architecture: IMHO it is still worth changing the specification, since with the update proposed here the implementation could be simplified to: if the |
Allow i16x8.relaxed_dot_i8x16_i7x16_s to use either wrapping or saturating pairwise addition via $R_idot instead of hardcoding saturating addition. ARM NEON uses smull/smull2/addp (signed multiply with wrapping i16 add), whereas x86-64 uses vpmaddubsw (signed/unsigned multiply with saturating i16 add).
Add a spec test covering both lowerings.