diff options
author | Tulio Magno Quites Machado Filho <tuliom@linux.ibm.com> | 2021-04-30 18:12:08 -0300 |
---|---|---|
committer | Tulio Magno Quites Machado Filho <tuliom@linux.ibm.com> | 2021-04-30 18:12:08 -0300 |
commit | e941e0ae80626b7661c1db8953a673cafd3b8b19 (patch) | |
tree | 42b3dcccfce69af0f7ffb0fa4ed2ed75734b82a2 /sysdeps/powerpc/powerpc64/multiarch/memcpy.c | |
parent | dd59655e9371af86043b97e38953f43bd9496699 (diff) | |
download | glibc-e941e0ae80626b7661c1db8953a673cafd3b8b19.tar glibc-e941e0ae80626b7661c1db8953a673cafd3b8b19.tar.gz glibc-e941e0ae80626b7661c1db8953a673cafd3b8b19.tar.bz2 glibc-e941e0ae80626b7661c1db8953a673cafd3b8b19.zip |
powerpc64le: Optimize memcpy for POWER10
This implementation is based on __memcpy_power8_cached and integrates
suggestions from Anton Blanchard.
It benefits from loads and stores with length for short lengths and for
tail code, simplifying the code.
All unaligned memory accesses use instructions that do not generate
alignment interrupts on POWER10, making it safe to use on
caching-inhibited memory.
The main loop has also been modified in order to increase instruction
throughput by reducing the dependency on updates from previous iterations.
On average, this implementation provides around 30% improvement when
compared to __memcpy_power7 and 10% improvement in comparison to
__memcpy_power8_cached.
Diffstat (limited to 'sysdeps/powerpc/powerpc64/multiarch/memcpy.c')
-rw-r--r-- | sysdeps/powerpc/powerpc64/multiarch/memcpy.c | 7 |
1 files changed, 7 insertions, 0 deletions
diff --git a/sysdeps/powerpc/powerpc64/multiarch/memcpy.c b/sysdeps/powerpc/powerpc64/multiarch/memcpy.c index 5733192932..53ab32ef26 100644 --- a/sysdeps/powerpc/powerpc64/multiarch/memcpy.c +++ b/sysdeps/powerpc/powerpc64/multiarch/memcpy.c @@ -36,8 +36,15 @@ extern __typeof (__redirect_memcpy) __memcpy_power6 attribute_hidden; extern __typeof (__redirect_memcpy) __memcpy_a2 attribute_hidden; extern __typeof (__redirect_memcpy) __memcpy_power7 attribute_hidden; extern __typeof (__redirect_memcpy) __memcpy_power8_cached attribute_hidden; +# if defined __LITTLE_ENDIAN__ +extern __typeof (__redirect_memcpy) __memcpy_power10 attribute_hidden; +# endif libc_ifunc (__libc_memcpy, +# if defined __LITTLE_ENDIAN__ + (hwcap2 & PPC_FEATURE2_ARCH_3_1 && hwcap & PPC_FEATURE_HAS_VSX) + ? __memcpy_power10 : +# endif ((hwcap2 & PPC_FEATURE2_ARCH_2_07) && use_cached_memopt) ? __memcpy_power8_cached : (hwcap & PPC_FEATURE_HAS_VSX) |