Home / Forums / SIMD Vectorization with AVX2 & SSE Intrinsics for High Performance C++

UnreliableCode Community

Developer Research, Reverse Engineering & Coding Community

Tutorial

SIMD Vectorization with AVX2 & SSE Intrinsics for High Performance C++

SigScannerPro
Pattern Scanning & IDA
MEMBER
Rep: 210
Join Date: Jul 2023
Posts: 25
Thanks: 52
4y ago · Jul 16, 2022 1:21 PM
#1
Using _mm256_loadu_ps, _mm256_fmadd_ps, and _mm256_storeu_ps to process 8 single-precision floats simultaneously.
SigScannerPro · Pattern scanning & Byte masking
The following users thanked SigScannerPro for this post:
DirectXRay
DirectX 12 / Vulkan Dev
MEMBER
Rep: 240
Join Date: Feb 2023
Posts: 31
Thanks: 60
4y ago · Jul 16, 2022 2:48 PM
#2
AVX2 FMA (Fused Multiply-Add) computes (a * b) + c in a single hardware cycle with higher floating point precision.
DirectXRay · DirectX 12 & Vulkan Command Lists
AsmDisasm
x86_64 Disassembler Dev
MEMBER
Rep: 145
Join Date: Feb 2025
Posts: 23
Thanks: 36
4y ago · Jul 16, 2022 5:48 PM
#3
Ensure target CPU flags are verified with __cpuid before executing AVX2 instructions to prevent illegal instruction crashes on older hardware.
AsmDisasm · Zydis & Opcode Length Disassembly