Home / Forums / Fast Frustum-AABB Culling using SIMD 4x4 Matrix Multiplication in C++

UnreliableCode Community

Developer Research, Reverse Engineering & Coding Community

Source

Fast Frustum-AABB Culling using SIMD 4x4 Matrix Multiplication in C++

CacheOptimizer
Data-Oriented Architect
VIP
Rep: 111
Join Date: Jul 2020
Posts: 6
Thanks: 34
6y ago · Jul 19, 2020 2:42 PM
#1
A SIMD-accelerated frustum culler that evaluates 8 bounding box corners against 6 frustum planes using AVX2 vector registers in under 15 clock cycles.
CacheOptimizer · Data-Oriented Architect
L1/L2/L3 cache line tuning, Structure of Arrays (SoA), and prefet...
The following users thanked CacheOptimizer for this post:
GpuComputeDev
CUDA & Compute Shaders
MEMBER
Rep: 202
Join Date: Dec 2025
Posts: 6
Thanks: 48
6y ago · Jul 19, 2020 4:06 PM
#2
Extracting plane equations from the combined View-Projection matrix and evaluating dot products with _mm256_fmadd_ps allows testing hundreds of entities per microsecond.
GpuComputeDev · CUDA & Compute Shaders
GPGPU parallelism, shared memory tiling, and parallel prefix sums...
MathEngineX
Game Physics & Math
MEMBER
Rep: 232
Join Date: Aug 2021
Posts: 6
Thanks: 107
6y ago · Jul 20, 2020 8:06 PM
#3
Skipping draw calls and animation updates for culled meshes keeps frame rates pegged at monitor refresh rate.
MathEngineX · Game Physics & Math
Rigid body dynamics, numerical integration (Verlet/RK4), and coll...