Special Relativity in Financial Modeling 1.0.0
Lorentz transforms, spacetime classification, and geodesic price paths for quantitative finance
Loading...
Searching...
No Matches
Namespaces | Functions | Variables
gamma_avx512.cpp File Reference

AVX-512F (512-bit, 8-wide) implementation of the gamma batch kernel. More...

#include "simd_batch_detail.hpp"
#include "momentum/momentum.hpp"
#include <immintrin.h>
#include <cmath>
#include <cstddef>

Go to the source code of this file.

Namespaces

namespace  srfm
 
namespace  srfm::simd
 
namespace  srfm::simd::detail
 

Functions

void srfm::simd::detail::compute_gamma_avx512 (const double *__restrict__ betas, std::size_t n, double *__restrict__ out) noexcept
 AVX-512F (512-bit, 8-wide) gamma batch kernel.
 

Variables

static constexpr double srfm::simd::detail::BETA_CLAMP_LIMIT
 

Detailed Description

AVX-512F (512-bit, 8-wide) implementation of the gamma batch kernel.

Module: src/simd/ Owner: AGT-08 — 2026-03-01

Responsibility

Vectorised gamma_i = 1.0 / sqrt(1.0 - betas[i]^2) for machines with AVX-512F. Processes 8 doubles per SIMD cycle using 512-bit ZMM registers.

Algorithm

For each 8-element chunk:

  1. Load 8 betas: _mm512_loadu_pd
  2. Clamp to limit: _mm512_min_pd(b, clamp_v)
  3. Square: _mm512_mul_pd(b, b) → b²
  4. Subtract from 1: _mm512_sub_pd(ones, b2) → 1 - b² ∈ (0,1]
  5. sqrt: _mm512_sqrt_pd(denom)
  6. Reciprocal exact: _mm512_div_pd(ones, sqrt_d) → γ
  7. Store: _mm512_storeu_pd Tail (n % 8 != 0): scalar fallback.

Performance Characteristics (Skylake-X / Ice Lake)

_mm512_sqrt_pd — throughput 1/cycle (reciprocal), latency 18 cycles _mm512_div_pd — throughput 1/cycle (reciprocal), latency 15 cycles Overall throughput goal: ~8 gammas per ~35 cycles ≈ 229M gammas/sec at 3.5 GHz with 2 AVX-512 FMA units.

Clamp-before-sqrt rationale

beta_i is pre-validated to [0, BETA_MAX_SAFE) by compute_beta_avx512. A second clamp here guards against any stale or externally constructed BetaVelocity value whose internal double was set to exactly BETA_MAX_SAFE via future refactoring. The overhead is one _mm512_min_pd per chunk.

Note: This file must be compiled with -mavx512f (GCC/Clang) or /arch:AVX512 (MSVC) so that the intrinsics are recognised.

Definition in file gamma_avx512.cpp.