Special Relativity in Financial Modeling 1.0.0
Lorentz transforms, spacetime classification, and geodesic price paths for quantitative finance
Loading...
Searching...
No Matches
gamma_avx512.cpp
Go to the documentation of this file.
1/**
2 * @file gamma_avx512.cpp
3 * @brief AVX-512F (512-bit, 8-wide) implementation of the gamma batch kernel.
4 *
5 * Module: src/simd/
6 * Owner: AGT-08 — 2026-03-01
7 *
8 * Responsibility
9 * --------------
10 * Vectorised gamma_i = 1.0 / sqrt(1.0 - betas[i]^2) for machines with
11 * AVX-512F. Processes 8 doubles per SIMD cycle using 512-bit ZMM registers.
12 *
13 * Algorithm
14 * ---------
15 * For each 8-element chunk:
16 * 1. Load 8 betas: _mm512_loadu_pd
17 * 2. Clamp to limit: _mm512_min_pd(b, clamp_v)
18 * 3. Square: _mm512_mul_pd(b, b) → b²
19 * 4. Subtract from 1: _mm512_sub_pd(ones, b2) → 1 - b² ∈ (0,1]
20 * 5. sqrt: _mm512_sqrt_pd(denom)
21 * 6. Reciprocal exact: _mm512_div_pd(ones, sqrt_d) → γ
22 * 7. Store: _mm512_storeu_pd
23 * Tail (n % 8 != 0): scalar fallback.
24 *
25 * Performance Characteristics (Skylake-X / Ice Lake)
26 * ---------------------------------------------------
27 * _mm512_sqrt_pd — throughput 1/cycle (reciprocal), latency 18 cycles
28 * _mm512_div_pd — throughput 1/cycle (reciprocal), latency 15 cycles
29 * Overall throughput goal: ~8 gammas per ~35 cycles ≈ 229M gammas/sec
30 * at 3.5 GHz with 2 AVX-512 FMA units.
31 *
32 * Clamp-before-sqrt rationale
33 * ---------------------------
34 * beta_i is pre-validated to [0, BETA_MAX_SAFE) by compute_beta_avx512.
35 * A second clamp here guards against any stale or externally constructed
36 * BetaVelocity value whose internal double was set to exactly BETA_MAX_SAFE
37 * via future refactoring. The overhead is one _mm512_min_pd per chunk.
38 *
39 * Note: This file must be compiled with -mavx512f (GCC/Clang) or
40 * /arch:AVX512 (MSVC) so that the intrinsics are recognised.
41 */
42
43#include "simd_batch_detail.hpp"
44#include "momentum/momentum.hpp"
45
46#include <immintrin.h>
47#include <cmath>
48#include <cstddef>
49
50namespace srfm::simd::detail {
51
52static constexpr double BETA_CLAMP_LIMIT =
54
56 const double* SRFM_RESTRICT betas,
57 std::size_t n,
58 double* SRFM_RESTRICT out) noexcept
59{
60 constexpr std::size_t LANE = 8;
61
62 const __m512d ones = _mm512_set1_pd(1.0);
63 const __m512d clamp_v = _mm512_set1_pd(BETA_CLAMP_LIMIT);
64
65 std::size_t i = 0;
66
67 // ── Vectorised main loop (8-wide) ─────────────────────────────────────────
68 for (; i + LANE <= n; i += LANE) {
69 // 1. Load 8 betas (unaligned).
70 __m512d b = _mm512_loadu_pd(betas + i);
71
72 // 2. Clamp betas to BETA_CLAMP_LIMIT to ensure sqrt argument > 0.
73 b = _mm512_min_pd(b, clamp_v);
74
75 // 3. b² = b * b.
76 __m512d b2 = _mm512_mul_pd(b, b);
77
78 // 4. denom = 1.0 - b² (strictly positive after clamping).
79 __m512d denom = _mm512_sub_pd(ones, b2);
80
81 // 5. sqrt(1 - b²).
82 __m512d sqrt_d = _mm512_sqrt_pd(denom);
83
84 // 6. gamma = 1.0 / sqrt(1 - b²).
85 __m512d gamma = _mm512_div_pd(ones, sqrt_d);
86
87 // 7. Store 8 results.
88 _mm512_storeu_pd(out + i, gamma);
89 }
90
91 // ── Scalar tail: n % 8 remaining elements ─────────────────────────────────
92 for (; i < n; ++i) {
93 const double b = (betas[i] > BETA_CLAMP_LIMIT) ? BETA_CLAMP_LIMIT : betas[i];
94 const double b2 = b * b;
95 out[i] = 1.0 / std::sqrt(1.0 - b2);
96 }
97}
98
99} // namespace srfm::simd::detail
constexpr double BETA_MAX_SAFE
Definition momentum.hpp:62
static constexpr double BETA_CLAMP_LIMIT
Definition beta_avx2.cpp:30
void compute_gamma_avx512(const double *__restrict__ betas, std::size_t n, double *__restrict__ out) noexcept
AVX-512F (512-bit, 8-wide) gamma batch kernel.
Internal raw-double compute signatures for SIMD batch kernels.
#define SRFM_RESTRICT
Momentum-Velocity Signal Processor (AGT-03 / SRFM)