pytc.kmat.calc_kmat_kernels_from_aux_streamed

pytc.kmat.calc_kmat_kernels_from_aux_streamed(xi_phi, xi_grad, weights, L_aux, H_aux, grid_block_size=4096, rank_block_size=128)[source]

Recover exact K1/K3 from streamed auxiliary-grid panels.

This is the production counterpart of calc_kmat_kernels_from_aux(). It accepts NumPy/JAX arrays or HDF5-style datasets, reads only a rank_block_size by grid_block_size slice of xi at a time, and keeps the output kernels on host RAM. The large L_aux/H_aux panels are transferred once per grid panel and reused for every left-rank block. Consequently no R-by-R temporary is materialised on the device.

The result is the same algebraic identity as the in-core routine:

K1[k,l,c] = sum_g xi_grad[k,g,c] w[g] L_aux[l,g,c] K3[k,l]    = sum_g xi_phi[k,g] w[g] H_aux[l,g].

Parameters are deliberately explicit rather than inferred from a device memory probe. The caller’s grid panel is the device-transfer limit and the rank panel bounds the GEMM output. This makes the out-of-core route predictable and testable on the same inputs as the direct kernel.