pytc.solver.isdf_xtc_ccsd._stream_partial_x_pipelined¶
- pytc.solver.isdf_xtc_ccsd._stream_partial_x_pipelined(panel_kernel, t2, left_out, left_inner, panel_source, nocc, *, occupied_pair_batch_size, rank_panel_size, panel_budget_bytes=None)[source]¶
Working-set X panel loop with double-buffered async prefetch.
Same math, padding, and accumulation semantics as
_stream_partial_x(), but the panel width is sized from the measured device working set and the next panel’s host read is issued (viaasync_read()) before the current panel’s kernel, so its H2Ddevice_putoverlaps compute (JAX async dispatch).panel_source(m0, m1, panel_size)returns one contiguous rank-leading(panel_size, nvir, nvir)float64 host panel, zero-padded on the rank tail exactly like_read_x_rank_panel(); every panel shares one shape, so each term compiles exactly once.noccis already baked intopanel_sourceand is accepted for symmetry with_stream_partial_x(); callers validate with_validate_x_stream().If a panel
device_putor kernel dispatch raises a device out-of-memory error anyway (the free-memory measurement cannot see BFC fragmentation), the whole loop retries from scratch at half the panel width, halving again on each failure down torank_panel_size.