pytc.solver.isdf_xtc_ccsd._stream_partial_x_pipelined

pytc.solver.isdf_xtc_ccsd._stream_partial_x_pipelined(panel_kernel, t2, left_out, left_inner, panel_source, nocc, *, occupied_pair_batch_size, rank_panel_size, panel_budget_bytes=None)[source]

Working-set X panel loop with double-buffered async prefetch.

Same math, padding, and accumulation semantics as _stream_partial_x(), but the panel width is sized from the measured device working set and the next panel’s host read is issued (via async_read()) before the current panel’s kernel, so its H2D device_put overlaps compute (JAX async dispatch).

panel_source(m0, m1, panel_size) returns one contiguous rank-leading (panel_size, nvir, nvir) float64 host panel, zero-padded on the rank tail exactly like _read_x_rank_panel(); every panel shares one shape, so each term compiles exactly once. nocc is already baked into panel_source and is accepted for symmetry with _stream_partial_x(); callers validate with _validate_x_stream().

If a panel device_put or kernel dispatch raises a device out-of-memory error anyway (the free-memory measurement cannot see BFC fragmentation), the whole loop retries from scratch at half the panel width, halving again on each failure down to rank_panel_size.