Overview
This post reviews HPC-vQPU: A Service-Export Architecture for Virtual QPUs on Batch-Scheduled HPC Systems (arXiv:2605.28845, 29 pages, 5 figures), a paper from the Pawsey Supercomputing Centre validated on the Setonix system with AMD MI250X GPUs and IBM Fez calibration data.
One-line takeaway: HPC-vQPU frees quantum simulators from the batch-job straitjacket and serves them externally like a real QPU.
Key points
The connectivity inversion problem
- HPC interfaces are scheduler-oriented: submit scripts, queue, run, exit. Quantum software wants a QPU — a device with submit/observe/query semantics like IBM Quantum or AWS Braket.
- Existing options fall short: cloud platforms scale poorly and cost much; science gateways don't handle calibration, topology semantics, or gate-set constraints; running quantum software inside batch systems turns every interaction into a job submission.
- Security is the deeper barrier: HPC compute nodes cannot open inbound ports. All coordination must be outbound and agent-initiated — hence "connectivity inversion": exporting the internal simulator rather than sending jobs in.
- Control plane (
vqpu_server): cloud-visible, speaks in devices/tasks/lifecycle states/results via REST API and SSE event streams; the authoritative source of truth. It knows nothing about partitions, module stacks, or batch scripts. - Execution plane (
vqpu_agent): an unprivileged process on HPC login nodes that CLAIMs tasks, translates them into scheduler jobs (Slurm/Dask), and REPORTs results. - Only three verbs cross the boundary: CLAIM, HEARTBEAT, REPORT. No callbacks, no RPC. A dead agent is detected by heartbeat expiry only.
- A snapshot Δ(𝒟, t) contains the directed coupling map G = (Q, E), native gate set G_native, per-qubit T1/T2, single-qubit error ε_1q, readout error r, and directional two-qubit errors ε_ij ≠ ε_ji.
- A 156-qubit virtual device
ibm-fez-0309was built from real IBM Fez calibration data, compared against a zero-erroribm-fez-ideal. - Snapshots bind at CLAIM time — not at submission (too early) nor on compute nodes (would break hermeticity). An experiment zeroed all noise parameters 1.2s after submission of 8 tasks: all 8 returned TV = 0, proving claim-time binding.
- D1 outbound-only connections; D2 terminal states (COMPLETED/FAILED/CANCELLED) are absorbing; D3 atomic exclusive CLAIM; D4 asymmetric heartbeat-driven failure detection; D5 immutable claim-bound snapshots; D6 graph-parametric topology validation; D7 TTL snapshot caching; D8 sealed execution (no network on compute nodes); D9 server-side observability via SSE.
- Environment: Setonix, AMD Instinct MI250X, Qiskit-Aer + cuQuantum, IBM Fez 2026-03-09 calibration — a production system, not a lab prototype.
- Experiment A (bounded overhead): 28–32 qubit random circuits, 1024 shots. At 32 qubits, T_exec = 23.77 s; service-layer overhead is only 3.7%, remaining constant-time while simulation cost grows exponentially.
- Experiment B (device-aware fidelity): noisy device gives TV = 0.03537 ± 0.00171 vs. deterministic |00⟩ on the ideal device, consistent with readout errors and the calibrated two-qubit channel on edge (0,1).
- Experiment C (claim-time binding): 8/8 tasks returned noiseless output after mid-flight calibration reset.
- Experiment D (crash recovery): killed agent leaves tasks RUNNING (no auto-fail, no auto-requeue); explicit admin requeue → 3/3 completed. The system deliberately avoids creating competing execution lineages.
- Single-qubit depolarizing (excluding virtual gates like Rz), symmetric readout matrices, and directional two-qubit depolarizing channels. The control-plane validator and compute-node runner use identical derivation logic, preventing semantic divergence between validation and execution paths.
- Complements cloud quantum platforms (same abstractions, HPC-scale backends), Qiskit Runtime (no assumption of a remotely reachable backend), simulators like Qiskit-Aer/cuQuantum (adds a service layer), science gateways (adds device contracts), and QPU time-multiplexing efforts like Quantum Brilliance (opposite direction: exporting simulators as QPUs).
- Exponential statevector simulation cost is isolated but not hidden — beyond ~32 qubits, simulation time dominates.
- Recovery is conservative; automated post-heartbeat-expiry policies are future work.
- Only Dask-backed Slurm integration; PBS/LSF/Kubernetes need adapters.
- Only simulators validated; physical QPU integration requires further work, though the architectural principles hold.
Two-plane architecture
Device snapshots as immutable contracts
Nine design invariants (D1–D9)
Experiments on Pawsey Setonix
Noise model consistency
Relation to existing work
Limitations
Commentary
The paper's value is not a novel algorithm but an engineering solution to a problem that has stalled quantum–HPC integration: HPC keeps doing what it does best (scheduling, isolation, GPU compute) while exposing a device interface instead of a job script. The pattern — dual planes, outbound-only, snapshot contracts, sealed execution — generalizes to any HPC resource needing interactive external service, from deep-learning inference to scientific visualization. The experiments are modest in scale but prove feasibility in a production environment: external clients submit circuits via REST, observe via SSE, and receive device-aware results driven by real calibration data — without a single inbound port opened into the HPC boundary.
> Reference: Shusen Liu et al., "HPC-vQPU: A Service-Export Architecture for Virtual QPUs on Batch-Scheduled HPC Systems", arXiv:2605.28845 [cs.DC], 2026.