Model & Topology Config
Dual RTX 3060
Transformer Layers (N)
28
VRAM per Layer (MB)
200 MB
GPU 0 Thermal Pressure
47°C
GPU 1 Thermal Pressure
43°C
PCIe Transfer Weight
10.0 MB/cut
⚛ Quantum Objective: Min H(x) = ∑ w_uv(x_u - x_v)² + λ(∑ x_i · v_i - C)² + Thermal_Bias
Real-Time Performance
QUBO Optimal
0.94 ms
Slicing Latency
632.2
Prompt Tok/s
65.76
Gen Tok/s
-8.75
QUBO Energy Cost
2D Hamiltonian Energy Landscape
Min Ground State H(x)
Inference Insight: GPU 1 is running 4°C cooler. QUBO dynamically shifted 28 layers (5.5GB) to GPU 1, eliminating cross-GPU PCIe bus transfer penalties.