⚛ Quantum GPU & LLM Lab
API Docs Whitepaper
Model & Topology Config Dual RTX 3060
Transformer Layers (N) 28
VRAM per Layer (MB) 200 MB
GPU 0 Thermal Pressure 47°C
GPU 1 Thermal Pressure 43°C
PCIe Transfer Weight 10.0 MB/cut
⚛ Quantum Objective: Min H(x) = ∑ w_uv(x_u - x_v)² + λ(∑ x_i · v_i - C)² + Thermal_Bias
📷 Isometric 3D View
GPU 0 (Primary)
GPU 1 (Cooler Target)
PCIe Cut Activation
Real-Time Performance QUBO Optimal
0.94 ms
Slicing Latency
632.2
Prompt Tok/s
65.76
Gen Tok/s
-8.75
QUBO Energy Cost
GPU 0 VRAM (0 layers) 0.0 / 12.2 GB
GPU 1 VRAM (28 layers) 5.5 / 12.2 GB
2D Hamiltonian Energy Landscape Min Ground State H(x)
Inference Insight: GPU 1 is running 4°C cooler. QUBO dynamically shifted 28 layers (5.5GB) to GPU 1, eliminating cross-GPU PCIe bus transfer penalties.