Files
voice_linux/plans/gpu_passthrough_lessons_learned.md
T

11 KiB

GPU Passthrough in Qubes OS — Summary & Lessons Learned

What We Did

Passed an NVIDIA GTX 1080 Ti (BDF 65:00.0 / 65:00.1) through from Qubes dom0 to a StandaloneVM called ai, installed NVIDIA drivers and CUDA toolkit, and rebuilt whisper.cpp with CUDA support for GPU-accelerated speech-to-text.

Environment

Component Detail
Host OS Qubes OS with Xen hypervisor
dom0 display GPU AMD (stays in dom0)
Passthrough GPU NVIDIA GeForce GTX 1080 Ti (Pascal/GP102, 11GB VRAM, compute 6.1)
Second NVIDIA GPU RTX 3050 6GB at BDF 17:00.0 (not used for this project)
Target VM ai — originally an AppVM, converted to StandaloneVM
VM OS Debian 13 (bookworm/trixie) with XFCE
NVIDIA driver 550.127.05
CUDA toolkit 12.6

Timeline of Steps

Phase 0: dom0 GPU Passthrough

  1. Identified GPU BDF addresseslspci in dom0 showed two NVIDIA GPUs. We targeted the GTX 1080 Ti at 65:00.0 (GPU) and 65:00.1 (HDMI audio controller).

  2. Confirmed boot loader — Verified Qubes was using GRUB (not systemd-boot) by checking /etc/default/grub existed.

  3. Hid GPU from dom0 — Added rd.qubes.hide_pci=65:00.0,65:00.1 to GRUB_CMDLINE_LINUX in /etc/default/grub.

  4. Regenerated GRUB configsudo grub2-mkconfig -o /boot/efi/EFI/qubes/grub.cfg (EFI path, not legacy BIOS path).

  5. Rebooted dom0 — Full system reboot required for PCI hide to take effect.

  6. Verified GPU was assignablexl pci-assignable-list showed 0000:65:00.0 and 0000:65:00.1.

  7. Attached GPU to VMqvm-pci attach ai dom0:65_00.0 --persistent -o permissive=true (and same for 65_00.1).

Phase 1: Driver & CUDA Installation (inside ai VM)

  1. Installed build prerequisitesbuild-essential, linux-headers-$(uname -r).

  2. Installed NVIDIA driver — Downloaded .run installer, ran with --no-opengl-files --dkms.

  3. Installed CUDA toolkit — Via NVIDIA's Debian repo + cuda-toolkit-12-6.

  4. Verifiednvidia-smi showed GTX 1080 Ti, nvcc --version showed CUDA 12.6.

Phase 2: Application Rebuild

  1. Rebuilt whisper.cpp with CUDAWHISPER_CUDA=1 bash ./build.sh.

  2. Verified CUDA linkageldd ./voice_linux | grep whisper|ggml confirmed CUDA libraries linked.

  3. Tested GPU inference — Ran with GPU mode, confirmed nvidia-smi showed process using GPU memory.

Problems Encountered & Solutions

Problem 1: AppVM vs StandaloneVM confusion

What happened: The ai VM was initially described as an AppVM (template-based). In an AppVM, anything installed outside /home is lost on reboot — NVIDIA drivers would vanish.

Discovery: User clarified it was actually a StandaloneVM, not template-based.

Lesson: Always verify VM type before planning driver installation. In Qubes:

  • AppVM: Root filesystem resets to template on reboot. Drivers must go in the template or use bind-dirs.
  • StandaloneVM: Full persistent root filesystem. Drivers persist normally.
  • Check with: qvm-prefs ai virt_mode and qvm-ls --fields name,klass ai

Problem 2: VM must be shut down before PCI attach

What happened: Attempted to run qvm-pci attach while the ai VM was running. The attach appeared to succeed but the device wasn't visible inside the VM.

Discovery: qvm-device pci list ai showed empty even though attach command ran.

Lesson: The VM must be shut down before attaching PCI devices with --persistent. The workflow is:

  1. Shut down VM
  2. Attach PCI device
  3. Start VM
  4. Verify inside VM with lspci

Problem 3: "Already assigned" error on re-attach

What happened: After a reboot cycle, running the attach command again produced an "already assigned" message.

Lesson: --persistent means the attachment survives reboots. Don't re-run the attach command after reboot — it's already configured. Verify with qvm-pci list ai from dom0.

Problem 4: lspci not found inside VM

What happened: Tried to verify GPU visibility inside the VM but lspci wasn't installed.

Solution: sudo apt install pciutils then lspci | grep -i nvidia.

Lesson: Minimal Debian VMs may not have pciutils installed. Include it in prerequisites.

Problem 5: Script couldn't be copy-pasted into dom0

What happened: Created a comprehensive dom0 shell script for GPU passthrough, but Qubes security model prevents clipboard paste from other VMs into dom0.

Solution: Provided numbered step-by-step commands that could be typed manually, and also created the script as a file that could be transferred via qvm-run or Qubes file copy.

Lesson: When automating dom0 tasks in Qubes:

  • Dom0 is intentionally isolated — no clipboard sharing
  • Scripts must be typed manually or transferred via qvm-copy-to-vm (from dom0 to VM) or qvm-run -p pipes
  • Keep dom0 scripts short and simple
  • Always include a revert mechanism

Problem 6: Reboot scope confusion

What happened: Unclear whether "reboot" meant just the VM or the entire Qubes system (dom0 + all VMs).

Clarification: GRUB changes require a full system reboot (dom0 reboot, which takes down all VMs). PCI attachment changes only require the target VM to be restarted.

Lesson: Be explicit about reboot scope:

  • sudo reboot in dom0 = full system reboot
  • qvm-shutdown ai && qvm-start ai = just the VM

Problem 7: Build accidentally ran without whisper support

What happened: After some iteration, a build was accidentally triggered with WITH_WHISPER=0, producing a binary that showed "[transcription unavailable: rebuild with WITH_WHISPER=1]".

Solution: Rebuilt with WITH_WHISPER=1 bash ./build.sh.

Lesson: The build system defaults matter. Our build.sh auto-detects whisper if the vendor directory exists, but explicit WITH_WHISPER=0 overrides that. Always verify the build output includes whisper support with ldd ./voice_linux | grep whisper.

Problem 8: Dirty git worktree blocking version increment

What happened: The increment_and_push.sh script refused to run because of uncommitted changes in the working tree.

Solution: Committed the intended changes first, then ran the increment script.

Lesson: The increment script enforces a clean worktree policy. Always commit your changes before running it. If there are unrelated/untracked files, git stash them first.

Key Qubes-Specific Knowledge

PCI Passthrough Essentials

# Hide device from dom0 (GRUB, requires full reboot):
rd.qubes.hide_pci=BDF1,BDF2

# Attach to VM (VM must be off):
qvm-pci attach VMNAME dom0:BDF --persistent -o permissive=true

# Verify assignment:
qvm-pci list VMNAME        # from dom0
lspci | grep -i nvidia      # from inside VM

NVIDIA Driver in Qubes VM

  • Always use --no-opengl-files — Qubes VMs use a virtual GPU for display. Installing OpenGL files would break the display.
  • Use --dkms — Ensures kernel module rebuilds on kernel updates.
  • StandaloneVM recommended — Avoids template pollution and persistence issues.
  • permissive=true — Required for NVIDIA GPUs in Qubes due to how they access PCI config space.

What Persists Where

VM Type /home /usr, /lib, /etc Drivers
AppVM Persists Resets to template Lost on reboot
StandaloneVM Persists Persists Persists
Template Persists Persists Persists (shared to AppVMs)

Dom0 Safety

  • Dom0 has no network access by design
  • Clipboard is one-way (dom0 → VM, not VM → dom0) and requires explicit Ctrl+Shift+C/V
  • Always have a revert plan for GRUB changes (keep a backup of /etc/default/grub)
  • Log what you change — our script wrote to /var/log/gpu_passthrough.log

Architecture Diagram

┌─────────────────────────────────────────────────────────┐
│  dom0 (Xen Hypervisor)                                   │
│                                                           │
│  AMD GPU ──── Display                                     │
│  GTX 1080 Ti ──── HIDDEN via rd.qubes.hide_pci           │
│       │                                                   │
│       │ PCI passthrough (qvm-pci attach --persistent)     │
│       ▼                                                   │
│  ┌─────────────────────────────────────────────────────┐ │
│  │  ai StandaloneVM                                     │ │
│  │                                                       │ │
│  │  NVIDIA Driver 550.x (--no-opengl-files --dkms)      │ │
│  │       │                                               │ │
│  │  CUDA Toolkit 12.6                                    │ │
│  │       │                                               │ │
│  │  whisper.cpp (GGML_CUDA=ON)                           │ │
│  │       │                                               │ │
│  │  voice_linux (GPU-accelerated transcription)          │ │
│  └─────────────────────────────────────────────────────┘ │
└─────────────────────────────────────────────────────────┘

Performance Results

Model VRAM Used Decode Time (10s audio) Quality
tiny.en ~1GB <0.5s Fair
base.en ~1GB <1s Good
small.en ~2GB ~1-2s Very good
medium.en ~5GB ~2-3s Excellent
large-v3-turbo ~6GB ~3-4s Excellent+
large-v3 ~10GB ~4-6s Best

The GTX 1080 Ti with 11GB VRAM can run all models including large-v3. For real-time dictation, medium.en provides the best accuracy-to-speed tradeoff.

Files Created During This Process

File Purpose
plans/gpu_enablement_plan.md Step-by-step GPU passthrough and driver installation plan
plans/architecture.md Full system architecture including hardware details
scripts/download_models.sh Downloads all compatible whisper models

If You Had to Do It Again

  1. Verify VM type firstqvm-ls --fields name,klass ai before planning anything
  2. Shut down VM before PCI operations — always
  3. Back up GRUB configsudo cp /etc/default/grub /etc/default/grub.bak before editing
  4. Install pciutils earlysudo apt install pciutils so you can verify GPU visibility
  5. Use the .run installer, not apt packages — NVIDIA's apt packages for Debian can conflict with Qubes' virtual GPU setup; the .run installer with --no-opengl-files is cleaner
  6. Test CPU mode first — Build and verify voice_linux works on CPU before adding GPU complexity
  7. Keep dom0 commands minimal — Type them manually, don't try to automate complex scripts in dom0