{
  "project": "gpu-pilot",
  "doc": "CONTINUE_6_28_2026.md",
  "title": "\u25b6 CONTINUE 6/28 (LOS AI-1003 moved to gtfs3's dedicated RTX 5060 Ti / qwen2.5-coder:14b \u2014 newest handoff)",
  "oneline": "Newest handoff: moved the mortgage LOS AI-1003 assembly off geo2's shared GPU onto gtfs3's new dedicated RTX 5060 Ti (16GB) running qwen2.5-coder:14b (an upgrade over the 7b) \u2014 always-on, quota-free, warm ~0.6s. Fixed the new GPU with no reboot / no commute-DB downtime (595-open driver hot-swap), installed Ollama LAN-bound, and freed geo2. macstudio was investigated + rejected (contended); geo<->Mac passwordless SSH staged.",
  "tldr": [
    "LOS AI host move (certihomes-los 4b11981 on geo2): OLLAMA_BASE=http://172.26.1.152:11434, OLLAMA_MODEL=qwen2.5-coder:14b. Verified live: /health=qwen2.5-coder:14b@ollama, 'assembled via ai:qwen2.5-coder:14b' (real URLA), warm ~0.6s, ~9GB/16GB resident. geo2's 7b removed -> its 32GB GPU freed for ComfyUI/avatar.",
    "gtfs3 GPU brought up with NO reboot / NO DB downtime: RTX 5060 Ti (10de:2d04) wasn't enumerated by driver 580 -> installed nvidia-driver-595-open (DKMS, kernel 6.8.0-124) -> hot-swapped modules (rmmod nvidia stack -> modprobe nvidia) -> nvidia-smi lists it. Ollama installed (systemd, 0.0.0.0:11434, KEEP_ALIVE=-1, CUDA 13). Persists across reboots.",
    "Model-fit decision: 16GB card -> qwen2.5-coder:14b (~9-12GB, fits fully = sweet spot). 32b (~20GB) does NOT fit 16GB (needs a 24GB+ card / RTX 5090); Q2/Q3-cramming loses more than it gains; CPU-offload too slow. 'A model that fits fully in VRAM beats a bigger one that spills to CPU.'",
    "GOTCHAS: new Blackwell 'No devices found' = driver too old (595-open for 5060 Ti); hot-swap GPU driver w/o reboot on headless boxes; curl|sh installers need run-as-root via 'sudo -S bash file < ~/secert.txt' over non-interactive SSH; Ollama LAN bind override; manage models via HTTP API (no SSH); geo's :11434 firewalled from peers; macstudio SSH locked + contended; 172.26.1.152 = gtfs3 (NOT macmini)."
  ],
  "updated": "2026-06-28 14:10",
  "sections": [
    {
      "level": 1,
      "title": "CONTINUE \u2014 2026-06-28 \u2014 Moved the LOS AI-1003 to gtfs3's dedicated RTX 5060 Ti (qwen2.5-coder:14b)"
    },
    {
      "level": 2,
      "title": "DONE this session"
    },
    {
      "level": 2,
      "title": "PENDING (precise)"
    },
    {
      "level": 2,
      "title": "DECISIONS (user chose + why)"
    },
    {
      "level": 2,
      "title": "GOTCHAS (don't relearn) \u2014 full detail in memory `fleet-llm-inference-hosts`"
    }
  ],
  "stats": {
    "lines": 33,
    "words": 659
  },
  "html": "https://board.certihomes.com/project-docs/gpu-pilot/CONTINUE_6_28_2026.html",
  "out_name": "CONTINUE_6_28_2026.md",
  "comments_api": "https://board.certihomes.com/api/doc-comments?project=gpu-pilot&doc=CONTINUE_6_28_2026.md"
}