Building a Dedicated Local AI Compute Node on an ASRock Z390 Platform

1

 Repurposing Retired Hardware: Building a Dedicated Local AI Compute Node on an ASRock Z390 Platform

When high-end hardware cycles out of the gaming rig rotation, it doesn't mean it belongs in a landfill. My retired ASRock Z390 Phantom Gaming 4 motherboard, paired with an Intel Core i7-9700K and 32GB of DDR4 memory, recently found a second life as a dedicated, headless local AI calculation engine for the workshop.
Instead of forcing a heavy virtualized storage server to handle machine learning inference, this build utilizes a clean, split-node architecture.

🧰 Hardware Specification & Architecture
  • Motherboard: ASRock Z390 Phantom Gaming 4 (Robust power delivery and PCIe lane flexibility).
  • CPU: Intel Core i7-9700K (8 Cores, 8 Threads; running stock JEDEC memory timings for absolute 24/7 reliability—no XMP aging degradation here).
  • RAM: 32GB Corsair DDR4 matched kit (Dual-channel configuration).
  • GPU: Gigabyte RTX 2080 Super 8GB VRAM.
  • Power Supply: Corsair RM850e (850W, 80 PLUS Gold, providing ample headroom for transient GPU spikes).
  • Storage: Dedicated 256GB NVMe boot and model weight cache drive.
🌐 The Split-Node Network Strategy
The core philosophy of this build is decoupling heavy mathematical compute from long-term data storage:
  1. The Compute Node (Bare-Metal Ubuntu Server): The ASRock rig runs Ollama natively on Ubuntu Server LTS. It is isolated on the network and acts strictly as a raw VRAM/RAM crunching engine. When the workshop is closed, a simple network command (sudo shutdown -h now) powers the machine completely down to zero watts.
  2. The Storage & Frontend Node (TrueNAS SCALE): A separate virtualization and storage server hosts Open WebUI (v0.11.3) inside a container. The TrueNAS instance acts as the persistent brain—storing chat histories, Thiele-Small driver datasheets, and REW acoustic measurement text files (.frd / .zma) on a resilient ZFS pool.
[ ASRock AI Compute Node ] <──( LAN Port 11434 )──> [ TrueNAS SCALE / Open WebUI ]
2
  - i7-9700K / 32GB / RTX 2080S                       - Persistent ZFS Storage Pool
3
  - Ollama Engine Hosting Qwen2.5:32B                 - Responsive Workbench Dashboard
4

🛠️ Overcoming the Setup Hurdles
Building a headless server out of retired desktop gaming parts is rarely a straight line. This deployment fought through a few classic Linux and container networking quirks:
  • The UEFI Storage Table Loop: The initial installation fought a glitched storage loop that dropped the primary user from the sudo administrative group. Resolved via a quick kernel drop-in and credential refresh.
  • The Web Redirect "Doom Loop": Direct curl and wget requests to generic download URLs kept grabbing 40KB–500KB HTML release webpage wrappers instead of the raw binary assets. Bypassing the web redirects and pulling the authentic compressed production archives directly solved the pathing errors.
  • Network Socket Binding (OLLAMA_HOST): By default, manual binary extractions bind strictly to loopback (127.0.0.1). Injecting a systemd drop-in override (Environment="OLLAMA_HOST=0.0.0.0:11434") and opening the UFW firewall rules bridged the ASRock hardware directly to the TrueNAS container.
  • The TrueNAS Protocol Syntax Trap: TrueNAS application environment variables require the explicit protocol header. Declaring OLLAMA_BASE_URL as a fully qualified string (http://192.168.1.50:11434) instantly brought the Qwen-2.5-32B model dropdown to life on the web dashboard.

🎯 The Result
The ASRock Z390 rig now sits quietly in the rack, pulling an 8GB VRAM / 11GB system RAM split to run the 32B model effortlessly. Coupled with custom engineering system prompts, the setup provides an instant, private, air-gapped loudspeaker design assistant right next to the workbench.
Technical Setup Guide: Multi-Node Local AI for Workshop & Engineering Workflows
🛠️ The Hardware Infrastructure
This environment operates across two distinct local server nodes pooled into a single frontend via Open WebUI:
  1. The Heavy Engine (Debian 13 Node): RTX 5070 (12GB VRAM) / 32GB DDR5 RAM. This machine handles the main engineering tasks, large context files, and native multimodal vision tasks.
  2. The Workbench Node (Ubuntu Server Node): ASRock Z390 / RTX 2080 Super (8GB VRAM) / 32GB DDR4 RAM. This machine handles quick reference lookup, code snippet generation, and troubleshooting directly at the physical workbench.

🎯 Model Allocation & VRAM Optimization Strategy
To eliminate the massive speed penalties caused by running models that spill into system RAM (like Qwen-2.5-32B on limited hardware), models are split strictly by VRAM availability:
1. 12GB VRAM Deployment (Debian Node)
  • Primary Engine: gemma4:12b
    • Footprint: ~8 GB VRAM.
    • Usage: Fits completely into memory with 4GB left for context. Elite coding logic and native vision.
  • Deep Reasoning Alternate: gemma4:26b (A4B Mixture of Experts)
    • Usage: Activates only 3.8B parameters per token. Runs drastically faster than traditional 26B+ dense models on limited hardware.
2. 8GB VRAM Deployment (Ubuntu Node)
  • Primary Workbench Tool: gemma4:e4b
    • Usage: Ultra-fast, natively multimodal, and leaves massive headroom for large text logs.
  • Text & Scripting Alternate: qwen2.5:7b
    • Usage: Blazing-fast generation for terminal commands, Proxmox scripts, and OpenWrt logs.

📁 Open WebUI Workspace & Data Management
To handle heavy speaker engineering data and networking configurations smoothly, data is structured into isolated Knowledge Bases:
  • #Master Driver Datasheets (Global Reference): A permanent Knowledge Base containing raw manufacturer specification text files, T/S parameters, and driver dimensions.
  • #Project_Name (Specific Workspaces): Isolated Knowledge Bases for active speaker builds. Holds exported Room EQ Wizard (REW) text curves (.txt/.frd/.zma) and active cross-over configuration plans.
  • Note on REW Files: Proprietary binary .mdat files are not uploaded directly. Data is fed to the AI either via text exports or by uploading graph screenshots using Gemma 4's native vision features.

🔄 Multi-Node Local Integration (Connections)
To link both independent Ollama backends into your single Open WebUI dashboard:
  1. Ensure both servers expose port 11434 externally (Environment="OLLAMA_HOST=0.0.0.0:11434").
  2. In Open WebUI, navigate to Settings > Admin Settings > Connections.
  3. Under the Ollama API configuration, click the Plus (+) icon and enter both explicit network endpoints:
    • http://<Debian_IP>:11434
    • http://<Ubuntu_IP>:11434

💬 Future Reference System Prompt (Save in Open WebUI)
Copy and paste the text below into your Open WebUI System Prompt settings for your workshop profile:
text
You are an expert workshop AI assistant specialized in loudspeaker crossover design, acoustics, network topology engineering, and Linux system administration (specifically Debian, Ubuntu Server, Proxmox, and OpenWrt). 
5
​
6
When analyzing uploaded Knowledge Base files (#Master Driver Datasheets or specific #Project Folders):
7
1. Read plain text acoustic (.frd) and impedance (.zma) tables column-by-column to identify resonant frequencies (Fs), breakup modes, and impedance minimums.
8
2. Provide clean, actionable CLI scripts for Proxmox container management and OpenWrt firewall rule manipulation.
9
3. Keep responses highly technical, precise, and concise. Your primary target hardware configuration includes an RTX 5070 (12GB) master node and an RTX 2080 Super (8GB) workbench node running optimized Gemma 4 and Qwen 2.5 variants. Focus on execution speed and context efficiency.
10
Use code with caution.


Page settings Options Reader comments Allow Do not allow, show existing Do not allow, hide existing Page moved to Trash

Comments

Popular posts from this blog

The DIY JBL Tribute "Wide-Baffle" Monitoring Towers

Project Master Blueprint: Hybrid MTM SLOB Tower Speaker