WarGreymon
A home server and network infrastructure design, named for WarGreymon — the Mega-level evolution of Agumon in the Digimon series, a defence-oriented form noted for its shield and its role protecting the Digital World. This document is the design reference: hardware, rack layout, network architecture, service interconnection, configuration procedure, and security posture for the system.
00Quick start — build order
Build the network before you build on the network. Everything downstream needs a gateway, an address, and a stable link before it can be configured — follow this order and each device is reachable, patched, and on the right subnet by the time you get to it. Tick boxes as you go; this is the "what's next" page — every step links to the full detail elsewhere on this page.
00 · Plan before you plug anything in15–30 min — cheapest mistakes are the ones fixed on paper
- Topology confirmed: Modem → OPNsense → MikroTik switch → NAS / Proxmox / everything else
- Subnet plan is set:
10.0.X.0/24per VLAN, 15 VLANs already designed in Section 05 (10 Mgmt, 15 Container, 20 Media, 25 Indexing, 30 Cloud, 35 Trading, 40 Photos, 45 Cameras, 50 Monitor, 60 DNS, 70 VPN, 80 Database, 90 Testing, 100 Home Devices, 999 NAS) - Static/reserved addresses noted: OPNsense LAN
10.0.1.1, switch mgmt10.0.1.5, NAS10.0.30.10+10.0.1.100, NanoKVM reservation, Technitiumdns01/dns02 - Cable + PDU labelling convention set —
SW-SFP1-OPNstyle, per Section 03
01 · Modem / ISP handoffConfirm raw internet before anything else joins the chain
- Laptop connected directly to the modem, confirmed it gets a public IP / internet access
- Connection type noted: DHCP, PPPoE, static, or VLAN-tagged (matters for OPNsense's WAN config next)
- Laptop disconnected — OPNsense's em1 (WAN) takes this port next
02 · OPNsense — before the switchEverything after this already has a gateway waiting for it
- Modem → OPNsense
em1(WAN) - Laptop → OPNsense LAN directly for initial setup (skip the switch for now)
- Firmware updated, root password changed immediately — full steps in Section 07 → Hardware → OPNsense, Step 1
- WAN configured per the connection type noted in Step 01
- LAN (SFP+ trunk) set up, all 15 VLAN interfaces created — Section 07 → Hardware → OPNsense, Steps 3–4
- Default-deny baseline + explicit allow rules in place — Section 07 → Hardware → OPNsense, Step 5
- Internet confirmed working through OPNsense from the laptop
Why not the switch first: a switch without a firewall behind it just forwards packets to a dead end — nothing gets an IP or a route out.
03 · MikroTik switchExtend the gateway cleanly to every port
- Switch uplink (SFP+1) → OPNsense SFP+1
- RouterOS firmware updated
- Switch given a static management IP on VLAN 10 (
10.0.1.5) — not on the isolated NAS VLAN - All 15 VLANs created to match the plan, trunk vs access ports set — Section 07 → Hardware → MikroTik switch
- Test device on a port confirmed reaching the internet
04 · NASConfigured once the network under it is stable — moving a NAS to a new subnet later is disruptive
- NAS NIC1 + NIC2 → switch ports 4 & 5
- TOS firmware updated before creating any storage pool
- Static/reserved IPs set on the correct VLANs (
10.0.30.10primary,10.0.1.100secondary) - RAID 5 array created — Section 07 → Hardware → NAS, Step 2
- Default admin credentials changed, unused default services disabled
- Only the shares actually needed are enabled (NFS to Proxmox, nothing extra left on)
05 · Other hosts — Proxmox, NanoKVMConnect, address, patch, harden — before exposing any service on top
- Proxmox (MS-01) connected to the switch, correct VLAN trunk confirmed — Section 07 → Hardware → Proxmox
- Static IP / reservation assigned per the plan
- OS patched, real credentials set before any service is exposed
- NanoKVM connected and secured — Section 07 → Hardware → NanoKVM
- Everything in Section 09 (Kubernetes, AI, Trading) waits until this phase is fully stable — see the resource budget there before starting it
06 · Wireless access pointNot part of the current hardware list — skip unless/until one is added
- N/A for now — no dedicated AP in the current build. If one is added later: connect it to a trunk port, map each SSID to a VLAN, same pattern as everything else here
07 · End-to-end testingProve the whole chain works before trusting it
- From a client on each VLAN, confirm internet access and expected inter-VLAN isolation against the traffic matrix in Section 05
- NAS and other hosts reachable by name, not just IP, once Technitium (Section 07 → Services → Monitoring & DNS) is live
- Deliberately reboot OPNsense and the switch once — confirm everything comes back up in the right order with no hand-holding
08 · HardeningFull detail in Section 07
- Every device audited for a changed default password — OPNsense, switch, NAS, NanoKVM
- NetBird preferred over any port-forward for remote access — Section 08
- Automatic security updates enabled where safe to do so
- UPS installed and confirmed on OPNsense, switch, and NAS at minimum — Section 04
09 · Backups & monitoring — last, on purposeNeeds everything above already stable to back up or monitor reliably
- Proxmox Backup Server jobs scheduled — Section 07 → Services → Backups
- Uptime Kuma monitoring OPNsense, switch, and NAS at minimum
- IPs, VLANs, and credentials documented somewhere outside the lab itself (this repo, plus a password manager for secrets — never commit credentials to git)
02Hardware inventory
Cabinet, UPS, and the 10GbE cables are now in hand. One optional addition below for easier remote hardware management.
| Role | Device | Form factor | Status |
|---|---|---|---|
| Enclosure | Tecmojo 12RU wall-mount cabinet, 450mm deep, 50kg rated, lockable glass door | 19" cabinet | Owned |
| UPS | APC Smart-UPS C SMC1500I-2UC — 1500VA/900W, line-interactive, 4× IEC C13 outlets | 2U rackmount, 439mm deep, 24.1kg | Owned |
| Firewall / router | OPNsense on SJRC N150 mini-ITX box — 2× 2.5GbE + 2× 10GbE SFP+ | Desktop, shelf-mount | Owned |
| Hypervisor | Proxmox VE on MS-01 (i9‑13900K, 32GB RAM, 2TB NVMe) — incl. 10GbE SFP+ | Desktop, shelf-mount | Owned |
| Switch | MikroTik CRS310‑8G+2S+in — 8×1GbE + 2×10GbE SFP+ | 1U, 483×88×288mm | Owned |
| Storage / NAS | Terramaster F4‑424 Pro, 4‑bay, RAID 5, ~4TB usable | Desktop, shelf-mount, ~18kg loaded | Owned |
| Power distribution | 6-outlet IEC PDU | 1U rackmount | Owned |
| 10GbE uplinks | 2× 10GbE DAC Twinax SFP+ cables, 0.5m | — | Owned — ready to install |
| Out-of-band mgmt | Sipeed NanoKVM — KVM-over-IP, HDMI capture + USB device emulation, remote power control | USB-C powered, 1GbE, 1080p60 capture | Owned — cables in hand |
| Home devices switch | Netgear ProSAFE Plus JGS524E — 24-port gigabit, 802.1Q VLAN tagging, satellite switch for household devices | Desktop, 330×173×43mm, 13.5W max | Owned |
| Dashboard display | Wall/desk monitor running a browser kiosk pointed at Grafana/Homarr | HDMI, needs a small feeder device | Planned — future |
| Security camera display | Second monitor for a future NVR (Frigate/ZoneMinder) camera wall | HDMI, needs a small feeder device | Planned — future |
JGS524E — role and physical placement
This is a Netgear ProSAFE Plus-series switch: 24 gigabit copper ports, a limited web-based/Windows-utility configuration interface (not RouterOS-class), supporting 802.1Q tagged VLANs, QoS, port mirroring, and broadcast storm control. It is not a rack-mount unit — at 330mm wide it is roughly two-thirds the width of a 19" rack — and functions as a satellite switch located wherever the household devices actually are (a media cabinet, a desk), rather than inside the enclosure documented in Section 03. A single Cat6 run carries a trunk link back to the MikroTik CRS310's port 7. Power draw is 13.5W maximum; since it sits outside the rack, it is not protected by the UPS in Section 04 unless a separate small UPS is added at its location, which is optional and not required for the lab itself to function.
Its function in this design is a single-purpose one: everything plugged into it — a television, a personal computer, other household devices — reaches the internet through the same OPNsense firewall, Technitium DNS (and its ad-blocking), and Suricata/CrowdSec inspection as every other device in the design, on a VLAN (100) that has no visibility into and no access from any lab VLAN. The remaining 23 ports and the switch's own configuration headroom leave capacity for future expansion — additional household devices, or a physical link to a second lab location, without any change to the design documented here.
NanoKVM — what it actually gives you
Three things bundled into one small device: video (HDMI capture, so you see exactly what a monitor would show — BIOS, boot menus, a crashed console), input (USB-C device emulation, so it behaves as a real keyboard/mouse to the host, no agent software needed on the host side), and optionally power control via the ATX breakout board if the host exposes accessible power/reset header pins. Combined with NetBird for the network path in, this is what makes "fix it from my phone while not at home" actually true — without NetBird it's still LAN-only, and without NanoKVM you still need physical access whenever the OS itself won't boot.
03Physical rack layout
Bottom-to-top: heaviest gear at the bottom. Cable routing between these devices is documented separately in the network topology diagram (Section 05) rather than overlaid here, so this stays a clean read of what occupies which U position.
Cable discipline
- 1GbE cables (switch ↔ OPNsense port 1, switch ↔ Proxmox port 3, switch ↔ NAS ×2) labelled both ends — kept as fallback
- 10GbE DAC cables (OPNsense SFP+1 ↔ switch SFP+1, switch SFP+2 ↔ Proxmox SFP+1) labelled
SW-SFP1-OPN/SW-SFP2-PRX— primary VLAN trunk path, already in hand - Power cables labelled to match PDU outlet number, e.g.
PDU-2-PRX - Network and power cables bundled and routed on opposite sides — never crossed
- DAC cables handled gently — no sharp bends
- NanoKVM: HDMI-in + USB-C emulation cable to Proxmox, Ethernet to switch port 6 (VLAN10), USB-C aux power from a spare Proxmox USB port — velcro or 3M tape it to the shelf, no dedicated RU needed
- Switch port 2 reserved and labelled as the admin laptop drop — run it out to wherever you'll actually sit with a laptop, not just to a wall plate nobody uses
- Every cable tug-tested before first power-on
04Power & UPS
Goal: if mains power drops, every device gets enough runtime to either ride out a short blip or shut down cleanly.
Idle draw
Typical draw
Peak draw
APC Smart-UPS C SMC1500I-2UC — owned
| Spec | Value |
|---|---|
| Capacity | 1500VA / 900W, line-interactive, AVR |
| Outlets | 4× IEC C13 (battery-backed) |
| Form factor | 2U rackmount, 439mm deep, 24.1kg |
| Management | USB + SmartConnect port, graphic LCD status panel |
| Est. runtime | ~25–35 min @ typical 160W, ~8–10 min @ 277W peak |
30% headroom over peak draw — sized for "shut everything down safely," not multi-hour outages. Right trade-off for a home lab.
Power path
Wall outlet → UPS input. UPS output → PDU (RU10) → switch, OPNsense, Proxmox, NAS individually. Never plug the UPS itself into the PDU it's powering.
Automated shutdown
Install and configure NUT (Network UPS Tools) on Proxmox
- Connect the UPS to Proxmox via USB
- Install:
apt install nut - Set
/etc/nut/ups.confdriver tousbhid-upsfor APC Smart-UPS C series - Configure
upsmon.confwith a shutdown threshold (e.g. trigger at 20% battery or 5 min remaining) - Test with a real unplug — confirm Proxmox begins an orderly VM shutdown before the battery is exhausted
- Expose UPS status to Grafana via the NUT exporter for Prometheus
05Network architecture
One firewall, one trunked switch, fifteen isolated VLANs, on a 10GbE backbone.
Logical network diagram — VLANs & services
Physical & core network topology
Client through to every core device, with the switch identified as the system's focal distribution point — every path to compute or storage passes through it. Remote access (NetBird, Cloudflare) is a separate diagram below, so this one stays about the physical/logical core, not VPNs.
Remote access (VPN) diagram
Two independent paths in, drawn as two independent lanes — nothing here shares a line with the physical diagram above, on purpose. NetBird is primary; Cloudflare Tunnel is the architecturally unrelated backup. The diagram below is a live embed of the real .drawio file — each lane is a real layer, and the small layers icon in the viewer's own toolbar (bottom-left of the embed) shows or hides NetBird and Cloudflare Tunnel independently, natively, without any custom button on this page.
VLANs
| VLAN | Name | Gateway | Purpose |
|---|---|---|---|
| 10 | Management | 10.0.1.1 | Proxmox, switch, OPNsense admin, dashboard, reverse proxy, backups, NanoKVM, laptop drop, Cloudflare Tunnel, Authentik, NetBox |
| 15 | Container | 10.0.15.1 | Docker, Talos K8s cluster, Argo CD, Traefik, Ollama + Open WebUI |
| 20 | Media | 10.0.20.1 | Jellyfin, Sonarr, Radarr, Jellyseerr, Bazarr |
| 25 | Indexing | 10.0.25.1 | Prowlarr |
| 30 | Cloud | 10.0.30.1 | Nextcloud (contacts/calendar/files), Vaultwarden — phone backup target |
| 35 | Trading | 10.0.35.1 | MetaTrader 5 (Windows VM) — isolated, outbound-only to broker + NTP + DNS |
| 40 | Photos | 10.0.40.1 | Immich — isolated, phone photo/video backup target |
| 45 | Cameras | 10.0.45.1 | Frigate NVR (planned) — switch-local, no lateral/internet access at all |
| 50 | Monitor | 10.0.50.1 | Prometheus, Grafana, InfluxDB, Uptime Kuma, Loki, Wazuh, ntfy, Beszel |
| 60 | DNS | 10.0.60.1 | Technitium DNS (primary + secondary, clustered), Unbound (recursive resolver) |
| 70 | VPN | 10.0.70.1 | NetBird — self-hosted mesh VPN, remote access |
| 80 | Database | 10.0.80.1 | PostgreSQL, MariaDB |
| 90 | Testing | 10.0.90.1 | Home Assistant, Homebridge, Gitea, Paperless-ngx, sandbox VMs, Flatcar test host |
| 100 | Home Devices | 10.0.100.1 | TV, personal PC, and other household devices — internet + DNS ad-blocking only, isolated from every lab VLAN |
| 999 | NAS | 10.0.99.1 | Storage only — fully isolated |
Switch port map
| Port | Connects to | Mode | Speed |
|---|---|---|---|
| SFP+ 1 | OPNsense SFP+ 1 | Trunk — all VLANs tagged | 10Gbps — primary |
| SFP+ 2 | Proxmox SFP+ 1 | Trunk — all VLANs tagged | 10Gbps — primary |
| 1 | OPNsense em0 (2.5GbE) | Trunk — all VLANs tagged | 1Gbps — fallback |
| 3 | Proxmox eth0 | Trunk — all VLANs tagged | 1Gbps — fallback |
| 4 | NAS NIC1 | Access — VLAN 30 untagged | 1Gbps |
| 5 | NAS NIC2 | Access — VLAN 10 untagged | 1Gbps |
| 2 | Admin laptop drop (room patch point) | Access — VLAN 10 untagged, occasional | 1Gbps |
| 6 | NanoKVM | Access — VLAN 10 untagged | 1Gbps |
| 7 | JGS524E home switch uplink | Trunk — VLAN 10 + 100 tagged only | 1Gbps |
| 8 | — | Reserved | — |
VLAN traffic matrix
| From \ To | 10 | 15 | 20 | 25 | 30 | 35 | 40 | 45 | 50 | 60 | 70 | 80 | 90 | 100 | 999 |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 10 Mgmt | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| 15 Container | ✗ | ✓ | ✓ | ✓ | ✓ | ✗ | ✓ | ✗ | ✗ | ✓ | ✗ | ✓ | ✗ | ✗ | ✗ |
| 20 Media | ✗ | ✗ | ✓ | ✗ | ✓ | ✗ | ✗ | ✗ | ✗ | ✓ | ✗ | ✓ | ✗ | ✗ | ✗ |
| 25 Indexing | ✗ | ✗ | ✓ | ✓ | ✗ | ✗ | ✗ | ✗ | ✗ | ✓ | ✗ | ✗ | ✗ | ✗ | ✗ |
| 30 Cloud | ✗ | ✗ | ✗ | ✗ | ✓ | ✗ | ✗ | ✗ | ✗ | ✓ | ✗ | ✓ | ✗ | ✗ | ✗ |
| 35 Trading | ✗ | ✗ | ✗ | ✗ | ✗ | ✓ | ✗ | ✗ | ✗ | ✓ | ✗ | ✗ | ✗ | ✗ | ✗ |
| 40 Photos | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✓ | ✗ | ✗ | ✓ | ✗ | ✗ | ✗ | ✗ | ✗ |
| 45 Cameras | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✓ | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ |
| 50 Monitor | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✗ | ✓ | ✓ | ✓ | ✓ | ✓ | ✗ | ✓ |
| 60 DNS | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✗ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| 70 VPN | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✗ | ✓ | ✓ | ✓ | ✓ | ✓ | ✗ | ✓ |
| 80 Database | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✓ | ✗ | ✓ | ✗ | ✗ | ✗ |
| 90 Testing | ✓ | ✓ | ✓ | ✓ | ✓ | ✗ | ✓ | ✗ | ✓ | ✓ | ✓ | ✓ | ✓ | ✗ | ✓ |
| 100 Home Devices | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✓ | ✗ | ✗ | ✗ | ✓ | ✗ |
| 999 NAS | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✓ |
Network boundary diagram — what crosses the edge
Four ways traffic can cross the OPNsense boundary, and only four. Direct inbound has nothing to land on — there are no port-forwards to receive it.
Data flow diagram — how information actually moves
Three representative journeys through the stack, left to right.
VLAN zone map
The same 15 VLANs as the swimlane diagram above, drawn instead as trust zones radiating from the firewall — every line is a boundary OPNsense enforces, solid where broadly reachable, dashed where isolated.
Switch port diagram
The MikroTik CRS310's actual front panel, port by port — what to plug into what while cabling the rack (Section 03).
06Service interconnection reference
Earlier sections describe how each service is deployed in isolation. This section describes how they operate together: four architectural patterns underpinning the system, and a consolidated reference table listing every service's purpose, internal connections, external network access, and authentication method.
Service discovery
Every service is addressed by name through Technitium DNS zones rather than a memorised IP address. Portainer additionally exposes container-to-container dependencies within a single Docker host. Name resolution is a one-time lookup; once resolved, services communicate directly and Technitium is no longer part of the data path.
Unified authentication
Authentik is positioned in front of proxied services as a forward-auth layer: a single login and MFA challenge cover every gated service, and every access attempt is logged in one place rather than scattered across separate per-service login systems.
Event-driven messaging
Mosquitto (MQTT) carries IoT device communication for Home Assistant. Sensors and switches publish state changes to topics; subscribers react to those topics without addressing the publisher directly. A dedicated message bus such as NATS is not required at the current service count, since inter-service communication elsewhere in the system runs over standard REST APIs.
Notification aggregation
ntfy is the single delivery path for every alert source — Uptime Kuma, Grafana, Wazuh. A plain SMTP relay serves as a secondary channel. A full mail-gateway appliance is not an appropriate substitute for this purpose, since that class of tool is built for spam filtering on a mail server, not for delivering occasional alert messages.
Service discovery — how a name becomes a connection
Every service in the system is reachable by a stable DNS name (service.lab.internal) rather than a memorised IP address, resolved by the internal Technitium DNS pair described in Section 05. Name resolution happens once per connection: a client asks Technitium for an address, receives it, and then communicates directly with the destination for the remainder of that session. This has two practical consequences. First, an IP address can change — a VM gets rebuilt, a container moves to a different host — without updating every service that talks to it, provided the DNS record is updated once. Second, Technitium sits outside the actual data path: it is consulted at the start of a connection, not on every packet, so it introduces negligible latency and is not a single point of failure for already-established connections if it becomes briefly unavailable.
The diagram below traces a concrete example: Jellyseerr resolving and reaching Radarr. The same three-step pattern — query, resolve, connect directly — applies to every internal service-to-service connection in the system.
End-to-end data flow across the homelab
The service-discovery diagram above traces one hop — a name becoming a connection. This diagram traces the complete path a request takes through the system, role by role: from the requesting device, through the edge (DNS resolution, authentication, reverse proxy), into the application service that handles it, and finally to where the resulting data is persisted.
Event-driven messaging — MQTT for IoT
Mosquitto is an MQTT message broker. IoT devices — Zigbee sensors via a Zigbee2MQTT bridge, DIY ESPHome sensors, Tasmota-flashed smart plugs — publish state changes to named topics on the broker (for example, home/livingroom/temperature) rather than being polled for their current state. Home Assistant, and any other interested service, subscribes to the relevant topics and receives updates the moment they are published. Publishers and subscribers never address each other directly; both only address the broker. This decouples the two sides: a new sensor can be added by subscribing Home Assistant to a new topic, without reconfiguring any existing device, and a battery-powered sensor can announce a change once and return to a low-power state rather than responding to repeated polling requests.
Mosquitto is required only once physical IoT devices using MQTT — Zigbee, Z-Wave via a bridge, or DIY sensors — are present on the network. A Home Assistant instance with no such devices connected has no MQTT traffic to broker and does not require it.
Deploy Mosquitto (MQTT broker)
- Deploy Mosquitto on VLAN 90, alongside Home Assistant.
- Enable authentication (username/password at minimum) — an open MQTT broker lets anyone on the VLAN publish fake sensor data or watch real device state.
- Connect Home Assistant's MQTT integration to it, confirm a test topic round-trips.
- Point Zigbee2MQTT or ESPHome devices at the same broker as they're added.
- Mosquitto deployed with authentication enabled, not anonymous access
- Home Assistant MQTT integration connected and tested
Current software versions
The version of every operating system and major platform component in the design, current as of the last review of this document. Software in active development moves continuously; this table is a reference point, not a pin. Each version should be re-confirmed against the project's own release page immediately before installation, and the deployed version recorded in NetBox (Section 07 → Group B → Proxy & dashboard) or the asset log (Section 14) once installed, rather than assumed to match this table indefinitely.
| Component | Current stable version | Base |
|---|---|---|
| Proxmox VE | 9.2 | Debian 13 "Trixie" |
| Proxmox Backup Server | 4.x | Debian 13 |
| OPNsense | 26.7 "Xenial Xenops" | FreeBSD 15.1 |
| MikroTik RouterOS | 7.x (Stable channel) | — |
| Talos Linux | 1.13.x | — |
| Flatcar Linux | Current stable channel | — |
| Technitium DNS Server | 15.4.x | .NET 10 |
| Debian (Docker/LXC hosts) | 13 "Trixie" | — |
| TrueNAS Community Edition | 25.10 "Goldeye" | Debian Linux |
TrueNAS is included for reference only — the NAS in this design runs Terramaster's own operating system (TOS), not TrueNAS, per Section 02. A fuller comparison of the two, including the ZFS data-integrity distinction, is in Section 07 → Group A → NAS, Step 1.
Application-layer services (Jellyfin, Nextcloud, Authentik, and the remainder of the master table below) are not tracked by version in this document individually — Watchtower (Section 07 → Group B → Proxy & dashboard) keeps container images current on an ongoing schedule, which is a better fit for a service catalogue this size than a manually maintained version table that goes stale between reviews.
Master service interconnection table
Every deployed service, what it's for, what it talks to, whether it reaches the internet, and how it authenticates. Cross-reference against the traffic matrix above for the exact firewall rule each connection depends on.
| Service | VLAN | Purpose | Talks to (internal) | Internet access | Auth |
|---|---|---|---|---|---|
| Proxmox | 10 | Hypervisor | PBS, NAS (NFS), every VM/LXC | Package updates only | Local + Authentik (RAC, optional) |
| PBS | 10 | Backup target | Proxmox | None | Local |
| OPNsense | 10 | Firewall/router | Everything (routes all VLANs) | WAN, updates | Local + MFA |
| MikroTik switch | 10 | VLAN trunk | All wired devices | None | Local |
| Homarr | 10 | Dashboard | Reads status from most services | None | Authentik (forward-auth) |
| NGINX Proxy Mgr | 10 | Reverse proxy | Every proxied service | Let's Encrypt DNS challenge | Local (admin only) |
| NanoKVM | 10 | Out-of-band mgmt | Proxmox (HDMI/USB, not IP) | None | Local + NetBird-only |
| Cloudflare Tunnel | 10 | Backup remote access | NPM | Outbound to Cloudflare edge | Cloudflare Access |
| Authentik | 10 | SSO / forward-auth | Own Postgres+Redis, every gated service | None required | Self (MFA) |
| NetBox | 10 | Infra documentation | Read-only reference, no live deps | None | Authentik |
| Portainer | 15 | Container mgmt / discovery | Docker socket on its host | None | Local + Authentik |
| Talos K8s cluster | 15 | Container orchestration | Argo CD, Traefik, Longhorn, NAS (storage) | Image pulls, Argo CD git sync | Kubernetes RBAC |
| Argo CD | 15 | GitOps sync | Git repo, Talos API | Git repo (outbound HTTPS) | Authentik (OIDC) |
| Traefik | 15 | K8s ingress | Every K8s-hosted service | None required | Passthrough to backend auth |
| Ollama + Open WebUI | 15 | Local AI | Standalone; optional NPM proxy | Model downloads only | Authentik (forward-auth) |
| Jellyfin | 20 | Media server | NAS (NFS media share) | Metadata lookups | Local + Authentik |
| Sonarr / Radarr | 20 | Media automation | Prowlarr (VLAN 25), qBittorrent, Jellyfin, NAS | Indexer APIs (via Prowlarr) | Local, API key |
| Jellyseerr | 20 | Request front-end | Jellyfin, Sonarr, Radarr | None required | Jellyfin account / Authentik |
| Bazarr | 20 | Subtitles | Sonarr, Radarr | Subtitle providers | Local, API key |
| Prowlarr | 25 | Indexer management | Sonarr, Radarr (push config) | Indexer sites | Local, API key |
| Nextcloud | 30 | Files/contacts/calendar | PostgreSQL (VLAN 80), NAS | App store, federation (optional) | Local + Authentik (OIDC) |
| Vaultwarden | 30 | Password manager | Standalone | None required | Self (master password + MFA) |
| MetaTrader 5 | 35 | Trading terminal | None internal — fully isolated | Broker servers, NTP only | Local Windows + broker login |
| Immich | 40 | Photo/video backup | Own Postgres, NAS storage | None required | Self, mobile app token |
| Frigate (planned) | 45 | Camera NVR | Cameras (RTSP, same VLAN only) | None | Local |
| Prometheus / Grafana | 50 | Metrics | Scrapes Proxmox, OPNsense, SNMP, most services | Community dashboard imports | Authentik (forward-auth) |
| Uptime Kuma | 50 | Uptime monitoring | Pings every monitored service | ntfy webhook | Local + Authentik |
| Wazuh | 50 | SIEM / logging | Agents on every host, OPNsense syslog, Authentik events | Threat intel feeds (optional) | Local + Authentik |
| ntfy | 50 | Notifications | Receives from Uptime Kuma, Wazuh, NUT, MT5 monitor | None required | Per-topic tokens, deny-all default |
| Beszel / Pulse | 50 | Host monitoring | Agents on Proxmox, LXCs, Talos nodes | None | Local |
| Technitium ×2 + Unbound | 60 | DNS | Answers every VLAN's queries; Technitium→Unbound→root servers | Root/TLD servers (Unbound only) | Local (admin) |
| NetBird | 70 | Remote access (primary) | Enrolled peers directly (mesh) | Coordination server only | Self (SSO/OIDC) |
| PostgreSQL / MariaDB | 80 | Shared databases | Nextcloud, Authentik, Immich, Grafana (as configured) | None | Per-service DB credentials |
| Home Assistant | 90 | Home automation | Mosquitto, Homebridge, IoT devices | Integration cloud APIs (optional) | Local + Authentik |
| Mosquitto | 90 | MQTT broker | Home Assistant, Zigbee2MQTT/ESPHome devices | None | Username/password |
| Homebridge | 90 | HomeKit bridge | Individual smart-home device APIs | Device cloud APIs (as needed) | Apple Home pairing code |
| Gitea | 90 | Git hosting | Argo CD (pulls manifests) | None required | Local + Authentik (OIDC) |
| Paperless-ngx | 90 | Document management | Own Postgres | None required | Local + Authentik |
| NAS shares | 999 | Bulk storage | Proxmox, PBS, Jellyfin, Immich (all via NFS) | None — fully isolated | Per-service scoped accounts |
| JGS524E + household devices | 100 | Household internet, ad-blocking, firewall inspection | None internal — Technitium and WAN only | Full outbound (WAN) | Switch: local admin password. Devices: none required. |
07Configuration steps
This section is written to teach, not just instruct — every step has a difficulty rating, a time estimate, a "why" explanation of the underlying concept, a numbered walkthrough, a checklist to confirm before moving on, and links to the real official docs for when you want more depth than fits here.
Virtualization strategy — how Proxmox, LXC, Docker, and Kubernetes fit together
Every deployment step in this section specifies one of four execution environments. This is the decision framework behind that choice, stated once rather than re-argued at every step.
| Layer | Used for | Reason |
|---|---|---|
| LXC (on Proxmox) | Single-purpose Linux services not tied to a Docker-specific ecosystem — Technitium, NPM, Homarr, Uptime Kuma, Vaultwarden | Near-native performance, no second kernel, trivial PBS snapshot/backup, lowest resource cost per service |
| Docker (inside a VM or LXC) | Anything distributed primarily as a container image or Compose stack, and multi-container applications — the *arr stack, Immich, Nextcloud+database, Authentik | Matches how the upstream project actually ships and documents itself; fighting an upstream Docker Compose file into a native LXC install usually costs more effort than it saves |
| VM (on Proxmox) | Anything needing its own full kernel or OS — OPNsense (FreeBSD), the MT5 Windows VM, each Talos node | Strong isolation and a genuinely different OS are the only two reasons to pay a VM's overhead over an LXC in this design |
| Kubernetes (Talos) | Automated scheduling, self-healing, and GitOps-managed rollout once manually tracking individual Compose stacks stops scaling | An orchestration layer above Docker, not a replacement for it — see Section 09 |
Proxmox is the layer underneath all four, and the reason to run it rather than any single layer alone: one backup system (PBS) covering VMs and LXCs alike, snapshots before any risky change, a live-migration path to the second node already planned in Section 11, and one place to see everything running rather than four disconnected tools.
What runs where — the complete breakdown
Every service named in the master table (Section 06) falls into exactly one of three categories. The deciding question for each is the same one asked in the table above, applied per service rather than left as a general principle: does the upstream project's own documentation assume Docker, does it run fine as a native package with no real benefit from containerising it, or does it need a genuinely separate OS?
Docker host operating system
Every Docker host in this design — mgmt01, media01, idx01, cloud01, photos01, ai01, test01 — runs the same base: Debian 13 "Trixie", the same OS already established for every Docker/LXC host in Section 06's version table. This is a deliberate default, not an arbitrary one:
- Matches Proxmox's own base OS — one set of
apthabits, one patch cadence, no second Linux distribution's quirks to remember - Minimal install footprint, with Docker's official
aptrepository well-supported and documented for Debian specifically - glibc-based — some upstream images and their bundled binaries assume glibc and behave unpredictably or fail outright on musl-based distributions such as Alpine used as a host OS (Alpine as a container base image is a different, unrelated question and is fine)
- Long, predictable support lifecycle — consistent with the patch-cadence expectations already set in Section 08
A dedicated container-optimised OS (Flatcar, Talos) is deliberately not used for these general Docker hosts, even though both already exist elsewhere in this design. Flatcar and Talos are immutable and API-managed by design — a good fit for a Kubernetes node that should never be hand-configured, and a poor fit for a host that still needs occasional native package installs, SSH-based troubleshooting, and ad hoc debugging the way these Docker hosts do.
How Docker actually runs inside an LXC
Docker needs kernel features an LXC does not expose by default. Getting this specific piece wrong is one of the most common home-lab stumbling blocks, so it is stated explicitly rather than assumed:
- Create the LXC as unprivileged — not privileged. Older guides frequently claim Docker needs a privileged container; current Proxmox versions do not require this, and running unprivileged keeps a container escape from directly compromising the Proxmox host, consistent with the least-privilege principle in Section 08.
- Enable the nesting feature on the LXC (
Options → Features → nesting: 1) — this is the specific setting that allows a container runtime to run inside the container at all. Without it, the Docker daemon fails to start. - Enable keyctl alongside nesting — several images rely on kernel keyring operations that are otherwise blocked in an unprivileged container.
- Use the overlay2 storage driver (Docker's default on modern kernels) — it works correctly inside an unprivileged nested LXC on the kernel versions Proxmox 9.x ships, with no special configuration required.
Once those three settings are in place, Docker installs and behaves inside the LXC exactly as it would on bare metal — the Compose file, the networking model in Section 10, and the image-trust checklist above all apply unchanged.
1 · Physical connection & first boot
- SFP+ 1 → Switch SFP+ 1 (10GbE, primary trunk). Optionally also connect em0 → Switch port 1 (1GbE fallback) — not required to boot, but handy from day one.
- em1 (WAN) → home router, using whatever cable is already running to your router.
- Power via PDU outlet, label
PDU-3-OPN. - Press power. Wait 30–60 seconds for boot — you'll see console text if a monitor is attached, or just wait if not.
- Default login is
root / opnsense. Log in once, then immediately go set a real password — don't leave this step for later.
- Power LED solid, no error beeps
- em1 shows link (WAN cable seated)
- Logged in once with default credentials and changed the password
2 · Identify and assign the SFP+ interface
igb0, em2...) based on hardware detection order, not what's printed on the case. You have to confirm which name maps to which physical port before you can safely configure it.- SSH in:
ssh admin@10.0.1.1(or use the console if SSH isn't enabled yet). - Run
ip link showand look for the interface reporting a 10000 Mbps link once the SFP+ cable is seated — that's your target. - In the web UI: System → Interfaces → Assignments → add that identified port as LAN.
ssh admin@10.0.1.1 ip link show # look for the interface reporting 10000 Mbps
- Correct 10GbE interface identified (not guessed)
- Assigned as LAN in System → Interfaces → Assignments
3 · Configure LAN (SFP+) as the VLAN trunk parent
10.0.1.1 — if that address ever changed, everything downstream would break.- Static IPv4:
10.0.1.1, prefix/24. - Enable DHCP on this interface: range
10.0.1.100–200. - Leave
em0configured too, as a secondary/fallback VLAN parent — this is what you manually switch to if the 10GbE link ever fails.
- LAN shows 10.0.1.1/24, interface up
- Can ping 10.0.1.1 from an admin PC
- https://10.0.1.1 loads the web UI
4 · Create the 15 VLAN interfaces
- Go to Interfaces → Other Types → VLAN, click add.
- Parent interface: LAN (the SFP+ trunk from Step 3). VLAN tag: the VLAN number (10, 15, 20…). Description: the VLAN's name.
- Go to Interfaces → Assignments, add the new VLAN as a real interface.
- Open that interface, set static
10.0.X.1/24, enable DHCP10.0.X.100–200. - Repeat for all 15: 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 999.
- All 15 VLAN interfaces created and visible under Interfaces → Assignments
- All show GREEN (up) status
- Each has the correct
10.0.X.1gateway and DHCP range
5 · Firewall rules — default deny, then explicit allows
Rule order per VLAN (first match wins, catch-all deny always last):
10 Management → ANY (allow) 20 Media → 30 Cloud, 80 DB, WAN (allow) 30 Cloud → 80 DB, WAN (allow) 35 Trading → 60 DNS, WAN (broker) (allow — outbound only, no lateral VLAN access) 40 Photos → 40 self, 60 DNS only (allow) 45 Cameras → nothing, not even DNS (isolated — Frigate reaches it via a dedicated NIC, not routing) 50 Monitor → ANY (allow) 60 DNS → ANY (allow, UDP/53) 70 VPN → ANY (allow) 80 Database → 60 DNS only (allow) 90 Testing → ANY (allow) 999 NAS → nothing outbound (isolated) * → * (BLOCK — catch-all)
Add outbound NAT for internet-needing VLANs (15, 20, 30, 35, 50, 70, 90) under Firewall → NAT → Outbound. VLAN 35 (Trading) should be restricted to just the broker's IP ranges/hostnames if your firewall supports FQDN aliases — not a blanket internet allow.
- Allow rules created per VLAN, in the order above
- Catch-all BLOCK rule sits at the bottom of every VLAN's rule list
- Outbound NAT configured for 20, 30, 50, 70, 90
6 · Verify & harden
- Ping each VLAN gateway (
10.0.X.1) and confirm DHCP hands out an address on each. - Test one allowed path (e.g. 20→30) and one blocked path (e.g. 40→internet) against the matrix in Section 05.
- Change the default password if you haven't, disable Telnet, force HTTPS-only, enable logging.
- Back up config: System → Configuration → Backup & Restore → download XML, store it off-box (in the GitHub repo's
/archiveor a password manager, not just on the NAS).
- All 14 gateways reachable, DHCP confirmed on at least 3 spot-checked VLANs
- One allow + one block tested and behaved as expected
- Config XML backed up off-box
6b · Perimeter hardening — UPnP, outbound blocklists, certificate exposure
- Disable UPnP and NAT-PMP on OPNsense entirely. Both protocols allow any device on the network to open an inbound port mapping without administrator approval, and were the root cause of the DeadBolt ransomware campaign against roughly 19,000 QNAP NAS devices via auto-forwarded UPnP mappings.
- Add an outbound firewall rule referencing a maintained threat-intelligence block list (Spamhaus DROP is a standard choice) as a destination alias. This does not affect legitimate traffic; its function is to prevent a compromised host from reaching a known command-and-control address if one is ever compromised regardless of the inbound protections in place.
- Optionally, enable GeoIP filtering on the WAN interface. This reduces log noise from high-volume scanning sources and is a genuine additional layer, though it does not stop a determined attacker routing through infrastructure inside an allowed country.
The wildcard certificate configured in Section 07 → Group B → Proxy & dashboard has a secondary benefit here: a wildcard entry in Certificate Transparency logs discloses that *.yourdomain.com exists, but does not enumerate individual subdomains the way a separate certificate per service would. This is a reason to prefer one wildcard certificate over per-service certificates beyond the operational convenience already described there.
- UPnP and NAT-PMP disabled
- Outbound block list rule active, referencing a maintained threat-intelligence feed
- GeoIP filtering considered and a decision made, whether enabled or not (optional)
1 · Enable and verify the SFP+ ports
ssh admin@10.0.1.5 (or console) /interface/ethernet set sfp-sfpplus1 disabled=no /interface/ethernet set sfp-sfpplus2 disabled=no /interface ethernet print # confirm both show speed=10Gbps once cables are seated
- Both SFP+ ports enabled
- Both show 10Gbps once DAC cables are seated on both ends
2 · Add SFP+ ports to the VLAN trunk bridge
/interface/bridge/port/add bridge=vlan-trunk interface=sfp-sfpplus1 /interface/bridge/port/add bridge=vlan-trunk interface=sfp-sfpplus2 /interface/bridge/vlan/print # both SFP+ ports should show all 15 VLAN IDs tagged
- Both SFP+ ports added to the trunk bridge
-
bridge/vlan/printshows all 15 VLAN IDs tagged on both
3 · Configure fallback, storage, and access ports
ether1— trunk, all 15 VLANs, → OPNsense em0 (1GbE fallback)ether3— trunk, all 15 VLANs, → Proxmox eth0 (1GbE fallback)ether4— access, untagged VLAN 30, → NAS NIC1ether5— access, untagged VLAN 10, → NAS NIC2ether2— access, untagged VLAN 10, → admin laptop drop (leave patched through to wherever you'll sit)ether6— access, untagged VLAN 10, → NanoKVM
- ether1 & ether3 confirmed trunk mode
- ether4/5/2/6 confirmed access mode with correct untagged VLAN
3b · Retire VLAN 1 — don't let it become the silent default
- Remove VLAN 1 from the allowed-VLAN list on both trunk ports (SFP+1, SFP+2, ether1, ether3).
- Set the switch's own management interface to VLAN 10 explicitly (already done in Step 3) rather than leaving it on VLAN 1 by default.
- Where the platform allows it, set an explicit native VLAN on trunk ports rather than leaving VLAN 1 as the implicit native — or disable native/untagged VLAN handling on trunks entirely if RouterOS supports it.
- VLAN 1 removed from both trunk ports' allowed lists
- A misconfigured/factory-reset port would now fail to pass traffic rather than silently landing on a working default
4 · Verify link status
/interface ethernet print # sfp-sfpplus1, sfp-sfpplus2, ether1-6 # should all show "R" (running)
- All expected ports show running ("R")
- Unused ports (7, 8) left disabled or unassigned
1 · Install Proxmox VE
- Flash the Proxmox VE ISO (version 9.2, built on Debian 13 "Trixie," current as of this writing — confirm the latest point release on the official downloads page before installing) to a USB stick.
- Boot the MS-01 from USB, install to the 2TB NVMe drive.
- First login:
https://<ip>:8006in a browser, accept the self-signed certificate warning.
- Proxmox installed, web UI reachable on port 8006
1b · PVE post-install script — repository and update configuration
apt update. The post-install script performs the standard set of first-run configuration changes — switching to the no-subscription repository, correcting the same issue for the Ceph repository if present, and optionally disabling the subscription nag dialog in the web UI — as a single reviewed, community-maintained script rather than manual editing of sources.list files.- Review the script's contents before running it — this applies to any script executed with root privileges from a third-party source, not a specific concern about this one.
- Run it from the Proxmox host shell:
bash -c "$(wget -qLO - https://raw.githubusercontent.com/community-scripts/ProxmoxVE/main/tools/pve/post-pve-install.sh)" - Select the no-subscription repository when prompted; the enterprise repository prompt can be declined unless a paid subscription is in use.
- No-subscription repository active,
apt updateruns without repository errors
1c · Proxmox VE Helper-Scripts — purpose and use throughout this build
The trust model is the same as any script executed with root privileges from a source not personally authored: the community-scripts organisation is the maintained successor to the original tteck repository, scripts are open source and reviewable before execution, and the general precaution in Step 1b — read a script before running it — applies uniformly across every application deployed this way in this document, not as a repeated caveat per service.
- Helper-Scripts site bookmarked as the first check before manually scripting a new LXC deployment
2 · Bridge the SFP+ trunk interface
auto enp2s0f0
iface enp2s0f0 inet manual
auto vmbr0
iface vmbr0 inet manual
bridge-ports enp2s0f0
bridge-stp off
bridge-fd 0
bridge-vlan-aware yes
bridge-vids 10,15,20,25,30,35,40,45,50,60,70,80,90,100,999
Apply with systemctl restart networking. Keep eth0 (the 1GbE fallback NIC) as a separate manual bridge you're not actively using, ready if the 10GbE link ever needs replacing.
- vmbr0 created, VLAN-aware, all 15 VLAN IDs listed
- Networking restarted without errors, SSH/web UI still reachable
3 · Tag each VM's NIC to its VLAN
Single VLAN-aware vmbr0, tag per VM in Hardware → Network Device → VLAN Tag. Jellyfin VM → 20, Nextcloud → 30, Postgres → 80, and so on per the VLAN table in Section 05.
vmbr0 itself or configure the switch's SFP+2 port as an access port for a specific VLAN — it's a trunk carrying all 14 tagged VLANs, and the switch port config in the MikroTik tab already reflects that. If a VM ever seems to land on the wrong network or nothing responds, check for exactly this: a tag applied at both the bridge and the VM level.
- Every VM's NIC has an explicit VLAN tag set (never left blank/untagged)
- vmbr0 itself carries no VLAN tag of its own — only individual VM NICs are tagged
4 · Mount NAS storage over NFS
Prerequisite: the NAS (Group A → Tab 3) must already be configured with its shares created before this step — this is why NAS precedes Proxmox in the tab order.
Datacenter → Storage → Add → NFS: server = NAS IP on VLAN 30, export = the share path from the NAS tab, content = Disk image / ISO / backup as needed.
- NFS storage added and shows green/active in Datacenter → Storage
- Throughput spot-checked (>100MB/s target on the 1GbE NAS link)
5 · First test VM
- Create a small VM (any lightweight Linux ISO works), tag its NIC VLAN 90 (Testing).
- Confirm it gets a DHCP lease from
10.0.90.100–200, and can ping its gateway10.0.90.1. - Confirm reachability matches the traffic matrix in Section 05 — it should reach the NAS if a rule allows it, and should not reach VLAN 40 (Photos).
- Test VM gets correct DHCP lease and gateway
- Allowed path confirmed, blocked path confirmed blocked
1 · Operating system — TOS, and how it compares to TrueNAS
This design uses the Terramaster appliance because the hardware was already owned at the point the network design began, and the appliance model trades some of ZFS's data-integrity guarantees for a simpler, faster initial setup. A TrueNAS deployment would require different hardware — a generic x86 host with sufficient SATA/NVMe connectivity, no Terramaster-specific components — and is noted here as an alternative rather than a planned change.
- Confirmed: this device runs TOS, and TrueNAS-specific instructions found elsewhere do not apply directly
2 · Physical install & initial setup
- Install all four drives into the bays.
- NIC1 → switch port 4 (VLAN 30), NIC2 → switch port 5 (VLAN 10).
- Complete initial setup via the management VLAN, using TerraMaster's discovery tool (TNAS PC/Mac, or the web-based discovery page) to locate the device's assigned address.
- Apply any pending TOS firmware update before creating the storage pool — a fresh appliance from stock may be several point releases behind.
- All four drives detected in the storage manager
- Both NICs show link, on the correct VLANs
- TOS updated to the current version before proceeding
3 · Configure RAID 5
- Storage Manager → create a RAID 5 array across the four bays. Approximately 4TB usable after parity overhead.
- Enable S.M.A.R.T. monitoring and a weekly scheduled scan.
- Enable TOS's own scheduled RAID scrub/data-integrity check if the model supports it, as a partial substitute for ZFS's continuous checksumming.
- RAID 5 array created, initial resync started
- S.M.A.R.T. monitoring enabled with a weekly scan scheduled
- Scheduled scrub/integrity check enabled if available
4 · Shares, NFS export, permissions
- Create shares:
/storage,/backups,/media. - Enable NFS, with the export scoped to the relevant VLAN subnets specifically — not an unrestricted "any" export.
- Create per-service scoped accounts rather than a single shared administrative login for every service that accesses the NAS.
- Shares created and exported
- NFS export restricted by subnet, not open to "any"
- Scoped accounts created; no shared administrative login in use by any service
5 · Snapshot schedule — TOS's native snapshot feature
- Enable scheduled snapshots on the shares created in Step 4 — a daily snapshot with a one- to two-week retention window is a reasonable starting point.
- Confirm the snapshot schedule does not depend on the RAID array being otherwise idle, and that it does not conflict with backup job timing (Section 07 → Group B → Backups).
- Scheduled snapshots enabled on primary shares
- Retention window set and confirmed not to fill available capacity
6 · Verify dual-NIC reachability
From Proxmox: ping 10.0.30.10 and ping 10.0.1.100. Disconnect each cable in turn to confirm the NAS remains reachable on the other NIC before relying on this configuration for redundancy.
- Both NICs individually confirmed reachable
1 · Physical connections
- HDMI-in ← Proxmox HDMI-out.
- USB-C (device emulation) ← Proxmox USB port.
- Ethernet → switch port 6 (VLAN 10, access, untagged).
- Aux power USB-C ← a spare USB port on Proxmox, or a small USB power adapter into a PDU-adjacent outlet if you'd rather it not depend on Proxmox's own power state.
- HDMI capture confirmed (see the Proxmox boot screen through NanoKVM)
- USB emulation confirmed (mouse cursor moves on the host from NanoKVM's web UI)
- Status LEDs normal, no error pattern
2 · First boot & firmware
- Power it on, wait for the status LED to settle.
- Find its IP (check your router's DHCP client list, or use the discovery method in the quick start guide).
- Log into the web UI, go to Settings → Check for Updates, apply if available.
- NanoKVM web UI reachable
- Firmware confirmed up to date
3 · Network configuration — VLAN 10, static reservation
- Confirm it pulled an address in the
10.0.1.100–200range from OPNsense. - In OPNsense, add a DHCP static mapping (by MAC address) so it always gets the same IP.
- Static DHCP reservation configured on OPNsense
- Address noted somewhere you'll actually find it during an outage (Homarr, a sticky note, wherever — just not "I'll remember")
4 · Web UI security — password, HTTPS, SSH
- Change the default password immediately.
- Enable HTTPS on the NanoKVM web UI if not already on.
- Disable SSH unless you specifically need it (Settings → Devices → SSH, per the user guide).
- Default password changed
- HTTPS enabled
- SSH disabled unless actively needed
5 · ATX power control (optional, hardware-dependent)
- Check whether the MS-01 exposes an accessible power/reset header (open the case, check Minisforum documentation, or search the model number plus "ATX header").
- If yes: wire the NanoKVM-B breakout board per Sipeed's wiring diagram.
- If no: skip this — you still have full video + keyboard/mouse. For hard power-cycling, a UPS-controlled smart outlet is the practical fallback.
- Verified whether MS-01 exposes ATX header pins before buying/wiring anything
- If wired: power on/off/reset tested remotely at least once
6 · Lock it down — NetBird-only access
- Confirm no port-forward exists on OPNsense exposing NanoKVM's port to WAN.
- Once NetBird is deployed (Services tab), scope access so only your "admin" peer group can reach VLAN 10 where NanoKVM lives.
- No WAN port-forward to NanoKVM exists
- Reachable only via NetBird admin group once that's deployed
7 · Set a BIOS password on Proxmox itself
- Set a BIOS/UEFI admin password on the MS-01, stored in Vaultwarden like every other credential.
- While in there, disable booting from any external/USB media by default, and disable any unused Thunderbolt/USB controllers not actually in use.
- BIOS password set and stored in Vaultwarden
- External boot disabled by default, unused USB/Thunderbolt controllers disabled
1 · Initial discovery and firmware check
- Connect the switch to power and to a temporary access-mode port on the MikroTik switch (not port 7, which is reserved for the trunk configured in Step 2).
- Run the ProSAFE Plus Configuration Utility from a Windows PC on the same network segment, or check the switch's DHCP lease on OPNsense if it obtained one automatically.
- Log in with the default password (
password) and change it immediately — this switch does not support the account-separation model used elsewhere in this design, so the single admin password is the only credential protecting it. - Check for and apply any available firmware update before proceeding; this hardware model dates to 2012 and a specific unit's firmware history is unknown until checked.
- Switch located and reachable via the configuration utility or web interface
- Default password changed
- Firmware checked and current
2 · VLAN configuration — 802.1Q tagging
- Create VLAN 10 (Management) and VLAN 100 (Home Devices) in the switch's VLAN configuration screen.
- Set the uplink port (the port that will connect to the MikroTik switch) to tag both VLAN 10 and VLAN 100.
- Set every other port intended for a household device to VLAN 100, untagged (PVID 100).
- Leave the switch's own management VLAN set to VLAN 10, matching every other managed device in this design.
- VLAN 10 and VLAN 100 created
- Uplink port tagging both VLANs confirmed
- Device-facing ports set to untagged VLAN 100
3 · Uplink to the MikroTik switch
- Move the uplink cable from the temporary access port used in Step 1 to MikroTik port 7.
- On the MikroTik switch, confirm port 7 is configured as a trunk carrying VLAN 10 and VLAN 100 only — not every VLAN, since this switch has no reason to see lab traffic.
- Confirm the switch's management IP (VLAN 10) is reachable from an admin device, and that a test device plugged into a VLAN 100 port receives a DHCP lease from OPNsense in the
10.0.100.0/24range.
- Port 7 confirmed as a two-VLAN trunk, not an all-VLANs trunk
- Switch management IP reachable on VLAN 10
- Test device on VLAN 100 receives a DHCP lease and reaches the internet
3b · Reserve one port for on-site troubleshooting access
- Set one spare port (port 24, the last physical port) to untagged VLAN 10, separate from the VLAN 100 access ports used for household devices.
- Label this port physically and in the switch's own port naming, since a household device plugged in here by mistake would receive a management-network address instead of an internet-only one.
- Confirm a laptop plugged into port 24 receives a VLAN 10 address from OPNsense and can reach lab services — the same reach as the standing Management-VLAN exception already applied everywhere else in this design, not a new or broader rule.
- Port 24 set to untagged VLAN 10, clearly labelled
- Confirmed a device on this port receives a VLAN 10 address and reaches lab services
- Confirmed the remaining household-device ports are unaffected — still VLAN 100 only
4 · OPNsense — VLAN 100 firewall rules
- Create the VLAN 100 interface on OPNsense (
10.0.100.1/24), following the same procedure as every other VLAN in Section 07 → OPNsense, Step 4. - Allow rule: VLAN 100 → Technitium (port 53) and → WAN, matching the DNS-enforcement pattern in Section 07 → Services → Monitoring & DNS.
- Explicit deny: VLAN 100 → every other VLAN. No exception rules — this VLAN has no legitimate reason to reach lab services.
- Confirm Suricata and CrowdSec (Section 07 → Group C → Security & identity), already inspecting all inter-VLAN and WAN-bound traffic, cover this VLAN without additional configuration — they operate on the trunk, not per-VLAN.
- VLAN 100 interface created, DHCP enabled
- Outbound allowed only to Technitium and WAN
- Explicit deny confirmed against every lab VLAN, tested from a device on VLAN 100
5 · Optional — outbound VPN egress for VLAN 100
- Configure a WireGuard client instance on OPNsense connecting to a chosen VPN provider's endpoint.
- Create a gateway using that WireGuard interface, and a policy-based routing rule scoped to VLAN 100 only, directing its traffic through that gateway instead of the default WAN gateway.
- Verify with an IP-address check site from a VLAN 100 device that egress traffic shows the VPN provider's address, and confirm DNS continues resolving through Technitium rather than leaking to the VPN provider's own resolver.
- Not configured by default — optional, add only if outbound privacy routing for household devices specifically is wanted
- If configured: egress IP and DNS resolution both verified from a VLAN 100 device
1 · Deploy the self-hosted NetBird server
This also sidesteps the classic reason people reach for a hosted mesh tool in the first place: a bare WireGuard server normally needs an inbound port forwarded on the router, which doesn't work at all behind carrier-grade NAT (CGNAT). NetBird's peers connect outbound to the coordination server and punch through NAT from there — no inbound port needed, which is also exactly why this build has zero port-forwards (Section 08).
Twingate, for context: a legitimate free-for-small-teams zero-trust option, but its control plane is Twingate's cloud — same category concern as Entra ID in Section 11. NetBird's self-hosted coordination server keeps that layer under this lab's own control instead, which is the deciding factor here, not a feature gap in Twingate itself.
- Create an LXC/VM on VLAN 70 (VPN).
- Follow NetBird's Docker Compose self-hosted quickstart to stand up the management server, signal server, and dashboard.
- Point a DNS name at it via NGINX Proxy Manager once that's deployed (Tab B), or access by IP for now.
- NetBird management dashboard reachable and logged in
2 · Enrol peers
- OPNsense: install via the OPNsense NetBird plugin (
os-netbird) — a native FreeBSD package, not Docker. - Docker-hosted peers (any Docker host on VLAN 15, e.g. the Talos/AI/media hosts): run NetBird as a container using the
netbirdio/netbirdimage withnetwork_mode: host— this lets it create the WireGuard interface (wt0) directly on the host's network stack rather than isolated inside Docker's bridge network. Mount/etc/netbirdand/var/lib/netbirdas volumes so the peer's identity survives a container restart. - Admin laptop/phone: native NetBird client from the official download page.
- Authenticate each against the self-hosted dashboard, confirm each shows "Connected" in the peers list.
# Docker-hosted peer example docker run -d --name netbird \ --network host \ --cap-add NET_ADMIN \ -v /etc/netbird:/etc/netbird \ -v /var/lib/netbird:/var/lib/netbird \ --restart unless-stopped \ netbirdio/netbird
- OPNsense enrolled and connected (via the OPNsense plugin)
- Proxmox / Docker hosts enrolled and connected (via the Docker container,
wt0interface confirmed withip a) - Admin devices enrolled and connected
3 · Access control groups
- Create an
admingroup containing your own admin devices, with access rules to VLAN 10 + VLAN 90. - Create a second, narrower
personal-devicesgroup for your phone (and any future phone/tablet), with access rules to VLAN 30 (Cloud) + VLAN 40 (Photos) only — everything it needs for Nextcloud, Vaultwarden, and Immich backup, nothing it doesn't. - Leave other VLANs unreachable via NetBird by default — add exceptions deliberately, not by accident.
- Admin group created, scoped to management + testing only
-
personal-devicesgroup created, scoped to Cloud + Photos only - No blanket "allow all" access rule left in place
4 · SSO / 2FA on the dashboard
Configure an OIDC identity provider (many free options work fine for a single-user homelab) instead of relying on local username/password login alone.
- SSO/OIDC or 2FA enabled on the NetBird dashboard login
5 · Cloudflare Tunnel — the backup path
cloudflared container on VLAN 10 makes an outbound-only connection to Cloudflare's edge network. There's nothing to port-forward and no inbound rule needed on OPNsense. Because it's architecturally unrelated to NetBird, an outage or misconfiguration in one doesn't take down the other — that's the entire point of having two paths.- Register a domain (or use one you already own) and add it to a free Cloudflare account.
- Deploy
cloudflaredas a container on VLAN 10, authenticate it to your Cloudflare account. - Create a tunnel and add public hostnames that map to internal services through NPM/Traefik, e.g.
jellyfin.yourdomain.com → 10.0.10.x:443.
- cloudflared deployed on VLAN 10, tunnel shows healthy in the Cloudflare dashboard
- At least one service reachable through the tunnel from outside the home network
6 · Cloudflare Access — Zero Trust policies
In the Cloudflare Zero Trust dashboard, create an Access application per exposed hostname, require authentication (email OTP is the simplest starting point), and scope each policy to just yourself unless you specifically intend to share access.
- Access application created for each exposed hostname
- Policy tested — an unauthenticated request is challenged, not served directly
6b · Geo-restrict Access and lock out repeated failures
- In each Access policy, add a country rule scoped to Australia (adjust temporarily if travelling).
- Enable Cloudflare's rate limiting / bot protection on the Access login page.
- Apply the same logic to NetBird and OPNsense's own admin login — fail2ban (or OPNsense's built-in lockout) on repeated auth failures, not just on the Cloudflare side.
- Country-restricted Access policy in place
- Rate limiting / auto-block confirmed active on Cloudflare Access, NetBird, and OPNsense admin login
7 · Optional: highly-available tunnels (later)
A single cloudflared replica is a fine backup path for a home lab. If this ever needs to be genuinely highly available, multiple cloudflared replicas behind Docker Swarm is a documented pattern — worth knowing exists, not worth building until the single-replica version has actually let you down.
1 · Deploy NGINX Proxy Manager
10.0.20.10:8096 for Jellyfin, 10.0.30.20:80 for Nextcloud...). A reverse proxy sits in front of all of them and lets you instead type jellyfin.lab.internal — it looks up which backend that name maps to and forwards the request there, transparently, with one HTTPS certificate covering everything instead of managing certs per-service.- Deploy as an LXC on VLAN 10.
- Only forward ports 80/443 to it from OPNsense for whatever you actually intend to expose externally — internal-only services don't need a WAN forward at all.
- NPM deployed, admin UI reachable
2 · Point subdomains at internal services, one wildcard cert for everything
*.yourdomain.com wildcard certificate, issued once via DNS challenge, covers every subdomain you add later — no per-service certificate requests, no renewal juggling across a dozen services.- Add a proxy host per service: domain name → internal IP:port.
- Issue one wildcard certificate (
*.yourdomain.com) via Let's Encrypt DNS challenge, and reuse it across every proxy host in NPM — this is the same domain and certificate strategy used for NetBird/Cloudflare Access below, so the whole lab shares one consistent chain of trust. - Confirm auto-renewal is actually configured, not a one-off manual issue.
- At least one service reachable via a clean HTTPS name, no port in the URL
- Wildcard certificate issued once and reused across services, renewing automatically
3 · Deploy Homarr dashboard
Deploy as an LXC/container, add links to every service, set as your browser home page.
- Homarr deployed with at least the core services linked
Alternative worth knowing: Homepage is configured almost exactly like Traefik — YAML/labels, not a GUI — which makes it genuinely simpler than Homarr if every service is already a pure Docker Compose stack. Homarr's GUI-first setup is why it's the default here; swap later if label-driven config starts feeling more natural than clicking.
4 · Wire up live status widgets
Connect Homarr's integrations to Proxmox, Jellyfin, and your download clients for live status/stats directly on the dashboard, not just static links.
- At least one live-status widget working (e.g. Proxmox resource usage)
5 · Watchtower or Dockcheck — automatic container updates
apt upgrade habit covering them. Watchtower checks for new images on a schedule and updates running containers automatically — pairs with the host-level patch cadence already covered in Section 07, so both halves of the stack stay current without manual tracking.- Deploy Watchtower (or Dockcheck) as a container with access to the Docker socket, scoped to check on a daily or weekly schedule.
- Exclude anything you'd rather update manually and review first (databases are a common exclusion — a surprise major-version bump is a worse outcome than a slightly stale image).
- Update checker deployed, schedule set
- Databases and anything version-sensitive explicitly excluded from auto-update
6 · NetBox — the actual source of truth for what's where
- Deploy via the official Docker Compose project on VLAN 10.
- Enter this build's actual inventory: the rack's devices, all 15 VLANs and their subnets, and current IP assignments — tedious once, but this is what turns it from a diagram into a live record.
- Retire the manual Section 09 table once NetBox reflects the same information — don't maintain both.
- NetBox deployed, admin account created
- Rack, devices, VLANs, and current IPs entered
1 · Deploy Proxmox Backup Server
Start as a VM/LXC on the same Proxmox host, its own storage target on the NAS. Upgrade to a dedicated small box later if you add a second hypervisor node.
- PBS deployed, storage datastore pointed at NAS
2 · Add PBS as a Proxmox backup target
Datacenter → Storage → Add → Proxmox Backup Server: enter the PBS host, datastore name, and credentials.
- PBS storage visible and green in Proxmox's Datacenter view
3 · Schedule jobs
Daily incremental backups, weekly full verification job. Start with everything except purely-disposable test VMs.
- Daily backup job scheduled and has run successfully at least once
- Weekly verify job scheduled
4 · Test a real restore
Restore a test VM to a new VM ID, boot it, confirm it actually works — not just that the restore operation completed without an error.
- Restore performed and the restored VM actually boots and functions
5 · Rclone — the offsite leg of 3-2-1
- Pick a destination: Backblaze B2 and Wasabi are the usual home-lab picks — S3-compatible, no egress fees on B2 up to your download volume, cheap per-GB storage.
- Deploy Rclone as a scheduled job (cron on a small LXC, or a Proxmox scheduled task) rather than installing it ad hoc — this needs to run unattended.
- Sync PBS's backup datastore (or just the most critical VMs/datasets, if full replication is more than the connection or budget can handle) on a nightly or weekly schedule.
- Confirm the destination bucket is encrypted at rest and that Rclone's own encryption (
rclone crypt) is layered on top if the destination provider itself isn't trusted with plaintext data.
- Offsite destination chosen and configured
- Rclone sync scheduled, not manual
- First sync completed and spot-checked — a file actually appears in the remote bucket
1 · Deploy Jellyfin
Deploy as a VM/LXC tagged VLAN 20, point its library at the NAS media share over the NFS mount from Proxmox (Hardware → Proxmox → Step 4).
- Jellyfin deployed, library scanning the NAS media share
2 · Deploy Prowlarr (indexers)
Deploy tagged VLAN 25, add your indexers, this becomes the single place Sonarr/Radarr pull indexer config from instead of configuring each app separately.
- Prowlarr deployed with at least one indexer configured
3 · Deploy Sonarr/Radarr, connect the pipeline
Both tagged VLAN 20. Connect each to Prowlarr for indexers, and to Jellyfin so newly-downloaded media appears in the library automatically.
- Sonarr connected to Prowlarr
- Radarr connected to Prowlarr
- Both connected to Jellyfin's library
4 · Deploy Bazarr (subtitles)
Deploy tagged VLAN 20, connect it to both Sonarr and Radarr as subtitle providers, pick your preferred subtitle languages/sources.
- Bazarr connected to Sonarr and Radarr
- At least one subtitle provider configured
5 · Deploy Jellyseerr (user requests)
Deploy tagged VLAN 20, connect it to Jellyfin (for library sync and login) and to Sonarr/Radarr (to actually place requests).
- Jellyseerr connected to Jellyfin, Sonarr, and Radarr
- One test request placed end-to-end and confirmed it reached Radarr/Sonarr
6 · Verify the full pipeline
Full chain per the standard automated-Jellyfin pattern: Jellyseerr request → Radarr/Sonarr → qBittorrent → Bazarr → Jellyfin. Trigger one real request end to end and confirm it lands correctly organised, subtitled, and playable — not just that individual apps look "connected" in their settings pages.
- One item requested via Jellyseerr, downloaded, subtitled, and confirmed playable in Jellyfin
1 · Deploy Prometheus + node exporters
Deploy on VLAN 50, install node_exporter on Proxmox and OPNsense (or use OPNsense's built-in Telegraf/Prometheus exporter plugin), configure Prometheus's scrape targets.
- Prometheus deployed, at least one target being scraped successfully
1b · Optional: SNMP on the switch, Pulse for Proxmox, Beszel for multi-host
ps and Pulse already cover that — but this build already spans Proxmox, several LXCs/VMs, and a Talos cluster with multiple nodes, which is exactly the multi-host case Beszel is built for: one lightweight view across every host instead of checking each individually.- Enable SNMP on the MikroTik CRS310, add the Prometheus SNMP exporter as a scrape target, alert on abnormal per-port traffic spikes.
- Deploy Pulse alongside Prometheus/Grafana for Proxmox-specific monitoring.
- Deploy Beszel's hub on VLAN 50, install its lightweight agent on Proxmox and each VM/LXC/Talos node worth watching.
- SNMP enabled on the switch, scraped by Prometheus (optional)
- Pulse deployed for Proxmox-specific views (optional)
- Beszel hub + agents deployed across multiple hosts (optional, most useful once the Talos cluster is live)
2 · Deploy Grafana, add data sources, import dashboards
Add Prometheus and InfluxDB as data sources, import a community Node Exporter dashboard (search grafana.com/dashboards by ID) rather than building one from scratch on day one.
- Grafana deployed, at least one imported dashboard showing live data
3 · Deploy Technitium DNS (clustered pair) + Unbound
- Deploy two Technitium instances on VLAN 60 as plain Docker containers or LXCs — not inside the Talos cluster (see the callout below for why that's deliberate).
- Give each a distinct
DNS_SERVER_DOMAINunder Settings → General before clustering — e.g.dns01anddns02— and make sure it's set correctly on both from the start; a stale or mismatched value here is the single most common cause of clustering looking "connected" one moment and "unreachable" the next. - On the primary: Administration → Cluster → Initialize → New Cluster. On the secondary: Administration → Cluster → Initialize → Join Cluster, pointing at the primary's cluster URL. Without a real certificate yet, check "ignore certificate validation errors" to join.
- Deploy Unbound (can sit alongside Technitium on the same host or its own small LXC), configure it as a pure recursive/validating resolver.
- Point Technitium's forwarder at Unbound instead of a public DNS provider.
- Set OPNsense's DHCP to hand out the Technitium primary as DNS server 1, secondary as DNS server 2.
Once this is solid, start giving services proper names instead of remembering IPs — grafana.lab.internal, ai.lab.internal, proxmox.lab.internal — via Technitium zones, matching the DNS-first organisation approach in Section 09.
- Distinct, correct
DNS_SERVER_DOMAINset on each node before clustering - Technitium primary + secondary clustered and confirmed syncing (refresh the page, not just the first "Connected" flash)
- Unbound deployed, Technitium forwarding to it
- DHCP handing out both Technitium instances, secondary tested by stopping the primary briefly
3b · Get the cluster domain right the first time
- The Cluster Domain must be a domain you don't already host authoritatively anywhere else. Don't reuse
lab.internal(already this build's general service-naming domain, per Section 09) — pick something separate and purely technical, e.g.dnscluster.local. - Never use
cluster.localspecifically — that's Kubernetes' own reserved internal service-discovery suffix. This build has a Talos cluster planned (Section 08); even though Technitium isn't running inside it, a colliding domain is exactly the kind of thing that causes confusing failures months later for no obvious reason. - Keep the cluster domain and this build's real internal naming domain (
lab.internal) conceptually separate forever — the cluster domain is plumbing between the two DNS nodes, not something you resolve services through.
0.0.0.0:53 rather than a MetalLB VIP, configure TLS passthrough (not termination) for port 53443 if anything like Traefik sits in front of cluster traffic, and add the Kubernetes pod CIDR to Technitium's zone-transfer allow list — the source IP the primary sees will be a pod address, not the node's external IP. None of this applies to the plain Docker/LXC deployment above; it's flagged here so it's not rediscovered the hard way if the architecture changes.
- Cluster domain confirmed distinct from both
lab.internalandcluster.local - Understood: replication is one-way, primary → secondary only — administrative changes must always be made on the primary node
- Understood: Technitium's own DHCP clustering isn't supported yet — not a gap for this build, since OPNsense handles DHCP per VLAN, not Technitium
4 · VLAN-wide DNS enforcement — make Technitium the only option
- OPNsense: Firewall → Aliases → create a Host(s) alias
Technitium_Serverscontaining both Technitium IPs (dns01, dns02). - Create a Port alias
DNS_Ports= 53 (standard DNS), and a second Port aliasDoH_DoT_Ports= 443, 853 (DNS-over-HTTPS / DNS-over-TLS) for use in the next step. - On every VLAN's rule set, add an explicit allow rule near the top: destination =
Technitium_Servers, port =DNS_Ports— this is what actually lets each VLAN reach DNS at all (already reflected in the traffic matrix in Section 05).
-
Technitium_Serversalias created with both IPs - Explicit allow-to-Technitium rule confirmed present on every VLAN
5 · Block third-party DNS & DoH bypass attempts
- On every VLAN except 10 (Management) and 70 (VPN), add a block rule below the Technitium-allow rule from Step 4: destination = any, port =
DNS_Ports(53) — this catches any device trying to reach a DNS server that isn't Technitium. - Add a second block rule: destination = WAN, port 853 (DNS-over-TLS) — there's rarely a legitimate reason for a LAN device to speak DoT directly to the internet.
- Create an alias
Known_Public_DoHwith the IP ranges of the most common public DoH providers (Cloudflare, Google, Quad9, NextDNS) and block outbound port 443 to that alias. This is necessarily incomplete — treat it as blocking the defaults, not every possible DoH provider. - Test from a client device: with a public resolver hardcoded, DNS lookups should now fail or time out, while normal browsing (through Technitium) keeps working.
- Port 53 to anywhere-but-Technitium blocked on every non-management/VPN VLAN
- Port 853 (DoT) blocked outbound to WAN
-
Known_Public_DoHalias created and blocked on port 443 - Tested: a hardcoded public resolver fails from a client device on a restricted VLAN
- Secure DNS / private DNS disabled in browser and OS settings on your own devices, as defence in depth
6 · Deploy Uptime Kuma
Add a monitor per key service (Jellyfin, Nextcloud, Technitium, NPM...), point its notification channel at ntfy (Step 7) rather than Discord — one self-hosted notification path for the whole lab, not a third-party service in the alert chain.
- Uptime Kuma deployed, core services monitored, at least one notification channel configured
7 · ntfy — one notification path for everything, no Discord required
curl command can push a notification straight to your phone, organised by topic (one for backups, one for security alerts, one for uptime), with no external dependency at all.- Deploy via the official Docker image on VLAN 50, alongside the rest of the monitoring stack.
- Set
NTFY_AUTH_DEFAULT_ACCESS=deny-alland create real users/tokens per topic — don't rely on an unguessable topic name as the only protection, which is a security-through-obscurity pattern this build avoids everywhere else (Section 07). - Put it behind NPM with the shared wildcard cert, like every other service —
ntfy.yourdomain.com, not plain HTTP. - Install the ntfy app (F-Droid or Google Play — GrapheneOS-friendly either way) and subscribe to your topics.
- Wire it in as the notification target for Uptime Kuma, Wazuh, the MT5 trading VM's fast-notification requirement (Section 07 → Group C → Trading), and the NUT UPS shutdown script — one app, every alert.
# quick test once it's running curl -H "Title: Test" -H "Priority: high" \ -d "ntfy is alive" \ https://ntfy.yourdomain.com/lab_alerts
- ntfy deployed with
deny-alldefault and real per-topic auth, not just an obscure topic name - Reachable over HTTPS through NPM with the shared wildcard cert
- App installed and subscribed, test notification received
- Uptime Kuma and Wazuh both pointed at it as their notification target
1 · Deploy Home Assistant
- Deploy via the community Proxmox VE Helper Script for Home Assistant OS (VM variant) — it automates what would otherwise be a manual ISO install and virtual disk setup.
- Tag VLAN 90 (Testing) for now, per the existing VLAN table — see the note below on a dedicated IoT VLAN if your device count grows.
- Give it real resources if you're a heavy user: 2+ cores, 4GB+ RAM. Light users can go smaller.
- If you have Zigbee/Z-Wave/Bluetooth USB dongles, pass them through to the VM individually (Proxmox → VM → Hardware → Add → USB Device) rather than sharing a USB controller.
- Home Assistant OS VM deployed via the helper script
- Onboarding wizard completed, admin account created
- Any USB radios (Zigbee/Z-Wave/Bluetooth) passed through and detected
2 · Multi-NIC for device discovery (optional)
Two options, in order of preference: (1) enable multicast/mDNS relay on OPNsense between the relevant VLANs, or (2) give the Home Assistant VM a second virtual NIC directly on the IoT VLAN if one exists, so discovery happens locally on that segment. Start with option 1 — it's less invasive.
- Not needed yet — VLAN 90 already reaches what HA currently manages (revisit only if a dedicated IoT VLAN is added)
3 · Deploy Homebridge (HomeKit bridge)
- Deploy via the community Proxmox VE Helper Script for Homebridge (LXC variant — it's lightweight, doesn't need VM-level USB passthrough for most plugins).
- Tag VLAN 90, same as Home Assistant.
- Install plugins for your specific devices through the Homebridge web UI's plugin search.
- Pair it to Apple Home by scanning the QR code / entering the setup code shown in the Homebridge UI.
- Homebridge LXC deployed
- At least one device plugin installed and configured
- Paired successfully in the Apple Home app
1 · NetBird on the phone — always-on connection
- Install NetBird — currently via Google Play or Aurora Store (Aurora works without a Google account). It isn't on F-Droid yet on GrapheneOS as of writing; there's an open community request to add it, worth checking again later.
- Open NetBird, enrol it against your self-hosted management server (same as the other peers in Section 07 → Group B → Remote access).
- Add it to a new, narrower NetBird access group —
personal-devices— scoped to just VLAN 30 (Cloud) and VLAN 40 (Photos), not the fulladmingroup's reach into VLAN 10/90. A phone doesn't need to manage OPNsense. - Android Settings → Network & Internet → VPN → NetBird (gear icon) → enable Always-on VPN.
- Settings → Apps → NetBird → Battery → set to Unrestricted — without this, Android's battery management will periodically kill the background connection regardless of the always-on setting.
- NetBird installed and enrolled
-
personal-devicesgroup created, scoped to VLAN 30 + 40 only - Always-on VPN enabled, battery optimisation disabled for the app
- Connection reliability observed over several days before treating it as fully dependable
2 · Immich — daily photo & video backup
- Install Immich from F-Droid (fully open-source build, GrapheneOS-friendly) or Google Play.
- Log in with the server URL — since the phone is always connected via NetBird, this can be the internal VLAN 40 address; it'll resolve whether you're home or out.
- Backup screen → select the Camera album (and Screenshots, Downloads, or anything else you want included).
- Enable background backup, and turn off the "Wi-Fi only" restriction if you want photos backing up over mobile data too, since the whole point is the phone is already tunnelled into the lab regardless of network.
- Settings → Apps → Immich → Battery → Unrestricted, same reasoning as NetBird.
- Immich installed, camera album selected for backup
- Background backup enabled, battery optimisation disabled
- Confirmed a photo taken today actually appears on the server without opening the app first — then repeat that check after a few days
2b · Deploy Nextcloud itself — the server side, plus Collabora
- Deploy Nextcloud (AIO or the official Docker image) on VLAN 30, backed by the PostgreSQL instance on VLAN 80.
- Deploy Collabora Online alongside it (its own container — it's a separate service, not a Nextcloud plugin) and connect it via the Nextcloud admin panel's Collabora app.
- Confirm a document actually opens and saves for in-browser editing, not just file storage.
- Nextcloud deployed, backed by PostgreSQL
- Collabora connected, a document edits and saves in-browser
Worth knowing, not building: a full self-hosted mail server (Mailcow, Mail-in-a-Box) is a real option for digital independence, but most residential ISPs block outbound port 25 and dynamic IPs get blacklisted quickly — it's a genuinely harder project than everything else in this build and not one to take on from a home connection without first confirming the ISP allows it.
CryptPad, if a specific need shows up: Collabora gives Nextcloud in-browser editing, but the server can still technically read document content. CryptPad is zero-knowledge — encrypted client-side, so even a compromised server can't read what's stored. Not worth running as a second, separate document platform for everything; worth adding later specifically for documents where that distinction actually matters (nothing in this build currently needs it).
3 · Nextcloud + DAVx5 — contacts, calendar, and files
- Install Nextcloud and DAVx5, both from F-Droid.
- Nextcloud app: add account with your server URL, enable auto-upload for the folders you want backed up (Downloads, Documents, Music — anything not already covered by Immich).
- DAVx5: add an account, point it at the same Nextcloud server's CardDAV/CalDAV endpoint, let it discover your contacts and calendar collections.
- In Android's own Contacts and Calendar apps, confirm the Nextcloud-synced account now shows up as a real sync source — this is what makes it "just work" like any other synced account on the phone.
- Nextcloud app syncing at least one folder
- DAVx5 configured, contacts and calendar visible in Android's native apps
- A test contact added on the phone confirmed to appear in Nextcloud's web UI (and vice versa)
4 · Seedvault — full app + app-data backup (GrapheneOS)
- Make sure the Nextcloud app is installed and logged in first (Step 3) — Seedvault backs up to any storage location exposed through Android's Storage Access Framework, and Nextcloud registers itself as one of those locations.
- Settings → System → Backup → set the backup location to Nextcloud via the SAF picker.
- Select which apps to include, confirm the encryption recovery code is generated, and store that code in Vaultwarden (Step 5) — without it, the backup is unrecoverable.
- Seedvault backup location set to Nextcloud
- Recovery code generated and stored in Vaultwarden, not just on the phone
- One test restore attempted (even to a spare/secondary device) to confirm the backup is actually usable
5 · Vaultwarden mobile — your passwords, from any device
- Install the official Bitwarden app — available on F-Droid, fully open-source.
- On the login screen, tap the gear/settings icon and enable Self-hosted, entering your Vaultwarden server URL.
- Log in, enable biometric unlock, and turn on Bitwarden as the phone's autofill service (Settings → Passwords & accounts, or similar depending on Android version).
- Bitwarden app pointed at the self-hosted Vaultwarden server, logged in
- Biometric unlock and autofill enabled
- Confirmed the same vault is reachable from a desktop browser too — one vault, every device
Architecture
The request path has two edges deliberately kept separate. NGINX Proxy Manager — already the TLS-terminating edge for every other service in this design (Section 07 → Group B → Proxy & dashboard) — keeps that same role here: it owns the wildcard certificate, listens on port 80 (redirected to 443) and port 443 (HTTPS), and forwards matching hostnames to DockerWakeUp's internal proxy port rather than directly to a container. DockerWakeUp itself never terminates TLS and never sees a certificate; its only job is deciding whether the target container is already running before forwarding the request on.
Two independent things happen on every request. First, the request path: DockerWakeUp checks whether the target's containers are running; if so, the request is forwarded immediately with no added latency worth measuring. If the target is stopped, DockerWakeUp runs docker compose up for that service's project, waits for it to start responding, then forwards the now-ready request — this is the "cold start" and its duration is entirely a function of the target application, not of DockerWakeUp itself. Second, and unrelated to any single request: a separate idle-checker process runs on a five-minute timer, independent of traffic, comparing each service's last-active time against its configured threshold, and runs docker compose down for anything that has exceeded it.
1 · Install prerequisites
sudo apt update sudo apt install -y nodejs npm nginx jq certbot sudo apt install -y python3-certbot-dns-cloudflare
Docker and Docker Compose are already present from earlier steps in this design; NGINX here refers to the underlying package DockerWakeUp's setup script expects to find — NPM already provides this role for every other service, so this instance stays internal-only (see Step 3) rather than becoming a second public-facing NGINX.
- Node.js, npm, nginx, jq, certbot installed
2 · Clone and configure
- Deploy on VLAN 15, alongside the other container-management tooling.
- Clone the repository and create the working config from the supplied example.
git clone https://github.com/jelliott2021/DockerWakeUp.git cd DockerWakeUp cp config.json.example config.json nano config.json
Each entry in services maps a route name to the Compose project that runs it. The route name and domain combine into the public hostname NPM forwards for — route: "gitea-test" under domain lab.internal becomes gitea-test.lab.internal.
{
"proxyPort": 8080,
"idleThreshold": 3600,
"domain": "lab.internal",
"services": [
{
"route": "gitea-test",
"target": "http://localhost:3001",
"composeDir": "/home/claude/homelabservices/gitea-test"
}
]
}
- config.json created, at least one service entry defined
3 · Wire NPM to DockerWakeUp instead of the container directly
- In NPM, create a proxy host for each on-demand hostname (e.g.
gitea-test.lab.internal), forwarding to DockerWakeUp's internal address on port 8080 — not to the target container's own port. - Enable the existing wildcard certificate on the proxy host, exactly as for every other NPM-fronted service.
- Leave DockerWakeUp's own NGINX config disabled or uninstalled — one TLS-terminating edge for the whole lab, not two.
- NPM proxy host forwards to DockerWakeUp's port, not the container's port
- Existing wildcard certificate applied — no separate Certbot instance created
4 · Enable at boot and test a cold start
- Run the setup script and enable the service.
chmod +x setup-service.sh ./setup-service.sh sudo systemctl enable docker-wakeup
To confirm it works: manually stop the test service (docker compose down) and request its hostname in a browser — the request should stall briefly while DockerWakeUp starts the stack, then complete once the target responds.
- docker-wakeup enabled at boot
- A manually stopped service starts on its next request and serves the page
Idle threshold policy
The idle checker runs every five minutes regardless of threshold value; the threshold only controls how long a service is allowed to sit idle before that check stops it. There is no single correct value — it is a direct trade-off against cold-start time, set per service.
| Workload | Suggested idle threshold |
|---|---|
| Frequently used utility | 12–24 hours |
| Occasionally used application | 1–4 hours |
| Development / test environment | 30–60 minutes |
| Demo environment | 15–30 minutes |
| Game server | 15–60 minutes |
Alternative worth knowing: Sablier solves the same problem and is the more established, better-documented project of the two — genuinely worth evaluating first if DockerWakeUp's rougher setup experience (Step 1 above exists because of it) becomes a real obstacle.
1 · Set up the repo structure
In soranthony5/homelab-deployment-v1.0, add:
/terraform/ - VM & LXC definitions (Proxmox provider)
/ansible/ - OS-level config & hardening playbooks
/kubernetes/ - Kustomize/Helm manifests, one folder per app
/docker/ - Compose stacks for non-k8s services, one folder per app
jellyfin/
docker-compose.yml
.env
README.md
/archive/ - old planning docs (already exists)
- Folder structure created and pushed
2 · Terraform (OpenTofu) + Proxmox provider
- Install OpenTofu (or Terraform) on your workstation or a dedicated LXC.
- Configure the
bpg/proxmoxorTelmate/proxmoxprovider with an API token from Proxmox (not your root password). - Write your first resource: recreate one existing LXC (e.g. Homarr) as code,
plan, thenapplyagainst a throwaway test VM ID first.
- Proxmox API token created, scoped (not root)
- One VM/LXC successfully created purely from Terraform code
3 · Ansible for OS-level configuration
- Install Ansible on your workstation or a dedicated LXC.
- Build an inventory file listing your hosts (pve, opn, docker hosts, etc.) grouped by role.
- Write a baseline playbook: SSH key auth, disable password login, install common packages, apply timezone/NTP.
- Run it against one host first, confirm idempotency (running it twice changes nothing the second time).
- Inventory file created, grouped by role
- Baseline playbook applied to at least one host, confirmed idempotent
4 · Terraform the Cloudflare Zero Trust config
- Add the
cloudflare/cloudflareprovider to your Terraform config, authenticate with an API token scoped to just the Tunnel/Access/DNS permissions it needs. - Define your tunnel, its public hostnames, and Access applications/policies as Terraform resources.
plan, review,apply— then make your next Cloudflare change by editing code and re-applying, not by clicking in the dashboard.
- Cloudflare API token scoped (not a global key)
- At least one tunnel hostname and Access policy defined and applied via Terraform
5 · Optional: Packer golden images
- Install Packer alongside Terraform and Ansible.
- Write a Packer template that builds a minimal, patched Ubuntu image on Proxmox and converts it to a VM template.
- Point your Terraform VM resources at that template instead of a raw ISO.
- One golden image built and confirmed usable as a Terraform VM source (optional — skip if the repo above already covers what you need)
1 · Bootstrap a single-node Talos cluster
- Download the Talos ISO, create a VM in Proxmox (suggested start: 4 vCPU, 8GB RAM, 60GB disk, tagged VLAN 15).
- Boot it — Talos comes up in maintenance mode with no config, listening on the API port.
- Generate a machine config with
talosctl gen config, edit it to set the single node as both controlplane and worker. - Apply the config, bootstrap etcd, pull down the kubeconfig.
- Confirm with
kubectl get nodes— one Ready node.
- Talos VM created and booted
- Machine config applied, cluster bootstrapped
-
kubectl get nodesshows one Ready node
2 · Deploy Argo CD, point it at the repo
/kubernetes repo folder, and reconciles any difference automatically. Push a change to git, Argo CD applies it — no kubectl apply from your laptop, no "which version is actually deployed" uncertainty.- Install Argo CD into the cluster via its official manifests.
- Point it at
soranthony5/homelab-deployment-v1.0, path/kubernetes. - Deploy one trivial test app purely by pushing a manifest to git — confirm Argo CD picks it up without manual intervention.
- Argo CD deployed and synced to the repo
- One app deployed purely via a git push, no manual kubectl
3 · Traefik as the cluster's ingress
Deploy via its official Helm chart or Argo CD manifest, expose it through a single OPNsense-forwarded port if you need external K8s ingress, or keep it internal-only.
- Traefik deployed, one test Ingress resource routing correctly
4 · Longhorn — read the caveat before deploying
- Deploy Longhorn via its Helm chart or Argo CD manifest.
- Use it for stateful workloads that benefit from snapshots (databases, Argo CD itself) — not for bulk media, which stays on the NAS over NFS regardless.
- Longhorn deployed, one PVC provisioned and tested
- Understood and accepted: single-node = no cross-node resilience yet
5 · Flatcar Linux — the testing counterpart
Deploy as a VM tagged VLAN 90 (Testing), named dockertest01 per the naming convention in Section 09. Use it for anything you're not ready to commit to the GitOps repo yet.
- Flatcar VM deployed on VLAN 90, named per convention
1 · Deploy Ollama
- Deploy Ollama as a container (VM/LXC tagged VLAN 15).
- Pull a realistic starting model:
llama3.2:3b,phi3.5, orqwen2.5:7b-q4_K_M— all usable on CPU without a long wait per response. - Test with a direct API call before adding a UI on top, so you know the model itself works.
- Ollama deployed, at least one small quantized model pulled and responding
2 · Deploy Open WebUI
ChatGPT-style web frontend for Ollama — chat history, multiple models, document upload for RAG-style question answering over your own files. Point it at the Ollama instance from Step 1, reverse-proxy it via NPM or Traefik as ai.lab.internal.
- Open WebUI deployed, connected to Ollama, reachable via a clean URL
3 · Optional: iGPU passthrough for a speed boost
Passthrough /dev/dri into the Ollama and Jellyfin LXCs (not full VM passthrough). Confirm with intel_gpu_top on the host that both containers can access it.
- iGPU passthrough configured on relevant LXCs
- Confirmed both Jellyfin transcoding and Ollama can use it
4 · Optional: self-hosted AI workflows (n8n)
n8n is a self-hosted workflow automation tool with native Ollama/OpenAI-compatible nodes — a natural fit for things like "summarise new Paperless-ngx documents" or "generate a daily digest for Homarr" using your local model instead of a cloud API.
- n8n deployed (optional), one simple workflow built as a test
1 · Create the isolated Windows VM
- Create a VM: suggested 4 vCPU, 8GB RAM, 80GB disk, tagged VLAN 35.
- Install Windows 10/11, apply all updates, keep Defender enabled.
- Install MetaTrader 5 from your broker's official installer only.
- VM created on VLAN 35, confirmed no lateral access to other VLANs
- Windows fully patched, Defender active
- MT5 installed from the broker's official source
2 · Lock down the firewall rule for this VLAN
- On OPNsense, restrict VLAN 35 outbound to: your broker's server hostnames (use FQDN aliases if supported), NTP, and DNS (VLAN 60) only.
- Confirm no rule allows VLAN 35 to reach any other internal VLAN, including Management.
- Admin access to this VM is RDP only, only reachable via the NetBird admin group — never a WAN port-forward.
- Outbound restricted to broker + NTP + DNS only
- No inter-VLAN rule exists for VLAN 35 in either direction except the exceptions above
- RDP reachable only via NetBird, never WAN-exposed
2b · Endpoint hardening — this VM is a Windows endpoint holding real money
- Enforce a USB device whitelist (Group Policy or a third-party tool) — block unknown USB devices by default rather than trusting anything plugged in.
- Apply restrictive Group Policy: no access to the C: drive directly, no ability to install new programs, no unnecessary system changes — reduces what a compromised EA or indicator could actually do even if it tried.
- Keep endpoint AV mandatory and active — treat AV being disabled or out of date as an automatic "something is wrong" signal, not a background nicety.
- Enable Windows Event Logging for USB connections and blocked actions, forwarded to Wazuh (Section 07 → Group C → Security & identity) alongside everything else.
- USB whitelist enforced
- Group Policy restrictions applied (no C: drive access, no arbitrary installs)
- AV active, logging forwarded to Wazuh
3 · Power & backup priority
- Confirm this VM is included in the NUT-triggered graceful shutdown sequence (Section 04), and check whether your specific EA has a documented safe-disconnect or pause behaviour.
- Add an Uptime Kuma monitor on this VM specifically, notification target set to ntfy (Section 07 → Group B → Monitoring & DNS, Step 7) with a dedicated high-priority topic — given real money is at stake, you want to know within minutes, not hours, and ntfy's priority flag can make that alert impossible to miss.
- Set a PBS backup schedule for this VM, and always take a manual snapshot before changing any EA or indicator.
- VM included in graceful shutdown sequence
- Fast-notification Uptime Kuma monitor configured
- PBS backup schedule set, manual pre-change snapshot habit established
4 · Optional: browser-based access via Guacamole
Apache Guacamole gives you RDP/VNC access to the MT5 VM through a plain browser tab (still tunneled through NetBird) — no native RDP client needed on whatever device you're checking from.
- Guacamole deployed (optional), RDP session to the MT5 VM tested through it
1 · Deploy Authentik — centralized identity & forward-auth
- Deploy Authentik (server + worker + its own dedicated Postgres + Redis — keep this separate from the shared Database VLAN 80 instance to avoid extra firewall exceptions) on VLAN 10.
- Complete the setup wizard, create your admin account with a strong unique password from Vaultwarden, enable MFA (TOTP or a hardware key such as a YubiKey) immediately, then disable or delete the wizard's default admin account if one remains.
- Create an
adminsgroup (just you, for now) and a separate lower-privilege group as a template for any future limited access — never grant a new user or service account admin-group membership by default.
- Authentik deployed, own admin account created with MFA enabled
- Default/wizard admin account disabled or removed
-
adminsgroup created, scoped to just you
2 · Wire NPM / Traefik to require Authentik before every request
- Configure Authentik's "Proxy Provider" (forward-auth mode) for each service you want gated.
- In NPM, add the forward-auth snippet to each proxy host's advanced config (or the equivalent Traefik forward-auth middleware for anything running in the Talos cluster).
- Start with low-risk services (Homarr, Grafana) before gating anything you rely on daily, in case the config needs tuning.
- At least one proxied service confirmed redirecting to Authentik login when unauthenticated
- Same service confirmed accessible once logged in
3 · Deploy Wazuh — logging & SIEM
- Deploy the Wazuh manager + indexer + dashboard on VLAN 50 (Monitor).
- Install the Wazuh agent on Proxmox and any Docker/Talos hosts that support it.
- Forward OPNsense's syslog output to the Wazuh manager.
- Forward Authentik's authentication events to Wazuh as well, so every login — successful or failed — lands in the same place as everything else.
- Wazuh manager/indexer/dashboard deployed on VLAN 50
- Agents installed on Proxmox and at least one other host
- OPNsense syslog and Authentik events both flowing into Wazuh
4 · Enable Suricata on OPNsense
- Install the
os-suricataplugin from OPNsense's plugin manager. - Enable it on the WAN interface first, subscribe to a free ruleset (Emerging Threats Open is the standard starting point).
- Run in IDS mode (alert-only, doesn't block anything) for at least a week to see what fires and tune out false positives.
- Only switch specific rule categories to IPS (blocking) mode once you've confirmed they're not flagging legitimate traffic.
- Forward Suricata's alerts to Wazuh (Step 3) for a single combined view.
- Suricata enabled on WAN, ET Open ruleset subscribed
- Run in IDS-only mode for at least a week before enabling any IPS blocking
- Alerts forwarded to Wazuh
4b · CrowdSec alongside Suricata — crowd-sourced blocking
- Install the CrowdSec OPNsense plugin, enrol it against the free community blocklist.
- Point it at OPNsense's own logs plus Suricata's alerts as scenarios to watch.
- Start in alert/log mode before enabling active blocking, same caution as Suricata's IPS step.
- CrowdSec deployed, community blocklist enrolled
- Confirmed logging real detections before switching to active blocking
4c · Optional: ClamAV filesystem scanning
Schedule a daily scan on Proxmox and/or the NAS shares, notify on detection via email or a Wazuh forward.
- ClamAV scheduled scan configured on at least one host (optional)
5 · Periodic vulnerability scanning — OpenVAS & Trivy, on-demand
- Deploy OpenVAS/Greenbone via Docker on VLAN 90 only when actively scanning.
- Run a full scan against your own VLANs from inside the lab — this is what Phase 1 of Section 14's Security Audit uses.
- Add Trivy as a pre-deploy check in the Argo CD/GitOps pipeline (Section 07 → Group C → GitOps foundation) to scan container images for known vulnerabilities before they ship.
- OpenVAS run at least once, results reviewed, container torn down afterward
- Trivy wired into the GitOps pipeline as an optional pre-deploy check
6 · Zeek — protocol-level network analysis, alongside Suricata
- Deploy Zeek on VLAN 50, with a network tap or mirrored port on the switch feeding it traffic to observe.
- Forward Zeek's logs to Wazuh, alongside Suricata's alerts and every other log source already centralised there.
- Zeek deployed, receiving mirrored traffic (optional — resource cost should be weighed against the budget in Section 09)
- Logs forwarded to Wazuh
7 · Atomic Red Team — testing whether the detections actually detect
- Run Atomic Red Team tests only against the Testing VLAN sandbox (VLAN 90), never against production services.
- Select a small number of tests relevant to the actual attack surface here — credential access, a simulated brute-force pattern against SSH, an unusual outbound connection — rather than the full library.
- For each test run, confirm the expected alert appeared in Wazuh and the expected notification arrived via ntfy. A test that produces no alert is a genuine finding — a detection gap to fix, not a step to skip past.
- At least one test run against the Testing VLAN, confirmed to produce a Wazuh alert
- Confirmed the alert produced a notification via ntfy, not just a log entry
08Security & remote access
Default-deny firewall
All inter-VLAN traffic blocked unless a rule explicitly allows it. Management (VLAN 10) reaches everything for administration; everything else is scoped to what it actually needs.
Isolation examples
Jellyfin → NAS: allowed
Photos → Internet: blocked
Media ↔ Photos: blocked
Trading → anything but broker/NTP/DNS: blocked
Management → everything: allowed
Zero trust — identity, privilege & logging
VLANs and firewall rules control which network a device sits on. This layer controls who is allowed to act, on top of that — the actual "zero trust" piece, since it doesn't matter which VLAN a request comes from if it can't authenticate.
Single identity, everywhere
Authentik sits in front of every proxied service (Section 07 → Group C → Security & identity). One real account — yours — with MFA required, instead of a dozen separate app logins with varying password strength.
Least privilege by default
An admins group exists containing only you. Any future account — family, guest, a service account — starts in a restricted group and is granted access explicitly, never inherits admin by default.
Full activity logging
Wazuh centralizes logs from every host, OPNsense's syslog, and every Authentik login attempt (successful or failed) into one searchable place — not scattered across a dozen hosts you'd have to SSH into individually during an incident.
Defence in depth on appliance UIs
OPNsense, Proxmox, and the NAS keep their own native logins + 2FA (forward-auth doesn't cleanly wrap appliance UIs) — but they're only reachable via NetBird or Cloudflare Access in the first place, so there are two layers to get through, not one.
Remote access — two independent paths
Primary — NetBird: replaces a plain point-to-point VPN or a third-party-hosted mesh. It's WireGuard under the hood (fast, modern crypto) with a self-hosted coordination server so peers connect directly to each other instead of routing through someone else's cloud. Access is grouped, not all-or-nothing — your admin laptop joins an admin group reaching VLAN 10 (management) and VLAN 90 (testing), while your phone joins a narrower personal-devices group reaching only VLAN 30 (Cloud — Nextcloud, Vaultwarden) and VLAN 40 (Photos — Immich) for always-on backup, per Section 07 → Group B → Phone & mobile backup. A phone doesn't need, and shouldn't have, the same reach as the device you administer the firewall from.
Backup — Cloudflare Tunnel: an architecturally unrelated path — outbound-only, no WireGuard, no inbound firewall rule at all. Cloudflare Access gates it with login before any request reaches your network. Kept deliberately separate from NetBird so one path's outage or misconfiguration never takes down the other.
Full config steps for both are in Section 07 → Group B → Remote access.
Hardening baseline
- SSH disabled by default on every device, enabled only when actually needed, key-only (no password) — a YubiKey FIDO2 resident key is a further step up, since the private key never touches disk at all
- MFA everywhere, admin logins first — Authentik, OPNsense, Proxmox, NAS, and NetBird all support it
- Default credentials changed on every device before go-live, default/wizard admin accounts disabled once your own account exists
- A named personal admin account for daily use on OPNsense, Proxmox, and the NAS — not
root/admindirectly — with elevation only when a task actually needs it - Where a device supports its own local firewall (Proxmox host, NAS), scope it to the specific VLAN subnets that need it — and narrower still, to specific source IPs for the most sensitive management ports, not the whole subnet
- HTTPS everywhere via NGINX Proxy Manager + internal CA or Let's Encrypt DNS challenge — one wildcard cert (Section 07 → Services → Proxy) covers everything rather than managing certs per-service
- fail2ban or equivalent on internet-facing services — extended explicitly to OPNsense's own admin login, NetBird, and Cloudflare Access, not just "the internet-facing stuff"
- UFW (or nftables) enabled on every Linux VM/LXC individually, not just relied on at the VLAN level — set default-deny inbound, allow only the specific LAN subnet each service actually needs. This is genuine defence in depth: VLAN isolation stops traffic between networks, a host firewall stops traffic from other devices on the same VLAN that shouldn't be talking to this one host specifically. Fail2ban belongs on the same hosts, watching SSH and any exposed login forms for repeated failures.
- Automatic security updates enabled where safe to do so; everything else reviewed manually against release notes before applying, so a known-bad update doesn't roll out unattended
- NetBird access groups reviewed whenever a new peer or service is added
- NanoKVM web UI never port-forwarded to WAN — reachable only via NetBird's admin group
- Every login — success or failure — lands in Wazuh; nothing authenticates silently
- Client-side ad/tracker blocking as a second layer on top of Technitium's network-level filtering — uBlock Origin in every browser, AdGuard on mobile devices. Malicious ads are a common entry point network-level DNS filtering alone doesn't fully close.
- Changing default ports is not treated as a security measure here — once a device is on the network, a port scan finds the real port trivially. VLANs, firewall rules, and authentication do the actual work; obscurity isn't a substitute for any of them. SSH is not reachable on any port from outside NetBird's tunnel, which is a stronger position than moving it off port 22 while still leaving it reachable.
09Kubernetes, GitOps & AI platform
This section is shaped by two real sources: Jim's Garage's GitOps-first homelab pattern (Terraform + Ansible + Talos + Argo CD), and Brandon Lee's VirtualizationHowto lab, which runs this exact stack — Talos, Flatcar, Argo CD, Traefik, Longhorn, Technitium, Unbound.
Resource budget — one 32GB host
| Workload | Realistic footprint | Notes |
|---|---|---|
| Proxmox host overhead | ~2GB | Fixed cost, always-on |
| Existing LXC services | ~6–8GB | Homarr, NPM, Technitium×2, Unbound, Uptime Kuma, Vaultwarden, etc. — individually light |
| Media + *arr stack | ~3–4GB | Jellyfin, Sonarr, Radarr, Prowlarr, Bazarr, Jellyseerr, qBittorrent |
| Single-node Talos cluster | ~8–10GB | One VM running control-plane + worker roles combined |
| Ollama + Open WebUI | ~6–10GB | Depends entirely on model size — budget for a 7B-class quantized model, not larger |
| MT5 Windows VM | ~6–8GB | Only when actively trading — see Section 07 → Group C → Trading |
| Security stack | ~6–8GB | Authentik + Postgres/Redis (~2GB), Wazuh manager/indexer (~4GB+, the heaviest single piece), Suricata (runs on OPNsense itself, minimal extra). OpenVAS deliberately excluded — on-demand only, not always-on. |
Add the middle columns up and you're already past 32GB before the AI stack, MT5 VM, or security stack even start — running everything simultaneously at generous allocations doesn't fit. Two honest paths forward, not mutually exclusive:
- Stay right-sized: single-node Talos (not a 3-node HA cluster), small quantized LLMs (3B–7B, not 70B), a minimally-sized Wazuh deployment, and accept that the AI stack, MT5 VM, and full security stack aren't all running flat-out at the same moment as a media transcode
- Accelerate the second Proxmox node from Section 10 specifically to carry the K8s cluster, the AI workload, or the security stack separately — this is the real, permanent fix, not a workaround. Wazuh in particular is a strong candidate to be the thing that justifies node 2.
Architecture decisions
| Decision | Choice | Why |
|---|---|---|
| Kubernetes distro | Talos Linux | Immutable, API-managed, no SSH — matches the GitOps philosophy directly. This is the production cluster. |
| Testing OS | Flatcar Linux | A second immutable OS, deliberately different from Talos, for trying new containers before they earn a place in the GitOps repo (VLAN 90) |
| Reverse proxy split | NGINX Proxy Manager (classic VM/LXC/Docker) + Traefik (inside K8s only) | Two clear territories rather than one tool awkwardly covering both |
| K8s persistent storage | Longhorn, with caveats | Useful for PVC provisioning/snapshots now; genuine cross-node resilience waits for node 2 |
| Bulk/media storage | Stays on the NAS over NFS | Longhorn is for stateful app data, not for the media library |
| GitOps engine | Argo CD, watching /kubernetes in the repo | Push to git, cluster state follows automatically |
| Infra provisioning | Terraform (OpenTofu) for VMs/LXCs, Ansible for OS config | Rebuild the whole lab from the repo if hardware is replaced |
| DNS platform | Technitium DNS (clustered pair) + Unbound | Native clustering instead of sync scripts; Unbound as a private validating resolver |
| Local AI | Ollama + Open WebUI, small quantized models | No discrete GPU — CPU inference is genuinely usable at the right model size |
| Identity / access control | Authentik (forward-auth), Wazuh (logging/SIEM), Suricata (IDS on OPNsense) | One identity everywhere, one place logs land, one layer inspecting traffic content — see Section 07 |
GitOps flow
Full step-by-step for all of this — bootstrapping Talos, deploying Argo CD/Traefik/Longhorn, standing up Ollama/Open WebUI, and the MT5 trading VM — lives in Section 07 → Group C.
10Home lab organisation & Docker best practices
Organisational structure and Docker networking practice, informed by established home lab documentation — the two primary factors determining whether a system of this scale remains manageable as it grows.
Organise by function, not technology
Devices and services are grouped by function rather than by underlying technology — not "Docker VMs" versus "Kubernetes VMs" versus "Windows VMs," but by what each one does. This grouping is applied as a Proxmox tag on every VM/LXC and as the top-level folder structure in the git repository:
| Role tag | What lives there |
|---|---|
| infrastructure | Proxmox, OPNsense, switch config, PBS |
| networking | Technitium, Unbound, NetBird, NPM, Traefik |
| monitoring | Prometheus, Grafana, InfluxDB, Uptime Kuma, Loki |
| ai | Ollama, Open WebUI, n8n |
| media | Jellyfin, Sonarr, Radarr, Prowlarr, Bazarr, Jellyseerr |
| storage | NAS shares, Longhorn, NFS/iSCSI exports |
| identity | Vaultwarden, Authentik |
| security | Wazuh, Suricata, OpenVAS (on-demand) |
| automation | Argo CD, Terraform, Ansible, Gitea |
| trading | MT5 VM — kept deliberately alone, nothing else shares this tag or this VLAN |
| testing | Flatcar host, sandbox VMs, anything not yet in the GitOps repo |
Predictable hostnames
Boring, functional names beat clever ones — you should know what a host does from its name alone, six months from now, half-asleep, during an outage.
| Host | Role |
|---|---|
| pve01 | The Proxmox host itself |
| opn01 | OPNsense |
| sw01 | MikroTik switch |
| nas01 | Terramaster NAS |
| kube-cp01 | Talos control-plane (+worker, single-node) |
| dockertest01 / test01 | Flatcar testing host / Gitea + Paperless-ngx Docker host |
| ai01 | Ollama + Open WebUI + NetBird (Docker) |
| mgmt01 | Homarr, NPM, Authentik, NetBox (Docker) |
| media01 / idx01 | Jellyfin + *arr stack / Prowlarr (Docker) |
| cloud01 / photos01 | Nextcloud + Vaultwarden / Immich (Docker) |
| trade01 | MT5 Windows VM |
| dns01 / dns02 | Technitium primary / secondary |
Git folder structure mirrors the infrastructure
Every Docker Compose service gets its own folder, same shape every time — migrate a service or rebuild a host by copying one folder:
/docker/
jellyfin/
docker-compose.yml
.env
README.md
config/
data/
backups/
homarr/
docker-compose.yml
.env
README.md
...
DNS-first naming
Once Technitium is live (Section 07 → Services → Monitoring & DNS), stop remembering IPs — create a zone and give every service a real name: grafana.lab.internal, ai.lab.internal, proxmox.lab.internal, argocd.lab.internal. NPM and Traefik both route on these names, and it's the single biggest daily-usability improvement in a lab this size.
Nine checks before trusting a Docker image
The vulnerability scanning in Section 07 (Trivy, OpenVAS) catches known CVEs in an image already pulled. These checks happen earlier — before an unfamiliar image is deployed at all, when the real question is not "is this patched" but "should this be trusted."
1. Publisher identity
Confirm the image is the project's own, not a look-alike fork — check the official docs for the canonical registry path before pulling. docker buildx imagetools inspect shows what a reference actually resolves to.
2. Active maintenance
Recent commits, a documented vulnerability-reporting process, reviewed dependency-update PRs, and images built automatically from tagged releases — an unmaintained image is a growing liability regardless of how clean it looks today.
3. The Dockerfile itself
When available, check the base image and watch for a bare curl | sh pulling and executing a remote script during build — not automatically malicious, but worth knowing it's there.
4. Image metadata and layers
docker image inspect and docker image history surface the default user, exposed ports, volumes, and what each layer actually did before a container is ever started from it.
5. Runtime privileges requested
Treat privileged: true, network_mode: host, or a mounted docker.sock in a Compose file as a question, not a default — does this specific application actually need it?
6. Vulnerability scan
Trivy or Docker Scout against the pulled image — already the default in this design (Section 07 → GitOps foundation) for anything going through the pipeline; worth running manually too for a one-off pull.
7. Signature verification
Where a project signs its images (cosign is the common tool), verifying the signature confirms the image matches what the developer actually published — closing the specific gap a scan alone can't: supply-chain tampering after the fact.
8. Tag vs. digest
A tag can be silently moved to point at a different image later; a digest (@sha256:...) identifies exact content. Pinning by digest for anything security-sensitive means what gets deployed tomorrow is provably what was tested today.
9. Test in isolation first
VLAN 90 (Testing) already exists for exactly this — run an unfamiliar image there first, optionally with --network none --read-only --cap-drop ALL, and watch what it actually tries to do before it reaches a real VLAN.
Docker networking & host mistakes to avoid
1. Publishing every port
Only publish the port users/browsers actually hit (the frontend, or Traefik/NPM). Backend containers reach each other over Docker's internal network — no port mapping needed.
2. One giant network
Give each app stack its own dedicated Docker network. Only the frontend container joins the shared proxy network too.
3. Exposing databases
Don't publish 5432/3306/6379 to the host. The app reaches postgres:5432 by service name on the shared internal network — no port needed at all.
4. Overusing host networking
Reserve it for the few tools that genuinely need to see the host's wire (e.g. network monitors). Everything else stays on bridge networking for real isolation.
5. Skipping the external proxy network
Create one dedicated external network for NPM/Traefik. App stacks connect their frontend to it and keep their own private network for backend traffic — the only way in is through the proxy.
6. Hardcoding IPs / subnets
Use Docker's built-in service-name DNS, never a container's IP. Let Docker auto-assign subnets — manually picking them is how you end up with silent overlaps down the line.
7. Trusting a host firewall to catch Docker traffic
Docker manipulates iptables directly and can silently bypass UFW or a host-level firewall rule — a ports: mapping in a compose file can expose a service even when the host firewall "should" have blocked it. VLAN isolation (Section 05) is the real control here, not a host firewall layered on top of Docker.
8. Running containers/LXCs privileged by default
Privileged mode grants root-level host access and skips normal isolation. Default to unprivileged/rootless for every container and LXC in this build, and only grant privileged mode to the specific few that genuinely need device passthrough (e.g. iGPU sharing for Jellyfin/Ollama) — never as a default habit to sidestep a permissions error.
Keep production and testing genuinely separate
This build already does this structurally: VLAN 90 (Testing) plus the Flatcar dockertest01 host is where anything new gets tried first. It only gets promoted into the Talos GitOps repo — becoming "production" — once it's actually proven out. Never test directly against the Talos cluster or the core LXC services.
Separate Linux users per service, not one account for everything
On any host running multiple services directly (rather than fully containerized), give each function its own unprivileged Linux user — one for media, one for backups, one for sandbox/testing work, and so on — rather than running everything as a single account. If one service gets compromised, the blast radius is whatever that one user can touch, not everything on the box. None of these service accounts should have sudo; elevate deliberately, per the named-admin-account principle in Section 07.
11Future expansion & HA readiness
Nothing below needs building today — it's what this design already leaves room for.
Second Proxmox node
RU9 is reserved specifically for this. A second node turns Proxmox from a single box into a real cluster capable of live-migrating VMs. Proxmox HA clustering wants an odd number of voting members for quorum — a 2-node cluster needs a lightweight third vote (a QDevice, which can be as small as a Raspberry Pi) rather than a full third server.
Shared storage for HA
Live migration and automatic VM failover need storage both nodes can see. The NAS's NFS/iSCSI exports (VLAN 999) already provide this path — no redesign needed when you add node 2. Ceph across nodes is the other common option if you later add local disks to each node instead of relying on the NAS.
10GbE headroom
The CRS310 has exactly 2 SFP+ ports, both spoken for (OPNsense + Proxmox). A second hypervisor node — or true LACP bonding — needs a switch with more 10GbE ports. Worth planning for when a second node becomes real rather than buying ahead of need.
3-2-1 backups
Proxmox Backup Server (Section 07) plus the NAS gives you 2 local copies. Add a third, offsite copy — a cloud target (Backblaze B2, etc.) or a drive rotated to another location — to complete a real 3-2-1 strategy rather than "two copies in the same rack."
Rough order of operations, when you're ready
- Go live on the current single-node design first — don't build for HA before you have a workload that needs it
- Add NanoKVM once NetBird is running, so remote hands-off recovery actually works end-to-end
- Bring up the single-node Talos cluster and AI stack only after the core services (Section 07 Groups A/B) are stable — Section 08 has the resource math on why not everything runs at once yet
- Add a QDevice (Raspberry Pi is plenty) before a second Proxmox node, so quorum works from day one of the cluster
- Add the second node into RU9, cable it the same way as Proxmox today (10GbE if the switch supports it by then, 1GbE otherwise) — this is also what unlocks genuine multi-node Talos and real Longhorn resilience
- Re-run the weight check in Section 02 every time something physical gets added
12Build checklist
Five phases, done in order. Expand a phase, check items off as you go by editing this file directly.
Phase 1 — PhysicalCabinet mounted, all hardware and UPS installed, cables verified, safe power-on
- Cabinet wall-mounted into studs, level, secure — confirm total load stays under 50kg
- UPS installed at RU 1–2 (bottom), switch/shelves/NAS installed per rack layout above
- All 1GbE fallback cables + 2× 10GbE DAC cables (owned) + PDU cabling installed and labelled both ends
- Cable verification pass (tug test, no pinches, DAC cables not sharply bent)
- Power-on sequence: switch → NAS → OPNsense → Proxmox (last — highest inrush)
- All LEDs verified green (incl. SFP+ links), console access confirmed on each device
Phase 2 — NetworkFirewall configured, 15 VLANs live on the 10GbE trunk, switch trunked correctly
- OPNsense SFP+ interface identified and set as VLAN trunk parent (em0 kept as fallback)
- 15 VLAN subinterfaces created with correct gateways
- DHCP scopes enabled per VLAN
- Switch SFP+1/SFP+2 added to trunk bridge, ports 4 & 5 set to access
- Default-deny rules in place, then explicit allow rules per VLAN
- Cross-VLAN traffic tested against the isolation matrix
Phase 3 — Compute & storageProxmox and NAS ready to host services
- Proxmox installed, VLAN-aware trunk bridge (vmbr0) created on the SFP+ NIC
- NAS RAID 5 configured, shares created, NFS export to Proxmox
- First test VM created, tagged VLAN 90, reaches its gateway and the NAS
- Storage throughput checked (>100MB/s target on 1GbE NAS link)
Phase 4 — Core services & add-onsMedia, cloud, DNS, monitoring, dashboard, and proxy stacks live
- Media stack: Jellyfin, Sonarr, Radarr, Prowlarr
- Cloud stack: Nextcloud, Vaultwarden
- DNS: Pi-hole + AdGuard Home (secondary)
- Monitoring: Prometheus, Grafana, InfluxDB, Uptime Kuma, Loki
- Backing databases: PostgreSQL, Redis/MariaDB on VLAN 80
- NGINX Proxy Manager + Homarr dashboard deployed on VLAN 10
- NetBird self-hosted server deployed on VLAN 70, admin devices enrolled
Phase 5 — Backup, UPS & hardeningProduction-ready and resilient to power loss
- UPS installed, NUT configured for automated graceful shutdown
- UPS battery test performed — confirm devices survive a simulated outage
- Proxmox Backup Server deployed, daily incremental + weekly full jobs scheduled
- Restore tested at least once, end to end
- Third, offsite backup copy configured (3-2-1)
- Security hardening checklist (Section 07) fully applied
- This document updated to reflect the final as-built state, changelog entry added
Phase 6 — Platform expansion (optional, later)GitOps foundation, Kubernetes, AI stack, and MT5 trading — after everything above is stable
- Repo restructured for Terraform/Ansible/Kubernetes per Section 08
- Terraform + Proxmox provider provisioning at least one real VM/LXC
- Ansible baseline playbook applied across existing hosts
- Single-node Talos cluster bootstrapped, Argo CD deployed and synced to the repo
- Traefik and Longhorn deployed inside the cluster (with the single-node caveat understood)
- Ollama + Open WebUI deployed, at least one small model tested end to end
- Jellyseerr + Bazarr added to complete the media automation pipeline
- MT5 VM deployed on isolated VLAN 35, firewall rules locked to broker + NTP + DNS only, included in backup and shutdown priority
- Resource budget in Section 08 re-checked against actual usage once everything is live
13Operations quick reference
Daily
Glance at Homarr/Grafana, check overnight backup succeeded, confirm no failed services, check Uptime Kuma status page.
Weekly
Review firewall logs, check pending updates, verify UPS self-test passed, review NetBird peer list.
Monthly
Full PBS restore drill, storage capacity review, rotate any credentials due for rotation.
On power loss
NUT should already be handling this automatically — check the shutdown log afterward to confirm it triggered cleanly.
14Security audit & recheck
This is not a one-time checklist — it's a cycle. Run it in full before this lab is ever exposed to the internet, and again on a quarterly cadence after that. Structured around AUDIT → HARDEN → SEGMENT → RECOVER, cross-checked against the homelabstarter network security audit and the community homelab-hardening-checklists repo — worth periodically diffing this section against that repo as it's updated.
AUDIT — run before going public, then quarterlyFind out what's actually true about the lab, not what you assume is true
-
nmapscan of the WAN IP from outside the LAN (mobile hotspot, not home wifi) — confirm only the Cloudflare Tunnel's outbound connection shows, no unexpected open ports - Cross-check with canyouseeme.org — a single-port web-based check independent of the device or network running the nmap scan above, useful as a quick sanity check without needing a second network to scan from
- Check shodan.io for the home WAN IP — confirm nothing is indexed
- Review every OPNsense port-forward rule — delete anything you don't immediately recognise the reason for
- Inventory every exposed service (Cloudflare Tunnel hostnames) — confirm each sits behind Cloudflare Access and Authentik, not just one
- Have I Been Pwned check for every admin email used across OPNsense, Proxmox, Authentik, NetBird, Cloudflare
- User/account audit — any accounts still active that shouldn't be? Any default/wizard accounts not yet disabled?
- Trivy scan of container images currently in use (Section 07 → Group C → Security & identity)
- OpenVAS scan run against the lab's own VLANs, results reviewed, container torn down afterward
- Firewall rule review — for each VLAN, does every allow rule still have a reason to exist?
- Every
.env/ compose secret for every deployed service checked against its default — a surprising number of self-hosted projects ship a working default password/key that's easy to miss on a quick setup -
netstat/ listening-socket check on Proxmox and any Docker hosts — confirm nothing is listening that you didn't intend to run
HARDEN — passwords, keys, patchingClose the gaps AUDIT just found
- Every credential is 20+ characters, generated and stored in Vaultwarden — no exceptions, no reused passwords
- MFA enabled everywhere it's supported, admin logins first (Authentik, OPNsense, Proxmox, NAS, NetBird, Cloudflare)
- SSH hardened: key-only auth, root login disabled, fail2ban (or equivalent) active — consider a YubiKey FIDO2 resident key so the private key never touches disk
- Patch review completed — Proxmox host, every VM/LXC OS, and container images all checked for pending updates
- Cloudflare Access confirmed in front of every admin UI that's reachable through the tunnel
- Unused services/plugins disabled — less running surface, less to keep patched
SEGMENT — VLANs, DNS, firewallConfirm the isolation design is still actually true, not just documented
- Traffic matrix in Section 05 re-tested against reality — pick 3 random VLAN pairs and confirm allowed/blocked matches what's documented
- DNS deny-by-default still enforced — a hardcoded public resolver still fails from a restricted VLAN (Section 07 → Services → Monitoring & DNS)
- Home Assistant / IoT devices confirmed still isolated to VLAN 90, no unexpected new devices present
- Any camera/RTSP traffic (Frigate, once added) confirmed on VLAN 45 only — zero routing to any other VLAN, dual-NIC pattern (like the NAS) used if Frigate itself needs a management-reachable interface
- VLAN 35 (Trading) re-confirmed: zero lateral access in either direction except broker/NTP/DNS
RECOVER — backups, tests, practiceAn unverified backup is a belief, not a fact
- Nightly PBS backup job verified — confirm the last 3 actually completed, not just scheduled
- Test restore performed — pull a real backup, mount it, actually read the data back
- 3-2-1 rule confirmed: 3 copies, at least 1 offsite
- At least one backup copy is immutable — can't be deleted or modified for a set retention window, so ransomware reaching the lab can't also delete the way back out
- Backups are encrypted, and the decryption key is stored somewhere other than the lab itself
- A recovery document exists that someone else (or a future, panicked you) could actually follow start to finish
- A simulated full host-loss drill practiced at least once — not just backups existing, but proof they'd actually get you back online
Asset hardening log
A lightweight record per device, filled in the first time it's hardened and updated whenever something changes — the difference between "I think I did this" and being able to check. Copy this table and extend it as devices are added. Once NetBox is deployed (Section 07 → Group B → Proxy & dashboard), it can take over as the living version of this same information — retire the manual table at that point rather than keeping both in parallel.
| Device | IP | MAC | Hardened by | Date | Notes |
|---|---|---|---|---|---|
| opn01 | 10.0.1.1 | — | — | — | — |
| sw01 | 10.0.1.5 | — | — | — | — |
| pve01 | 10.0.1.x | — | — | — | — |
| nas01 | 10.0.30.10 | — | — | — | — |
Last full audit: 20XX-XX-XX · Findings: ... · Fixed: ... — keep the history in the changelog below so the audit trail is as real as the lab itself.
REFExternal references
Every official documentation link used throughout this document, grouped by subsystem, with the tab or section each is cited from. Each entry opens the source directly, for verification or for going deeper than this document's own summary.
| Link | Cited in | |
|---|---|---|
| Firewall & perimeter (OPNsense) | ||
| CrowdSec OPNsense integration | 07 → Security & identity | |
| NAT | 07 → OPNsense | |
| OPNsense Suricata / IDS-IPS | 07 → Security & identity | |
| OPNsense VLAN configuration | 07 → OPNsense | |
| OPNsense firewall rules | 07 → OPNsense | |
| OPNsense install & first boot | 07 → OPNsense | |
| OPNsense interface assignment | 07 → OPNsense | |
| Diagram design system | ||
| diagram-design — architecture & data flow visual system | 05 → Architecture view · 06 → Data flow | |
| Switch (MikroTik) | ||
| MikroTik bridge VLAN filtering | 07 → MikroTik switch | |
| MikroTik interfaces | 07 → MikroTik switch | |
| Home devices switch (Netgear JGS524E) | ||
| JGS524E installation guide | 07 → Home switch (JGS524E) | |
| Netgear support — JGS524E | 07 → Home switch (JGS524E) | |
| Hypervisor & Proxmox tooling | ||
| Original tteck repository (archival reference) | 07 → Proxmox | |
| PVE Post Install script | 07 → Proxmox | |
| Proxmox Backup Server docs | 07 → Backups | |
| Proxmox Helper Script — HA OS VM | 07 → Home automation | |
| Proxmox Helper Script — Homebridge LXC | 07 → Home automation | |
| Proxmox VE Helper-Scripts index | 07 → Proxmox | |
| Proxmox VE installation | 07 → Proxmox | |
| Proxmox network configuration | 07 → Proxmox | |
| NAS (TOS / TrueNAS reference) | ||
| TOS downloads & firmware | 07 → NAS | |
| TerraMaster TOS overview | 07 → NAS | |
| TrueNAS Community Edition documentation (reference/comparison) | 07 → NAS | |
| Out-of-band management (NanoKVM) | ||
| NanoKVM ATX wiring guide | 07 → NanoKVM | |
| NanoKVM quick start | 07 → NanoKVM | |
| NanoKVM user guide | 07 → NanoKVM | |
| Remote access (NetBird / Cloudflare) | ||
| Cloudflare Access policies | 07 → Remote access | |
| Cloudflare Tunnel docs | 07 → Remote access | |
| Cloudflare Tunnel homelab walkthrough | 07 → Remote access | |
| Load-balanced Cloudflare Tunnels with Docker Swarm | 07 → Remote access | |
| NetBird Android install | 07 → Phone & mobile backup | |
| NetBird identity providers | 07 → Remote access | |
| NetBird self-hosted guide | 07 → Remote access | |
| Reverse proxy & dashboard | ||
| Homarr getting started | 07 → Proxy & dashboard | |
| NGINX Proxy Manager guide | 07 → Proxy & dashboard | |
| NetBox Docker | 07 → Proxy & dashboard | |
| Traefik quick start | 07 → Kubernetes (Talos) | |
| Watchtower documentation | 07 → Proxy & dashboard | |
| DNS (Technitium / Unbound) | ||
| Technitium DNS Server | 07 → Monitoring & DNS | |
| Unbound documentation | 07 → Monitoring & DNS | |
| a follow-up troubleshooting post | 07 → Monitoring & DNS | |
| the original clustering setup | 07 → Monitoring & DNS | |
| Media automation | ||
| Bazarr wiki | 07 → Media stack | |
| Jellyfin documentation | 07 → Media stack | |
| Jellyseerr documentation | 07 → Media stack | |
| Prowlarr wiki | 07 → Media stack | |
| Radarr wiki | 07 → Media stack | |
| Sonarr wiki | 07 → Media stack | |
| Home automation | ||
| Home Assistant installation | 07 → Home automation | |
| Homebridge documentation | 07 → Home automation | |
| Mosquitto documentation | 06 Service interconnection | |
| Mobile backup & Vaultwarden | ||
| DAVx5 | 07 → Phone & mobile backup | |
| GrapheneOS backup (Seedvault) | 07 → Phone & mobile backup | |
| Immich mobile backup | 07 → Phone & mobile backup | |
| Nextcloud Android app | 07 → Phone & mobile backup | |
| Nextcloud server install | 07 → Phone & mobile backup | |
| Vaultwarden wiki | 07 → Phone & mobile backup | |
| Trading VM access | ||
| Apache Guacamole docs | 07 → Trading (MT5) | |
| Kubernetes platform | ||
| Argo CD getting started | 07 → Kubernetes (Talos) | |
| Flatcar installation | 07 → Kubernetes (Talos) | |
| Longhorn installation | 07 → Kubernetes (Talos) | |
| Talos getting started | 07 → Kubernetes (Talos) | |
| AI stack | ||
| Ollama Docker guide | 07 → AI stack | |
| Open WebUI documentation | 07 → AI stack | |
| n8n documentation | 07 → AI stack | |
| Security & identity | ||
| Atomic Red Team | 07 → Security & identity | |
| Authentik Docker Compose install | 07 → Security & identity | |
| Authentik proxy provider / forward-auth | 07 → Security & identity | |
| ClamAV documentation | 07 → Security & identity | |
| Greenbone/OpenVAS Docker | 07 → Security & identity | |
| Trivy installation | 07 → Security & identity | |
| Wazuh Docker deployment | 07 → Security & identity | |
| Zeek documentation | 07 → Security & identity | |
| Monitoring & notifications | ||
| Beszel getting started | 07 → Monitoring & DNS | |
| Grafana getting started | 07 → Monitoring & DNS | |
| Prometheus overview | 07 → Monitoring & DNS | |
| Uptime Kuma wiki | 07 → Monitoring & DNS | |
| ntfy documentation | 07 → Monitoring & DNS | |
| GitOps & infrastructure-as-code | ||
| Ansible getting started | 07 → GitOps foundation | |
| Cloudflare Terraform provider | 07 → GitOps foundation | |
| Collabora Online Docker | 07 → Phone & mobile backup | |
| Homelab as Code — Merox | 07 → GitOps foundation | |
| Packer documentation | 07 → GitOps foundation | |
| bpg/proxmox Terraform provider | 07 → GitOps foundation | |
| homelab-as-code | 07 → GitOps foundation | |
| Security audit tools | ||
| AUDIT → HARDEN → SEGMENT → RECOVER | 14 Security audit | |
| Have I Been Pwned | 14 Security audit | |
| canyouseeme.org | 14 Security audit | |
| homelab-hardening-checklists | 14 Security audit | |
| homelabstarter network security audit | 14 Security audit | |
| shodan.io | 14 Security audit | |
| External design references | ||
| 5 Home Lab VLAN Mistakes | 07 → Kubernetes (Talos); 07 → MikroTik switch | |
| Jim's Garage | 09 Kubernetes/GitOps/AI | |
| VirtualizationHowto lab | 09 Kubernetes/GitOps/AI | |
| Other | ||
| Rclone documentation | 07 → Backups | |
15Changelog
Add a row here every time hardware, VLANs, or major config actually changes. Newest on top.
/diagrams/* path is deliberately exempted from that gate: initial research proposed relying on Cloudflare's authentication cookie to also cover the diagram embeds, but this was checked against Cloudflare's own community forum before being implemented and found to be a documented failure mode — the cookie does not reliably reach a resource loaded inside a cross-origin iframe, which is exactly how the diagrams.net viewer fetches each file. Gating that path would either break every diagram or require real engineering effort with no guaranteed outcome, so it stays public by explicit, stated design — the diagram files contain no credentials, only architecture information already fully described in the gated text. Also corrected the diagram embed URL format itself: verified against three independent official draw.io sources that the target file belongs after a #U hash fragment, not a ?url= query parameter as the previous version used — the earlier format would not have rendered any diagram at all. Added a Gemfile so Cloudflare's build environment installs Jekyll correctly, and rewrote Section 16 and the accompanying HOW-TO-RUN-THIS-SITE.txt to document the actual current setup rather than the superseded GitHub-Pages-only instructions.
.drawio file via the public diagrams.net viewer — the paired .svg export files (16 of them, added in the previous diagram-editability pass) are deleted entirely. Editing the .drawio now updates the page with no export step, ever. The trade-off is stated directly in script.js where the base URL lives: this only renders once the page is actually hosted (GitHub Pages already covers this) — it will not render when index.html is opened locally by double-click, since the viewer is a remote service fetching the file over the internet. The Remote Access diagram's two custom "toggle path" buttons — broken by this change, since JavaScript can no longer reach inside a cross-origin viewer iframe — are removed; the diagram was rebuilt with two real draw.io layers instead, toggled natively from the viewer's own layers panel. Second: the single 3,553-line index.html is split into a 46-line Jekyll shell plus 21 _includes/ files (one per section, with the largest section — Configuration Steps — further split into intro/Group A/Group B/Group C so no single file exceeds ~900 lines). GitHub Pages runs Jekyll automatically; no build step was added. Verified by writing an include-simulator and diffing its output against the pre-split original: every difference was whitespace or a restored organisational comment, confirmed by a full tag-balance and tab-wiring check against the reconstructed page, not just the source files.
architecture.png (client → edge → focal distribution point → grouped compute/storage zone), applied to OPNsense, the MikroTik switch, Proxmox, and the NAS. Both are self-contained SVG with proper title/desc accessibility metadata per that system's contract. Scope note: this covers the two diagrams requested; the remaining 13 existing diagrams remain in the document's original visual language and were not converted in this pass.
deny-all default), not an obscure-topic-name-as-password pattern. Added Beszel for multi-host system monitoring, with the critique's own point stated directly: it adds little on a single host, but this build already spans Proxmox, several LXCs, and a multi-node Talos cluster. Addressed the rest of the critique honestly rather than silently: sharpened the NetBird step to state plainly that Tailscale is a WireGuard implementation, not an alternative to it; added tradeoff callouts on NPM-vs-Traefik, Homarr-vs-Homepage, and Authentik-vs-Authelia rather than pretending the criticism doesn't apply.
lab.internal naming domain was a plausible, wrong answer that both source articles independently identify as the top cause of clusters that show "Connected" and then flip to "Unreachable." Added a dedicated new step (3b) covering the fix (a distinct, unused cluster domain like dnscluster.local; never cluster.local, which collides with Kubernetes), DNS_SERVER_DOMAIN consistency, one-way replication direction, and a forward-looking callout on the specific Kubernetes traps (VIP binding, TLS passthrough on 53443, pod-CIDR zone-transfer allow-listing) that only apply if Technitium is ever moved into the Talos cluster later.
personal-devices access group (Cloud + Photos VLANs only) narrower than the admin group, so the phone gets exactly what it needs and nothing more.
.env/compose secrets and open listening sockets to AUDIT, an immutable-backup-copy check to RECOVER, and a lightweight asset hardening log table (device/IP/MAC/who/when) to track what's actually been hardened over time. NanoKVM tab: added a BIOS password + disabled external boot as a second gate behind the KVM access path itself. Consciously did not adopt one piece of advice from the dev.to source (blanket-disabling ICMP) — that's dated guidance that breaks path MTU discovery and diagnostics for little real benefit; noting the omission rather than silently skipping it.
network_mode: host, wt0 interface). Added optional Packer golden-image step to the GitOps foundation tab. Added a new Home Automation services tab — Home Assistant (VM, USB passthrough for Zigbee/Z-Wave/Bluetooth) and Homebridge (LXC, HomeKit bridge) — plus a Homebridge chip in the logical diagram. Fixed: switch management IP was previously on the isolated NAS VLAN.
10.0.99.254) — moved to the Management VLAN (10.0.1.5) where it belongs.
.changelog-row block, bump the version, set today's date, describe what changed. Keep newest at the top.
16Hosting this page & keeping it updated
Served from Cloudflare Pages, built from a private GitHub repository, gated by Cloudflare Access — real authentication on the page content, not just an unlisted link.
Why not plain GitHub Pages
GitHub Pages served from a public repository has no access control at all — anyone with the URL sees everything, and the repository itself is browsable. GitHub Pages from a private repository requires a paid GitHub plan and still does not provide simple per-person link sharing. Cloudflare Pages was chosen specifically because it builds directly from a private repository on the free tier, and pairs natively with Cloudflare Access for authentication — the same access-control product already used for the Cloudflare Tunnel backup path in Section 08.
One-time setup
- Make the GitHub repository private (Settings → General → Danger Zone → Change visibility), if it is not already. This alone is what stops anyone but explicitly-added collaborators from editing anything — a repository permissions question, independent of hosting.
- Create a free Cloudflare account if one does not already exist, and add a zone (a domain — Cloudflare offers free subdomains for testing, or use any domain already owned).
- Cloudflare dashboard → Workers & Pages → Create → Pages → Connect to Git, authorise Cloudflare's GitHub App, and select the private repository.
- Build settings: framework preset Jekyll, build command
bundle exec jekyll build, output directory_site. TheGemfileat the repository root tells Cloudflare's build image to install Jekyll — no further Ruby configuration needed. - Deploy. Cloudflare assigns a
*.pages.devURL immediately; a custom domain can be attached afterward under the project's Custom domains tab if wanted. - Update
DIAGRAMS_BASE_URLinscript.jsto that URL, commit, push — this is the one line the whole live-diagram system depends on (Section 06).
Cloudflare Access — gating the page
- Zero Trust dashboard → Access → Applications → Add an application → Self-hosted.
- Domain: the Pages project's
*.pages.devURL (or the custom domain, once attached). - Policy: allow rule listing specific email addresses — every person who should be able to view this, and no one else. One-time email code is the simplest identity provider for a personal project; no separate account creation needed for anyone on the list.
- Add a second, separate application scoped to the path
/diagrams/*, policy: Bypass — public, no login required for that path specifically.
/diagrams/* is deliberately excluded from the login requirement: the diagrams render via a third-party viewer (viewer.diagrams.net) loading each file inside a cross-origin iframe. Cloudflare's own authentication cookie does not reliably reach a resource fetched from inside a cross-origin iframe — a documented browser cookie-scoping limitation, not a configuration mistake — so attempting to gate this path either breaks every diagram embed or requires real engineering effort with no guaranteed outcome. The trade-off actually taken: every paragraph, table, and configuration step on this page requires login; the diagram XML files themselves are reachable only by someone with the exact file URL, which is linked nowhere public and contains no credentials — only architecture information already fully described in the gated text regardless.
Day-to-day editing
- Prose: edit the relevant file directly under
_includes/sections/(seeHOW-TO-RUN-THIS-SITE.txtfor which file covers which section), commit, push. Cloudflare rebuilds automatically within about a minute. - Diagrams: open the
.drawiofile directly in app.diagrams.net, edit, save. No export step — the live embed reads the same file. - Both routes require push access to the private repository — the same permission boundary that keeps editing restricted, regardless of who can view the published page.
index.html, _config.yml, and Gemfile at the root, _includes/ and diagrams/ alongside them — exactly the structure already in this repository.