Status: rebuilding from scratch

WarGreymon

A home server and network infrastructure design, named for WarGreymon — the Mega-level evolution of Agumon in the Digimon series, a defence-oriented form noted for its shield and its role protecting the Digital World. This document is the design reference: hardware, rack layout, network architecture, service interconnection, configuration procedure, and security posture for the system.

19" · 12RU wall cabinet — owned APC SMC1500I-2UC UPS — owned 10GbE SFP+ DAC cables — owned NanoKVM out-of-band mgmt — owned NetBird remote access Dashboard + security monitors — planned

00Quick start — build order

Build the network before you build on the network. Everything downstream needs a gateway, an address, and a stable link before it can be configured — follow this order and each device is reachable, patched, and on the right subnet by the time you get to it. Tick boxes as you go; this is the "what's next" page — every step links to the full detail elsewhere on this page.

How to use this: work top to bottom, don't skip ahead. Each phase below is collapsed — expand the one you're on, leave the rest closed so this stays a fast "where am I" glance rather than another wall of text.
00 · Plan before you plug anything in15–30 min — cheapest mistakes are the ones fixed on paper
  • Topology confirmed: Modem → OPNsense → MikroTik switch → NAS / Proxmox / everything else
  • Subnet plan is set: 10.0.X.0/24 per VLAN, 15 VLANs already designed in Section 05 (10 Mgmt, 15 Container, 20 Media, 25 Indexing, 30 Cloud, 35 Trading, 40 Photos, 45 Cameras, 50 Monitor, 60 DNS, 70 VPN, 80 Database, 90 Testing, 100 Home Devices, 999 NAS)
  • Static/reserved addresses noted: OPNsense LAN 10.0.1.1, switch mgmt 10.0.1.5, NAS 10.0.30.10 + 10.0.1.100, NanoKVM reservation, Technitium dns01/dns02
  • Cable + PDU labelling convention set — SW-SFP1-OPN style, per Section 03
01 · Modem / ISP handoffConfirm raw internet before anything else joins the chain
  • Laptop connected directly to the modem, confirmed it gets a public IP / internet access
  • Connection type noted: DHCP, PPPoE, static, or VLAN-tagged (matters for OPNsense's WAN config next)
  • Laptop disconnected — OPNsense's em1 (WAN) takes this port next
02 · OPNsense — before the switchEverything after this already has a gateway waiting for it

Why not the switch first: a switch without a firewall behind it just forwards packets to a dead end — nothing gets an IP or a route out.

03 · MikroTik switchExtend the gateway cleanly to every port
  • Switch uplink (SFP+1) → OPNsense SFP+1
  • RouterOS firmware updated
  • Switch given a static management IP on VLAN 10 (10.0.1.5) — not on the isolated NAS VLAN
  • All 15 VLANs created to match the plan, trunk vs access ports set — Section 07 → Hardware → MikroTik switch
  • Test device on a port confirmed reaching the internet
04 · NASConfigured once the network under it is stable — moving a NAS to a new subnet later is disruptive
  • NAS NIC1 + NIC2 → switch ports 4 & 5
  • TOS firmware updated before creating any storage pool
  • Static/reserved IPs set on the correct VLANs (10.0.30.10 primary, 10.0.1.100 secondary)
  • RAID 5 array created — Section 07 → Hardware → NAS, Step 2
  • Default admin credentials changed, unused default services disabled
  • Only the shares actually needed are enabled (NFS to Proxmox, nothing extra left on)
05 · Other hosts — Proxmox, NanoKVMConnect, address, patch, harden — before exposing any service on top
  • Proxmox (MS-01) connected to the switch, correct VLAN trunk confirmed — Section 07 → Hardware → Proxmox
  • Static IP / reservation assigned per the plan
  • OS patched, real credentials set before any service is exposed
  • NanoKVM connected and secured — Section 07 → Hardware → NanoKVM
  • Everything in Section 09 (Kubernetes, AI, Trading) waits until this phase is fully stable — see the resource budget there before starting it
06 · Wireless access pointNot part of the current hardware list — skip unless/until one is added
  • N/A for now — no dedicated AP in the current build. If one is added later: connect it to a trunk port, map each SSID to a VLAN, same pattern as everything else here
07 · End-to-end testingProve the whole chain works before trusting it
08 · HardeningFull detail in Section 07
  • Every device audited for a changed default password — OPNsense, switch, NAS, NanoKVM
  • NetBird preferred over any port-forward for remote access — Section 08
  • Automatic security updates enabled where safe to do so
  • UPS installed and confirmed on OPNsense, switch, and NAS at minimum — Section 04
09 · Backups & monitoring — last, on purposeNeeds everything above already stable to back up or monitor reliably
  • Proxmox Backup Server jobs scheduled — Section 07 → Services → Backups
  • Uptime Kuma monitoring OPNsense, switch, and NAS at minimum
  • IPs, VLANs, and credentials documented somewhere outside the lab itself (this repo, plus a password manager for secrets — never commit credentials to git)

02Hardware inventory

Cabinet, UPS, and the 10GbE cables are now in hand. One optional addition below for easier remote hardware management.

RoleDeviceForm factorStatus
EnclosureTecmojo 12RU wall-mount cabinet, 450mm deep, 50kg rated, lockable glass door19" cabinetOwned
UPSAPC Smart-UPS C SMC1500I-2UC — 1500VA/900W, line-interactive, 4× IEC C13 outlets2U rackmount, 439mm deep, 24.1kgOwned
Firewall / routerOPNsense on SJRC N150 mini-ITX box — 2× 2.5GbE + 2× 10GbE SFP+Desktop, shelf-mountOwned
HypervisorProxmox VE on MS-01 (i9‑13900K, 32GB RAM, 2TB NVMe) — incl. 10GbE SFP+Desktop, shelf-mountOwned
SwitchMikroTik CRS310‑8G+2S+in — 8×1GbE + 2×10GbE SFP+1U, 483×88×288mmOwned
Storage / NASTerramaster F4‑424 Pro, 4‑bay, RAID 5, ~4TB usableDesktop, shelf-mount, ~18kg loadedOwned
Power distribution6-outlet IEC PDU1U rackmountOwned
10GbE uplinks2× 10GbE DAC Twinax SFP+ cables, 0.5mOwned — ready to install
Out-of-band mgmtSipeed NanoKVM — KVM-over-IP, HDMI capture + USB device emulation, remote power controlUSB-C powered, 1GbE, 1080p60 captureOwned — cables in hand
Home devices switchNetgear ProSAFE Plus JGS524E — 24-port gigabit, 802.1Q VLAN tagging, satellite switch for household devicesDesktop, 330×173×43mm, 13.5W maxOwned
Dashboard displayWall/desk monitor running a browser kiosk pointed at Grafana/HomarrHDMI, needs a small feeder devicePlanned — future
Security camera displaySecond monitor for a future NVR (Frigate/ZoneMinder) camera wallHDMI, needs a small feeder devicePlanned — future
Weight check (unchanged): switch + NAS + UPS + PDU + two mini-ITX boxes ≈ ~49.6kg against the cabinet's 50kg rating. NanoKVM adds negligible weight (<100g), doesn't move the needle. The JGS524E does not affect this figure at all — it is not rack-mounted; see below. Weigh the real boxes on arrival and don't add extra gear into this cabinet without re-checking.

JGS524E — role and physical placement

This is a Netgear ProSAFE Plus-series switch: 24 gigabit copper ports, a limited web-based/Windows-utility configuration interface (not RouterOS-class), supporting 802.1Q tagged VLANs, QoS, port mirroring, and broadcast storm control. It is not a rack-mount unit — at 330mm wide it is roughly two-thirds the width of a 19" rack — and functions as a satellite switch located wherever the household devices actually are (a media cabinet, a desk), rather than inside the enclosure documented in Section 03. A single Cat6 run carries a trunk link back to the MikroTik CRS310's port 7. Power draw is 13.5W maximum; since it sits outside the rack, it is not protected by the UPS in Section 04 unless a separate small UPS is added at its location, which is optional and not required for the lab itself to function.

Its function in this design is a single-purpose one: everything plugged into it — a television, a personal computer, other household devices — reaches the internet through the same OPNsense firewall, Technitium DNS (and its ad-blocking), and Suricata/CrowdSec inspection as every other device in the design, on a VLAN (100) that has no visibility into and no access from any lab VLAN. The remaining 23 ports and the switch's own configuration headroom leave capacity for future expansion — additional household devices, or a physical link to a second lab location, without any change to the design documented here.

NanoKVM — what it actually gives you

Three things bundled into one small device: video (HDMI capture, so you see exactly what a monitor would show — BIOS, boot menus, a crashed console), input (USB-C device emulation, so it behaves as a real keyboard/mouse to the host, no agent software needed on the host side), and optionally power control via the ATX breakout board if the host exposes accessible power/reset header pins. Combined with NetBird for the network path in, this is what makes "fix it from my phone while not at home" actually true — without NetBird it's still LAN-only, and without NanoKVM you still need physical access whenever the OS itself won't boot.

MS-01 caveat: the ATX breakout board needs an accessible power/reset header on the host motherboard. The Proxmox MS-01 is a sealed mini-PC, not a bare motherboard in a case — check Minisforum's documentation for whether it exposes header pins before assuming remote power-cycling will work. If it doesn't, you still get full video + keyboard/mouse control, just not a remote power button; a UPS-controlled smart outlet is a reasonable fallback for hard power cycling.

03Physical rack layout

Bottom-to-top: heaviest gear at the bottom. Cable routing between these devices is documented separately in the network topology diagram (Section 05) rather than overlaid here, so this stays a clean read of what occupies which U position.

Front elevation. The switch is called out as the focal distribution point, consistent with the topology diagram in Section 05. The JGS524E is not rack-mounted and lives at its own satellite location.

Cable discipline

  • 1GbE cables (switch ↔ OPNsense port 1, switch ↔ Proxmox port 3, switch ↔ NAS ×2) labelled both ends — kept as fallback
  • 10GbE DAC cables (OPNsense SFP+1 ↔ switch SFP+1, switch SFP+2 ↔ Proxmox SFP+1) labelled SW-SFP1-OPN / SW-SFP2-PRX — primary VLAN trunk path, already in hand
  • Power cables labelled to match PDU outlet number, e.g. PDU-2-PRX
  • Network and power cables bundled and routed on opposite sides — never crossed
  • DAC cables handled gently — no sharp bends
  • NanoKVM: HDMI-in + USB-C emulation cable to Proxmox, Ethernet to switch port 6 (VLAN10), USB-C aux power from a spare Proxmox USB port — velcro or 3M tape it to the shelf, no dedicated RU needed
  • Switch port 2 reserved and labelled as the admin laptop drop — run it out to wherever you'll actually sit with a laptop, not just to a wall plate nobody uses
  • Every cable tug-tested before first power-on

04Power & UPS

Goal: if mains power drops, every device gets enough runtime to either ride out a short blip or shut down cleanly.

Idle draw

~50W
UPS standby + switch + NAS idle

Typical draw

~160W
all services running

Peak draw

~277W
worst case, all devices busy

APC Smart-UPS C SMC1500I-2UC — owned

SpecValue
Capacity1500VA / 900W, line-interactive, AVR
Outlets4× IEC C13 (battery-backed)
Form factor2U rackmount, 439mm deep, 24.1kg
ManagementUSB + SmartConnect port, graphic LCD status panel
Est. runtime~25–35 min @ typical 160W, ~8–10 min @ 277W peak

30% headroom over peak draw — sized for "shut everything down safely," not multi-hour outages. Right trade-off for a home lab.

Power path

Wall outlet → UPS input. UPS output → PDU (RU10) → switch, OPNsense, Proxmox, NAS individually. Never plug the UPS itself into the PDU it's powering.

Easy power-efficiency win on the MS-01: Core Performance Boost (Intel Turbo Boost) can be toggled in BIOS independently of everything else in this build. Disabling it trades peak single-core burst speed for meaningfully lower sustained power draw — reportedly close to half, in some mini-PC reports — with little practical impact on a host that's mostly running steady background services rather than bursty single-threaded workloads. Worth testing once the lab is stable: disable it, watch the UPS's reported load for a week, and decide whether the trade-off is worth it for your actual usage pattern.

Automated shutdown

Install and configure NUT (Network UPS Tools) on Proxmox
  • Connect the UPS to Proxmox via USB
  • Install: apt install nut
  • Set /etc/nut/ups.conf driver to usbhid-ups for APC Smart-UPS C series
  • Configure upsmon.conf with a shutdown threshold (e.g. trigger at 20% battery or 5 min remaining)
  • Test with a real unplug — confirm Proxmox begins an orderly VM shutdown before the battery is exhausted
  • Expose UPS status to Grafana via the NUT exporter for Prometheus

05Network architecture

One firewall, one trunked switch, fifteen isolated VLANs, on a 10GbE backbone.

Correction from the reference diagram you provided: it described Proxmox using a 20GbE LACP bond across both switch SFP+ ports. The CRS310 only has two SFP+ ports total, and one of them is needed for OPNsense's trunk link — there aren't three SFP+ ports available, so a 2-port bond on Proxmox alone isn't physically possible with this switch. The design below keeps the two real 10GbE links: SFP+1 → OPNsense and SFP+2 → Proxmox, each a single 10Gbps trunk. If you want true LACP redundancy later, that requires a switch with more SFP+ ports (see Section 08).
15 VLANs isn't sprawl, checked against the actual test: the usual home-lab VLAN mistake is creating networks without a reason — more segments than anyone can document or reason about. Every VLAN here exists because a specific pair of things needed to be kept apart (Trading from everything, Cameras from everything including each other's blast radius, Photos from Media, NAS from all compute). None were added "because more VLANs seems more secure" — each has a one-line justification in the table below, and that's the actual bar worth holding this design to over time, not a target VLAN count.

Logical network diagram — VLANs & services

Each lane is a VLAN — colour marks its trust tier: teal = broadly trusted, amber = scoped, red = isolated, grey = internal-only

Physical & core network topology

Client through to every core device, with the switch identified as the system's focal distribution point — every path to compute or storage passes through it. Remote access (NetBird, Cloudflare) is a separate diagram below, so this one stays about the physical/logical core, not VPNs.

The JGS524E's connection is drawn routing around the lab zone rather than through it, matching its actual isolation — it has no path into any lab VLAN.

Remote access (VPN) diagram

Two independent paths in, drawn as two independent lanes — nothing here shares a line with the physical diagram above, on purpose. NetBird is primary; Cloudflare Tunnel is the architecturally unrelated backup. The diagram below is a live embed of the real .drawio file — each lane is a real layer, and the small layers icon in the viewer's own toolbar (bottom-left of the embed) shows or hides NetBird and Cloudflare Tunnel independently, natively, without any custom button on this page.

NetBird (left): full mesh, scoped by access group. Cloudflare Tunnel (right): single services only, gated by a login before the request ever reaches the tunnel.

VLANs

VLANNameGatewayPurpose
10Management10.0.1.1Proxmox, switch, OPNsense admin, dashboard, reverse proxy, backups, NanoKVM, laptop drop, Cloudflare Tunnel, Authentik, NetBox
15Container10.0.15.1Docker, Talos K8s cluster, Argo CD, Traefik, Ollama + Open WebUI
20Media10.0.20.1Jellyfin, Sonarr, Radarr, Jellyseerr, Bazarr
25Indexing10.0.25.1Prowlarr
30Cloud10.0.30.1Nextcloud (contacts/calendar/files), Vaultwarden — phone backup target
35Trading10.0.35.1MetaTrader 5 (Windows VM) — isolated, outbound-only to broker + NTP + DNS
40Photos10.0.40.1Immich — isolated, phone photo/video backup target
45Cameras10.0.45.1Frigate NVR (planned) — switch-local, no lateral/internet access at all
50Monitor10.0.50.1Prometheus, Grafana, InfluxDB, Uptime Kuma, Loki, Wazuh, ntfy, Beszel
60DNS10.0.60.1Technitium DNS (primary + secondary, clustered), Unbound (recursive resolver)
70VPN10.0.70.1NetBird — self-hosted mesh VPN, remote access
80Database10.0.80.1PostgreSQL, MariaDB
90Testing10.0.90.1Home Assistant, Homebridge, Gitea, Paperless-ngx, sandbox VMs, Flatcar test host
100Home Devices10.0.100.1TV, personal PC, and other household devices — internet + DNS ad-blocking only, isolated from every lab VLAN
999NAS10.0.99.1Storage only — fully isolated

Switch port map

PortConnects toModeSpeed
SFP+ 1OPNsense SFP+ 1Trunk — all VLANs tagged10Gbps — primary
SFP+ 2Proxmox SFP+ 1Trunk — all VLANs tagged10Gbps — primary
1OPNsense em0 (2.5GbE)Trunk — all VLANs tagged1Gbps — fallback
3Proxmox eth0Trunk — all VLANs tagged1Gbps — fallback
4NAS NIC1Access — VLAN 30 untagged1Gbps
5NAS NIC2Access — VLAN 10 untagged1Gbps
2Admin laptop drop (room patch point)Access — VLAN 10 untagged, occasional1Gbps
6NanoKVMAccess — VLAN 10 untagged1Gbps
7JGS524E home switch uplinkTrunk — VLAN 10 + 100 tagged only1Gbps
8Reserved

VLAN traffic matrix

From \ To10152025303540455060708090100999
10 Mgmt
15 Container
20 Media
25 Indexing
30 Cloud
35 Trading
40 Photos
45 Cameras
50 Monitor
60 DNS
70 VPN
80 Database
90 Testing
100 Home Devices
999 NAS

Network boundary diagram — what crosses the edge

Four ways traffic can cross the OPNsense boundary, and only four. Direct inbound has nothing to land on — there are no port-forwards to receive it.

The only two paths in are both outbound-initiated from inside the lab — nothing on the internet can open a connection to this network unprompted.

Data flow diagram — how information actually moves

Three representative journeys through the stack, left to right.

Media requests, remote access, and DNS resolution — the three flows that touch the most services in this build.

VLAN zone map

The same 15 VLANs as the swimlane diagram above, drawn instead as trust zones radiating from the firewall — every line is a boundary OPNsense enforces, solid where broadly reachable, dashed where isolated.

broad access scoped isolated (solid line = still monitored/testing-reachable) isolated (dashed = no lateral access)
A shape-at-a-glance view — for exact rules, see the traffic matrix above; for what's running where, see the swimlane diagram.

Switch port diagram

The MikroTik CRS310's actual front panel, port by port — what to plug into what while cabling the rack (Section 03).

Matches the switch port map table above exactly — this is the same information, drawn as the physical panel you're looking at while cabling.

06Service interconnection reference

Earlier sections describe how each service is deployed in isolation. This section describes how they operate together: four architectural patterns underpinning the system, and a consolidated reference table listing every service's purpose, internal connections, external network access, and authentication method.

Service discovery

Every service is addressed by name through Technitium DNS zones rather than a memorised IP address. Portainer additionally exposes container-to-container dependencies within a single Docker host. Name resolution is a one-time lookup; once resolved, services communicate directly and Technitium is no longer part of the data path.

Unified authentication

Authentik is positioned in front of proxied services as a forward-auth layer: a single login and MFA challenge cover every gated service, and every access attempt is logged in one place rather than scattered across separate per-service login systems.

Event-driven messaging

Mosquitto (MQTT) carries IoT device communication for Home Assistant. Sensors and switches publish state changes to topics; subscribers react to those topics without addressing the publisher directly. A dedicated message bus such as NATS is not required at the current service count, since inter-service communication elsewhere in the system runs over standard REST APIs.

Notification aggregation

ntfy is the single delivery path for every alert source — Uptime Kuma, Grafana, Wazuh. A plain SMTP relay serves as a secondary channel. A full mail-gateway appliance is not an appropriate substitute for this purpose, since that class of tool is built for spam filtering on a mail server, not for delivering occasional alert messages.

Service discovery — how a name becomes a connection

Every service in the system is reachable by a stable DNS name (service.lab.internal) rather than a memorised IP address, resolved by the internal Technitium DNS pair described in Section 05. Name resolution happens once per connection: a client asks Technitium for an address, receives it, and then communicates directly with the destination for the remainder of that session. This has two practical consequences. First, an IP address can change — a VM gets rebuilt, a container moves to a different host — without updating every service that talks to it, provided the DNS record is updated once. Second, Technitium sits outside the actual data path: it is consulted at the start of a connection, not on every packet, so it introduces negligible latency and is not a single point of failure for already-established connections if it becomes briefly unavailable.

The diagram below traces a concrete example: Jellyseerr resolving and reaching Radarr. The same three-step pattern — query, resolve, connect directly — applies to every internal service-to-service connection in the system.

Purpose: illustrates that DNS resolution is a discrete, one-time step preceding a connection, not an intermediary the connection continues to depend on.

End-to-end data flow across the homelab

The service-discovery diagram above traces one hop — a name becoming a connection. This diagram traces the complete path a request takes through the system, role by role: from the requesting device, through the edge (DNS resolution, authentication, reverse proxy), into the application service that handles it, and finally to where the resulting data is persisted.

Purpose: shows who does what at each stage of a request — not just which services exist, but the order data actually moves through them, and where the one authenticated handoff into the application layer occurs.

Event-driven messaging — MQTT for IoT

Mosquitto is an MQTT message broker. IoT devices — Zigbee sensors via a Zigbee2MQTT bridge, DIY ESPHome sensors, Tasmota-flashed smart plugs — publish state changes to named topics on the broker (for example, home/livingroom/temperature) rather than being polled for their current state. Home Assistant, and any other interested service, subscribes to the relevant topics and receives updates the moment they are published. Publishers and subscribers never address each other directly; both only address the broker. This decouples the two sides: a new sensor can be added by subscribing Home Assistant to a new topic, without reconfiguring any existing device, and a battery-powered sensor can announce a change once and return to a low-power state rather than responding to repeated polling requests.

Mosquitto is required only once physical IoT devices using MQTT — Zigbee, Z-Wave via a bridge, or DIY sensors — are present on the network. A Home Assistant instance with no such devices connected has no MQTT traffic to broker and does not require it.

Purpose: illustrates the publish/subscribe pattern — publishers and subscribers address only the broker, never each other. Adding a new sensor requires subscribing Home Assistant to a new topic; it does not require reconfiguring any existing device.
Deploy Mosquitto (MQTT broker)
Intermediate~25 min
Rationale: IoT devices are frequently battery-powered and communicate in short bursts. Polling such devices over REST at regular intervals drains battery life and adds latency. MQTT's publish/subscribe model allows a device to announce a state change once and return to a low-power state; any subscriber receives the update immediately, with substantially less overhead than repeated polling.
  1. Deploy Mosquitto on VLAN 90, alongside Home Assistant.
  2. Enable authentication (username/password at minimum) — an open MQTT broker lets anyone on the VLAN publish fake sensor data or watch real device state.
  3. Connect Home Assistant's MQTT integration to it, confirm a test topic round-trips.
  4. Point Zigbee2MQTT or ESPHome devices at the same broker as they're added.
  • Mosquitto deployed with authentication enabled, not anonymous access
  • Home Assistant MQTT integration connected and tested

Current software versions

The version of every operating system and major platform component in the design, current as of the last review of this document. Software in active development moves continuously; this table is a reference point, not a pin. Each version should be re-confirmed against the project's own release page immediately before installation, and the deployed version recorded in NetBox (Section 07 → Group B → Proxy & dashboard) or the asset log (Section 14) once installed, rather than assumed to match this table indefinitely.

ComponentCurrent stable versionBase
Proxmox VE9.2Debian 13 "Trixie"
Proxmox Backup Server4.xDebian 13
OPNsense26.7 "Xenial Xenops"FreeBSD 15.1
MikroTik RouterOS7.x (Stable channel)
Talos Linux1.13.x
Flatcar LinuxCurrent stable channel
Technitium DNS Server15.4.x.NET 10
Debian (Docker/LXC hosts)13 "Trixie"
TrueNAS Community Edition25.10 "Goldeye"Debian Linux

TrueNAS is included for reference only — the NAS in this design runs Terramaster's own operating system (TOS), not TrueNAS, per Section 02. A fuller comparison of the two, including the ZFS data-integrity distinction, is in Section 07 → Group A → NAS, Step 1.

Application-layer services (Jellyfin, Nextcloud, Authentik, and the remainder of the master table below) are not tracked by version in this document individually — Watchtower (Section 07 → Group B → Proxy & dashboard) keeps container images current on an ongoing schedule, which is a better fit for a service catalogue this size than a manually maintained version table that goes stale between reviews.

Master service interconnection table

Every deployed service, what it's for, what it talks to, whether it reaches the internet, and how it authenticates. Cross-reference against the traffic matrix above for the exact firewall rule each connection depends on.

ServiceVLANPurposeTalks to (internal)Internet accessAuth
Proxmox10HypervisorPBS, NAS (NFS), every VM/LXCPackage updates onlyLocal + Authentik (RAC, optional)
PBS10Backup targetProxmoxNoneLocal
OPNsense10Firewall/routerEverything (routes all VLANs)WAN, updatesLocal + MFA
MikroTik switch10VLAN trunkAll wired devicesNoneLocal
Homarr10DashboardReads status from most servicesNoneAuthentik (forward-auth)
NGINX Proxy Mgr10Reverse proxyEvery proxied serviceLet's Encrypt DNS challengeLocal (admin only)
NanoKVM10Out-of-band mgmtProxmox (HDMI/USB, not IP)NoneLocal + NetBird-only
Cloudflare Tunnel10Backup remote accessNPMOutbound to Cloudflare edgeCloudflare Access
Authentik10SSO / forward-authOwn Postgres+Redis, every gated serviceNone requiredSelf (MFA)
NetBox10Infra documentationRead-only reference, no live depsNoneAuthentik
Portainer15Container mgmt / discoveryDocker socket on its hostNoneLocal + Authentik
Talos K8s cluster15Container orchestrationArgo CD, Traefik, Longhorn, NAS (storage)Image pulls, Argo CD git syncKubernetes RBAC
Argo CD15GitOps syncGit repo, Talos APIGit repo (outbound HTTPS)Authentik (OIDC)
Traefik15K8s ingressEvery K8s-hosted serviceNone requiredPassthrough to backend auth
Ollama + Open WebUI15Local AIStandalone; optional NPM proxyModel downloads onlyAuthentik (forward-auth)
Jellyfin20Media serverNAS (NFS media share)Metadata lookupsLocal + Authentik
Sonarr / Radarr20Media automationProwlarr (VLAN 25), qBittorrent, Jellyfin, NASIndexer APIs (via Prowlarr)Local, API key
Jellyseerr20Request front-endJellyfin, Sonarr, RadarrNone requiredJellyfin account / Authentik
Bazarr20SubtitlesSonarr, RadarrSubtitle providersLocal, API key
Prowlarr25Indexer managementSonarr, Radarr (push config)Indexer sitesLocal, API key
Nextcloud30Files/contacts/calendarPostgreSQL (VLAN 80), NASApp store, federation (optional)Local + Authentik (OIDC)
Vaultwarden30Password managerStandaloneNone requiredSelf (master password + MFA)
MetaTrader 535Trading terminalNone internal — fully isolatedBroker servers, NTP onlyLocal Windows + broker login
Immich40Photo/video backupOwn Postgres, NAS storageNone requiredSelf, mobile app token
Frigate (planned)45Camera NVRCameras (RTSP, same VLAN only)NoneLocal
Prometheus / Grafana50MetricsScrapes Proxmox, OPNsense, SNMP, most servicesCommunity dashboard importsAuthentik (forward-auth)
Uptime Kuma50Uptime monitoringPings every monitored servicentfy webhookLocal + Authentik
Wazuh50SIEM / loggingAgents on every host, OPNsense syslog, Authentik eventsThreat intel feeds (optional)Local + Authentik
ntfy50NotificationsReceives from Uptime Kuma, Wazuh, NUT, MT5 monitorNone requiredPer-topic tokens, deny-all default
Beszel / Pulse50Host monitoringAgents on Proxmox, LXCs, Talos nodesNoneLocal
Technitium ×2 + Unbound60DNSAnswers every VLAN's queries; Technitium→Unbound→root serversRoot/TLD servers (Unbound only)Local (admin)
NetBird70Remote access (primary)Enrolled peers directly (mesh)Coordination server onlySelf (SSO/OIDC)
PostgreSQL / MariaDB80Shared databasesNextcloud, Authentik, Immich, Grafana (as configured)NonePer-service DB credentials
Home Assistant90Home automationMosquitto, Homebridge, IoT devicesIntegration cloud APIs (optional)Local + Authentik
Mosquitto90MQTT brokerHome Assistant, Zigbee2MQTT/ESPHome devicesNoneUsername/password
Homebridge90HomeKit bridgeIndividual smart-home device APIsDevice cloud APIs (as needed)Apple Home pairing code
Gitea90Git hostingArgo CD (pulls manifests)None requiredLocal + Authentik (OIDC)
Paperless-ngx90Document managementOwn PostgresNone requiredLocal + Authentik
NAS shares999Bulk storageProxmox, PBS, Jellyfin, Immich (all via NFS)None — fully isolatedPer-service scoped accounts
JGS524E + household devices100Household internet, ad-blocking, firewall inspectionNone internal — Technitium and WAN onlyFull outbound (WAN)Switch: local admin password. Devices: none required.

07Configuration steps

This section is written to teach, not just instruct — every step has a difficulty rating, a time estimate, a "why" explanation of the underlying concept, a numbered walkthrough, a checklist to confirm before moving on, and links to the real official docs for when you want more depth than fits here.

How to use this section: work through Hardware & Core Devices first, in tab order (OPNsense → MikroTik → NAS → Proxmox → NanoKVM) — that is the dependency order these devices need to come up in, matching Section 00. The NAS is configured before Proxmox specifically because Proxmox's own storage step (Group A → Proxmox → Step 4) mounts NAS storage over NFS, which requires the NAS to already exist. Services & Add-ons follows, once the network and hypervisor exist for services to run on. Every step is collapsed by default so the page stays scannable — expand only what is being worked on, or use "Expand all" to read a whole tab start to finish.

Virtualization strategy — how Proxmox, LXC, Docker, and Kubernetes fit together

Every deployment step in this section specifies one of four execution environments. This is the decision framework behind that choice, stated once rather than re-argued at every step.

LayerUsed forReason
LXC (on Proxmox)Single-purpose Linux services not tied to a Docker-specific ecosystem — Technitium, NPM, Homarr, Uptime Kuma, VaultwardenNear-native performance, no second kernel, trivial PBS snapshot/backup, lowest resource cost per service
Docker (inside a VM or LXC)Anything distributed primarily as a container image or Compose stack, and multi-container applications — the *arr stack, Immich, Nextcloud+database, AuthentikMatches how the upstream project actually ships and documents itself; fighting an upstream Docker Compose file into a native LXC install usually costs more effort than it saves
VM (on Proxmox)Anything needing its own full kernel or OS — OPNsense (FreeBSD), the MT5 Windows VM, each Talos nodeStrong isolation and a genuinely different OS are the only two reasons to pay a VM's overhead over an LXC in this design
Kubernetes (Talos)Automated scheduling, self-healing, and GitOps-managed rollout once manually tracking individual Compose stacks stops scalingAn orchestration layer above Docker, not a replacement for it — see Section 09

Proxmox is the layer underneath all four, and the reason to run it rather than any single layer alone: one backup system (PBS) covering VMs and LXCs alike, snapshots before any risky change, a live-migration path to the second node already planned in Section 11, and one place to see everything running rather than four disconnected tools.

Should Docker run on a separate box instead of inside Proxmox? Not for this build. Docker running inside a Proxmox VM or LXC has negligible overhead against bare metal for the workloads here — self-hosted web apps are not CPU-bound in a way virtualization meaningfully taxes. Moving Docker to its own bare-metal box would trade away PBS-unified backup, snapshotting, and the live-migration path for no real performance gain. A separate box becomes justified specifically when PCIe device passthrough is cleaner bare-metal than through Proxmox (uncommon at this stage), or when building toward a genuinely independent multi-node cluster for real high availability — which is exactly what the second Proxmox node in Section 11 already is. That planned second node is the answer to "more Docker capacity," not an unmanaged box outside this design.

What runs where — the complete breakdown

Every service named in the master table (Section 06) falls into exactly one of three categories. The deciding question for each is the same one asked in the table above, applied per service rather than left as a general principle: does the upstream project's own documentation assume Docker, does it run fine as a native package with no real benefit from containerising it, or does it need a genuinely separate OS?

Purpose: shows the actual placement decision made for every service category in this design, not just the abstract rule — including that VLAN 50's monitoring stack runs natively with no Docker host at all.

Docker host operating system

Every Docker host in this design — mgmt01, media01, idx01, cloud01, photos01, ai01, test01 — runs the same base: Debian 13 "Trixie", the same OS already established for every Docker/LXC host in Section 06's version table. This is a deliberate default, not an arbitrary one:

  • Matches Proxmox's own base OS — one set of apt habits, one patch cadence, no second Linux distribution's quirks to remember
  • Minimal install footprint, with Docker's official apt repository well-supported and documented for Debian specifically
  • glibc-based — some upstream images and their bundled binaries assume glibc and behave unpredictably or fail outright on musl-based distributions such as Alpine used as a host OS (Alpine as a container base image is a different, unrelated question and is fine)
  • Long, predictable support lifecycle — consistent with the patch-cadence expectations already set in Section 08

A dedicated container-optimised OS (Flatcar, Talos) is deliberately not used for these general Docker hosts, even though both already exist elsewhere in this design. Flatcar and Talos are immutable and API-managed by design — a good fit for a Kubernetes node that should never be hand-configured, and a poor fit for a host that still needs occasional native package installs, SSH-based troubleshooting, and ad hoc debugging the way these Docker hosts do.

How Docker actually runs inside an LXC

Docker needs kernel features an LXC does not expose by default. Getting this specific piece wrong is one of the most common home-lab stumbling blocks, so it is stated explicitly rather than assumed:

  1. Create the LXC as unprivileged — not privileged. Older guides frequently claim Docker needs a privileged container; current Proxmox versions do not require this, and running unprivileged keeps a container escape from directly compromising the Proxmox host, consistent with the least-privilege principle in Section 08.
  2. Enable the nesting feature on the LXC (Options → Features → nesting: 1) — this is the specific setting that allows a container runtime to run inside the container at all. Without it, the Docker daemon fails to start.
  3. Enable keyctl alongside nesting — several images rely on kernel keyring operations that are otherwise blocked in an unprivileged container.
  4. Use the overlay2 storage driver (Docker's default on modern kernels) — it works correctly inside an unprivileged nested LXC on the kernel versions Proxmox 9.x ships, with no special configuration required.

Once those three settings are in place, Docker installs and behaves inside the LXC exactly as it would on bare metal — the Compose file, the networking model in Section 10, and the image-trust checklist above all apply unchanged.

Group A — Hardware & core devices
1 · Physical connection & first boot
Beginner~20 min
Why: OPNsense is the box everything else trusts to decide what traffic is allowed. Getting the physical cabling right before touching any software means every later step is testing config, not chasing a loose cable.
  1. SFP+ 1 → Switch SFP+ 1 (10GbE, primary trunk). Optionally also connect em0 → Switch port 1 (1GbE fallback) — not required to boot, but handy from day one.
  2. em1 (WAN) → home router, using whatever cable is already running to your router.
  3. Power via PDU outlet, label PDU-3-OPN.
  4. Press power. Wait 30–60 seconds for boot — you'll see console text if a monitor is attached, or just wait if not.
  5. Default login is root / opnsense. Log in once, then immediately go set a real password — don't leave this step for later.
  • Power LED solid, no error beeps
  • em1 shows link (WAN cable seated)
  • Logged in once with default credentials and changed the password
2 · Identify and assign the SFP+ interface
Intermediate~15 min
Why: OPNsense doesn't know which physical port you mean by "the 10GbE one" — Linux/BSD network interfaces get auto-named (igb0, em2...) based on hardware detection order, not what's printed on the case. You have to confirm which name maps to which physical port before you can safely configure it.
  1. SSH in: ssh admin@10.0.1.1 (or use the console if SSH isn't enabled yet).
  2. Run ip link show and look for the interface reporting a 10000 Mbps link once the SFP+ cable is seated — that's your target.
  3. In the web UI: System → Interfaces → Assignments → add that identified port as LAN.
ssh admin@10.0.1.1
ip link show
# look for the interface reporting 10000 Mbps
  • Correct 10GbE interface identified (not guessed)
  • Assigned as LAN in System → Interfaces → Assignments
3 · Configure LAN (SFP+) as the VLAN trunk parent
Beginner~10 min
Why: this interface is the "parent" every VLAN in Step 4 will attach to. Give it a stable static IP now (rather than DHCP) because every other device on VLAN 10 is going to be configured to look for OPNsense at exactly 10.0.1.1 — if that address ever changed, everything downstream would break.
  1. Static IPv4: 10.0.1.1, prefix /24.
  2. Enable DHCP on this interface: range 10.0.1.100–200.
  3. Leave em0 configured too, as a secondary/fallback VLAN parent — this is what you manually switch to if the 10GbE link ever fails.
  • LAN shows 10.0.1.1/24, interface up
  • Can ping 10.0.1.1 from an admin PC
  • https://10.0.1.1 loads the web UI
4 · Create the 15 VLAN interfaces
Intermediate~45 min (mechanical, repeated 15×)
Why a VLAN interface at all? Physically there's one cable (the SFP+ trunk) carrying all 12 networks at once. A VLAN tag is a small marker stamped on each Ethernet frame saying "I belong to network 20" or "network 30" — OPNsense reads that tag and treats the traffic as if it arrived on a completely separate wire. That's what makes 12 isolated networks possible over 1 physical link instead of needing 12 physical cables.
  1. Go to Interfaces → Other Types → VLAN, click add.
  2. Parent interface: LAN (the SFP+ trunk from Step 3). VLAN tag: the VLAN number (10, 15, 20…). Description: the VLAN's name.
  3. Go to Interfaces → Assignments, add the new VLAN as a real interface.
  4. Open that interface, set static 10.0.X.1/24, enable DHCP 10.0.X.100–200.
  5. Repeat for all 15: 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 999.
  • All 15 VLAN interfaces created and visible under Interfaces → Assignments
  • All show GREEN (up) status
  • Each has the correct 10.0.X.1 gateway and DHCP range
5 · Firewall rules — default deny, then explicit allows
Advanced~60 min
The core security concept of this whole build: a default-deny firewall starts from "nothing is allowed" and you punch specific holes for traffic you've decided is legitimate — the opposite of a typical home router, which allows everything on the LAN by default. It's more setup work up front, but it means a compromised device on, say, the Media VLAN physically cannot reach your Database VLAN unless you explicitly said it could — there's no rule to exploit because none exists.

Rule order per VLAN (first match wins, catch-all deny always last):

10 Management  → ANY                 (allow)
20 Media       → 30 Cloud, 80 DB, WAN (allow)
30 Cloud       → 80 DB, WAN           (allow)
35 Trading     → 60 DNS, WAN (broker) (allow — outbound only, no lateral VLAN access)
40 Photos      → 40 self, 60 DNS only (allow)
45 Cameras     → nothing, not even DNS   (isolated — Frigate reaches it via a dedicated NIC, not routing)
50 Monitor     → ANY                 (allow)
60 DNS         → ANY                 (allow, UDP/53)
70 VPN         → ANY                 (allow)
80 Database    → 60 DNS only         (allow)
90 Testing     → ANY                 (allow)
999 NAS        → nothing outbound    (isolated)
*  → *                               (BLOCK — catch-all)

Add outbound NAT for internet-needing VLANs (15, 20, 30, 35, 50, 70, 90) under Firewall → NAT → Outbound. VLAN 35 (Trading) should be restricted to just the broker's IP ranges/hostnames if your firewall supports FQDN aliases — not a blanket internet allow.

  • Allow rules created per VLAN, in the order above
  • Catch-all BLOCK rule sits at the bottom of every VLAN's rule list
  • Outbound NAT configured for 20, 30, 50, 70, 90
6 · Verify & harden
Intermediate~30 min
Why: a rule that "should" work and a rule that's confirmed working are different things — test both the allow and the deny sides against the traffic matrix in Section 05 before trusting this firewall with real services behind it.
  1. Ping each VLAN gateway (10.0.X.1) and confirm DHCP hands out an address on each.
  2. Test one allowed path (e.g. 20→30) and one blocked path (e.g. 40→internet) against the matrix in Section 05.
  3. Change the default password if you haven't, disable Telnet, force HTTPS-only, enable logging.
  4. Back up config: System → Configuration → Backup & Restore → download XML, store it off-box (in the GitHub repo's /archive or a password manager, not just on the NAS).
  • All 14 gateways reachable, DHCP confirmed on at least 3 spot-checked VLANs
  • One allow + one block tested and behaved as expected
  • Config XML backed up off-box
6b · Perimeter hardening — UPnP, outbound blocklists, certificate exposure
Intermediate~20 min
Rationale: automated internet-wide scanning is continuous and effectively instantaneous — a routable IPv4 address can be swept by tools such as masscan in minutes, and a newly issued TLS certificate has been observed being probed within one second of appearing in public Certificate Transparency logs. This is a baseline condition of operating anything with a public IP, not a targeted event. The design in this document already accounts for it structurally: zero inbound port-forwards mean there is nothing for a scanner to connect to regardless of whether the address is known. The steps below close the remaining gaps — an unmanaged UPnP mapping, or an unrestricted outbound path for anything that does get in.
  1. Disable UPnP and NAT-PMP on OPNsense entirely. Both protocols allow any device on the network to open an inbound port mapping without administrator approval, and were the root cause of the DeadBolt ransomware campaign against roughly 19,000 QNAP NAS devices via auto-forwarded UPnP mappings.
  2. Add an outbound firewall rule referencing a maintained threat-intelligence block list (Spamhaus DROP is a standard choice) as a destination alias. This does not affect legitimate traffic; its function is to prevent a compromised host from reaching a known command-and-control address if one is ever compromised regardless of the inbound protections in place.
  3. Optionally, enable GeoIP filtering on the WAN interface. This reduces log noise from high-volume scanning sources and is a genuine additional layer, though it does not stop a determined attacker routing through infrastructure inside an allowed country.

The wildcard certificate configured in Section 07 → Group B → Proxy & dashboard has a secondary benefit here: a wildcard entry in Certificate Transparency logs discloses that *.yourdomain.com exists, but does not enumerate individual subdomains the way a separate certificate per service would. This is a reason to prefer one wildcard certificate over per-service certificates beyond the operational convenience already described there.

  • UPnP and NAT-PMP disabled
  • Outbound block list rule active, referencing a maintained threat-intelligence feed
  • GeoIP filtering considered and a decision made, whether enabled or not (optional)
1 · Enable and verify the SFP+ ports
Beginner~10 min
Why: SFP+ ports ship disabled by default on some RouterOS configs. This is the first thing to check if your 10GbE link doesn't come up — before assuming the cable or the other end is the problem.
ssh admin@10.0.1.5 (or console)

/interface/ethernet set sfp-sfpplus1 disabled=no
/interface/ethernet set sfp-sfpplus2 disabled=no

/interface ethernet print
# confirm both show speed=10Gbps once cables are seated
  • Both SFP+ ports enabled
  • Both show 10Gbps once DAC cables are seated on both ends
2 · Add SFP+ ports to the VLAN trunk bridge
Intermediate~15 min
Trunk vs. access, visually: a trunk port carries every VLAN's tagged traffic at once — think of it as several separate lanes squeezed onto one road, each car (frame) still labelled with which lane it belongs to. An access port carries exactly one VLAN, untagged — the switch strips the label before handing it to a simple device (like the NAS) that doesn't understand VLAN tags at all.
/interface/bridge/port/add bridge=vlan-trunk interface=sfp-sfpplus1
/interface/bridge/port/add bridge=vlan-trunk interface=sfp-sfpplus2

/interface/bridge/vlan/print
# both SFP+ ports should show all 15 VLAN IDs tagged
  • Both SFP+ ports added to the trunk bridge
  • bridge/vlan/print shows all 15 VLAN IDs tagged on both
3 · Configure fallback, storage, and access ports
Beginner~15 min
Why: not every device needs to see all 15 VLANs — the NAS, the laptop drop, and NanoKVM each only need one, so they get simple access ports rather than trunk ports. Fewer things a misconfiguration can go wrong with.
  1. ether1 — trunk, all 15 VLANs, → OPNsense em0 (1GbE fallback)
  2. ether3 — trunk, all 15 VLANs, → Proxmox eth0 (1GbE fallback)
  3. ether4 — access, untagged VLAN 30, → NAS NIC1
  4. ether5 — access, untagged VLAN 10, → NAS NIC2
  5. ether2 — access, untagged VLAN 10, → admin laptop drop (leave patched through to wherever you'll sit)
  6. ether6 — access, untagged VLAN 10, → NanoKVM
  • ether1 & ether3 confirmed trunk mode
  • ether4/5/2/6 confirmed access mode with correct untagged VLAN
3b · Retire VLAN 1 — don't let it become the silent default
Intermediate~10 min
Why this matters even though nothing here intentionally uses VLAN 1: switches ship with VLAN 1 as the default — untagged traffic lands there, it's implicitly allowed on every trunk, and a factory-reset or misconfigured port silently falls back to it. That makes VLAN 1 a quiet catch-all for mistakes instead of a network you designed. The fix is to make mistakes fail loudly instead of landing somewhere that "just works."
  1. Remove VLAN 1 from the allowed-VLAN list on both trunk ports (SFP+1, SFP+2, ether1, ether3).
  2. Set the switch's own management interface to VLAN 10 explicitly (already done in Step 3) rather than leaving it on VLAN 1 by default.
  3. Where the platform allows it, set an explicit native VLAN on trunk ports rather than leaving VLAN 1 as the implicit native — or disable native/untagged VLAN handling on trunks entirely if RouterOS supports it.
  • VLAN 1 removed from both trunk ports' allowed lists
  • A misconfigured/factory-reset port would now fail to pass traffic rather than silently landing on a working default
4 · Verify link status
Beginner~5 min
/interface ethernet print
# sfp-sfpplus1, sfp-sfpplus2, ether1-6
# should all show "R" (running)
  • All expected ports show running ("R")
  • Unused ports (7, 8) left disabled or unassigned
1 · Install Proxmox VE
Beginner~40 min
What Proxmox is, functionally: it is a hypervisor — an operating system whose function is to run other operating systems inside it, as isolated virtual machines (VMs) or lighter-weight containers (LXCs). One physical host can run twenty or more separate guest systems concurrently, each operating as if it has dedicated hardware.
  1. Flash the Proxmox VE ISO (version 9.2, built on Debian 13 "Trixie," current as of this writing — confirm the latest point release on the official downloads page before installing) to a USB stick.
  2. Boot the MS-01 from USB, install to the 2TB NVMe drive.
  3. First login: https://<ip>:8006 in a browser, accept the self-signed certificate warning.
  • Proxmox installed, web UI reachable on port 8006
1b · PVE post-install script — repository and update configuration
Beginner~10 min
Purpose: a fresh Proxmox installation is configured to use the enterprise package repository, which requires a paid subscription and otherwise produces update errors on every apt update. The post-install script performs the standard set of first-run configuration changes — switching to the no-subscription repository, correcting the same issue for the Ceph repository if present, and optionally disabling the subscription nag dialog in the web UI — as a single reviewed, community-maintained script rather than manual editing of sources.list files.
  1. Review the script's contents before running it — this applies to any script executed with root privileges from a third-party source, not a specific concern about this one.
  2. Run it from the Proxmox host shell: bash -c "$(wget -qLO - https://raw.githubusercontent.com/community-scripts/ProxmoxVE/main/tools/pve/post-pve-install.sh)"
  3. Select the no-subscription repository when prompted; the enterprise repository prompt can be declined unless a paid subscription is in use.
  • No-subscription repository active, apt update runs without repository errors
1c · Proxmox VE Helper-Scripts — purpose and use throughout this build
Beginnerreference
What these are: the Proxmox VE Helper-Scripts project (originally maintained by tteck, now continued by the community-scripts organisation) is a library of shell scripts that automate LXC or VM creation for several hundred common self-hosted applications — provisioning the container, installing the application and its dependencies, and applying sane defaults, in place of a manual multi-step installation. Each application-specific deployment step in this document that references a Proxmox Helper Script (Home Assistant and Homebridge in Section 07 → Group B → Home automation are two examples already in use) is invoking one of these scripts rather than a fully manual procedure.

The trust model is the same as any script executed with root privileges from a source not personally authored: the community-scripts organisation is the maintained successor to the original tteck repository, scripts are open source and reviewable before execution, and the general precaution in Step 1b — read a script before running it — applies uniformly across every application deployed this way in this document, not as a repeated caveat per service.

  • Helper-Scripts site bookmarked as the first check before manually scripting a new LXC deployment
2 · Bridge the SFP+ trunk interface
Intermediate~15 min
Why a bridge? A Proxmox "bridge" is a virtual switch living inside the host — every VM plugs into it the same way a physical device plugs into a physical switch port. Making it "VLAN-aware" means one bridge can carry all 12 tagged VLANs, and each VM just declares which VLAN tag its virtual network card uses — you don't need 12 separate bridges.
auto enp2s0f0
iface enp2s0f0 inet manual

auto vmbr0
iface vmbr0 inet manual
    bridge-ports enp2s0f0
    bridge-stp off
    bridge-fd 0
    bridge-vlan-aware yes
    bridge-vids 10,15,20,25,30,35,40,45,50,60,70,80,90,100,999

Apply with systemctl restart networking. Keep eth0 (the 1GbE fallback NIC) as a separate manual bridge you're not actively using, ready if the 10GbE link ever needs replacing.

  • vmbr0 created, VLAN-aware, all 15 VLAN IDs listed
  • Networking restarted without errors, SSH/web UI still reachable
3 · Tag each VM's NIC to its VLAN
Beginner~5 min per VM
Why: this is the step that actually puts a VM "on" a VLAN — everything before this just built the road, this is choosing which lane each car drives in.

Single VLAN-aware vmbr0, tag per VM in Hardware → Network Device → VLAN Tag. Jellyfin VM → 20, Nextcloud → 30, Postgres → 80, and so on per the VLAN table in Section 05.

Avoid the most common VLAN mistake — double tagging: the tag belongs on the VM's NIC only. Never also set a VLAN tag on vmbr0 itself or configure the switch's SFP+2 port as an access port for a specific VLAN — it's a trunk carrying all 14 tagged VLANs, and the switch port config in the MikroTik tab already reflects that. If a VM ever seems to land on the wrong network or nothing responds, check for exactly this: a tag applied at both the bridge and the VM level.
  • Every VM's NIC has an explicit VLAN tag set (never left blank/untagged)
  • vmbr0 itself carries no VLAN tag of its own — only individual VM NICs are tagged
4 · Mount NAS storage over NFS
Beginner~10 min
Why NFS instead of a local disk: it lets every VM share one pool of storage on the NAS rather than each VM carrying its own separate disk image — media, backups, and config all live in one place, growable independently of Proxmox's own NVMe.

Prerequisite: the NAS (Group A → Tab 3) must already be configured with its shares created before this step — this is why NAS precedes Proxmox in the tab order.

Datacenter → Storage → Add → NFS: server = NAS IP on VLAN 30, export = the share path from the NAS tab, content = Disk image / ISO / backup as needed.

  • NFS storage added and shows green/active in Datacenter → Storage
  • Throughput spot-checked (>100MB/s target on the 1GbE NAS link)
5 · First test VM
Beginner~20 min
Why bother with a throwaway VM: this is your first end-to-end proof that networking, VLAN tagging, and the firewall rules all agree with each other, before you build anything you actually care about on top of a possibly-broken foundation.
  1. Create a small VM (any lightweight Linux ISO works), tag its NIC VLAN 90 (Testing).
  2. Confirm it gets a DHCP lease from 10.0.90.100–200, and can ping its gateway 10.0.90.1.
  3. Confirm reachability matches the traffic matrix in Section 05 — it should reach the NAS if a rule allows it, and should not reach VLAN 40 (Photos).
  • Test VM gets correct DHCP lease and gateway
  • Allowed path confirmed, blocked path confirmed blocked
1 · Operating system — TOS, and how it compares to TrueNAS
Beginnerreference
Clarification, since the two are frequently discussed together: the Terramaster F4-424 Pro ships with TerraMaster's own operating system, TOS — not TrueNAS. The two solve the same problem from different directions. TOS is a proprietary, appliance-style NAS operating system that comes pre-installed on purpose-built Terramaster hardware; setup is closer to configuring a consumer appliance than building a server. TrueNAS Community Edition is free, open-source, and Debian-based (following the 2022 merger of TrueNAS Scale into the mainline product), built on OpenZFS — but it is software installed onto generic x86 hardware assembled for the purpose, not an appliance. TrueNAS's OpenZFS foundation provides automatic checksumming and self-healing against data corruption ("bit rot") as a built-in filesystem property; TOS's RAID implementation does not provide the equivalent guarantee at the filesystem level, which is why S.M.A.R.T. monitoring (Step 3) and Proxmox Backup Server (Section 07 → Group B → Backups) carry more of that responsibility in this design than they would on a ZFS-based system.

This design uses the Terramaster appliance because the hardware was already owned at the point the network design began, and the appliance model trades some of ZFS's data-integrity guarantees for a simpler, faster initial setup. A TrueNAS deployment would require different hardware — a generic x86 host with sufficient SATA/NVMe connectivity, no Terramaster-specific components — and is noted here as an alternative rather than a planned change.

  • Confirmed: this device runs TOS, and TrueNAS-specific instructions found elsewhere do not apply directly
2 · Physical install & initial setup
Beginner~20 min
Rationale for two NICs: one carries the storage path other hosts use for actual file traffic (VLAN 30); the other is a management-only path (VLAN 10), so the NAS's admin panel remains reachable even if something is wrong with the storage VLAN specifically.
  1. Install all four drives into the bays.
  2. NIC1 → switch port 4 (VLAN 30), NIC2 → switch port 5 (VLAN 10).
  3. Complete initial setup via the management VLAN, using TerraMaster's discovery tool (TNAS PC/Mac, or the web-based discovery page) to locate the device's assigned address.
  4. Apply any pending TOS firmware update before creating the storage pool — a fresh appliance from stock may be several point releases behind.
  • All four drives detected in the storage manager
  • Both NICs show link, on the correct VLANs
  • TOS updated to the current version before proceeding
3 · Configure RAID 5
Beginner~15 min active, hours passive resync
What RAID 5 provides: data striped across all four drives with a parity calculation, so any single drive failure is recoverable from the survivors with no data loss. RAID is not a backup — it protects against a drive failing, not against accidental deletion or ransomware, which is what Proxmox Backup Server (Section 07 → Group B → Backups) addresses. As noted in Step 1, TOS's RAID does not carry ZFS's automatic checksumming, so S.M.A.R.T. monitoring here is doing work that would otherwise be partly automatic on a ZFS-based system.
  1. Storage Manager → create a RAID 5 array across the four bays. Approximately 4TB usable after parity overhead.
  2. Enable S.M.A.R.T. monitoring and a weekly scheduled scan.
  3. Enable TOS's own scheduled RAID scrub/data-integrity check if the model supports it, as a partial substitute for ZFS's continuous checksumming.
  • RAID 5 array created, initial resync started
  • S.M.A.R.T. monitoring enabled with a weekly scan scheduled
  • Scheduled scrub/integrity check enabled if available
4 · Shares, NFS export, permissions
Intermediate~20 min
Rationale for scoping NFS exports: without restriction, any device able to reach the NAS's IP address can mount the share. Scoping the export to only the VLANs that legitimately need it means a compromised device elsewhere on the network cannot mount the storage regardless of what else it can reach.
  1. Create shares: /storage, /backups, /media.
  2. Enable NFS, with the export scoped to the relevant VLAN subnets specifically — not an unrestricted "any" export.
  3. Create per-service scoped accounts rather than a single shared administrative login for every service that accesses the NAS.
  • Shares created and exported
  • NFS export restricted by subnet, not open to "any"
  • Scoped accounts created; no shared administrative login in use by any service
5 · Snapshot schedule — TOS's native snapshot feature
Beginner~15 min
Where this sits relative to Proxmox Backup Server: TOS's built-in snapshot feature provides fast, space-efficient point-in-time recovery for accidental deletion or modification directly on the NAS, distinct from PBS's role as the actual off-device backup copy. A snapshot on the same physical device is not a backup by itself — it does not survive drive failure at the array level or the appliance itself failing — but it is a useful first line of recovery for the common case of a file being overwritten or deleted in error, with much faster recovery than restoring from PBS for that specific scenario.
  1. Enable scheduled snapshots on the shares created in Step 4 — a daily snapshot with a one- to two-week retention window is a reasonable starting point.
  2. Confirm the snapshot schedule does not depend on the RAID array being otherwise idle, and that it does not conflict with backup job timing (Section 07 → Group B → Backups).
  • Scheduled snapshots enabled on primary shares
  • Retention window set and confirmed not to fill available capacity
6 · Verify dual-NIC reachability
Beginner~5 min

From Proxmox: ping 10.0.30.10 and ping 10.0.1.100. Disconnect each cable in turn to confirm the NAS remains reachable on the other NIC before relying on this configuration for redundancy.

  • Both NICs individually confirmed reachable
1 · Physical connections
Beginner~15 min
What each port does: HDMI-in captures whatever a monitor would show from the host. USB-C device emulation makes NanoKVM present itself to the host as a real keyboard/mouse — the host has no idea it isn't a physical human typing. Ethernet is how you reach NanoKVM's own web UI. The auxiliary USB-C power port keeps it running independent of the host's own power state, which matters if you ever need to power-cycle the host itself.
  1. HDMI-in ← Proxmox HDMI-out.
  2. USB-C (device emulation) ← Proxmox USB port.
  3. Ethernet → switch port 6 (VLAN 10, access, untagged).
  4. Aux power USB-C ← a spare USB port on Proxmox, or a small USB power adapter into a PDU-adjacent outlet if you'd rather it not depend on Proxmox's own power state.
  • HDMI capture confirmed (see the Proxmox boot screen through NanoKVM)
  • USB emulation confirmed (mouse cursor moves on the host from NanoKVM's web UI)
  • Status LEDs normal, no error pattern
2 · First boot & firmware
Beginner~20 min
Why check firmware first: NanoKVM gets frequent updates and the web UI changes between versions — starting on the latest firmware means the rest of these steps' menu paths are more likely to match what you actually see on screen.
  1. Power it on, wait for the status LED to settle.
  2. Find its IP (check your router's DHCP client list, or use the discovery method in the quick start guide).
  3. Log into the web UI, go to Settings → Check for Updates, apply if available.
  • NanoKVM web UI reachable
  • Firmware confirmed up to date
3 · Network configuration — VLAN 10, static reservation
Beginner~10 min
Why a static reservation, not just DHCP: you'll want to bookmark or hard-code NanoKVM's address for emergency access — you don't want to be hunting for its current IP in a DHCP lease table during the exact moment something else is broken.
  1. Confirm it pulled an address in the 10.0.1.100–200 range from OPNsense.
  2. In OPNsense, add a DHCP static mapping (by MAC address) so it always gets the same IP.
  • Static DHCP reservation configured on OPNsense
  • Address noted somewhere you'll actually find it during an outage (Homarr, a sticky note, wherever — just not "I'll remember")
4 · Web UI security — password, HTTPS, SSH
Intermediate~10 min
Why this device is a high-value target: anyone who can reach NanoKVM's web UI has full keyboard/video/power access to whatever it's plugged into — effectively physical access without being in the room. It deserves the same seriousness as your firewall's admin login, arguably more.
  1. Change the default password immediately.
  2. Enable HTTPS on the NanoKVM web UI if not already on.
  3. Disable SSH unless you specifically need it (Settings → Devices → SSH, per the user guide).
  • Default password changed
  • HTTPS enabled
  • SSH disabled unless actively needed
5 · ATX power control (optional, hardware-dependent)
Advanced~30 min, if possible at all
Why this one's marked "if possible": remote power-cycling needs NanoKVM's ATX breakout board wired to the host motherboard's power/reset header pins. That assumes the header is physically exposed — true for a bare motherboard in a case, not guaranteed for a sealed mini-PC like the MS-01. Check Minisforum's own teardown/documentation before buying or wiring the breakout board.
  1. Check whether the MS-01 exposes an accessible power/reset header (open the case, check Minisforum documentation, or search the model number plus "ATX header").
  2. If yes: wire the NanoKVM-B breakout board per Sipeed's wiring diagram.
  3. If no: skip this — you still have full video + keyboard/mouse. For hard power-cycling, a UPS-controlled smart outlet is the practical fallback.
  • Verified whether MS-01 exposes ATX header pins before buying/wiring anything
  • If wired: power on/off/reset tested remotely at least once
6 · Lock it down — NetBird-only access
Intermediate~10 min
Why: this device should never be reachable directly from the internet, and ideally not even casually from every device on your home network — only from your own admin devices, over the NetBird tunnel. Full remote KVM access is exactly the kind of thing you don't want discoverable.
  1. Confirm no port-forward exists on OPNsense exposing NanoKVM's port to WAN.
  2. Once NetBird is deployed (Services tab), scope access so only your "admin" peer group can reach VLAN 10 where NanoKVM lives.
  • No WAN port-forward to NanoKVM exists
  • Reachable only via NetBird admin group once that's deployed
7 · Set a BIOS password on Proxmox itself
Beginner~10 min
Why this matters specifically because NanoKVM exists: NanoKVM gives BIOS-level video and keyboard access to Proxmox remotely — genuinely useful, but it also means Proxmox's BIOS is now reachable by anything that reaches NanoKVM's web UI. A BIOS password is a second, independent gate: even in the worst case where NanoKVM itself were somehow compromised, boot order, secure boot settings, and other BIOS-level changes still require a credential NanoKVM access alone doesn't grant.
  1. Set a BIOS/UEFI admin password on the MS-01, stored in Vaultwarden like every other credential.
  2. While in there, disable booting from any external/USB media by default, and disable any unused Thunderbolt/USB controllers not actually in use.
  • BIOS password set and stored in Vaultwarden
  • External boot disabled by default, unused USB/Thunderbolt controllers disabled
Purpose: a second, physically separate switch dedicated to household devices — a television, a personal computer, and similar — that routes their traffic through the same firewall, DNS ad-blocking, and intrusion detection as the lab, on a VLAN with no visibility into or access from any lab VLAN. A fault or misconfiguration on the lab side does not affect these devices' internet access either, beyond the DNS/firewall path both share. One exception exists by design, not by accident: a single dedicated port (Step 3b) carries VLAN 10 instead of VLAN 100, giving an admin device plugged in at this switch's location the same lab-troubleshooting reach as the existing admin laptop drop elsewhere — every other port on this switch has no path into the lab at all.
1 · Initial discovery and firmware check
Beginner~20 min
Configuration model: the JGS524E is a Netgear ProSAFE Plus-series switch, not a fully managed switch. Configuration is performed either through its onboard web interface once its IP address is known, or through the ProSAFE Plus Configuration Utility — a Windows-only discovery and configuration tool, supplied on the original resource CD and also available from Netgear's support site, that locates Plus switches on the local network by broadcast, independent of what IP address they have picked up.
  1. Connect the switch to power and to a temporary access-mode port on the MikroTik switch (not port 7, which is reserved for the trunk configured in Step 2).
  2. Run the ProSAFE Plus Configuration Utility from a Windows PC on the same network segment, or check the switch's DHCP lease on OPNsense if it obtained one automatically.
  3. Log in with the default password (password) and change it immediately — this switch does not support the account-separation model used elsewhere in this design, so the single admin password is the only credential protecting it.
  4. Check for and apply any available firmware update before proceeding; this hardware model dates to 2012 and a specific unit's firmware history is unknown until checked.
  • Switch located and reachable via the configuration utility or web interface
  • Default password changed
  • Firmware checked and current
2 · VLAN configuration — 802.1Q tagging
Intermediate~20 min
What this switch can and cannot do: the ProSAFE Plus feature set supports 802.1Q tagged VLANs, which is sufficient to carry the two VLANs this design requires — it does not support inter-VLAN routing, DHCP, or firewall rules, all of which remain OPNsense's responsibility exactly as for every other VLAN in this design. This switch's only function is delivering tagged and untagged Ethernet frames to the correct ports.
  1. Create VLAN 10 (Management) and VLAN 100 (Home Devices) in the switch's VLAN configuration screen.
  2. Set the uplink port (the port that will connect to the MikroTik switch) to tag both VLAN 10 and VLAN 100.
  3. Set every other port intended for a household device to VLAN 100, untagged (PVID 100).
  4. Leave the switch's own management VLAN set to VLAN 10, matching every other managed device in this design.
  • VLAN 10 and VLAN 100 created
  • Uplink port tagging both VLANs confirmed
  • Device-facing ports set to untagged VLAN 100
3 · Uplink to the MikroTik switch
Beginner~10 min
Reference: this step connects the configuration above to the trunk already reserved on the MikroTik switch in Section 05's port map, completing the physical path from a household device to OPNsense.
  1. Move the uplink cable from the temporary access port used in Step 1 to MikroTik port 7.
  2. On the MikroTik switch, confirm port 7 is configured as a trunk carrying VLAN 10 and VLAN 100 only — not every VLAN, since this switch has no reason to see lab traffic.
  3. Confirm the switch's management IP (VLAN 10) is reachable from an admin device, and that a test device plugged into a VLAN 100 port receives a DHCP lease from OPNsense in the 10.0.100.0/24 range.
  • Port 7 confirmed as a two-VLAN trunk, not an all-VLANs trunk
  • Switch management IP reachable on VLAN 10
  • Test device on VLAN 100 receives a DHCP lease and reaches the internet
3b · Reserve one port for on-site troubleshooting access
Beginner~10 min
Rationale: the household devices on VLAN 100 have no path into the lab, by design — but that also means diagnosing a lab issue from the JGS524E's physical location would otherwise require walking to the rack. One dedicated port carrying VLAN 10 instead of VLAN 100 gives an admin device plugged in at that location the same reach as the existing admin laptop drop on the MikroTik switch (Section 05's port map, port 2) — full troubleshooting access to the lab, without granting that access to anything else connected to this switch.
  1. Set one spare port (port 24, the last physical port) to untagged VLAN 10, separate from the VLAN 100 access ports used for household devices.
  2. Label this port physically and in the switch's own port naming, since a household device plugged in here by mistake would receive a management-network address instead of an internet-only one.
  3. Confirm a laptop plugged into port 24 receives a VLAN 10 address from OPNsense and can reach lab services — the same reach as the standing Management-VLAN exception already applied everywhere else in this design, not a new or broader rule.
  • Port 24 set to untagged VLAN 10, clearly labelled
  • Confirmed a device on this port receives a VLAN 10 address and reaches lab services
  • Confirmed the remaining household-device ports are unaffected — still VLAN 100 only
4 · OPNsense — VLAN 100 firewall rules
Intermediate~20 min
Rule intent: VLAN 100 should behave like any household internet connection with DNS-level ad-blocking and firewall inspection applied — full outbound internet access, forced through Technitium exactly as every other VLAN, and zero access to or from any lab VLAN beyond the standing Management-VLAN exception applied everywhere in this design.
  1. Create the VLAN 100 interface on OPNsense (10.0.100.1/24), following the same procedure as every other VLAN in Section 07 → OPNsense, Step 4.
  2. Allow rule: VLAN 100 → Technitium (port 53) and → WAN, matching the DNS-enforcement pattern in Section 07 → Services → Monitoring & DNS.
  3. Explicit deny: VLAN 100 → every other VLAN. No exception rules — this VLAN has no legitimate reason to reach lab services.
  4. Confirm Suricata and CrowdSec (Section 07 → Group C → Security & identity), already inspecting all inter-VLAN and WAN-bound traffic, cover this VLAN without additional configuration — they operate on the trunk, not per-VLAN.
  • VLAN 100 interface created, DHCP enabled
  • Outbound allowed only to Technitium and WAN
  • Explicit deny confirmed against every lab VLAN, tested from a device on VLAN 100
5 · Optional — outbound VPN egress for VLAN 100
Advanced~45 min
Distinction from NetBird: NetBird (Section 07 → Services → Remote access) provides inbound remote access into this network from elsewhere — it is not related to routing this network's own outbound traffic through a VPN provider for privacy. That is a separate, optional configuration: OPNsense can run a WireGuard client to a third-party VPN provider and, via policy-based routing, send only VLAN 100's egress traffic through that tunnel, leaving every lab VLAN's traffic on its normal path.
  1. Configure a WireGuard client instance on OPNsense connecting to a chosen VPN provider's endpoint.
  2. Create a gateway using that WireGuard interface, and a policy-based routing rule scoped to VLAN 100 only, directing its traffic through that gateway instead of the default WAN gateway.
  3. Verify with an IP-address check site from a VLAN 100 device that egress traffic shows the VPN provider's address, and confirm DNS continues resolving through Technitium rather than leaking to the VPN provider's own resolver.
  • Not configured by default — optional, add only if outbound privacy routing for household devices specifically is wanted
  • If configured: egress IP and DNS resolution both verified from a VLAN 100 device
Group B — Services & add-ons
Two independent remote-access paths, on purpose. NetBird is the primary — a real mesh VPN, full access, lowest latency. Cloudflare Tunnel is the backup — a completely different mechanism (outbound-only tunnel to Cloudflare's edge, no WireGuard involved) so that if NetBird ever has an outage or a peer won't connect, you still have a second, unrelated way in.
1 · Deploy the self-hosted NetBird server
Intermediate~45 min
What NetBird is doing under the hood: WireGuard handles the actual encrypted tunnels — it's the protocol, not a competing product. NetBird and Tailscale are both implementations built on WireGuard, not alternatives to it; the real choice is which coordination layer introduces peers to each other and enforces policy. NetBird's is self-hostable, which is why it's the pick here — the WireGuard tunnels themselves are identical either way.

This also sidesteps the classic reason people reach for a hosted mesh tool in the first place: a bare WireGuard server normally needs an inbound port forwarded on the router, which doesn't work at all behind carrier-grade NAT (CGNAT). NetBird's peers connect outbound to the coordination server and punch through NAT from there — no inbound port needed, which is also exactly why this build has zero port-forwards (Section 08).

Twingate, for context: a legitimate free-for-small-teams zero-trust option, but its control plane is Twingate's cloud — same category concern as Entra ID in Section 11. NetBird's self-hosted coordination server keeps that layer under this lab's own control instead, which is the deciding factor here, not a feature gap in Twingate itself.

  1. Create an LXC/VM on VLAN 70 (VPN).
  2. Follow NetBird's Docker Compose self-hosted quickstart to stand up the management server, signal server, and dashboard.
  3. Point a DNS name at it via NGINX Proxy Manager once that's deployed (Tab B), or access by IP for now.
  • NetBird management dashboard reachable and logged in
2 · Enrol peers
Beginner~20 min
Why "mesh" matters: once enrolled, peers connect directly to each other (peer-to-peer) rather than all traffic funnelling through one central VPN gateway — lower latency, and no single choke point if you're streaming Jellyfin remotely while also managing OPNsense.
  1. OPNsense: install via the OPNsense NetBird plugin (os-netbird) — a native FreeBSD package, not Docker.
  2. Docker-hosted peers (any Docker host on VLAN 15, e.g. the Talos/AI/media hosts): run NetBird as a container using the netbirdio/netbird image with network_mode: host — this lets it create the WireGuard interface (wt0) directly on the host's network stack rather than isolated inside Docker's bridge network. Mount /etc/netbird and /var/lib/netbird as volumes so the peer's identity survives a container restart.
  3. Admin laptop/phone: native NetBird client from the official download page.
  4. Authenticate each against the self-hosted dashboard, confirm each shows "Connected" in the peers list.
# Docker-hosted peer example
docker run -d --name netbird \
  --network host \
  --cap-add NET_ADMIN \
  -v /etc/netbird:/etc/netbird \
  -v /var/lib/netbird:/var/lib/netbird \
  --restart unless-stopped \
  netbirdio/netbird
  • OPNsense enrolled and connected (via the OPNsense plugin)
  • Proxmox / Docker hosts enrolled and connected (via the Docker container, wt0 interface confirmed with ip a)
  • Admin devices enrolled and connected
3 · Access control groups
Intermediate~20 min
Why: "connected to the VPN" and "allowed to reach everything" shouldn't be the same thing. Groups let you say your admin laptop can reach VLAN 10 + 90 for administration, while your phone — or a family member's device — gets scoped to only what it actually needs.
  1. Create an admin group containing your own admin devices, with access rules to VLAN 10 + VLAN 90.
  2. Create a second, narrower personal-devices group for your phone (and any future phone/tablet), with access rules to VLAN 30 (Cloud) + VLAN 40 (Photos) only — everything it needs for Nextcloud, Vaultwarden, and Immich backup, nothing it doesn't.
  3. Leave other VLANs unreachable via NetBird by default — add exceptions deliberately, not by accident.
  • Admin group created, scoped to management + testing only
  • personal-devices group created, scoped to Cloud + Photos only
  • No blanket "allow all" access rule left in place
4 · SSO / 2FA on the dashboard
Intermediate~20 min
Why: this dashboard controls who can reach your entire lab remotely — it deserves stronger login protection than a shared username/password, the same logic as the NanoKVM lockdown in the Hardware tab.

Configure an OIDC identity provider (many free options work fine for a single-user homelab) instead of relying on local username/password login alone.

  • SSO/OIDC or 2FA enabled on the NetBird dashboard login
5 · Cloudflare Tunnel — the backup path
Intermediate~30 min
How this differs from NetBird: Cloudflare Tunnel doesn't use WireGuard or peer-to-peer connections at all — a small cloudflared container on VLAN 10 makes an outbound-only connection to Cloudflare's edge network. There's nothing to port-forward and no inbound rule needed on OPNsense. Because it's architecturally unrelated to NetBird, an outage or misconfiguration in one doesn't take down the other — that's the entire point of having two paths.
  1. Register a domain (or use one you already own) and add it to a free Cloudflare account.
  2. Deploy cloudflared as a container on VLAN 10, authenticate it to your Cloudflare account.
  3. Create a tunnel and add public hostnames that map to internal services through NPM/Traefik, e.g. jellyfin.yourdomain.com → 10.0.10.x:443.
  • cloudflared deployed on VLAN 10, tunnel shows healthy in the Cloudflare dashboard
  • At least one service reachable through the tunnel from outside the home network
6 · Cloudflare Access — Zero Trust policies
Intermediate~20 min
Why the tunnel alone isn't enough: a tunnel just gets a request to your service — it doesn't decide who's allowed to make that request. Cloudflare Access sits in front of the tunnel and requires login (email OTP, or an identity provider) before traffic ever reaches your network, so an internet-exposed hostname isn't the same as an internet-exposed service.

In the Cloudflare Zero Trust dashboard, create an Access application per exposed hostname, require authentication (email OTP is the simplest starting point), and scope each policy to just yourself unless you specifically intend to share access.

  • Access application created for each exposed hostname
  • Policy tested — an unauthenticated request is challenged, not served directly
6b · Geo-restrict Access and lock out repeated failures
Intermediate~15 min
Why narrow the country as well as require login: unless remote access is ever needed from outside Australia, there's no reason to accept login attempts from anywhere else — it shrinks the pool of anyone who can even attempt to authenticate, before MFA is even considered. Combined with an auto-block on repeated failures, this mirrors the layered approach most seasoned homelabbers converge on for their one exposed entry point.
  1. In each Access policy, add a country rule scoped to Australia (adjust temporarily if travelling).
  2. Enable Cloudflare's rate limiting / bot protection on the Access login page.
  3. Apply the same logic to NetBird and OPNsense's own admin login — fail2ban (or OPNsense's built-in lockout) on repeated auth failures, not just on the Cloudflare side.
  • Country-restricted Access policy in place
  • Rate limiting / auto-block confirmed active on Cloudflare Access, NetBird, and OPNsense admin login
7 · Optional: highly-available tunnels (later)
Advancednot needed yet

A single cloudflared replica is a fine backup path for a home lab. If this ever needs to be genuinely highly available, multiple cloudflared replicas behind Docker Swarm is a documented pattern — worth knowing exists, not worth building until the single-replica version has actually let you down.

1 · Deploy NGINX Proxy Manager
Beginner~15 min
What a reverse proxy does, concretely: right now every service lives at its own IP and port (10.0.20.10:8096 for Jellyfin, 10.0.30.20:80 for Nextcloud...). A reverse proxy sits in front of all of them and lets you instead type jellyfin.lab.internal — it looks up which backend that name maps to and forwards the request there, transparently, with one HTTPS certificate covering everything instead of managing certs per-service.
  1. Deploy as an LXC on VLAN 10.
  2. Only forward ports 80/443 to it from OPNsense for whatever you actually intend to expose externally — internal-only services don't need a WAN forward at all.
  • NPM deployed, admin UI reachable
Honest tradeoff, worth knowing: Traefik (already in this build for the Talos cluster, Section 07 → Group C → Kubernetes) is genuinely simpler than NPM once you're comfortable with Docker Compose — it auto-discovers services from labels in the compose file itself, no separate GUI clicking per host. NPM's GUI is the more beginner-friendly starting point precisely because it doesn't require learning label syntax first, which is why it's the default here. If you end up preferring label-driven config once you're comfortable, migrating a service from NPM to Traefik is a reasonable later move — not a sign the original choice was wrong, just a different point on the same tradeoff.
2 · Point subdomains at internal services, one wildcard cert for everything
Intermediate~30 min
Why the DNS challenge specifically: Let's Encrypt's DNS challenge proves you own a domain by creating a DNS record, not by exposing a public web server — meaning you can get real, trusted certificates for purely internal services that are never reachable from the internet at all. A single *.yourdomain.com wildcard certificate, issued once via DNS challenge, covers every subdomain you add later — no per-service certificate requests, no renewal juggling across a dozen services.
  1. Add a proxy host per service: domain name → internal IP:port.
  2. Issue one wildcard certificate (*.yourdomain.com) via Let's Encrypt DNS challenge, and reuse it across every proxy host in NPM — this is the same domain and certificate strategy used for NetBird/Cloudflare Access below, so the whole lab shares one consistent chain of trust.
  3. Confirm auto-renewal is actually configured, not a one-off manual issue.
  • At least one service reachable via a clean HTTPS name, no port in the URL
  • Wildcard certificate issued once and reused across services, renewing automatically
3 · Deploy Homarr dashboard
Beginner~15 min
Why bother: with 20+ services this is the difference between remembering "was Grafana on .50 or .51" and just clicking a tile. It also gives you live up/down status at a glance.

Deploy as an LXC/container, add links to every service, set as your browser home page.

  • Homarr deployed with at least the core services linked

Alternative worth knowing: Homepage is configured almost exactly like Traefik — YAML/labels, not a GUI — which makes it genuinely simpler than Homarr if every service is already a pure Docker Compose stack. Homarr's GUI-first setup is why it's the default here; swap later if label-driven config starts feeling more natural than clicking.

4 · Wire up live status widgets
Beginner~15 min

Connect Homarr's integrations to Proxmox, Jellyfin, and your download clients for live status/stats directly on the dashboard, not just static links.

  • At least one live-status widget working (e.g. Proxmox resource usage)
5 · Watchtower or Dockcheck — automatic container updates
Beginner~15 min
What this closes: patching gets skipped more often on containers than on the host OS, precisely because there's no apt upgrade habit covering them. Watchtower checks for new images on a schedule and updates running containers automatically — pairs with the host-level patch cadence already covered in Section 07, so both halves of the stack stay current without manual tracking.
Watchtower's maintenance status: the original project is now deprecated, though community forks keep it functioning. Dockcheck is the actively maintained alternative worth using instead for a new deployment — comparable functionality, runs fully containerized, and adds notification integrations (including Home Assistant events) that Watchtower never had. Watchtower is left as the documented option below since it still works and many existing labs already run it; treat Dockcheck as the default choice starting today.
  1. Deploy Watchtower (or Dockcheck) as a container with access to the Docker socket, scoped to check on a daily or weekly schedule.
  2. Exclude anything you'd rather update manually and review first (databases are a common exclusion — a surprise major-version bump is a worse outcome than a slightly stale image).
  • Update checker deployed, schedule set
  • Databases and anything version-sensitive explicitly excluded from auto-update
6 · NetBox — the actual source of truth for what's where
Intermediate~45 min
What this replaces: Section 09's asset hardening log is a static table you edit by hand — fine as a starting point, but it drifts the moment you forget to update it. NetBox is purpose-built infrastructure documentation: devices, IP addresses, VLANs, and their relationships as a real, queryable database rather than a table that's only as current as your last edit. It's also the natural next step if the GitOps work in Section 07 → Group C ever grows — NetBox can act as the source of truth Terraform and Ansible read from, instead of hardcoding IPs into playbooks.
  1. Deploy via the official Docker Compose project on VLAN 10.
  2. Enter this build's actual inventory: the rack's devices, all 15 VLANs and their subnets, and current IP assignments — tedious once, but this is what turns it from a diagram into a live record.
  3. Retire the manual Section 09 table once NetBox reflects the same information — don't maintain both.
  • NetBox deployed, admin account created
  • Rack, devices, VLANs, and current IPs entered
1 · Deploy Proxmox Backup Server
Intermediate~20 min
Why PBS over a raw copy script: PBS deduplicates at the chunk level — if only 2% of a VM's disk changed since yesterday, it only stores that 2% again. That makes daily full-VM backups actually affordable on storage, instead of needing 20× the space for 20 days of history.

Start as a VM/LXC on the same Proxmox host, its own storage target on the NAS. Upgrade to a dedicated small box later if you add a second hypervisor node.

  • PBS deployed, storage datastore pointed at NAS
2 · Add PBS as a Proxmox backup target
Beginner~10 min

Datacenter → Storage → Add → Proxmox Backup Server: enter the PBS host, datastore name, and credentials.

  • PBS storage visible and green in Proxmox's Datacenter view
3 · Schedule jobs
Beginner~10 min

Daily incremental backups, weekly full verification job. Start with everything except purely-disposable test VMs.

  • Daily backup job scheduled and has run successfully at least once
  • Weekly verify job scheduled
4 · Test a real restore
Intermediate~20 min
Why this is not optional: a backup you haven't restored from is a theory, not a backup. This is the single most skipped step in every homelab guide, and the one that matters most the day you actually need it.

Restore a test VM to a new VM ID, boot it, confirm it actually works — not just that the restore operation completed without an error.

  • Restore performed and the restored VM actually boots and functions
5 · Rclone — the offsite leg of 3-2-1
Intermediate~30 min
What this actually closes: every earlier mention of "3-2-1 backup" and "at least one offsite copy" (Sections 04, 08, 14) has assumed something does the actual offsite copying — this is that something. PBS gives you deduplicated local backups; Rclone is what gets a copy of them physically out of the building, which is the part that protects against fire, theft, or the whole rack failing at once, not just a single failed drive.
  1. Pick a destination: Backblaze B2 and Wasabi are the usual home-lab picks — S3-compatible, no egress fees on B2 up to your download volume, cheap per-GB storage.
  2. Deploy Rclone as a scheduled job (cron on a small LXC, or a Proxmox scheduled task) rather than installing it ad hoc — this needs to run unattended.
  3. Sync PBS's backup datastore (or just the most critical VMs/datasets, if full replication is more than the connection or budget can handle) on a nightly or weekly schedule.
  4. Confirm the destination bucket is encrypted at rest and that Rclone's own encryption (rclone crypt) is layered on top if the destination provider itself isn't trusted with plaintext data.
  • Offsite destination chosen and configured
  • Rclone sync scheduled, not manual
  • First sync completed and spot-checked — a file actually appears in the remote bucket
1 · Deploy Jellyfin
Beginner~20 min
Why Jellyfin over Plex here: fully open-source, no account or cloud dependency to stream on your own network, and no paywalled features. Plex remains an easy add-on later if you specifically want its easier external-sharing/login experience for family.

Deploy as a VM/LXC tagged VLAN 20, point its library at the NAS media share over the NFS mount from Proxmox (Hardware → Proxmox → Step 4).

  • Jellyfin deployed, library scanning the NAS media share
2 · Deploy Prowlarr (indexers)
Intermediate~20 min
Why Prowlarr sits on its own VLAN 25: it's the one component in this stack that talks to external indexers directly — keeping it separate makes the firewall rule for "who's allowed outbound to random indexer sites" narrow and explicit, rather than bundled in with everything else on the Media VLAN.

Deploy tagged VLAN 25, add your indexers, this becomes the single place Sonarr/Radarr pull indexer config from instead of configuring each app separately.

  • Prowlarr deployed with at least one indexer configured
3 · Deploy Sonarr/Radarr, connect the pipeline
Intermediate~30 min

Both tagged VLAN 20. Connect each to Prowlarr for indexers, and to Jellyfin so newly-downloaded media appears in the library automatically.

  • Sonarr connected to Prowlarr
  • Radarr connected to Prowlarr
  • Both connected to Jellyfin's library
4 · Deploy Bazarr (subtitles)
Beginner~15 min
Where it sits in the pipeline: Bazarr watches Sonarr/Radarr's libraries and automatically searches for and downloads matching subtitles after a file lands — the last step before Jellyfin serves it, per the standard request → grab → download → subtitles → serve pipeline.

Deploy tagged VLAN 20, connect it to both Sonarr and Radarr as subtitle providers, pick your preferred subtitle languages/sources.

  • Bazarr connected to Sonarr and Radarr
  • At least one subtitle provider configured
5 · Deploy Jellyseerr (user requests)
Beginner~15 min
What this actually buys you: this is the piece that turns "I have to SSH in and add things to Radarr myself" into a simple search-and-request page — the front door of the whole automation stack. A request in Jellyseerr triggers Radarr/Sonarr to search, grab, and once it lands, it shows up in Jellyfin automatically.

Deploy tagged VLAN 20, connect it to Jellyfin (for library sync and login) and to Sonarr/Radarr (to actually place requests).

  • Jellyseerr connected to Jellyfin, Sonarr, and Radarr
  • One test request placed end-to-end and confirmed it reached Radarr/Sonarr
6 · Verify the full pipeline
Beginner~15 min

Full chain per the standard automated-Jellyfin pattern: Jellyseerr request → Radarr/Sonarr → qBittorrent → Bazarr → Jellyfin. Trigger one real request end to end and confirm it lands correctly organised, subtitled, and playable — not just that individual apps look "connected" in their settings pages.

  • One item requested via Jellyseerr, downloaded, subtitled, and confirmed playable in Jellyfin
1 · Deploy Prometheus + node exporters
Intermediate~30 min
What's actually happening: Prometheus periodically "scrapes" (pulls) metrics from small agent processes (exporters) running on each host — it doesn't push data itself, hosts just expose a metrics endpoint and Prometheus polls it on a schedule.

Deploy on VLAN 50, install node_exporter on Proxmox and OPNsense (or use OPNsense's built-in Telegraf/Prometheus exporter plugin), configure Prometheus's scrape targets.

  • Prometheus deployed, at least one target being scraped successfully
1b · Optional: SNMP on the switch, Pulse for Proxmox, Beszel for multi-host
Beginner~25 min
Three small, high-value additions: enabling SNMP on the MikroTik switch and scraping it with Prometheus's SNMP exporter surfaces unusual traffic patterns per-port. Pulse is a lightweight, Proxmox-specific monitoring tool that gives a cleaner at-a-glance view of host/VM health than a generic Grafana dashboard alone. Beszel is genuinely skippable on a single host — ps and Pulse already cover that — but this build already spans Proxmox, several LXCs/VMs, and a Talos cluster with multiple nodes, which is exactly the multi-host case Beszel is built for: one lightweight view across every host instead of checking each individually.
  1. Enable SNMP on the MikroTik CRS310, add the Prometheus SNMP exporter as a scrape target, alert on abnormal per-port traffic spikes.
  2. Deploy Pulse alongside Prometheus/Grafana for Proxmox-specific monitoring.
  3. Deploy Beszel's hub on VLAN 50, install its lightweight agent on Proxmox and each VM/LXC/Talos node worth watching.
  • SNMP enabled on the switch, scraped by Prometheus (optional)
  • Pulse deployed for Proxmox-specific views (optional)
  • Beszel hub + agents deployed across multiple hosts (optional, most useful once the Talos cluster is live)
2 · Deploy Grafana, add data sources, import dashboards
Beginner~20 min

Add Prometheus and InfluxDB as data sources, import a community Node Exporter dashboard (search grafana.com/dashboards by ID) rather than building one from scratch on day one.

  • Grafana deployed, at least one imported dashboard showing live data
3 · Deploy Technitium DNS (clustered pair) + Unbound
Intermediate~30 min
Why Technitium clustering over separate synced instances: two independent Pi-hole-style servers need a sync tool to keep block lists and settings aligned, and can drift out of sync. Technitium has native primary/secondary clustering built in (v14+) — configure the primary once, and it replicates to the secondary automatically via a catalogue zone over an authenticated, encrypted channel. Combined with Unbound as the upstream recursive resolver, DNS queries get validated (DNSSEC) and resolved without depending on a third-party resolver like 1.1.1.1 or 8.8.8.8.
  1. Deploy two Technitium instances on VLAN 60 as plain Docker containers or LXCs — not inside the Talos cluster (see the callout below for why that's deliberate).
  2. Give each a distinct DNS_SERVER_DOMAIN under Settings → General before clustering — e.g. dns01 and dns02 — and make sure it's set correctly on both from the start; a stale or mismatched value here is the single most common cause of clustering looking "connected" one moment and "unreachable" the next.
  3. On the primary: Administration → Cluster → Initialize → New Cluster. On the secondary: Administration → Cluster → Initialize → Join Cluster, pointing at the primary's cluster URL. Without a real certificate yet, check "ignore certificate validation errors" to join.
  4. Deploy Unbound (can sit alongside Technitium on the same host or its own small LXC), configure it as a pure recursive/validating resolver.
  5. Point Technitium's forwarder at Unbound instead of a public DNS provider.
  6. Set OPNsense's DHCP to hand out the Technitium primary as DNS server 1, secondary as DNS server 2.

Once this is solid, start giving services proper names instead of remembering IPs — grafana.lab.internal, ai.lab.internal, proxmox.lab.internal — via Technitium zones, matching the DNS-first organisation approach in Section 09.

  • Distinct, correct DNS_SERVER_DOMAIN set on each node before clustering
  • Technitium primary + secondary clustered and confirmed syncing (refresh the page, not just the first "Connected" flash)
  • Unbound deployed, Technitium forwarding to it
  • DHCP handing out both Technitium instances, secondary tested by stopping the primary briefly
3b · Get the cluster domain right the first time
Intermediate~10 min, but read before Step 3, not after
Why this gets its own step: two separate Brandon Lee (VirtualizationHowto) write-ups — the original clustering setup and a follow-up troubleshooting post — both converge on the same root cause for clustering that looks fine for a moment and then goes "unreachable": the Cluster Domain field in the setup wizard is easy to get wrong, and the wrong value is genuinely hard to diagnose after the fact because the symptoms (DANE errors, intermittent NOTIFY failures, zone transfer refusals) all look unrelated to the actual cause.
  1. The Cluster Domain must be a domain you don't already host authoritatively anywhere else. Don't reuse lab.internal (already this build's general service-naming domain, per Section 09) — pick something separate and purely technical, e.g. dnscluster.local.
  2. Never use cluster.local specifically — that's Kubernetes' own reserved internal service-discovery suffix. This build has a Talos cluster planned (Section 08); even though Technitium isn't running inside it, a colliding domain is exactly the kind of thing that causes confusing failures months later for no obvious reason.
  3. Keep the cluster domain and this build's real internal naming domain (lab.internal) conceptually separate forever — the cluster domain is plumbing between the two DNS nodes, not something you resolve services through.
If you ever move Technitium into the Talos cluster later (matching the JimsGarage-style "everything in K8s" pattern), re-read both source articles in full first. The second one documents a specific, non-obvious set of Kubernetes traps: bind to 0.0.0.0:53 rather than a MetalLB VIP, configure TLS passthrough (not termination) for port 53443 if anything like Traefik sits in front of cluster traffic, and add the Kubernetes pod CIDR to Technitium's zone-transfer allow list — the source IP the primary sees will be a pod address, not the node's external IP. None of this applies to the plain Docker/LXC deployment above; it's flagged here so it's not rediscovered the hard way if the architecture changes.
  • Cluster domain confirmed distinct from both lab.internal and cluster.local
  • Understood: replication is one-way, primary → secondary only — administrative changes must always be made on the primary node
  • Understood: Technitium's own DHCP clustering isn't supported yet — not a gap for this build, since OPNsense handles DHCP per VLAN, not Technitium
4 · VLAN-wide DNS enforcement — make Technitium the only option
Intermediate~25 min
Why DHCP alone isn't enough: handing out Technitium via DHCP is a suggestion, not a rule — plenty of devices and apps hardcode a public resolver (8.8.8.8, 1.1.1.1) regardless of what DHCP says, quietly bypassing every block list and local name you've set up. This step turns the suggestion into an enforced rule at the firewall.
  1. OPNsense: Firewall → Aliases → create a Host(s) alias Technitium_Servers containing both Technitium IPs (dns01, dns02).
  2. Create a Port alias DNS_Ports = 53 (standard DNS), and a second Port alias DoH_DoT_Ports = 443, 853 (DNS-over-HTTPS / DNS-over-TLS) for use in the next step.
  3. On every VLAN's rule set, add an explicit allow rule near the top: destination = Technitium_Servers, port = DNS_Ports — this is what actually lets each VLAN reach DNS at all (already reflected in the traffic matrix in Section 05).
  • Technitium_Servers alias created with both IPs
  • Explicit allow-to-Technitium rule confirmed present on every VLAN
5 · Block third-party DNS & DoH bypass attempts
Advanced~30 min
What this actually stops, and what it doesn't: browsers (Firefox defaults to Cloudflare's DoH) and some OSes/apps will silently switch to an encrypted third-party resolver the moment plain port-53 DNS is blocked, unless that's blocked too. This is a genuine arms race — new DoH endpoints appear over time — so treat this as raising the bar significantly, not achieving a perfect seal. Pair it with turning off "secure DNS" / "private DNS" in browser and OS settings on devices you control, which is more reliable than any firewall rule.
  1. On every VLAN except 10 (Management) and 70 (VPN), add a block rule below the Technitium-allow rule from Step 4: destination = any, port = DNS_Ports (53) — this catches any device trying to reach a DNS server that isn't Technitium.
  2. Add a second block rule: destination = WAN, port 853 (DNS-over-TLS) — there's rarely a legitimate reason for a LAN device to speak DoT directly to the internet.
  3. Create an alias Known_Public_DoH with the IP ranges of the most common public DoH providers (Cloudflare, Google, Quad9, NextDNS) and block outbound port 443 to that alias. This is necessarily incomplete — treat it as blocking the defaults, not every possible DoH provider.
  4. Test from a client device: with a public resolver hardcoded, DNS lookups should now fail or time out, while normal browsing (through Technitium) keeps working.
  • Port 53 to anywhere-but-Technitium blocked on every non-management/VPN VLAN
  • Port 853 (DoT) blocked outbound to WAN
  • Known_Public_DoH alias created and blocked on port 443
  • Tested: a hardcoded public resolver fails from a client device on a restricted VLAN
  • Secure DNS / private DNS disabled in browser and OS settings on your own devices, as defence in depth
6 · Deploy Uptime Kuma
Beginner~10 min

Add a monitor per key service (Jellyfin, Nextcloud, Technitium, NPM...), point its notification channel at ntfy (Step 7) rather than Discord — one self-hosted notification path for the whole lab, not a third-party service in the alert chain.

  • Uptime Kuma deployed, core services monitored, at least one notification channel configured
7 · ntfy — one notification path for everything, no Discord required
Beginner~20 min
What this replaces: routing homelab alerts through Discord webhooks means a third-party chat platform sits in the middle of your monitoring, and its own outages or rate limits become your outages. ntfy is a small self-hosted push service — anything that can run a single curl command can push a notification straight to your phone, organised by topic (one for backups, one for security alerts, one for uptime), with no external dependency at all.
  1. Deploy via the official Docker image on VLAN 50, alongside the rest of the monitoring stack.
  2. Set NTFY_AUTH_DEFAULT_ACCESS=deny-all and create real users/tokens per topic — don't rely on an unguessable topic name as the only protection, which is a security-through-obscurity pattern this build avoids everywhere else (Section 07).
  3. Put it behind NPM with the shared wildcard cert, like every other service — ntfy.yourdomain.com, not plain HTTP.
  4. Install the ntfy app (F-Droid or Google Play — GrapheneOS-friendly either way) and subscribe to your topics.
  5. Wire it in as the notification target for Uptime Kuma, Wazuh, the MT5 trading VM's fast-notification requirement (Section 07 → Group C → Trading), and the NUT UPS shutdown script — one app, every alert.
# quick test once it's running
curl -H "Title: Test" -H "Priority: high" \
  -d "ntfy is alive" \
  https://ntfy.yourdomain.com/lab_alerts
  • ntfy deployed with deny-all default and real per-topic auth, not just an obscure topic name
  • Reachable over HTTPS through NPM with the shared wildcard cert
  • App installed and subscribed, test notification received
  • Uptime Kuma and Wazuh both pointed at it as their notification target
1 · Deploy Home Assistant
Intermediate~45 min
Why the VM install over the container/add-on route: Home Assistant OS as a VM (rather than Home Assistant Container) gives you the full add-on ecosystem plus clean USB device passthrough — important the moment you're plugging in a physical Zigbee, Z-Wave, or Bluetooth USB antenna, which a container can't easily reach.
  1. Deploy via the community Proxmox VE Helper Script for Home Assistant OS (VM variant) — it automates what would otherwise be a manual ISO install and virtual disk setup.
  2. Tag VLAN 90 (Testing) for now, per the existing VLAN table — see the note below on a dedicated IoT VLAN if your device count grows.
  3. Give it real resources if you're a heavy user: 2+ cores, 4GB+ RAM. Light users can go smaller.
  4. If you have Zigbee/Z-Wave/Bluetooth USB dongles, pass them through to the VM individually (Proxmox → VM → Hardware → Add → USB Device) rather than sharing a USB controller.
  • Home Assistant OS VM deployed via the helper script
  • Onboarding wizard completed, admin account created
  • Any USB radios (Zigbee/Z-Wave/Bluetooth) passed through and detected
2 · Multi-NIC for device discovery (optional)
Advanced~20 min, only if needed
The problem this solves: many smart-home discovery protocols (mDNS, SSDP/UPnP) don't cross VLAN boundaries by default, since routing between VLANs doesn't automatically forward multicast traffic. If you later add a dedicated IoT VLAN full of smart plugs and bulbs, Home Assistant may not "see" them automatically even though the firewall allows the traffic.

Two options, in order of preference: (1) enable multicast/mDNS relay on OPNsense between the relevant VLANs, or (2) give the Home Assistant VM a second virtual NIC directly on the IoT VLAN if one exists, so discovery happens locally on that segment. Start with option 1 — it's less invasive.

  • Not needed yet — VLAN 90 already reaches what HA currently manages (revisit only if a dedicated IoT VLAN is added)
3 · Deploy Homebridge (HomeKit bridge)
Beginner~20 min
What this is for specifically: if you're in the Apple Home/HomeKit ecosystem, Homebridge presents non-HomeKit-native devices (Switchbot, Elgato lights, LG TVs, and hundreds of others via community plugins) to Apple Home as if they were natively supported — a translation layer, not a replacement for Home Assistant. The two can run side by side; some people even bridge Home Assistant's own entities into HomeKit through Homebridge for a single unified Home app view.
  1. Deploy via the community Proxmox VE Helper Script for Homebridge (LXC variant — it's lightweight, doesn't need VM-level USB passthrough for most plugins).
  2. Tag VLAN 90, same as Home Assistant.
  3. Install plugins for your specific devices through the Homebridge web UI's plugin search.
  4. Pair it to Apple Home by scanning the QR code / entering the setup code shown in the Homebridge UI.
  • Homebridge LXC deployed
  • At least one device plugin installed and configured
  • Paired successfully in the Apple Home app
The pieces for this already exist in the design — Immich (VLAN 40), Nextcloud (VLAN 30), and Vaultwarden (VLAN 30, a self-hosted Bitwarden-compatible server — you already have this, just under its real name) are all already planned. What's new here is the phone-side setup: getting an Android phone — Nothing 2a today, a future Pixel or Motorola GrapheneOS phone later — actually talking to them, continuously, over NetBird.
1 · NetBird on the phone — always-on connection
Intermediate~20 min
Why "always on" is a real Android setting, not just leaving the app open: Android has a dedicated VPN mode that keeps a chosen VPN connected continuously and can restart it automatically — this is what makes the phone reachable from the lab and the lab reachable from the phone without you remembering to open anything first.
  1. Install NetBird — currently via Google Play or Aurora Store (Aurora works without a Google account). It isn't on F-Droid yet on GrapheneOS as of writing; there's an open community request to add it, worth checking again later.
  2. Open NetBird, enrol it against your self-hosted management server (same as the other peers in Section 07 → Group B → Remote access).
  3. Add it to a new, narrower NetBird access group — personal-devices — scoped to just VLAN 30 (Cloud) and VLAN 40 (Photos), not the full admin group's reach into VLAN 10/90. A phone doesn't need to manage OPNsense.
  4. Android Settings → Network & Internet → VPN → NetBird (gear icon) → enable Always-on VPN.
  5. Settings → Apps → NetBird → Battery → set to Unrestricted — without this, Android's battery management will periodically kill the background connection regardless of the always-on setting.
Known quirk, worth testing rather than assuming: there are scattered reports of NetBird's Android client occasionally dropping and needing a reconnect on GrapheneOS specifically, even with always-on enabled. Run it for a few days and check the connection state periodically before fully relying on it — and hold off on enabling Android's stricter "Block connections without VPN" option until you've confirmed reconnects are reliable, since that setting will cut all connectivity if the VPN service hiccups.
  • NetBird installed and enrolled
  • personal-devices group created, scoped to VLAN 30 + 40 only
  • Always-on VPN enabled, battery optimisation disabled for the app
  • Connection reliability observed over several days before treating it as fully dependable
2 · Immich — daily photo & video backup
Beginner~15 min
Why this is the right tool for photos specifically: Immich's mobile app does one-way, continuous camera-roll backup to your own server — the closest self-hosted equivalent to Google Photos' auto-backup, without your photos ever leaving your network.
  1. Install Immich from F-Droid (fully open-source build, GrapheneOS-friendly) or Google Play.
  2. Log in with the server URL — since the phone is always connected via NetBird, this can be the internal VLAN 40 address; it'll resolve whether you're home or out.
  3. Backup screen → select the Camera album (and Screenshots, Downloads, or anything else you want included).
  4. Enable background backup, and turn off the "Wi-Fi only" restriction if you want photos backing up over mobile data too, since the whole point is the phone is already tunnelled into the lab regardless of network.
  5. Settings → Apps → Immich → Battery → Unrestricted, same reasoning as NetBird.
Set the honest expectation here: Immich's own docs and issue tracker both note that background backup on Android can quietly stop after some days on certain phones and needs the app reopened to "wake" it back up — this is a widely reported Android background-execution limitation, not unique to Immich. Get in the habit of opening the app every day or two, and don't treat this as a zero-touch system the way Google Photos is.
  • Immich installed, camera album selected for backup
  • Background backup enabled, battery optimisation disabled
  • Confirmed a photo taken today actually appears on the server without opening the app first — then repeat that check after a few days
2b · Deploy Nextcloud itself — the server side, plus Collabora
Intermediate~40 min
A real gap worth closing: Nextcloud has been referenced throughout this build (VLAN 30, mobile backup target) without ever covering the server deployment itself, and without Collabora it's file sync only — not a real Google Drive/Docs replacement. Collabora Online adds in-browser document editing on top of Nextcloud's storage, which is what actually makes it a full office-suite alternative rather than just a folder that syncs.
  1. Deploy Nextcloud (AIO or the official Docker image) on VLAN 30, backed by the PostgreSQL instance on VLAN 80.
  2. Deploy Collabora Online alongside it (its own container — it's a separate service, not a Nextcloud plugin) and connect it via the Nextcloud admin panel's Collabora app.
  3. Confirm a document actually opens and saves for in-browser editing, not just file storage.
  • Nextcloud deployed, backed by PostgreSQL
  • Collabora connected, a document edits and saves in-browser

Worth knowing, not building: a full self-hosted mail server (Mailcow, Mail-in-a-Box) is a real option for digital independence, but most residential ISPs block outbound port 25 and dynamic IPs get blacklisted quickly — it's a genuinely harder project than everything else in this build and not one to take on from a home connection without first confirming the ISP allows it.

CryptPad, if a specific need shows up: Collabora gives Nextcloud in-browser editing, but the server can still technically read document content. CryptPad is zero-knowledge — encrypted client-side, so even a compromised server can't read what's stored. Not worth running as a second, separate document platform for everything; worth adding later specifically for documents where that distinction actually matters (nothing in this build currently needs it).

3 · Nextcloud + DAVx5 — contacts, calendar, and files
Intermediate~25 min
Why two apps, not one: the Nextcloud app itself handles file/folder sync (documents, music, anything that isn't a photo you're already routing through Immich). Contacts and calendar syncing into Android's actual native Contacts and Calendar apps needs a CardDAV/CalDAV sync client — DAVx5 is the standard companion for this, both pointed at the same Nextcloud server.
  1. Install Nextcloud and DAVx5, both from F-Droid.
  2. Nextcloud app: add account with your server URL, enable auto-upload for the folders you want backed up (Downloads, Documents, Music — anything not already covered by Immich).
  3. DAVx5: add an account, point it at the same Nextcloud server's CardDAV/CalDAV endpoint, let it discover your contacts and calendar collections.
  4. In Android's own Contacts and Calendar apps, confirm the Nextcloud-synced account now shows up as a real sync source — this is what makes it "just work" like any other synced account on the phone.
  • Nextcloud app syncing at least one folder
  • DAVx5 configured, contacts and calendar visible in Android's native apps
  • A test contact added on the phone confirmed to appear in Nextcloud's web UI (and vice versa)
4 · Seedvault — full app + app-data backup (GrapheneOS)
Intermediate~20 min
What this covers that Immich and Nextcloud don't: installed apps and their internal data/settings — the "everything" in your ask that isn't a photo, a contact, or a file. Seedvault is GrapheneOS's built-in backup system; on a stock/Motorola Android phone this specific tool isn't available, and app-level backup there is more limited without root — worth knowing that this step is GrapheneOS-specific, which is one more reason to prioritise that upgrade if full-device backup matters to you.
  1. Make sure the Nextcloud app is installed and logged in first (Step 3) — Seedvault backs up to any storage location exposed through Android's Storage Access Framework, and Nextcloud registers itself as one of those locations.
  2. Settings → System → Backup → set the backup location to Nextcloud via the SAF picker.
  3. Select which apps to include, confirm the encryption recovery code is generated, and store that code in Vaultwarden (Step 5) — without it, the backup is unrecoverable.
  • Seedvault backup location set to Nextcloud
  • Recovery code generated and stored in Vaultwarden, not just on the phone
  • One test restore attempted (even to a spare/secondary device) to confirm the backup is actually usable
5 · Vaultwarden mobile — your passwords, from any device
Beginner~15 min
This is the "host my own Bitwarden" part, already answered: Vaultwarden is a lightweight, fully API-compatible reimplementation of the Bitwarden server — the official Bitwarden apps work against it unmodified, you just point them at your own server instead of Bitwarden's cloud. Nothing to build here that isn't already planned; this step is just wiring the phone to it.
  1. Install the official Bitwarden app — available on F-Droid, fully open-source.
  2. On the login screen, tap the gear/settings icon and enable Self-hosted, entering your Vaultwarden server URL.
  3. Log in, enable biometric unlock, and turn on Bitwarden as the phone's autofill service (Settings → Passwords & accounts, or similar depending on Android version).
  • Bitwarden app pointed at the self-hosted Vaultwarden server, logged in
  • Biometric unlock and autofill enabled
  • Confirmed the same vault is reachable from a desktop browser too — one vault, every device
Purpose: most services in this design run continuously because something else depends on them being available (DNS, the reverse proxy, monitoring). A smaller set — an occasional admin tool, a demo environment, a game server, a rarely-used dashboard — does not need to run between uses. DockerWakeUp sits in front of that second group specifically: it starts a container's Compose stack on the first request that arrives for it, and stops it again after a configured idle period, without changing how the service itself is built or deployed.

Architecture

The request path has two edges deliberately kept separate. NGINX Proxy Manager — already the TLS-terminating edge for every other service in this design (Section 07 → Group B → Proxy & dashboard) — keeps that same role here: it owns the wildcard certificate, listens on port 80 (redirected to 443) and port 443 (HTTPS), and forwards matching hostnames to DockerWakeUp's internal proxy port rather than directly to a container. DockerWakeUp itself never terminates TLS and never sees a certificate; its only job is deciding whether the target container is already running before forwarding the request on.

Purpose: traces a single request through the state check, the cold-start path when the target is stopped, and the separate idle-checker loop that stops it again later — three logically distinct paths through the same proxy.

Two independent things happen on every request. First, the request path: DockerWakeUp checks whether the target's containers are running; if so, the request is forwarded immediately with no added latency worth measuring. If the target is stopped, DockerWakeUp runs docker compose up for that service's project, waits for it to start responding, then forwards the now-ready request — this is the "cold start" and its duration is entirely a function of the target application, not of DockerWakeUp itself. Second, and unrelated to any single request: a separate idle-checker process runs on a five-minute timer, independent of traffic, comparing each service's last-active time against its configured threshold, and runs docker compose down for anything that has exceeded it.

1 · Install prerequisites
Intermediate~20 min
Note on the upstream documentation: the project's quick-start script implies it installs its own dependencies; in practice it does not fully do so. Installing everything below first avoids the setup script failing partway through.
sudo apt update
sudo apt install -y nodejs npm nginx jq certbot
sudo apt install -y python3-certbot-dns-cloudflare

Docker and Docker Compose are already present from earlier steps in this design; NGINX here refers to the underlying package DockerWakeUp's setup script expects to find — NPM already provides this role for every other service, so this instance stays internal-only (see Step 3) rather than becoming a second public-facing NGINX.

  • Node.js, npm, nginx, jq, certbot installed
2 · Clone and configure
Intermediate~20 min
  1. Deploy on VLAN 15, alongside the other container-management tooling.
  2. Clone the repository and create the working config from the supplied example.
git clone https://github.com/jelliott2021/DockerWakeUp.git
cd DockerWakeUp
cp config.json.example config.json
nano config.json

Each entry in services maps a route name to the Compose project that runs it. The route name and domain combine into the public hostname NPM forwards for — route: "gitea-test" under domain lab.internal becomes gitea-test.lab.internal.

{
  "proxyPort": 8080,
  "idleThreshold": 3600,
  "domain": "lab.internal",
  "services": [
    {
      "route": "gitea-test",
      "target": "http://localhost:3001",
      "composeDir": "/home/claude/homelabservices/gitea-test"
    }
  ]
}
  • config.json created, at least one service entry defined
3 · Wire NPM to DockerWakeUp instead of the container directly
Advanced~20 min
Why this differs from the upstream install guide: the project's own documentation has its bundled NGINX terminate TLS directly via Certbot, running a second independent certificate-management path alongside the wildcard certificate NPM already manages for everything else. That duplication is avoidable here specifically because NPM already exists. Point NPM's proxy host at DockerWakeUp's internal port instead, and DockerWakeUp's own NGINX/Certbot setup in the upstream quick-start can be skipped entirely.
  1. In NPM, create a proxy host for each on-demand hostname (e.g. gitea-test.lab.internal), forwarding to DockerWakeUp's internal address on port 8080 — not to the target container's own port.
  2. Enable the existing wildcard certificate on the proxy host, exactly as for every other NPM-fronted service.
  3. Leave DockerWakeUp's own NGINX config disabled or uninstalled — one TLS-terminating edge for the whole lab, not two.
  • NPM proxy host forwards to DockerWakeUp's port, not the container's port
  • Existing wildcard certificate applied — no separate Certbot instance created
4 · Enable at boot and test a cold start
Beginner~15 min
  1. Run the setup script and enable the service.
chmod +x setup-service.sh
./setup-service.sh
sudo systemctl enable docker-wakeup

To confirm it works: manually stop the test service (docker compose down) and request its hostname in a browser — the request should stall briefly while DockerWakeUp starts the stack, then complete once the target responds.

  • docker-wakeup enabled at boot
  • A manually stopped service starts on its next request and serves the page

Idle threshold policy

The idle checker runs every five minutes regardless of threshold value; the threshold only controls how long a service is allowed to sit idle before that check stops it. There is no single correct value — it is a direct trade-off against cold-start time, set per service.

WorkloadSuggested idle threshold
Frequently used utility12–24 hours
Occasionally used application1–4 hours
Development / test environment30–60 minutes
Demo environment15–30 minutes
Game server15–60 minutes
What this is not a good fit for: Technitium, NPM itself, Authentik, and anything else something else depends on being reachable at all times. Waking DNS on demand means the request asking to wake it can't resolve the hostname to reach it in the first place — a genuine deadlock, not a hypothetical one. DockerWakeUp belongs in front of standalone, occasionally-used services only.

Alternative worth knowing: Sablier solves the same problem and is the more established, better-documented project of the two — genuinely worth evaluating first if DockerWakeUp's rougher setup experience (Step 1 above exists because of it) becomes a real obstacle.

Group C — Platform & advanced (Kubernetes, GitOps, AI, Trading, Security)
Read Section 08 before starting this group. These tabs are meaningfully heavier than Groups A/B — they cover things that genuinely stretch a single 32GB host. Section 08 has the resource budget and the honest reality check on what fits today versus what waits for a second node.
1 · Set up the repo structure
Beginner~20 min
Why this comes first, before any tool: Terraform, Ansible, and Argo CD are all just automation that reads a git repo and makes reality match it — none of them help if the repo itself is a mess. Get the folder structure right once, and everything after this tab just slots into it.

In soranthony5/homelab-deployment-v1.0, add:

/terraform/       - VM & LXC definitions (Proxmox provider)
/ansible/         - OS-level config & hardening playbooks
/kubernetes/      - Kustomize/Helm manifests, one folder per app
/docker/          - Compose stacks for non-k8s services, one folder per app
  jellyfin/
    docker-compose.yml
    .env
    README.md
/archive/         - old planning docs (already exists)
  • Folder structure created and pushed
2 · Terraform (OpenTofu) + Proxmox provider
Intermediate~45 min
What this replaces: clicking "Create VM" in the Proxmox web UI every time, and forgetting exactly what settings you used last time. A Terraform file describing a VM is the same VM every time you apply it — rebuild the whole lab from git if a disk dies.
  1. Install OpenTofu (or Terraform) on your workstation or a dedicated LXC.
  2. Configure the bpg/proxmox or Telmate/proxmox provider with an API token from Proxmox (not your root password).
  3. Write your first resource: recreate one existing LXC (e.g. Homarr) as code, plan, then apply against a throwaway test VM ID first.
  • Proxmox API token created, scoped (not root)
  • One VM/LXC successfully created purely from Terraform code
3 · Ansible for OS-level configuration
Intermediate~45 min
Division of labor: Terraform's job ends once the VM exists and boots. Ansible's job is everything after that — users, SSH keys, packages, hardening — applied consistently across every host instead of manually SSHing into each one and typing the same commands from memory.
  1. Install Ansible on your workstation or a dedicated LXC.
  2. Build an inventory file listing your hosts (pve, opn, docker hosts, etc.) grouped by role.
  3. Write a baseline playbook: SSH key auth, disable password login, install common packages, apply timezone/NTP.
  4. Run it against one host first, confirm idempotency (running it twice changes nothing the second time).
  • Inventory file created, grouped by role
  • Baseline playbook applied to at least one host, confirmed idempotent
4 · Terraform the Cloudflare Zero Trust config
Intermediate~30 min
Why put Cloudflare Tunnel/Access in code too: click-ops in the Cloudflare dashboard works fine to get started, but once you have several exposed hostnames and Access policies, tracking what changed and when — and being able to tear it all down and rebuild it identically — is exactly what Terraform is for. Same philosophy as the rest of this repo: infrastructure state lives in git, not just in your memory of what you clicked.
  1. Add the cloudflare/cloudflare provider to your Terraform config, authenticate with an API token scoped to just the Tunnel/Access/DNS permissions it needs.
  2. Define your tunnel, its public hostnames, and Access applications/policies as Terraform resources.
  3. plan, review, apply — then make your next Cloudflare change by editing code and re-applying, not by clicking in the dashboard.
  • Cloudflare API token scoped (not a global key)
  • At least one tunnel hostname and Access policy defined and applied via Terraform
5 · Optional: Packer golden images
Advanced~60 min
What this adds on top of Terraform alone: Terraform provisions a VM from some base image — Packer's job is building that base image itself, pre-patched and pre-configured, so every new VM starts from the same known-good template instead of a stock cloud image plus a long bootstrap script. This is the pattern the homelab-as-code project documents end-to-end: Packer builds an Ubuntu template on Proxmox → Terraform provisions VMs from it → Ansible configures the roles on top (Docker, Portainer, Traefik).
  1. Install Packer alongside Terraform and Ansible.
  2. Write a Packer template that builds a minimal, patched Ubuntu image on Proxmox and converts it to a VM template.
  3. Point your Terraform VM resources at that template instead of a raw ISO.
  • One golden image built and confirmed usable as a Terraform VM source (optional — skip if the repo above already covers what you need)
1 · Bootstrap a single-node Talos cluster
Advanced~60 min
What Talos is, and why start single-node: Talos is a minimal Linux built specifically to run Kubernetes and nothing else — no SSH, no shell, entirely managed through an API and a declarative config file (a genuine fit for the GitOps philosophy). A "single-node cluster" runs the control-plane and worker role on one VM — it's a completely valid, real Kubernetes cluster for learning, just without the multi-node fault tolerance you'd get later. See Section 08 for why this is the right starting size on this hardware.
  1. Download the Talos ISO, create a VM in Proxmox (suggested start: 4 vCPU, 8GB RAM, 60GB disk, tagged VLAN 15).
  2. Boot it — Talos comes up in maintenance mode with no config, listening on the API port.
  3. Generate a machine config with talosctl gen config, edit it to set the single node as both controlplane and worker.
  4. Apply the config, bootstrap etcd, pull down the kubeconfig.
  5. Confirm with kubectl get nodes — one Ready node.
  • Talos VM created and booted
  • Machine config applied, cluster bootstrapped
  • kubectl get nodes shows one Ready node
2 · Deploy Argo CD, point it at the repo
Intermediate~30 min
The GitOps loop, concretely: Argo CD continuously compares what's actually running in the cluster against what's declared in your /kubernetes repo folder, and reconciles any difference automatically. Push a change to git, Argo CD applies it — no kubectl apply from your laptop, no "which version is actually deployed" uncertainty.
  1. Install Argo CD into the cluster via its official manifests.
  2. Point it at soranthony5/homelab-deployment-v1.0, path /kubernetes.
  3. Deploy one trivial test app purely by pushing a manifest to git — confirm Argo CD picks it up without manual intervention.
  • Argo CD deployed and synced to the repo
  • One app deployed purely via a git push, no manual kubectl
3 · Traefik as the cluster's ingress
Intermediate~30 min
Traefik vs. NGINX Proxy Manager: NPM (Section 07 → Services B) stays as the reverse proxy for your "classic" VM/LXC/Docker services. Traefik does the equivalent job specifically for things running inside Kubernetes — it watches the cluster's API and automatically creates routes as you deploy new apps, which a manually-configured NPM host entry can't do. Two proxies, two clearly separate territories: outside the cluster = NPM, inside the cluster = Traefik.

Deploy via its official Helm chart or Argo CD manifest, expose it through a single OPNsense-forwarded port if you need external K8s ingress, or keep it internal-only.

  • Traefik deployed, one test Ingress resource routing correctly
4 · Longhorn — read the caveat before deploying
Advanced~30 min
Be honest about what this buys you on one physical host: Longhorn replicates persistent volumes across nodes for resilience — on a single-node cluster, "replicating across nodes" is meaningless, since there's only one node. It still gives you automatic PVC provisioning and easy snapshots/backups of stateful workloads, which is genuinely useful, but it does not protect you if the physical MS-01 dies. True Longhorn resilience is a good reason to prioritise the second Proxmox node in Section 10.
  1. Deploy Longhorn via its Helm chart or Argo CD manifest.
  2. Use it for stateful workloads that benefit from snapshots (databases, Argo CD itself) — not for bulk media, which stays on the NAS over NFS regardless.
  • Longhorn deployed, one PVC provisioned and tested
  • Understood and accepted: single-node = no cross-node resilience yet
Flagging this now, for when node 2 arrives: once Longhorn is actually replicating across multiple nodes, its replication traffic is exactly the kind of chatty east-west storage traffic that should not share a VLAN with light management or general container traffic — mixing them is one of the most common home-lab VLAN mistakes (bursty storage traffic causing latency spikes on whatever else shares the network). At that point, budget a dedicated, unrouted storage VLAN for Longhorn specifically. If jumbo frames (MTU 9000) are used to speed it up, that VLAN must stay entirely within one layer-2 segment — the moment jumbo-frame traffic gets routed through OPNsense (a layer-3 device that may not support the larger MTU), it silently fragments or drops. Neither of these applies yet on a single node; this is purely a note for the future.
5 · Flatcar Linux — the testing counterpart
Intermediate~30 min
Where this fits: per the production-vs-testing separation principle in Section 09, this is where you try a new container image or a new k3s/Docker experiment before it earns a place in the Talos GitOps repo. Flatcar is another immutable, container-focused Linux — a reasonable choice specifically because it's different enough from Talos to genuinely test portability, not just a second copy of the same thing.

Deploy as a VM tagged VLAN 90 (Testing), named dockertest01 per the naming convention in Section 09. Use it for anything you're not ready to commit to the GitOps repo yet.

  • Flatcar VM deployed on VLAN 90, named per convention
1 · Deploy Ollama
Beginner~20 min
Set expectations on hardware: the MS-01 has no discrete GPU — inference runs on the i9-13900K's CPU cores (genuinely capable for this) or lightly accelerated via the iGPU. That's fine for small, modern quantized models; it is not going to match a machine with a real GPU for large models. Right-size the model to the hardware, not the other way around.
  1. Deploy Ollama as a container (VM/LXC tagged VLAN 15).
  2. Pull a realistic starting model: llama3.2:3b, phi3.5, or qwen2.5:7b-q4_K_M — all usable on CPU without a long wait per response.
  3. Test with a direct API call before adding a UI on top, so you know the model itself works.
  • Ollama deployed, at least one small quantized model pulled and responding
2 · Deploy Open WebUI
Beginner~15 min

ChatGPT-style web frontend for Ollama — chat history, multiple models, document upload for RAG-style question answering over your own files. Point it at the Ollama instance from Step 1, reverse-proxy it via NPM or Traefik as ai.lab.internal.

  • Open WebUI deployed, connected to Ollama, reachable via a clean URL
3 · Optional: iGPU passthrough for a speed boost
Advanced~45 min
Why this is worth the effort: unlike a full discrete GPU passthrough (which locks the GPU to one VM exclusively), the MS-01's integrated GPU can be shared across multiple LXCs at once via cgroup device passthrough — meaning the same iGPU can accelerate both Jellyfin's hardware transcoding and lend a hand to Ollama, without picking one or the other.

Passthrough /dev/dri into the Ollama and Jellyfin LXCs (not full VM passthrough). Confirm with intel_gpu_top on the host that both containers can access it.

  • iGPU passthrough configured on relevant LXCs
  • Confirmed both Jellyfin transcoding and Ollama can use it
4 · Optional: self-hosted AI workflows (n8n)
Intermediate~30 min

n8n is a self-hosted workflow automation tool with native Ollama/OpenAI-compatible nodes — a natural fit for things like "summarise new Paperless-ngx documents" or "generate a daily digest for Homarr" using your local model instead of a cloud API.

  • n8n deployed (optional), one simple workflow built as a test
1 · Create the isolated Windows VM
Intermediate~60 min
Why Windows, and why so isolated: most MetaTrader Expert Advisors and custom indicators still assume a genuine Windows runtime — a Wine-based Linux container saves resources but risks subtle EA compatibility bugs, which is a bad trade-off when real money is involved. VLAN 35 exists specifically so this VM has zero lateral access to anything else in the lab — a compromise here can't reach your NAS, your database, or anything else, and nothing else can reach it except for administration.
  1. Create a VM: suggested 4 vCPU, 8GB RAM, 80GB disk, tagged VLAN 35.
  2. Install Windows 10/11, apply all updates, keep Defender enabled.
  3. Install MetaTrader 5 from your broker's official installer only.
  • VM created on VLAN 35, confirmed no lateral access to other VLANs
  • Windows fully patched, Defender active
  • MT5 installed from the broker's official source
2 · Lock down the firewall rule for this VLAN
Advanced~20 min
Why outbound needs restricting too, not just inbound: a default "allow all outbound" rule on this VLAN would still let a compromised EA or a malicious indicator phone home anywhere. Scoping outbound to just your broker's known hostnames/IP ranges and NTP is the difference between "isolated" and "isolated except for the one thing that matters."
  1. On OPNsense, restrict VLAN 35 outbound to: your broker's server hostnames (use FQDN aliases if supported), NTP, and DNS (VLAN 60) only.
  2. Confirm no rule allows VLAN 35 to reach any other internal VLAN, including Management.
  3. Admin access to this VM is RDP only, only reachable via the NetBird admin group — never a WAN port-forward.
  • Outbound restricted to broker + NTP + DNS only
  • No inter-VLAN rule exists for VLAN 35 in either direction except the exceptions above
  • RDP reachable only via NetBird, never WAN-exposed
2b · Endpoint hardening — this VM is a Windows endpoint holding real money
Intermediate~30 min
Why this gets more attention than any other VM in the build: network isolation stops lateral movement, but the VM itself is still a full Windows install running third-party Expert Advisors and indicators — code you may not have written. Treat it the way a security-conscious admin treats any endpoint handling money: locked down, watched, and quarantined automatically on unusual behaviour rather than trusted by default.
  1. Enforce a USB device whitelist (Group Policy or a third-party tool) — block unknown USB devices by default rather than trusting anything plugged in.
  2. Apply restrictive Group Policy: no access to the C: drive directly, no ability to install new programs, no unnecessary system changes — reduces what a compromised EA or indicator could actually do even if it tried.
  3. Keep endpoint AV mandatory and active — treat AV being disabled or out of date as an automatic "something is wrong" signal, not a background nicety.
  4. Enable Windows Event Logging for USB connections and blocked actions, forwarded to Wazuh (Section 07 → Group C → Security & identity) alongside everything else.
  • USB whitelist enforced
  • Group Policy restrictions applied (no C: drive access, no arbitrary installs)
  • AV active, logging forwarded to Wazuh
3 · Power & backup priority
Intermediate~20 min
Why this VM deserves special treatment in your shutdown plan: an ungraceful power loss mid-trade is a worse outcome than almost anything else in this lab — an open position left in an unknown state is a real financial risk, not just an inconvenience. A clean shutdown, or better, EA logic that closes/pauses on connection loss, matters more here than for any other VM.
  1. Confirm this VM is included in the NUT-triggered graceful shutdown sequence (Section 04), and check whether your specific EA has a documented safe-disconnect or pause behaviour.
  2. Add an Uptime Kuma monitor on this VM specifically, notification target set to ntfy (Section 07 → Group B → Monitoring & DNS, Step 7) with a dedicated high-priority topic — given real money is at stake, you want to know within minutes, not hours, and ntfy's priority flag can make that alert impossible to miss.
  3. Set a PBS backup schedule for this VM, and always take a manual snapshot before changing any EA or indicator.
  • VM included in graceful shutdown sequence
  • Fast-notification Uptime Kuma monitor configured
  • PBS backup schedule set, manual pre-change snapshot habit established
4 · Optional: browser-based access via Guacamole
Intermediate~30 min

Apache Guacamole gives you RDP/VNC access to the MT5 VM through a plain browser tab (still tunneled through NetBird) — no native RDP client needed on whatever device you're checking from.

  • Guacamole deployed (optional), RDP session to the MT5 VM tested through it
This tab is what makes "zero trust" real rather than aspirational. Everything else in this document controls which network a device is on. This tab controls who is allowed to do anything, requires them to prove it (MFA), and writes down that they did it (logging) — regardless of which VLAN they're coming from.
1 · Deploy Authentik — centralized identity & forward-auth
Advanced~60 min
What this replaces: right now, "logged into NPM" and "logged into Jellyfin" and "logged into Grafana" are all separate logins with separate passwords, some possibly shared or reused. Authentik sits in front of every proxied service as a single identity provider — one real account (yours), with MFA, and every login attempt against every service logged in one place. This is the difference between a home network and an actual access-controlled system.
  1. Deploy Authentik (server + worker + its own dedicated Postgres + Redis — keep this separate from the shared Database VLAN 80 instance to avoid extra firewall exceptions) on VLAN 10.
  2. Complete the setup wizard, create your admin account with a strong unique password from Vaultwarden, enable MFA (TOTP or a hardware key such as a YubiKey) immediately, then disable or delete the wizard's default admin account if one remains.
  3. Create an admins group (just you, for now) and a separate lower-privilege group as a template for any future limited access — never grant a new user or service account admin-group membership by default.
  • Authentik deployed, own admin account created with MFA enabled
  • Default/wizard admin account disabled or removed
  • admins group created, scoped to just you
Honest tradeoff, worth stating plainly: Authelia is genuinely lighter than Authentik — no Postgres/Redis backing, a single small binary, lower resource draw. It's a fair criticism that Authentik is heavyweight for a home lab this size, and Section 08's resource budget already reflects that cost. It stayed the choice here for two concrete reasons that matter for this build specifically: a proper web GUI for managing users/groups/events (Authelia's config is YAML-file-based, a steeper ask given the "explain everything" learning goal from Section 07's intro), and the RAC provider, which gives browser-based RDP/SSH/VNC to hosts like the MT5 trading VM without a separate tool. If the resource budget ever gets genuinely tight, Authelia is the reasonable downgrade path — not a sign this choice was wrong, just a different point on the weight-vs-features tradeoff.
2 · Wire NPM / Traefik to require Authentik before every request
Advanced~45 min
Forward-auth, explained: before NPM or Traefik forwards a request to the actual backend service, it first asks Authentik "has this browser session already authenticated, and are they allowed here?" — only a yes lets the request through. An unauthenticated visitor never even reaches Jellyfin's or Grafana's own login page; they're stopped at the proxy.
  1. Configure Authentik's "Proxy Provider" (forward-auth mode) for each service you want gated.
  2. In NPM, add the forward-auth snippet to each proxy host's advanced config (or the equivalent Traefik forward-auth middleware for anything running in the Talos cluster).
  3. Start with low-risk services (Homarr, Grafana) before gating anything you rely on daily, in case the config needs tuning.
  • At least one proxied service confirmed redirecting to Authentik login when unauthenticated
  • Same service confirmed accessible once logged in
3 · Deploy Wazuh — logging & SIEM
Advanced~90 min
Why one central place for logs matters: right now, "what happened and when" is scattered across a dozen hosts' local logs — useless during an actual incident when you need answers in minutes, not after SSHing into six different machines. Wazuh centralizes logs and alerts from every host into one searchable, correlated view.
Resource note: this is the single heaviest addition in this update — budget 4GB+ RAM for the manager/indexer alone. Re-check the resource budget in Section 08 before deploying this alongside everything else already running.
  1. Deploy the Wazuh manager + indexer + dashboard on VLAN 50 (Monitor).
  2. Install the Wazuh agent on Proxmox and any Docker/Talos hosts that support it.
  3. Forward OPNsense's syslog output to the Wazuh manager.
  4. Forward Authentik's authentication events to Wazuh as well, so every login — successful or failed — lands in the same place as everything else.
  • Wazuh manager/indexer/dashboard deployed on VLAN 50
  • Agents installed on Proxmox and at least one other host
  • OPNsense syslog and Authentik events both flowing into Wazuh
4 · Enable Suricata on OPNsense
Intermediate~30 min
What this adds above the firewall rules already in place: your firewall rules answer "is this destination allowed" — Suricata answers "does this traffic look malicious even if it's technically allowed," by matching packets against known attack signatures as they cross the firewall. It runs as a plugin directly on OPNsense, so there's no separate VM to provision.
  1. Install the os-suricata plugin from OPNsense's plugin manager.
  2. Enable it on the WAN interface first, subscribe to a free ruleset (Emerging Threats Open is the standard starting point).
  3. Run in IDS mode (alert-only, doesn't block anything) for at least a week to see what fires and tune out false positives.
  4. Only switch specific rule categories to IPS (blocking) mode once you've confirmed they're not flagging legitimate traffic.
  5. Forward Suricata's alerts to Wazuh (Step 3) for a single combined view.
  • Suricata enabled on WAN, ET Open ruleset subscribed
  • Run in IDS-only mode for at least a week before enabling any IPS blocking
  • Alerts forwarded to Wazuh
4b · CrowdSec alongside Suricata — crowd-sourced blocking
Intermediate~30 min
How this complements Suricata rather than duplicating it: Suricata matches known attack signatures in traffic content. CrowdSec instead watches behaviour (repeated failed logins, scanning patterns) and cross-references against a crowd-sourced reputation list of IPs currently misbehaving elsewhere on the internet — the two catch different things, and running both is a genuinely common OPNsense pairing rather than redundant.
  1. Install the CrowdSec OPNsense plugin, enrol it against the free community blocklist.
  2. Point it at OPNsense's own logs plus Suricata's alerts as scenarios to watch.
  3. Start in alert/log mode before enabling active blocking, same caution as Suricata's IPS step.
  • CrowdSec deployed, community blocklist enrolled
  • Confirmed logging real detections before switching to active blocking
4c · Optional: ClamAV filesystem scanning
Beginner~20 min
Where this fits: Suricata and CrowdSec watch network traffic; ClamAV instead scans the filesystem itself for known malware signatures — catching something that already landed (a bad download, an infected file synced in) rather than something crossing the wire. Not real-time, but a daily scheduled pass costs little and closes a gap the network-level tools don't cover.

Schedule a daily scan on Proxmox and/or the NAS shares, notify on detection via email or a Wazuh forward.

  • ClamAV scheduled scan configured on at least one host (optional)
5 · Periodic vulnerability scanning — OpenVAS & Trivy, on-demand
Intermediate~45 min per scan cycle
Why on-demand, not always-on: OpenVAS is resource-heavy and only earns its keep during an active scan — leaving it running 24/7 burns RAM for no ongoing benefit. Spin it up on VLAN 90 for a scan, review results, tear it back down.
  1. Deploy OpenVAS/Greenbone via Docker on VLAN 90 only when actively scanning.
  2. Run a full scan against your own VLANs from inside the lab — this is what Phase 1 of Section 14's Security Audit uses.
  3. Add Trivy as a pre-deploy check in the Argo CD/GitOps pipeline (Section 07 → Group C → GitOps foundation) to scan container images for known vulnerabilities before they ship.
  • OpenVAS run at least once, results reviewed, container torn down afterward
  • Trivy wired into the GitOps pipeline as an optional pre-deploy check
6 · Zeek — protocol-level network analysis, alongside Suricata
Advanced~60 min
How this differs from Suricata: Suricata matches traffic content against known attack signatures. Zeek instead produces structured, protocol-level logs of network activity — DNS queries, HTTP requests, TLS handshakes, connection metadata — regardless of whether anything matched a signature. The two are commonly deployed together rather than as alternatives: Zeek's logs provide the historical record to investigate an incident after the fact, including activity that did not trigger any signature at the time.
  1. Deploy Zeek on VLAN 50, with a network tap or mirrored port on the switch feeding it traffic to observe.
  2. Forward Zeek's logs to Wazuh, alongside Suricata's alerts and every other log source already centralised there.
  • Zeek deployed, receiving mirrored traffic (optional — resource cost should be weighed against the budget in Section 09)
  • Logs forwarded to Wazuh
7 · Atomic Red Team — testing whether the detections actually detect
Advanced~30 min per test cycle
The gap this closes: Wazuh, Suricata, and CrowdSec being deployed is not evidence that they would catch a real attack technique — a rule can be present and still fail to match, or an alert can fire without a notification path behind it. Atomic Red Team is a library of small, individually-scoped tests, each mapped to a specific MITRE ATT&CK technique, that safely reproduce the behaviour a real technique would exhibit without carrying out an actual attack. Running one and confirming it produced a Wazuh alert and an ntfy notification is a direct test of the detection pipeline end to end, rather than an assumption that it works.
  1. Run Atomic Red Team tests only against the Testing VLAN sandbox (VLAN 90), never against production services.
  2. Select a small number of tests relevant to the actual attack surface here — credential access, a simulated brute-force pattern against SSH, an unusual outbound connection — rather than the full library.
  3. For each test run, confirm the expected alert appeared in Wazuh and the expected notification arrived via ntfy. A test that produces no alert is a genuine finding — a detection gap to fix, not a step to skip past.
  • At least one test run against the Testing VLAN, confirmed to produce a Wazuh alert
  • Confirmed the alert produced a notification via ntfy, not just a log entry

08Security & remote access

Default-deny firewall

All inter-VLAN traffic blocked unless a rule explicitly allows it. Management (VLAN 10) reaches everything for administration; everything else is scoped to what it actually needs.

Isolation examples

Jellyfin → NAS: allowed
Photos → Internet: blocked
Media ↔ Photos: blocked
Trading → anything but broker/NTP/DNS: blocked
Management → everything: allowed

Zero trust — identity, privilege & logging

VLANs and firewall rules control which network a device sits on. This layer controls who is allowed to act, on top of that — the actual "zero trust" piece, since it doesn't matter which VLAN a request comes from if it can't authenticate.

Single identity, everywhere

Authentik sits in front of every proxied service (Section 07 → Group C → Security & identity). One real account — yours — with MFA required, instead of a dozen separate app logins with varying password strength.

Least privilege by default

An admins group exists containing only you. Any future account — family, guest, a service account — starts in a restricted group and is granted access explicitly, never inherits admin by default.

Full activity logging

Wazuh centralizes logs from every host, OPNsense's syslog, and every Authentik login attempt (successful or failed) into one searchable place — not scattered across a dozen hosts you'd have to SSH into individually during an incident.

Defence in depth on appliance UIs

OPNsense, Proxmox, and the NAS keep their own native logins + 2FA (forward-auth doesn't cleanly wrap appliance UIs) — but they're only reachable via NetBird or Cloudflare Access in the first place, so there are two layers to get through, not one.

Remote access — two independent paths

Primary — NetBird: replaces a plain point-to-point VPN or a third-party-hosted mesh. It's WireGuard under the hood (fast, modern crypto) with a self-hosted coordination server so peers connect directly to each other instead of routing through someone else's cloud. Access is grouped, not all-or-nothing — your admin laptop joins an admin group reaching VLAN 10 (management) and VLAN 90 (testing), while your phone joins a narrower personal-devices group reaching only VLAN 30 (Cloud — Nextcloud, Vaultwarden) and VLAN 40 (Photos — Immich) for always-on backup, per Section 07 → Group B → Phone & mobile backup. A phone doesn't need, and shouldn't have, the same reach as the device you administer the firewall from.

Backup — Cloudflare Tunnel: an architecturally unrelated path — outbound-only, no WireGuard, no inbound firewall rule at all. Cloudflare Access gates it with login before any request reaches your network. Kept deliberately separate from NetBird so one path's outage or misconfiguration never takes down the other.

Full config steps for both are in Section 07 → Group B → Remote access.

Hardening baseline

  • SSH disabled by default on every device, enabled only when actually needed, key-only (no password) — a YubiKey FIDO2 resident key is a further step up, since the private key never touches disk at all
  • MFA everywhere, admin logins first — Authentik, OPNsense, Proxmox, NAS, and NetBird all support it
  • Default credentials changed on every device before go-live, default/wizard admin accounts disabled once your own account exists
  • A named personal admin account for daily use on OPNsense, Proxmox, and the NAS — not root/admin directly — with elevation only when a task actually needs it
  • Where a device supports its own local firewall (Proxmox host, NAS), scope it to the specific VLAN subnets that need it — and narrower still, to specific source IPs for the most sensitive management ports, not the whole subnet
  • HTTPS everywhere via NGINX Proxy Manager + internal CA or Let's Encrypt DNS challenge — one wildcard cert (Section 07 → Services → Proxy) covers everything rather than managing certs per-service
  • fail2ban or equivalent on internet-facing services — extended explicitly to OPNsense's own admin login, NetBird, and Cloudflare Access, not just "the internet-facing stuff"
  • UFW (or nftables) enabled on every Linux VM/LXC individually, not just relied on at the VLAN level — set default-deny inbound, allow only the specific LAN subnet each service actually needs. This is genuine defence in depth: VLAN isolation stops traffic between networks, a host firewall stops traffic from other devices on the same VLAN that shouldn't be talking to this one host specifically. Fail2ban belongs on the same hosts, watching SSH and any exposed login forms for repeated failures.
  • Automatic security updates enabled where safe to do so; everything else reviewed manually against release notes before applying, so a known-bad update doesn't roll out unattended
  • NetBird access groups reviewed whenever a new peer or service is added
  • NanoKVM web UI never port-forwarded to WAN — reachable only via NetBird's admin group
  • Every login — success or failure — lands in Wazuh; nothing authenticates silently
  • Client-side ad/tracker blocking as a second layer on top of Technitium's network-level filtering — uBlock Origin in every browser, AdGuard on mobile devices. Malicious ads are a common entry point network-level DNS filtering alone doesn't fully close.
  • Changing default ports is not treated as a security measure here — once a device is on the network, a port scan finds the real port trivially. VLANs, firewall rules, and authentication do the actual work; obscurity isn't a substitute for any of them. SSH is not reachable on any port from outside NetBird's tunnel, which is a stronger position than moving it off port 22 while still leaving it reachable.
Discoverability and compromise are separate problems. A public IP address or a certificate in a transparency log will be found by automated scanning within minutes to hours, as a baseline condition of being connected to the internet, regardless of how obscure the domain or how small the network. That fact alone does not constitute a security failure. Documented incidents such as the DeadBolt ransomware campaign against QNAP devices were caused by a specific combination — UPnP auto-forwarding an admin interface, plus unpatched software behind it — not by the existence of a public-facing service in general. This build addresses both root causes structurally: UPnP is disabled (Section 07 → OPNsense, Step 6b), so nothing is auto-forwarded; Watchtower and the patch-review cadence (Section 07) keep exposed software current; and there are zero inbound port-forwards in the first place, so the question of what an unpatched exposed service could do does not arise for anything other than the two deliberately exposed, authenticated remote-access paths.
Full periodic re-check of all of this — an actual AUDIT → HARDEN → SEGMENT → RECOVER cycle, not a one-time checklist — lives in Section 14. Run it before this lab is ever exposed to the internet, and again on a quarterly cadence after that.

09Kubernetes, GitOps & AI platform

This section is shaped by two real sources: Jim's Garage's GitOps-first homelab pattern (Terraform + Ansible + Talos + Argo CD), and Brandon Lee's VirtualizationHowto lab, which runs this exact stack — Talos, Flatcar, Argo CD, Traefik, Longhorn, Technitium, Unbound.

Read this before Section 07 → Group C. Brandon Lee's lab running this stack is a 5-node cluster with 480GB of combined RAM across dedicated MS-01 nodes. This build has one MS-01 with 32GB. Everything below is still genuinely achievable — Kubernetes, GitOps, and local AI are all real and valuable at this scale — but it has to be right-sized, not copied 1:1. The honest budget is below.

Resource budget — one 32GB host

WorkloadRealistic footprintNotes
Proxmox host overhead~2GBFixed cost, always-on
Existing LXC services~6–8GBHomarr, NPM, Technitium×2, Unbound, Uptime Kuma, Vaultwarden, etc. — individually light
Media + *arr stack~3–4GBJellyfin, Sonarr, Radarr, Prowlarr, Bazarr, Jellyseerr, qBittorrent
Single-node Talos cluster~8–10GBOne VM running control-plane + worker roles combined
Ollama + Open WebUI~6–10GBDepends entirely on model size — budget for a 7B-class quantized model, not larger
MT5 Windows VM~6–8GBOnly when actively trading — see Section 07 → Group C → Trading
Security stack~6–8GBAuthentik + Postgres/Redis (~2GB), Wazuh manager/indexer (~4GB+, the heaviest single piece), Suricata (runs on OPNsense itself, minimal extra). OpenVAS deliberately excluded — on-demand only, not always-on.

Add the middle columns up and you're already past 32GB before the AI stack, MT5 VM, or security stack even start — running everything simultaneously at generous allocations doesn't fit. Two honest paths forward, not mutually exclusive:

  • Stay right-sized: single-node Talos (not a 3-node HA cluster), small quantized LLMs (3B–7B, not 70B), a minimally-sized Wazuh deployment, and accept that the AI stack, MT5 VM, and full security stack aren't all running flat-out at the same moment as a media transcode
  • Accelerate the second Proxmox node from Section 10 specifically to carry the K8s cluster, the AI workload, or the security stack separately — this is the real, permanent fix, not a workaround. Wazuh in particular is a strong candidate to be the thing that justifies node 2.

Architecture decisions

DecisionChoiceWhy
Kubernetes distroTalos LinuxImmutable, API-managed, no SSH — matches the GitOps philosophy directly. This is the production cluster.
Testing OSFlatcar LinuxA second immutable OS, deliberately different from Talos, for trying new containers before they earn a place in the GitOps repo (VLAN 90)
Reverse proxy splitNGINX Proxy Manager (classic VM/LXC/Docker) + Traefik (inside K8s only)Two clear territories rather than one tool awkwardly covering both
K8s persistent storageLonghorn, with caveatsUseful for PVC provisioning/snapshots now; genuine cross-node resilience waits for node 2
Bulk/media storageStays on the NAS over NFSLonghorn is for stateful app data, not for the media library
GitOps engineArgo CD, watching /kubernetes in the repoPush to git, cluster state follows automatically
Infra provisioningTerraform (OpenTofu) for VMs/LXCs, Ansible for OS configRebuild the whole lab from the repo if hardware is replaced
DNS platformTechnitium DNS (clustered pair) + UnboundNative clustering instead of sync scripts; Unbound as a private validating resolver
Local AIOllama + Open WebUI, small quantized modelsNo discrete GPU — CPU inference is genuinely usable at the right model size
Identity / access controlAuthentik (forward-auth), Wazuh (logging/SIEM), Suricata (IDS on OPNsense)One identity everywhere, one place logs land, one layer inspecting traffic content — see Section 07

GitOps flow

Push to git — two paths reconcile automatically: infrastructure (Terraform/Ansible → Proxmox) and workloads (Argo CD → Talos)

Full step-by-step for all of this — bootstrapping Talos, deploying Argo CD/Traefik/Longhorn, standing up Ollama/Open WebUI, and the MT5 trading VM — lives in Section 07 → Group C.

10Home lab organisation & Docker best practices

Organisational structure and Docker networking practice, informed by established home lab documentation — the two primary factors determining whether a system of this scale remains manageable as it grows.

Organise by function, not technology

Devices and services are grouped by function rather than by underlying technology — not "Docker VMs" versus "Kubernetes VMs" versus "Windows VMs," but by what each one does. This grouping is applied as a Proxmox tag on every VM/LXC and as the top-level folder structure in the git repository:

Role tagWhat lives there
infrastructureProxmox, OPNsense, switch config, PBS
networkingTechnitium, Unbound, NetBird, NPM, Traefik
monitoringPrometheus, Grafana, InfluxDB, Uptime Kuma, Loki
aiOllama, Open WebUI, n8n
mediaJellyfin, Sonarr, Radarr, Prowlarr, Bazarr, Jellyseerr
storageNAS shares, Longhorn, NFS/iSCSI exports
identityVaultwarden, Authentik
securityWazuh, Suricata, OpenVAS (on-demand)
automationArgo CD, Terraform, Ansible, Gitea
tradingMT5 VM — kept deliberately alone, nothing else shares this tag or this VLAN
testingFlatcar host, sandbox VMs, anything not yet in the GitOps repo

Predictable hostnames

Boring, functional names beat clever ones — you should know what a host does from its name alone, six months from now, half-asleep, during an outage.

HostRole
pve01The Proxmox host itself
opn01OPNsense
sw01MikroTik switch
nas01Terramaster NAS
kube-cp01Talos control-plane (+worker, single-node)
dockertest01 / test01Flatcar testing host / Gitea + Paperless-ngx Docker host
ai01Ollama + Open WebUI + NetBird (Docker)
mgmt01Homarr, NPM, Authentik, NetBox (Docker)
media01 / idx01Jellyfin + *arr stack / Prowlarr (Docker)
cloud01 / photos01Nextcloud + Vaultwarden / Immich (Docker)
trade01MT5 Windows VM
dns01 / dns02Technitium primary / secondary

Git folder structure mirrors the infrastructure

Every Docker Compose service gets its own folder, same shape every time — migrate a service or rebuild a host by copying one folder:

/docker/
  jellyfin/
    docker-compose.yml
    .env
    README.md
    config/
    data/
    backups/
  homarr/
    docker-compose.yml
    .env
    README.md
    ...

DNS-first naming

Once Technitium is live (Section 07 → Services → Monitoring & DNS), stop remembering IPs — create a zone and give every service a real name: grafana.lab.internal, ai.lab.internal, proxmox.lab.internal, argocd.lab.internal. NPM and Traefik both route on these names, and it's the single biggest daily-usability improvement in a lab this size.

Nine checks before trusting a Docker image

The vulnerability scanning in Section 07 (Trivy, OpenVAS) catches known CVEs in an image already pulled. These checks happen earlier — before an unfamiliar image is deployed at all, when the real question is not "is this patched" but "should this be trusted."

1. Publisher identity

Confirm the image is the project's own, not a look-alike fork — check the official docs for the canonical registry path before pulling. docker buildx imagetools inspect shows what a reference actually resolves to.

2. Active maintenance

Recent commits, a documented vulnerability-reporting process, reviewed dependency-update PRs, and images built automatically from tagged releases — an unmaintained image is a growing liability regardless of how clean it looks today.

3. The Dockerfile itself

When available, check the base image and watch for a bare curl | sh pulling and executing a remote script during build — not automatically malicious, but worth knowing it's there.

4. Image metadata and layers

docker image inspect and docker image history surface the default user, exposed ports, volumes, and what each layer actually did before a container is ever started from it.

5. Runtime privileges requested

Treat privileged: true, network_mode: host, or a mounted docker.sock in a Compose file as a question, not a default — does this specific application actually need it?

6. Vulnerability scan

Trivy or Docker Scout against the pulled image — already the default in this design (Section 07 → GitOps foundation) for anything going through the pipeline; worth running manually too for a one-off pull.

7. Signature verification

Where a project signs its images (cosign is the common tool), verifying the signature confirms the image matches what the developer actually published — closing the specific gap a scan alone can't: supply-chain tampering after the fact.

8. Tag vs. digest

A tag can be silently moved to point at a different image later; a digest (@sha256:...) identifies exact content. Pinning by digest for anything security-sensitive means what gets deployed tomorrow is provably what was tested today.

9. Test in isolation first

VLAN 90 (Testing) already exists for exactly this — run an unfamiliar image there first, optionally with --network none --read-only --cap-drop ALL, and watch what it actually tries to do before it reaches a real VLAN.

Docker networking & host mistakes to avoid

1. Publishing every port

Only publish the port users/browsers actually hit (the frontend, or Traefik/NPM). Backend containers reach each other over Docker's internal network — no port mapping needed.

2. One giant network

Give each app stack its own dedicated Docker network. Only the frontend container joins the shared proxy network too.

3. Exposing databases

Don't publish 5432/3306/6379 to the host. The app reaches postgres:5432 by service name on the shared internal network — no port needed at all.

4. Overusing host networking

Reserve it for the few tools that genuinely need to see the host's wire (e.g. network monitors). Everything else stays on bridge networking for real isolation.

5. Skipping the external proxy network

Create one dedicated external network for NPM/Traefik. App stacks connect their frontend to it and keep their own private network for backend traffic — the only way in is through the proxy.

6. Hardcoding IPs / subnets

Use Docker's built-in service-name DNS, never a container's IP. Let Docker auto-assign subnets — manually picking them is how you end up with silent overlaps down the line.

7. Trusting a host firewall to catch Docker traffic

Docker manipulates iptables directly and can silently bypass UFW or a host-level firewall rule — a ports: mapping in a compose file can expose a service even when the host firewall "should" have blocked it. VLAN isolation (Section 05) is the real control here, not a host firewall layered on top of Docker.

8. Running containers/LXCs privileged by default

Privileged mode grants root-level host access and skips normal isolation. Default to unprivileged/rootless for every container and LXC in this build, and only grant privileged mode to the specific few that genuinely need device passthrough (e.g. iGPU sharing for Jellyfin/Ollama) — never as a default habit to sidestep a permissions error.

Keep production and testing genuinely separate

This build already does this structurally: VLAN 90 (Testing) plus the Flatcar dockertest01 host is where anything new gets tried first. It only gets promoted into the Talos GitOps repo — becoming "production" — once it's actually proven out. Never test directly against the Talos cluster or the core LXC services.

Separate Linux users per service, not one account for everything

On any host running multiple services directly (rather than fully containerized), give each function its own unprivileged Linux user — one for media, one for backups, one for sandbox/testing work, and so on — rather than running everything as a single account. If one service gets compromised, the blast radius is whatever that one user can touch, not everything on the box. None of these service accounts should have sudo; elevate deliberately, per the named-admin-account principle in Section 07.

11Future expansion & HA readiness

Nothing below needs building today — it's what this design already leaves room for.

Second Proxmox node

RU9 is reserved specifically for this. A second node turns Proxmox from a single box into a real cluster capable of live-migrating VMs. Proxmox HA clustering wants an odd number of voting members for quorum — a 2-node cluster needs a lightweight third vote (a QDevice, which can be as small as a Raspberry Pi) rather than a full third server.

Shared storage for HA

Live migration and automatic VM failover need storage both nodes can see. The NAS's NFS/iSCSI exports (VLAN 999) already provide this path — no redesign needed when you add node 2. Ceph across nodes is the other common option if you later add local disks to each node instead of relying on the NAS.

10GbE headroom

The CRS310 has exactly 2 SFP+ ports, both spoken for (OPNsense + Proxmox). A second hypervisor node — or true LACP bonding — needs a switch with more 10GbE ports. Worth planning for when a second node becomes real rather than buying ahead of need.

3-2-1 backups

Proxmox Backup Server (Section 07) plus the NAS gives you 2 local copies. Add a third, offsite copy — a cloud target (Backblaze B2, etc.) or a drive rotated to another location — to complete a real 3-2-1 strategy rather than "two copies in the same rack."

Deliberately not in this build — optional career-skills sandboxes, not needs: a Windows Server + Active Directory domain (domain join, GPOs, hybrid identity with Entra ID) is genuinely valuable if the goal is CV-building enterprise IT experience, but it doesn't serve this lab's actual purpose and would add a Windows attack surface with no functional payoff here. Same logic for a full self-hosted mail server (Mailcow) — real digital independence, but residential ISPs blocking port 25 and dynamic-IP blacklisting make it a genuinely harder project than everything else in this build. Worth knowing both exist as legitimate paths; neither is a gap in this design.

Rough order of operations, when you're ready

  • Go live on the current single-node design first — don't build for HA before you have a workload that needs it
  • Add NanoKVM once NetBird is running, so remote hands-off recovery actually works end-to-end
  • Bring up the single-node Talos cluster and AI stack only after the core services (Section 07 Groups A/B) are stable — Section 08 has the resource math on why not everything runs at once yet
  • Add a QDevice (Raspberry Pi is plenty) before a second Proxmox node, so quorum works from day one of the cluster
  • Add the second node into RU9, cable it the same way as Proxmox today (10GbE if the switch supports it by then, 1GbE otherwise) — this is also what unlocks genuine multi-node Talos and real Longhorn resilience
  • Re-run the weight check in Section 02 every time something physical gets added

12Build checklist

Five phases, done in order. Expand a phase, check items off as you go by editing this file directly.

Phase 1 — PhysicalCabinet mounted, all hardware and UPS installed, cables verified, safe power-on
  • Cabinet wall-mounted into studs, level, secure — confirm total load stays under 50kg
  • UPS installed at RU 1–2 (bottom), switch/shelves/NAS installed per rack layout above
  • All 1GbE fallback cables + 2× 10GbE DAC cables (owned) + PDU cabling installed and labelled both ends
  • Cable verification pass (tug test, no pinches, DAC cables not sharply bent)
  • Power-on sequence: switch → NAS → OPNsense → Proxmox (last — highest inrush)
  • All LEDs verified green (incl. SFP+ links), console access confirmed on each device
Phase 2 — NetworkFirewall configured, 15 VLANs live on the 10GbE trunk, switch trunked correctly
  • OPNsense SFP+ interface identified and set as VLAN trunk parent (em0 kept as fallback)
  • 15 VLAN subinterfaces created with correct gateways
  • DHCP scopes enabled per VLAN
  • Switch SFP+1/SFP+2 added to trunk bridge, ports 4 & 5 set to access
  • Default-deny rules in place, then explicit allow rules per VLAN
  • Cross-VLAN traffic tested against the isolation matrix
Phase 3 — Compute & storageProxmox and NAS ready to host services
  • Proxmox installed, VLAN-aware trunk bridge (vmbr0) created on the SFP+ NIC
  • NAS RAID 5 configured, shares created, NFS export to Proxmox
  • First test VM created, tagged VLAN 90, reaches its gateway and the NAS
  • Storage throughput checked (>100MB/s target on 1GbE NAS link)
Phase 4 — Core services & add-onsMedia, cloud, DNS, monitoring, dashboard, and proxy stacks live
  • Media stack: Jellyfin, Sonarr, Radarr, Prowlarr
  • Cloud stack: Nextcloud, Vaultwarden
  • DNS: Pi-hole + AdGuard Home (secondary)
  • Monitoring: Prometheus, Grafana, InfluxDB, Uptime Kuma, Loki
  • Backing databases: PostgreSQL, Redis/MariaDB on VLAN 80
  • NGINX Proxy Manager + Homarr dashboard deployed on VLAN 10
  • NetBird self-hosted server deployed on VLAN 70, admin devices enrolled
Phase 5 — Backup, UPS & hardeningProduction-ready and resilient to power loss
  • UPS installed, NUT configured for automated graceful shutdown
  • UPS battery test performed — confirm devices survive a simulated outage
  • Proxmox Backup Server deployed, daily incremental + weekly full jobs scheduled
  • Restore tested at least once, end to end
  • Third, offsite backup copy configured (3-2-1)
  • Security hardening checklist (Section 07) fully applied
  • This document updated to reflect the final as-built state, changelog entry added
Phase 6 — Platform expansion (optional, later)GitOps foundation, Kubernetes, AI stack, and MT5 trading — after everything above is stable
  • Repo restructured for Terraform/Ansible/Kubernetes per Section 08
  • Terraform + Proxmox provider provisioning at least one real VM/LXC
  • Ansible baseline playbook applied across existing hosts
  • Single-node Talos cluster bootstrapped, Argo CD deployed and synced to the repo
  • Traefik and Longhorn deployed inside the cluster (with the single-node caveat understood)
  • Ollama + Open WebUI deployed, at least one small model tested end to end
  • Jellyseerr + Bazarr added to complete the media automation pipeline
  • MT5 VM deployed on isolated VLAN 35, firewall rules locked to broker + NTP + DNS only, included in backup and shutdown priority
  • Resource budget in Section 08 re-checked against actual usage once everything is live

13Operations quick reference

Daily

Glance at Homarr/Grafana, check overnight backup succeeded, confirm no failed services, check Uptime Kuma status page.

Weekly

Review firewall logs, check pending updates, verify UPS self-test passed, review NetBird peer list.

Monthly

Full PBS restore drill, storage capacity review, rotate any credentials due for rotation.

On power loss

NUT should already be handling this automatically — check the shutdown log afterward to confirm it triggered cleanly.

14Security audit & recheck

This is not a one-time checklist — it's a cycle. Run it in full before this lab is ever exposed to the internet, and again on a quarterly cadence after that. Structured around AUDIT → HARDEN → SEGMENT → RECOVER, cross-checked against the homelabstarter network security audit and the community homelab-hardening-checklists repo — worth periodically diffing this section against that repo as it's updated.

Untested = not a backup, unaudited = not secure. Every item below is a real check, not a box to tick from memory. If you can't currently verify an item, that's the honest finding — write it down and fix it, don't check it anyway.
AUDIT — run before going public, then quarterlyFind out what's actually true about the lab, not what you assume is true
  • nmap scan of the WAN IP from outside the LAN (mobile hotspot, not home wifi) — confirm only the Cloudflare Tunnel's outbound connection shows, no unexpected open ports
  • Cross-check with canyouseeme.org — a single-port web-based check independent of the device or network running the nmap scan above, useful as a quick sanity check without needing a second network to scan from
  • Check shodan.io for the home WAN IP — confirm nothing is indexed
  • Review every OPNsense port-forward rule — delete anything you don't immediately recognise the reason for
  • Inventory every exposed service (Cloudflare Tunnel hostnames) — confirm each sits behind Cloudflare Access and Authentik, not just one
  • Have I Been Pwned check for every admin email used across OPNsense, Proxmox, Authentik, NetBird, Cloudflare
  • User/account audit — any accounts still active that shouldn't be? Any default/wizard accounts not yet disabled?
  • Trivy scan of container images currently in use (Section 07 → Group C → Security & identity)
  • OpenVAS scan run against the lab's own VLANs, results reviewed, container torn down afterward
  • Firewall rule review — for each VLAN, does every allow rule still have a reason to exist?
  • Every .env / compose secret for every deployed service checked against its default — a surprising number of self-hosted projects ship a working default password/key that's easy to miss on a quick setup
  • netstat / listening-socket check on Proxmox and any Docker hosts — confirm nothing is listening that you didn't intend to run
HARDEN — passwords, keys, patchingClose the gaps AUDIT just found
  • Every credential is 20+ characters, generated and stored in Vaultwarden — no exceptions, no reused passwords
  • MFA enabled everywhere it's supported, admin logins first (Authentik, OPNsense, Proxmox, NAS, NetBird, Cloudflare)
  • SSH hardened: key-only auth, root login disabled, fail2ban (or equivalent) active — consider a YubiKey FIDO2 resident key so the private key never touches disk
  • Patch review completed — Proxmox host, every VM/LXC OS, and container images all checked for pending updates
  • Cloudflare Access confirmed in front of every admin UI that's reachable through the tunnel
  • Unused services/plugins disabled — less running surface, less to keep patched
SEGMENT — VLANs, DNS, firewallConfirm the isolation design is still actually true, not just documented
  • Traffic matrix in Section 05 re-tested against reality — pick 3 random VLAN pairs and confirm allowed/blocked matches what's documented
  • DNS deny-by-default still enforced — a hardcoded public resolver still fails from a restricted VLAN (Section 07 → Services → Monitoring & DNS)
  • Home Assistant / IoT devices confirmed still isolated to VLAN 90, no unexpected new devices present
  • Any camera/RTSP traffic (Frigate, once added) confirmed on VLAN 45 only — zero routing to any other VLAN, dual-NIC pattern (like the NAS) used if Frigate itself needs a management-reachable interface
  • VLAN 35 (Trading) re-confirmed: zero lateral access in either direction except broker/NTP/DNS
RECOVER — backups, tests, practiceAn unverified backup is a belief, not a fact
  • Nightly PBS backup job verified — confirm the last 3 actually completed, not just scheduled
  • Test restore performed — pull a real backup, mount it, actually read the data back
  • 3-2-1 rule confirmed: 3 copies, at least 1 offsite
  • At least one backup copy is immutable — can't be deleted or modified for a set retention window, so ransomware reaching the lab can't also delete the way back out
  • Backups are encrypted, and the decryption key is stored somewhere other than the lab itself
  • A recovery document exists that someone else (or a future, panicked you) could actually follow start to finish
  • A simulated full host-loss drill practiced at least once — not just backups existing, but proof they'd actually get you back online

Asset hardening log

A lightweight record per device, filled in the first time it's hardened and updated whenever something changes — the difference between "I think I did this" and being able to check. Copy this table and extend it as devices are added. Once NetBox is deployed (Section 07 → Group B → Proxy & dashboard), it can take over as the living version of this same information — retire the manual table at that point rather than keeping both in parallel.

DeviceIPMACHardened byDateNotes
opn0110.0.1.1
sw0110.0.1.5
pve0110.0.1.x
nas0110.0.30.10
Tracking runs: copy this line and update it after each full pass — Last full audit: 20XX-XX-XX · Findings: ... · Fixed: ... — keep the history in the changelog below so the audit trail is as real as the lab itself.

REFExternal references

Every official documentation link used throughout this document, grouped by subsystem, with the tab or section each is cited from. Each entry opens the source directly, for verification or for going deeper than this document's own summary.

LinkCited in
Firewall & perimeter (OPNsense)
CrowdSec OPNsense integration07 → Security & identity
NAT07 → OPNsense
OPNsense Suricata / IDS-IPS07 → Security & identity
OPNsense VLAN configuration07 → OPNsense
OPNsense firewall rules07 → OPNsense
OPNsense install & first boot07 → OPNsense
OPNsense interface assignment07 → OPNsense
Diagram design system
diagram-design — architecture & data flow visual system05 → Architecture view · 06 → Data flow
Switch (MikroTik)
MikroTik bridge VLAN filtering07 → MikroTik switch
MikroTik interfaces07 → MikroTik switch
Home devices switch (Netgear JGS524E)
JGS524E installation guide07 → Home switch (JGS524E)
Netgear support — JGS524E07 → Home switch (JGS524E)
Hypervisor & Proxmox tooling
Original tteck repository (archival reference)07 → Proxmox
PVE Post Install script07 → Proxmox
Proxmox Backup Server docs07 → Backups
Proxmox Helper Script — HA OS VM07 → Home automation
Proxmox Helper Script — Homebridge LXC07 → Home automation
Proxmox VE Helper-Scripts index07 → Proxmox
Proxmox VE installation07 → Proxmox
Proxmox network configuration07 → Proxmox
NAS (TOS / TrueNAS reference)
TOS downloads & firmware07 → NAS
TerraMaster TOS overview07 → NAS
TrueNAS Community Edition documentation (reference/comparison)07 → NAS
Out-of-band management (NanoKVM)
NanoKVM ATX wiring guide07 → NanoKVM
NanoKVM quick start07 → NanoKVM
NanoKVM user guide07 → NanoKVM
Remote access (NetBird / Cloudflare)
Cloudflare Access policies07 → Remote access
Cloudflare Tunnel docs07 → Remote access
Cloudflare Tunnel homelab walkthrough07 → Remote access
Load-balanced Cloudflare Tunnels with Docker Swarm07 → Remote access
NetBird Android install07 → Phone & mobile backup
NetBird identity providers07 → Remote access
NetBird self-hosted guide07 → Remote access
Reverse proxy & dashboard
Homarr getting started07 → Proxy & dashboard
NGINX Proxy Manager guide07 → Proxy & dashboard
NetBox Docker07 → Proxy & dashboard
Traefik quick start07 → Kubernetes (Talos)
Watchtower documentation07 → Proxy & dashboard
DNS (Technitium / Unbound)
Technitium DNS Server07 → Monitoring & DNS
Unbound documentation07 → Monitoring & DNS
a follow-up troubleshooting post07 → Monitoring & DNS
the original clustering setup07 → Monitoring & DNS
Media automation
Bazarr wiki07 → Media stack
Jellyfin documentation07 → Media stack
Jellyseerr documentation07 → Media stack
Prowlarr wiki07 → Media stack
Radarr wiki07 → Media stack
Sonarr wiki07 → Media stack
Home automation
Home Assistant installation07 → Home automation
Homebridge documentation07 → Home automation
Mosquitto documentation06 Service interconnection
Mobile backup & Vaultwarden
DAVx507 → Phone & mobile backup
GrapheneOS backup (Seedvault)07 → Phone & mobile backup
Immich mobile backup07 → Phone & mobile backup
Nextcloud Android app07 → Phone & mobile backup
Nextcloud server install07 → Phone & mobile backup
Vaultwarden wiki07 → Phone & mobile backup
Trading VM access
Apache Guacamole docs07 → Trading (MT5)
Kubernetes platform
Argo CD getting started07 → Kubernetes (Talos)
Flatcar installation07 → Kubernetes (Talos)
Longhorn installation07 → Kubernetes (Talos)
Talos getting started07 → Kubernetes (Talos)
AI stack
Ollama Docker guide07 → AI stack
Open WebUI documentation07 → AI stack
n8n documentation07 → AI stack
Security & identity
Atomic Red Team07 → Security & identity
Authentik Docker Compose install07 → Security & identity
Authentik proxy provider / forward-auth07 → Security & identity
ClamAV documentation07 → Security & identity
Greenbone/OpenVAS Docker07 → Security & identity
Trivy installation07 → Security & identity
Wazuh Docker deployment07 → Security & identity
Zeek documentation07 → Security & identity
Monitoring & notifications
Beszel getting started07 → Monitoring & DNS
Grafana getting started07 → Monitoring & DNS
Prometheus overview07 → Monitoring & DNS
Uptime Kuma wiki07 → Monitoring & DNS
ntfy documentation07 → Monitoring & DNS
GitOps & infrastructure-as-code
Ansible getting started07 → GitOps foundation
Cloudflare Terraform provider07 → GitOps foundation
Collabora Online Docker07 → Phone & mobile backup
Homelab as Code — Merox07 → GitOps foundation
Packer documentation07 → GitOps foundation
bpg/proxmox Terraform provider07 → GitOps foundation
homelab-as-code07 → GitOps foundation
Security audit tools
AUDIT → HARDEN → SEGMENT → RECOVER14 Security audit
Have I Been Pwned14 Security audit
canyouseeme.org14 Security audit
homelab-hardening-checklists14 Security audit
homelabstarter network security audit14 Security audit
shodan.io14 Security audit
External design references
5 Home Lab VLAN Mistakes07 → Kubernetes (Talos); 07 → MikroTik switch
Jim's Garage09 Kubernetes/GitOps/AI
VirtualizationHowto lab09 Kubernetes/GitOps/AI
Other
Rclone documentation07 → Backups

15Changelog

Add a row here every time hardware, VLANs, or major config actually changes. Newest on top.

v9.12026‑08‑31 Replaced public GitHub Pages hosting with Cloudflare Pages, built from a private GitHub repository and gated by Cloudflare Access — the page now requires an authenticated login from an explicit allow-list, not just an unlisted URL. The /diagrams/* path is deliberately exempted from that gate: initial research proposed relying on Cloudflare's authentication cookie to also cover the diagram embeds, but this was checked against Cloudflare's own community forum before being implemented and found to be a documented failure mode — the cookie does not reliably reach a resource loaded inside a cross-origin iframe, which is exactly how the diagrams.net viewer fetches each file. Gating that path would either break every diagram or require real engineering effort with no guaranteed outcome, so it stays public by explicit, stated design — the diagram files contain no credentials, only architecture information already fully described in the gated text. Also corrected the diagram embed URL format itself: verified against three independent official draw.io sources that the target file belongs after a #U hash fragment, not a ?url= query parameter as the previous version used — the earlier format would not have rendered any diagram at all. Added a Gemfile so Cloudflare's build environment installs Jekyll correctly, and rewrote Section 16 and the accompanying HOW-TO-RUN-THIS-SITE.txt to document the actual current setup rather than the superseded GitHub-Pages-only instructions.
v9.02026‑08‑30 Two structural changes, both reversing earlier approaches that weren't working in practice. First: every diagram is now live-embedded directly from its single .drawio file via the public diagrams.net viewer — the paired .svg export files (16 of them, added in the previous diagram-editability pass) are deleted entirely. Editing the .drawio now updates the page with no export step, ever. The trade-off is stated directly in script.js where the base URL lives: this only renders once the page is actually hosted (GitHub Pages already covers this) — it will not render when index.html is opened locally by double-click, since the viewer is a remote service fetching the file over the internet. The Remote Access diagram's two custom "toggle path" buttons — broken by this change, since JavaScript can no longer reach inside a cross-origin viewer iframe — are removed; the diagram was rebuilt with two real draw.io layers instead, toggled natively from the viewer's own layers panel. Second: the single 3,553-line index.html is split into a 46-line Jekyll shell plus 21 _includes/ files (one per section, with the largest section — Configuration Steps — further split into intro/Group A/Group B/Group C so no single file exceeds ~900 lines). GitHub Pages runs Jekyll automatically; no build step was added. Verified by writing an include-simulator and diffing its output against the pre-split original: every difference was whitespace or a restored organisational comment, confirmed by a full tag-balance and tab-wiring check against the reconstructed page, not just the source files.
v8.22026‑08‑27 Expanded the virtualization strategy subsection into a complete placement breakdown, classifying every service named in Section 06's master table into native LXC, Docker (grouped by seven new named Docker hosts — mgmt01, media01, idx01, cloud01, photos01, ai01, test01 — each tagged to the same VLAN as the services it runs, consistent with every other per-service VLAN assignment already documented), or full VM. Added a diagram showing the three categories side by side, deliberately including that VLAN 50's monitoring stack (Prometheus, Grafana, Wazuh, Beszel) runs natively with no Docker host at all — not every service benefits from containerising. Documented the Docker host OS decision explicitly (Debian 13, matching every other host, glibc-based, not the immutable Flatcar/Talos pattern used elsewhere since these hosts need ad hoc native package installs and SSH-based troubleshooting). Documented the specific mechanics of running Docker inside an unprivileged Proxmox LXC — the nesting and keyctl features, and why privileged mode is unnecessary on current Proxmox versions despite older guides claiming otherwise. Updated the Section 10 hostname table to include the newly-named Docker hosts.
v8.12026‑08‑27 Added a virtualization strategy subsection to Section 07, stating explicitly (rather than leaving implicit across a dozen scattered per-service choices) the decision framework for LXC versus Docker versus VM versus Kubernetes, and the reasoning for why Docker runs inside Proxmox rather than on a separate bare-metal box — the planned second node in Section 11 is the answer to more Docker capacity, not an unmanaged box outside this design. Added a new Group B tab, on-demand containers, covering DockerWakeUp: a diagram of the request path (NPM remains the single TLS-terminating edge; DockerWakeUp only decides whether to cold-start a container before forwarding), the config format, an idle-threshold policy table by workload type, and an explicit warning against using it on anything else depends on being always-up (DNS, the reverse proxy, Authentik). Sablier noted as the more mature alternative worth evaluating first. Updated the Watchtower step: the original project is now deprecated, with Dockcheck (mag37/dockcheck) noted as the actively maintained replacement — the correct repository was verified by search before linking, after an initial guess at the wrong one was caught in review. Added a nine-point Docker image trust checklist to Section 10, distinct from the vulnerability scanning already in Section 07 — publisher identity, maintenance signals, Dockerfile review, runtime privilege requests, signature verification, and digest pinning, closing the "should this be trusted" question a CVE scan alone doesn't answer.
v8.02026‑08‑23 Redesigned five diagrams to the diagram-design system introduced in v7.9 — physical rack layout, physical & core network topology, service discovery, event-driven MQTT messaging, and GitOps flow — consolidating the topology diagram with the now-redundant "architecture view" companion into one. Every diagram was rendered to a static image and visually inspected before being accepted, not just assumed correct from the generating code, which caught four genuine bugs: (1) the data flow diagram's reported text overflow, caused by fixed 100px node widths that didn't fit real content — replaced with content-measured auto-sizing and a hard pre-render assertion that halts generation if any label wouldn't fit its box; (2) a z-order bug present in three diagrams where arrow labels were drawn before their nodes, letting a node's opaque fill paint over part of the label text; (3) a duplicated "DATA" prefix on every lane label in the data flow diagram, left over from copying the type spec's own example; (4) a genuine layout bug in the rack diagram where the frame's top/bottom boundary variables were swapped, producing a viewBox a quarter of the height actually needed and silently clipping most of the diagram. The topology diagram's connection to the JGS524E is now routed around the lab zone rather than through it, matching its actual isolation. Structural validation (tag balance, marker/pattern ID uniqueness, zero orphaned references, zero broken anchors) confirmed clean after every change.
v7.92026‑08‑23 Added two diagrams built to the uploaded diagram-design system's exact specification rather than an approximation of its look — dark-skin tokens, Instrument Serif/Geist/Geist Mono typography, kind-tag chips, single-bend orthogonal connectors, and a bottom legend strip, all per that project's own style guide and type references. Section 06 gained an end-to-end data flow diagram (the explicit ask): built to the Data Flow type's parametric grid — four role lanes (Client, Edge, Service, Data) across six stages, with exactly one focal step, node, and labelled arrow as that type's rules require. Section 05 gained an architecture-view diagram, a direct structural reskin of the referenced architecture.png (client → edge → focal distribution point → grouped compute/storage zone), applied to OPNsense, the MikroTik switch, Proxmox, and the NAS. Both are self-contained SVG with proper title/desc accessibility metadata per that system's contract. Scope note: this covers the two diagrams requested; the remaining 13 existing diagrams remain in the document's original visual language and were not converted in this pass.
v7.82026‑08‑19 The JGS524E (added in v7.7) was present in the VLAN table, traffic matrix, logical swimlane diagram, and configuration steps, but absent from three other diagrams — corrected. Physical & core network topology: added as a node connected from MikroTik port 7, with a downstream "Household devices" node representing the 23 ports serving the television, personal computer, and similar. VLAN zone map: added as the 15th trust zone. Switch port diagram: port 7 updated from "reserved" to show the JGS524E trunk. Added a genuinely new capability rather than just a diagram fix: a dedicated port on the JGS524E (Step 3b) carrying VLAN 10 instead of VLAN 100, giving an admin device plugged in at that switch's physical location the same lab-troubleshooting reach as the existing admin laptop drop on the MikroTik switch — every other port remains VLAN 100 only, with no path into the lab. Verified zero duplicate SVG marker IDs and zero orphaned marker references across all 13 diagrams after these edits.
v7.72026‑08‑19 Added the Netgear ProSAFE Plus JGS524E (24-port gigabit, owned hardware) as a satellite switch for household devices — television, personal computer, and similar — isolated from the lab on a new VLAN 100 (Home Devices). 14→15 VLANs throughout: table, 15×15 traffic matrix, and logical diagram updated, with VLAN 100 forced through Technitium DNS (inheriting ad-blocking) and Suricata/CrowdSec inspection identically to every other VLAN, and given zero access to or from any lab VLAN beyond the standing Management-VLAN exception. Added a new Hardware configuration tab (Section 07 → Group A) covering the switch's ProSAFE Plus configuration model (Windows utility or web interface, not RouterOS-class), 802.1Q trunk setup to MikroTik port 7, the corresponding OPNsense firewall rules, and an optional outbound-VPN policy-routing step for the household devices specifically, explicitly distinguished from NetBird's unrelated inbound-access role. Noted the switch's physical placement (not rack-mounted, satellite location) and that it is not covered by the rack's UPS. Updated the master service interconnection table and References appendix accordingly.
v7.62026‑08‑19 Full audit pass. Fixed a genuine HTML validity issue found by an automated check: SVG marker IDs (arrowheads used across the diagrams) were duplicated across all 13 diagrams, since each was authored independently with the same marker names — every marker ID is now unique per diagram, with all references updated to match, verified with zero duplicates remaining. Corrected a factual inconsistency in the stated hardware build order: Section 07's own guidance previously listed Proxmox before the NAS, contradicting both Section 00's build order and Proxmox's own NFS-mount step, which depends on the NAS already existing — the NAS tab now precedes Proxmox, with an explicit prerequisite note added to the NFS-mount step itself. Expanded the NAS configuration procedure substantially: an explicit TOS-versus-TrueNAS comparison (appliance model versus self-built ZFS system, and the resulting difference in built-in data-integrity guarantees), a native snapshot step distinct from Proxmox Backup Server's role, and official TerraMaster documentation links. Added a new References appendix cataloguing all 92 external documentation links used throughout the document, grouped by subsystem with the specific tab or section each is cited from, addressable independently of the changelog. Verified zero broken internal anchors, zero duplicate element IDs, and correct section-heading sequencing across the full document.
v7.52026‑08‑17 Added a current software versions reference table to Section 06 (Proxmox VE 9.2, OPNsense 26.7, Talos 1.13.x, Technitium 15.4.x, and related base operating systems), sourced from each project's own release channel rather than DistroWatch directly, since DistroWatch's own catalogue covers operating systems and not most of the application-layer services in this design — TrueNAS is documented there for reference despite not being in use, since the NAS in this design runs Terramaster's own OS. Added canyouseeme.org as a secondary, independent check alongside the existing nmap step in the security audit procedure. Added a PVE post-install script step and a general explanation of the Proxmox VE Helper-Scripts project (community-scripts, successor to tteck's original repository) to the Proxmox configuration procedure, with the trust model — review before running — stated once rather than repeated per script. Reviewed a curated blue-team home lab resource list; added Zeek (protocol-level network logging, complementing Suricata's signature matching) and Atomic Red Team (safe, MITRE ATT&CK-mapped tests to verify the existing detection stack actually produces an alert and a notification, not just that it is deployed) to the security and identity configuration group. The majority of the reviewed list — Active Directory-adjacent tooling, MISP, TheHive/Cortex — was assessed as enterprise-SOC-scale tooling disproportionate to this design and not added.
v7.42026‑08‑17 Renamed to WarGreymon. Checked against an XDA report on automated scanning of newly exposed home servers, and its top comment's critique that discoverability alone is not a failure mode. Added a perimeter hardening step to the OPNsense configuration procedure: UPnP and NAT-PMP disabled (the specific mechanism behind the DeadBolt ransomware campaign against QNAP devices), an outbound threat-intelligence block list (Spamhaus DROP) added so a compromised host cannot reach a command-and-control address, and optional WAN GeoIP filtering. Documented the wildcard certificate's secondary benefit against Certificate Transparency log enumeration. Added an explicit statement distinguishing discoverability from compromise, and the specific root causes this design addresses structurally. Corrected three stale section cross-references left over from the v7.0 renumbering, found by systematically checking every in-document section reference against its linked anchor.
v7.32026‑08‑15 Renamed to Gardudan. Layout: removed fixed-pixel width caps on body paragraphs across the document, which were preventing text from filling the available content column. Rebuilt the VLAN zone map — the centre firewall node was undersized for its label and connection lines terminated at the node's centre point rather than its edge; the node is now sized to its content with an opaque fill, and every line is drawn to the boundary of its target box. Expanded the service discovery explanation with the resolution mechanism described in full, and both diagrams in that section now state their purpose explicitly in their captions. Applied Australian English spelling throughout the document's prose content (CSS and code samples deliberately excluded, since "color" and equivalent terms are literal syntax there, not a spelling choice). Revised the document's voice: removed second-person address, conversational asides, and references to the document's own revision history from technical description text, in favour of declarative third-person statements of fact and design rationale. This pass covered Section 06 in full and representative sections elsewhere; a complete pass across every configuration step remains outstanding.
v7.22026‑08‑15 Closed the last two items from the 10-layer stack list. CryptPad noted alongside Nextcloud/Collabora — zero-knowledge client-side encryption that Collabora doesn't provide, worth adding later for specific documents where that distinction matters, not as a second general document platform. Twingate addressed the same way as Entra ID: a legitimate free-tier zero-trust option, passed over because its control plane is Twingate's cloud rather than self-hosted, which is the actual deciding factor over NetBird — not a feature gap.
v7.12026‑08‑15 Cross-checked against ReadTheManual's "10-Layer Stack." Closed two real gaps that had been referenced but never actually built: added a proper Nextcloud server deployment step with Collabora Online (previously only the mobile-client side existed — Nextcloud without Collabora is file sync, not an office-suite replacement), and Rclone as the actual offsite leg of 3-2-1 backups (previously mentioned in three places as a requirement with no tool to do it). Section 08 hardening baseline: added explicit UFW/nftables + fail2ban guidance for every Linux VM/LXC individually — genuine defence in depth underneath VLAN isolation, not a replacement for it. Section 11: added an explicit "deliberately not in this build" callout for Windows Server/Active Directory and Mailcow — both legitimate for CV-building or mail independence, neither serving this lab's actual purpose. Nessus Essentials considered and passed over in favor of the already-planned OpenVAS, which has no 16-host cap. Fixed ~25 stale "Section 06" cross-references left over from the v7.0 renumbering (historical changelog entries correctly left untouched).
v7.02026‑08‑15 New Section 06 — Service Interconnection Reference. Explains four architectural patterns already in use — service discovery (Technitium as a phonebook, not a proxy), unified authentication (Authentik forward-auth), event-driven messaging, and notification aggregation — with two new diagrams (DNS-resolution flow, MQTT publish/subscribe). Added Mosquitto as a genuinely new service for IoT/Home Assistant, deliberately framed as needed only once real Zigbee/ESPHome devices exist, not on day one. Explained why a dedicated message bus (NATS) isn't warranted at this service count, and why Proxmox Mail Gateway would be the wrong tool for simple alert emails (it's an anti-spam appliance, not a notification channel). Closes with a full reference table covering every deployed service's purpose, internal connections, internet access, and auth method in one place. All subsequent sections renumbered (07–16).
v6.92026‑08‑14 Checked against the XDA "free tools" article and a sharp reader critique of it. Added NetBox (Section 06 → Group B → Proxy & dashboard) as the living replacement for the manual asset log in Section 13 — device/IP/VLAN documentation as a real database, not a table that drifts. Added ntfy as the self-hosted notification path for Uptime Kuma, Wazuh, and the MT5 trading VM, replacing the generic "Discord webhook" placeholder — deployed with real per-topic auth (deny-all default), not an obscure-topic-name-as-password pattern. Added Beszel for multi-host system monitoring, with the critique's own point stated directly: it adds little on a single host, but this build already spans Proxmox, several LXCs, and a multi-node Talos cluster. Addressed the rest of the critique honestly rather than silently: sharpened the NetBird step to state plainly that Tailscale is a WireGuard implementation, not an alternative to it; added tradeoff callouts on NPM-vs-Traefik, Homarr-vs-Homepage, and Authentik-vs-Authelia rather than pretending the criticism doesn't apply.
v6.82026‑08‑14 Checked both Technitium clustering sources against the design. Neither article changed the architecture (plain Docker/LXC clustering was already right) but exposed a real gap: the Technitium deployment step never specified what to actually type into the Cluster Domain field — and this build's existing lab.internal naming domain was a plausible, wrong answer that both source articles independently identify as the top cause of clusters that show "Connected" and then flip to "Unreachable." Added a dedicated new step (3b) covering the fix (a distinct, unused cluster domain like dnscluster.local; never cluster.local, which collides with Kubernetes), DNS_SERVER_DOMAIN consistency, one-way replication direction, and a forward-looking callout on the specific Kubernetes traps (VIP binding, TLS passthrough on 53443, pod-CIDR zone-transfer allow-listing) that only apply if Technitium is ever moved into the Talos cluster later.
v6.72026‑08‑13 Phone & mobile backup added — new Section 06 → Group B → G tab covering the Nothing 2a today, future Pixel/Motorola GrapheneOS phones later. NetBird always-on VPN on Android (with the GrapheneOS-specific disconnect caveat noted honestly, plus the battery-optimisation exemption that actually makes "always on" work), Immich for daily photo/video backup, Nextcloud + DAVx5 for contacts/calendar/files, Seedvault for full app+data backup (GrapheneOS-only, backing up to Nextcloud via SAF), and Bitwarden pointed at the already-planned self-hosted Vaultwarden — which turns out to already answer the "host my own password manager" ask, just under its real name. Added a new NetBird personal-devices access group (Cloud + Photos VLANs only) narrower than the admin group, so the phone gets exactly what it needs and nothing more.
v6.62026‑08‑13 Rack diagram: fixed the actual cause of text getting cut by lines. Cable paths and their labels were interleaved in drawing order — a label defined early (like "SFP+2 primary") could still get painted over by a different cable's line defined later, since SVG draws in document order. Restructured so every cable path renders first and every label renders last, guaranteeing labels always sit on top regardless of definition order. Also nudged the SFP+1/SFP+2/main-feed labels further from their lines' bend points and widened the HDMI+USB label's clearance. MAINS text changed from dim grey to bold amber for real contrast against the dark background.
v6.52026‑08‑13 Diagram legibility fixes. Logical VLAN diagram: the lane label (VLAN name/subtitle) was overlapping into the first row of service chips on the longer entries (Trading, Photos, Cameras, NAS) — reserved a fixed, width-verified label column with a divider line so no label can ever run into the chip area again. Physical rack diagram: rebuilt on a much larger canvas with real spacing — every network/power cable now runs in its own vertical lane instead of several sharing near-identical coordinates, the NanoKVM label (previously overflowing its box) now sits in a properly sized two-line chip, and cable/device connection points are spread across each device's full height instead of clustering at one spot.
v6.42026‑08‑13 Three more hardening sources reviewed (Pen Test Partners, XDA, dev.to). Section 09: added two more Docker/host mistake cards (Docker silently bypassing a host firewall like UFW — VLAN isolation is the real control, not a host firewall layered on top; default to unprivileged/rootless containers and LXCs) and a note on separating Linux users per service rather than one account for everything. Security & Identity tab: added CrowdSec alongside Suricata (behaviour/reputation-based, complements signature-based detection) and optional ClamAV filesystem scanning. Proxy & Dashboard tab: added Watchtower for automatic container updates, with databases explicitly excluded. Section 13: added checks for default .env/compose secrets and open listening sockets to AUDIT, an immutable-backup-copy check to RECOVER, and a lightweight asset hardening log table (device/IP/MAC/who/when) to track what's actually been hardened over time. NanoKVM tab: added a BIOS password + disabled external boot as a second gate behind the KVM access path itself. Consciously did not adopt one piece of advice from the dev.to source (blanket-disabling ICMP) — that's dated guidance that breaks path MTU discovery and diagnostics for little real benefit; noting the omission rather than silently skipping it.
v6.32026‑08‑12 Best-practice audit pass against a Fortigate-based homelab security writeup and VirtualizationHowto's "5 VLAN Mistakes" article. Added VLAN 45 (Cameras) — fully isolated, separated out from the more permissive VLAN 90 Testing where Frigate originally sat (table, matrix, both VLAN diagrams updated, 13→14 VLANs throughout). MikroTik tab: added a step to retire VLAN 1 as the implicit native/catch-all. Proxmox tab: added a double-tagging avoidance callout (the #1 VLAN mistake in the source article). Cloudflare Tunnel: added country-restricted Access policy and explicit brute-force lockout across Cloudflare/NetBird/OPNsense. NPM: consolidated to one wildcard certificate. Section 07 hardening baseline expanded — SSH off by default, named daily-admin accounts instead of root, device-level IP scoping, client-side ad-blocking as a second layer, an explicit note that port-obscurity isn't a control, and a callout confirming this build has zero inbound port-forwards. MT5 Trading tab: added Windows endpoint hardening (USB whitelist, GPO lockdown, mandatory AV, Wazuh logging). Monitoring tab: added optional SNMP switch monitoring and Pulse. Added a Core Performance Boost power-efficiency note (Section 04) and a forward-looking Longhorn/jumbo-frames warning for when a second node arrives (Section 08).
v6.22026‑08‑12 Diagram fixes. The network boundary diagram was genuinely broken (labels wider than their lane, arrows not touching their nodes) — rebuilt with four dedicated full-width lanes so nothing overlaps. Split the single, overloaded topology diagram into two: a Physical & core network diagram (wired path only — Internet through every rack device, bigger, with fallback links and planned devices now hidden behind toggles instead of drawn as a permanent tangle) and a standalone Remote access (VPN) diagram (NetBird and Cloudflare Tunnel as two independent, individually-toggleable lanes, sharing no lines with the physical diagram at all).
v6.12026‑08‑12 Diagram overhaul. Rebuilt the network connection topology diagram from scratch with generous node spacing, background-boxed labels so crossing lines never obscure text, the NAS's two 1GbE cables drawn as one bundled pair instead of two overlapping lines, bidirectional arrows on the WAN link, and a clear outbound-initiated marker on the NetBird and Cloudflare Tunnel overlay paths. Added four new diagrams to Section 05: a network boundary diagram (the four — and only four — ways traffic crosses the OPNsense edge), a data flow diagram (three representative request journeys through the stack), a VLAN zone map (trust-tier radial view, distinct from the swimlane diagram's service-listing view), and a switch port diagram (the CRS310's actual front panel, port by port). All diagrams regenerated programmatically this round for consistent spacing rather than hand-placed coordinates.
v6.02026‑08‑11 Zero trust identity layer + recurring security audit. Added a new Security & Identity tab (Section 06 → Group C → V): Authentik as a single identity provider with forward-auth in front of NPM/Traefik, Wazuh as a centralized SIEM/logging platform, Suricata enabled on OPNsense for signature-based traffic inspection, and OpenVAS/Trivy for on-demand vulnerability scanning. Section 07 gained a "Zero trust — identity, privilege & logging" block explaining how these fit together (least privilege by default, full activity logging, defence in depth on appliance UIs). Added new Section 13 — Security Audit & Recheck, a recurring (not one-time) AUDIT → HARDEN → SEGMENT → RECOVER checklist cross-referenced against the sebasantana/homelab-hardening-checklists repo, meant to be run before going live and quarterly after. Resource budget in Section 08 updated — the security stack is the second-heaviest addition after Wazuh's own footprint, reinforcing the case for a second Proxmox node.
v5.22026‑08‑11 DNS privacy hardening, dual remote access, and home automation added. Applied the Pi-hole-style DNS privacy pattern (VLAN-wide enforcement + third-party DNS/DoH blocking) to the existing Technitium setup rather than adding a redundant Pi-hole instance. Added Cloudflare Tunnel + Access as a second, architecturally independent remote-access path alongside NetBird (Section 06 → Services → Remote access), including a Terraform step for managing it as code. Refined NetBird's Docker-hosted peer deployment (network_mode: host, wt0 interface). Added optional Packer golden-image step to the GitOps foundation tab. Added a new Home Automation services tab — Home Assistant (VM, USB passthrough for Zigbee/Z-Wave/Bluetooth) and Homebridge (LXC, HomeKit bridge) — plus a Homebridge chip in the logical diagram. Fixed: switch management IP was previously on the isolated NAS VLAN.
v5.12026‑08‑06 Quick Start added — new Section 00 at the very top of the page: a sequential, tick-box build order (plan → ISP handoff → OPNsense → switch → NAS → hosts → AP → testing → hardening → backups/monitoring), mapped to this build's actual devices and cross-linked to the detailed section for each step. Also fixed an inconsistency: the MikroTik switch's management IP was set on the isolated NAS VLAN (10.0.99.254) — moved to the Management VLAN (10.0.1.5) where it belongs.
v5.02026‑08‑06 Platform expansion based on Jim's Garage and VirtualizationHowto's published homelab patterns: added VLAN 35 (Trading, isolated), swapped Pi-hole/AdGuard for a clustered Technitium DNS pair + Unbound, completed the media pipeline with Jellyseerr + Bazarr. Added two new top-level sections — 08 Kubernetes, GitOps & AI Platform (resource budget, architecture decisions, GitOps flow diagram) and 09 Home Lab Organisation & Docker Best Practices (role-based tagging, hostname convention, git folder structure, the six Docker networking mistakes, DNS-first naming, prod/test separation). Added Group C to Section 06 — GitOps foundation (Terraform/Ansible), Kubernetes (Talos, Argo CD, Traefik, Longhorn, Flatcar), AI stack (Ollama, Open WebUI, iGPU sharing, n8n), and Trading (MT5) — full step-cards throughout. Added a Phase 6 (optional, later) to the build checklist. Flagged honestly: the requested stack mirrors a published 5-node/480GB-RAM lab; this build has one 32GB host, so Section 08 right-sizes it (single-node Talos, small quantized models) rather than pretending it all fits at once.
v4.02026‑08‑04 NanoKVM confirmed owned and fully worked into the build: hardware inventory, physical rack diagram (shelf-mounted, cable links), logical VLAN diagram, topology diagram, switch port map, and a full 6-step configuration tab covering physical setup, firmware, network, web UI hardening, optional ATX power control, and NetBird-only access lockdown. Diagrams also gained an explicit 1GbE ISP uplink label, an admin laptop drop (switch port 2, VLAN10), and planned (not-yet-owned) dashboard and security-camera monitor nodes. Section 06 fully rebuilt into two tab groups — Hardware & Core Devices and Services & Add-ons — with every step now carrying a difficulty rating, time estimate, a "why" explanation, numbered instructions, a checkbox task list, and links to official docs. Added two original concept diagrams (default-deny firewall logic, VLAN trunk vs. access) for learning purposes.
v3.02026‑08‑04 Diagrams added: physical rack (with toggleable cable routing), logical VLAN swimlane diagram, and network connection topology (with toggleable fallback/VPN overlay). Cabinet, UPS, and 10GbE cables marked owned. Stack upgraded: Jellyfin (was Plex), Homarr dashboard, NGINX Proxy Manager, AdGuard Home secondary DNS, InfluxDB + Uptime Kuma monitoring, Proxmox Backup Server, and NetBird self-hosted remote access (replacing Tailscale). Added Future Expansion & HA section. Corrected an inconsistency in the supplied reference diagram (20GbE LACP bond isn't physically possible — switch only has 2 SFP+ ports). Section 06 gained expand-all/collapse-all controls.
v2.12026‑08‑04 UPS confirmed: APC Smart-UPS C SMC1500I-2UC added to hardware list and rack layout at RU 1–2. Weight and depth margins flagged.
v2.02026‑08‑04 10GbE backbone added: SFP+ trunk between OPNsense ↔ switch ↔ Proxmox, 1GbE ports kept as fallback. Full device configuration steps added.
v1.12026‑08‑04 Rack corrected from a 10" mini rack (didn't fit existing 19" gear) to the Tecmojo 12RU / 450mm / 50kg wall-mount cabinet.
v1.02026‑01‑29 Original plan: 12RU cabinet, 12 VLANs, single Proxmox hypervisor, 1GbE backbone, no UPS specified yet.
Adding a new entry: copy one .changelog-row block, bump the version, set today's date, describe what changed. Keep newest at the top.

16Hosting this page & keeping it updated

Served from Cloudflare Pages, built from a private GitHub repository, gated by Cloudflare Access — real authentication on the page content, not just an unlisted link.

Why not plain GitHub Pages

GitHub Pages served from a public repository has no access control at all — anyone with the URL sees everything, and the repository itself is browsable. GitHub Pages from a private repository requires a paid GitHub plan and still does not provide simple per-person link sharing. Cloudflare Pages was chosen specifically because it builds directly from a private repository on the free tier, and pairs natively with Cloudflare Access for authentication — the same access-control product already used for the Cloudflare Tunnel backup path in Section 08.

One-time setup

  1. Make the GitHub repository private (Settings → General → Danger Zone → Change visibility), if it is not already. This alone is what stops anyone but explicitly-added collaborators from editing anything — a repository permissions question, independent of hosting.
  2. Create a free Cloudflare account if one does not already exist, and add a zone (a domain — Cloudflare offers free subdomains for testing, or use any domain already owned).
  3. Cloudflare dashboard → Workers & Pages → Create → Pages → Connect to Git, authorise Cloudflare's GitHub App, and select the private repository.
  4. Build settings: framework preset Jekyll, build command bundle exec jekyll build, output directory _site. The Gemfile at the repository root tells Cloudflare's build image to install Jekyll — no further Ruby configuration needed.
  5. Deploy. Cloudflare assigns a *.pages.dev URL immediately; a custom domain can be attached afterward under the project's Custom domains tab if wanted.
  6. Update DIAGRAMS_BASE_URL in script.js to that URL, commit, push — this is the one line the whole live-diagram system depends on (Section 06).

Cloudflare Access — gating the page

  1. Zero Trust dashboard → Access → Applications → Add an application → Self-hosted.
  2. Domain: the Pages project's *.pages.dev URL (or the custom domain, once attached).
  3. Policy: allow rule listing specific email addresses — every person who should be able to view this, and no one else. One-time email code is the simplest identity provider for a personal project; no separate account creation needed for anyone on the list.
  4. Add a second, separate application scoped to the path /diagrams/*, policy: Bypass — public, no login required for that path specifically.
Why /diagrams/* is deliberately excluded from the login requirement: the diagrams render via a third-party viewer (viewer.diagrams.net) loading each file inside a cross-origin iframe. Cloudflare's own authentication cookie does not reliably reach a resource fetched from inside a cross-origin iframe — a documented browser cookie-scoping limitation, not a configuration mistake — so attempting to gate this path either breaks every diagram embed or requires real engineering effort with no guaranteed outcome. The trade-off actually taken: every paragraph, table, and configuration step on this page requires login; the diagram XML files themselves are reachable only by someone with the exact file URL, which is linked nowhere public and contains no credentials — only architecture information already fully described in the gated text regardless.

Day-to-day editing

  • Prose: edit the relevant file directly under _includes/sections/ (see HOW-TO-RUN-THIS-SITE.txt for which file covers which section), commit, push. Cloudflare rebuilds automatically within about a minute.
  • Diagrams: open the .drawio file directly in app.diagrams.net, edit, save. No export step — the live embed reads the same file.
  • Both routes require push access to the private repository — the same permission boundary that keeps editing restricted, regardless of who can view the published page.
Suggested repository layout: index.html, _config.yml, and Gemfile at the root, _includes/ and diagrams/ alongside them — exactly the structure already in this repository.