#hardware

Hello, tok/s

What this blog is, what's running in the closet today, and why local LLMs are next.

This is a blog about the server in my closet and the local LLMs I’m working toward running on it. My day job is integrations and solution architecture, so I spend a lot of time drawing boxes and arrows. This is where I build some of them for myself and write down what happened.

The box

The whole homelab is one refurbished Lenovo ThinkStation P330, bought on eBay for about $300 in September 2026. It’s quiet, it sips power, and it has full-height PCIe slots, which matters later.

CPUIntel i7-8700, 6C/12TQuick Sync for transcoding
RAM16 GB DDR4Room to grow
Boot512 GB NVMeProxmox + guests
Media4 TB USB drive18 TB upgrade planned
GPUNone (yet)See below

What runs on it today

Proxmox VE is the hypervisor. On top of it:

  • A media container running a media server and a reverse proxy that handles HTTPS.
  • An apps container with the usual download-automation stack behind a VPN, plus uptime monitoring.
  • A Home Assistant VM with a Zigbee coordinator passed through, running a few smart plugs, door sensors and too many light bulbs.

Admin access is private over Tailscale. Config lives in git, and backups hold the state. Keeping those two separate has already paid off once.

the closet, roughly
P330 (Proxmox VE)
├── media container     media server, reverse proxy
├── apps container      downloads behind a VPN, monitoring
└── Home Assistant VM   Zigbee, lights, automations

Why “tokens per second”

Tokens per second is the number you watch when you run a language model on your own hardware. It tells you whether a model is something you’ll actually use or just a demo. Right now the P330 has no GPU and 16 GB of RAM, so the honest answer for anything big is “not many”.

That’s the point of this blog. I want to find out what’s actually useful at home:

  • which models are worth running on modest hardware, and at what quantization
  • how to serve them so the rest of the homelab can use them
  • how to wire them into the integration tools I use at work

What’s next

First, some unglamorous groundwork: ad-blocking DNS, proper off-site backups, and a UPS so a power cut doesn’t take out a drive. After that comes the fun part: picking hardware for inference and measuring it honestly.