- My Forums
- Tiger Rant
- LSU Recruiting
- SEC Rant
- Saints Talk
- Pelicans Talk
- More Sports Board
- Fantasy Sports
- Golf Board
- Soccer Board
- O-T Lounge
- Tech Board
- Home/Garden Board
- Outdoor Board
- Health/Fitness Board
- Movie/TV Board
- Book Board
- Music Board
- Political Talk
- Money Talk
- Fark Board
- Gaming Board
- Travel Board
- Food/Drink Board
- Ticket Exchange
- TD Help Board
Customize My Forums- View All Forums
- Show Left Links
- Topic Sort Options
- Trending Topics
- Recent Topics
- Active Topics
Started By
Message
Official LLM discussion thread
Posted on 7/23/26 at 8:31 am
Posted on 7/23/26 at 8:31 am
It's time baws.
Who has done it?
What is your setup?
Pros and cons?
Would you do it again?
Are you pairing it with local voice devices?
Is your electric bill affected?
Running on dedicated hardware or mixed in with your existing virtualizations?
What model are you using?
How does it figure/dovetail with your usage of Claude or chatgpt?
How big of a pita is care and feeding of it?
My current architecture plan
Proxmox
|__ HAOS VM
|__ Windows VM
|__ Plex LXC
|__ AI & Video Debian VM
---->|__ Frigate container
---->|__ Ollama container
-------> |__ NVIDIA GPU passthrough
Who has done it?
What is your setup?
Pros and cons?
Would you do it again?
Are you pairing it with local voice devices?
Is your electric bill affected?
Running on dedicated hardware or mixed in with your existing virtualizations?
What model are you using?
How does it figure/dovetail with your usage of Claude or chatgpt?
How big of a pita is care and feeding of it?
My current architecture plan
Proxmox
|__ HAOS VM
|__ Windows VM
|__ Plex LXC
|__ AI & Video Debian VM
---->|__ Frigate container
---->|__ Ollama container
-------> |__ NVIDIA GPU passthrough
This post was edited on 7/23/26 at 8:48 am
Posted on 7/23/26 at 10:21 am to CAD703X
Novice here…what is the benefit of running an LLM locally? What does it achieve versus accessing LLM through the cloud?
Posted on 7/23/26 at 10:32 am to CAD703X
explain this to me like i'm a 5 year old
Posted on 7/23/26 at 10:44 am to lynxcat
quote:Most of the same benefits of running anything locally. Same cons as well.
Novice here…what is the benefit of running an LLM locally? What does it achieve versus accessing LLM through the cloud?
Mostly boils down to data ownership and control.
Honestly this is a question I'd send to a cloud LLM.
Posted on 7/23/26 at 10:46 am to BankLSU
quote:local language model (?) i think thats what it stands for.
explain this to me like i'm a 5 year old
basically you're taking 'chatgpt' (the opensource stuff anyway) and installing it on your local server so you can interact with a 'chatbot' that lives with you as opposed to the cloud avoiding
* forced changes and updates that break what you working on
* paywall changes and price increases
* giving Big Tech insight into everything you are asking about.
for me, i can imagine everything from 'learning' things about you and your family based on interactions to make relevant gags or observations "CAD left the wet clothes in the washing machine AGAIN" or if you go with local voice control, it can expand from simple "turn light x on/off" to something more nuanced like “What happened while I was gone?” or “Why is the house using so much electricity?” or “Summarize anything unusual today.”
im a hack and basically starting from ground zero on this stuff so some of the smarter posters on here can weigh in with better use cases but it never hurts to think about doing things like this and less dependent on our benevolent tech overlords to always have our best interests in mind.
eta this is most relevant for new 'home labbers' who are spec-ing out servers to plan on purchasing one with a GPU that is capable of handling some modest LLM as well. it never hurts to think ahead.
This post was edited on 7/23/26 at 10:49 am
Posted on 7/23/26 at 11:32 am to CAD703X
Ubuntu/ollama 16gb card giving ollama all the access it wants.
Dell 5820 128gb ram 1tb ssd
Building containers to satisfy everything else. I found the speed loss through proxmox was noticeable so that’s why I pulled it.
You can get way more uncensored models with local llm.
edited: Adding the use of cloudflare free tunnel to create an outbound connection only to cloudflare that you can use for remote access.
LINK
Dell 5820 128gb ram 1tb ssd
Building containers to satisfy everything else. I found the speed loss through proxmox was noticeable so that’s why I pulled it.
You can get way more uncensored models with local llm.
edited: Adding the use of cloudflare free tunnel to create an outbound connection only to cloudflare that you can use for remote access.
LINK
This post was edited on 7/23/26 at 11:43 am
Posted on 7/23/26 at 11:50 am to XanderCrews
quote:explain this architecture a bit to me. interested in all options since i haven't begun to provision the PC yet.
Building containers to satisfy everything else. I found the speed loss through proxmox was noticeable so that’s why I pulled it.
Posted on 7/23/26 at 11:53 am to j1897
quote:
It's free
Minus the hardware.
I was looking into this yesterday. I'd like to get a Mac Studio m3 Ultra, but they stopped production and they were about $7,000 at that point I think. That's 5.8 years of Claude Max subscriptions (and no maintenance). I'm gonna wait until they can get the quantization down without losing accuracy and hardware costs come down, hopefully.
What is Quantization?
It is essentially file compression for AI models. Under the hood, an AI is just a massive spreadsheet of highly precise numbers. Quantization rounds those numbers down to take up far less space. Think of it like taking a massive, uncompressed RAW photo and saving it as a smaller JPEG.
Why is it done?
By default, powerful AI models are so huge they require expensive data center supercomputers just to turn on. Quantization shrinks the model's footprint by up to 75% so it can run directly on everyday consumer hardware—like a standard PC, laptop, or smartphone—without noticeably losing its "smartness."
This post was edited on 7/23/26 at 12:00 pm
Posted on 7/23/26 at 11:57 am to AaronDeTiger
quote:damn baw
I was looking into this yesterday. I'd like to get a Mac Studio m3 Ultra, but they stopped production and they were about $7,000 at that point I think. That's 5.8 years of Claude Max subscriptions. I'm gonna wait until they can get the quantization down without losing accuracy and hardware costs come down.
i'm looking at an i7 dell 3660 with maybe a 20 gb RTX 4000 SFF Ada for a fraction of that cost.
Posted on 7/23/26 at 12:03 pm to CAD703X
That mac had 512 GB of unified memory. You can put a cutting edge 4-bit model on it.
Posted on 7/23/26 at 12:14 pm to AaronDeTiger
quote:
Minus the hardware.
I already have the hardware to play factorio, so free to me.
Posted on 7/23/26 at 12:27 pm to j1897
quote:from what i understand; you need low power GPUs for AI/LLMs that can spin up on demand and gaming GPUs are typically higher power draw all the time and more suitable as game engines but again, i know just enough about this stuff to sound like an idiot
I already have the hardware to play factorio, so free to me.
Posted on 7/23/26 at 1:37 pm to CAD703X
for 24/7 use yea, it's a waste of power. I generally use mine for agentic coding, so I can boot everything up with the right OS, run my tasks from laptop for a few hours or overnight, and switch my gaming pc back to windows gaming.
I use n8n for automation and it can run local llm's but they are very slow and not very good for a low power pc, and I've not had the best results using it in those workflows.
I use n8n for automation and it can run local llm's but they are very slow and not very good for a low power pc, and I've not had the best results using it in those workflows.
Posted on 7/23/26 at 2:25 pm to CAD703X
Let me know if you want more specifics, but the mgmt was just way easier through base ubuntu as opposed to PM vms. I liked learning and messing with PM but now i was managing a hypervisor on top of other things.
Pulled my LLM stuff off Proxmox and moved to bare-metal Ubuntu with a GPU passthrough-free setup, then containerized everything else around it. Here's the rundown:
**Hardware/OS**
- Bare Ubuntu install (no hypervisor) on a Xeon box, dedicated GPU for local LLM inference
- Reasoning: skipping the VM layer meant the GPU was just... there, no passthrough config, no driver headaches inside a guest
**Container stack (all Docker Compose, managed through Dockge)**
- Open WebUI in front of local models (Ollama) — this is the "LLM out of Proxmox" piece
- Media: Jellyfin, Navidrome, Beets for tagging
- Photos: Immich (server + ML + Postgres/pgvector + Redis)
- Docs/wiki: Wiki.js (Postgres-backed)
- CRM: EspoCRM (MariaDB-backed)
- Offline knowledge: Kiwix serving Wikipedia/Khan Academy ZIMs
- Dashboards: Homepage + Homarr as front doors, Glances for system monitoring
- Filebrowser for poking around the NAS mounts from a browser
- SearXNG for private search
- All NAS storage comes in over CIFS mounts from a separate box, with a systemd override so Docker waits for the mounts before starting containers (otherwise anything touching `/mnt/*` fails hard on reboot)
**Local DNS + routing**
- Avahi for mDNS so everything resolves as `service.local` on the LAN
- One nginx container doing simple 301 redirects from `service.local` ? `ip:port` per container, so nobody has to remember port numbers
**Remote access**
- Cloudflare Tunnel (`cloudflared`) — outbound-only connection from the box to Cloudflare's edge, so there's nothing listening on the router at all. One tunnel, one config file, a hostname-per-service ingress list.
- Cloudflare Access sits in front of the whole wildcard domain and requires login (email one-time-code) before any request even reaches the tunnel — so even the services with zero built-in auth (looking at you, Glances) aren't just sitting open on the internet.
- SSH rides the same tunnel as a raw TCP service, wrapped through `cloudflared access ssh` client-side instead of opening port 22 externally.
**Backups**
- Monthly cron job tars up all the stack config directories + relevant Docker volumes, keeps the last 6, prunes the rest
Happy to go deeper on any piece of this if useful — the tunnel/Access setup in particular was the part that took the most care to get right (mostly making sure nothing was reachable *before* the auth policy was actually in place, not after).
Forgot to add you will need your own domain name to pull off the cloudflare piece. Those are pretty cheap and easy these days though.
Pulled my LLM stuff off Proxmox and moved to bare-metal Ubuntu with a GPU passthrough-free setup, then containerized everything else around it. Here's the rundown:
**Hardware/OS**
- Bare Ubuntu install (no hypervisor) on a Xeon box, dedicated GPU for local LLM inference
- Reasoning: skipping the VM layer meant the GPU was just... there, no passthrough config, no driver headaches inside a guest
**Container stack (all Docker Compose, managed through Dockge)**
- Open WebUI in front of local models (Ollama) — this is the "LLM out of Proxmox" piece
- Media: Jellyfin, Navidrome, Beets for tagging
- Photos: Immich (server + ML + Postgres/pgvector + Redis)
- Docs/wiki: Wiki.js (Postgres-backed)
- CRM: EspoCRM (MariaDB-backed)
- Offline knowledge: Kiwix serving Wikipedia/Khan Academy ZIMs
- Dashboards: Homepage + Homarr as front doors, Glances for system monitoring
- Filebrowser for poking around the NAS mounts from a browser
- SearXNG for private search
- All NAS storage comes in over CIFS mounts from a separate box, with a systemd override so Docker waits for the mounts before starting containers (otherwise anything touching `/mnt/*` fails hard on reboot)
**Local DNS + routing**
- Avahi for mDNS so everything resolves as `service.local` on the LAN
- One nginx container doing simple 301 redirects from `service.local` ? `ip:port` per container, so nobody has to remember port numbers
**Remote access**
- Cloudflare Tunnel (`cloudflared`) — outbound-only connection from the box to Cloudflare's edge, so there's nothing listening on the router at all. One tunnel, one config file, a hostname-per-service ingress list.
- Cloudflare Access sits in front of the whole wildcard domain and requires login (email one-time-code) before any request even reaches the tunnel — so even the services with zero built-in auth (looking at you, Glances) aren't just sitting open on the internet.
- SSH rides the same tunnel as a raw TCP service, wrapped through `cloudflared access ssh` client-side instead of opening port 22 externally.
**Backups**
- Monthly cron job tars up all the stack config directories + relevant Docker volumes, keeps the last 6, prunes the rest
Happy to go deeper on any piece of this if useful — the tunnel/Access setup in particular was the part that took the most care to get right (mostly making sure nothing was reachable *before* the auth policy was actually in place, not after).
Forgot to add you will need your own domain name to pull off the cloudflare piece. Those are pretty cheap and easy these days though.
This post was edited on 7/23/26 at 3:19 pm
Posted on 7/23/26 at 2:58 pm to AaronDeTiger
quote:And the power.quote:Minus the hardware.
It's free
My math says a modest machine and modest usage would cost ~$100/year to run at 15 cents/kwh.
Posted on 7/23/26 at 4:51 pm to Korkstand
Is that power cost 24/7 usage on the LLM though?
Posted on 7/23/26 at 4:52 pm to CAD703X
quote:
local language model (?)
Large language model
Yours would be called local LLM
Posted on 7/23/26 at 4:57 pm to bluebarracuda
I tried to roughly estimate 24/7 always-on with a couple hours of active LLM use per day. If you're idling at 50 watts that'll burn 438kWh/year which is probably $50 minimum anywhere. Maybe you can tune idle consumption lower than that.
Popular
Back to top


5






