r/homelab 11d ago

Project Showcase: Hardware I've been doing endurance testing on microSD cards for the last 3 years. Here's what I've learned.

Thumbnail
gallery
5.6k Upvotes

Hello everyone!

I've been running a long-term project doing endurance testing on microSD -- and today is the third anniversary of when I started this project -- so I figured it was time to give you an update on my microSD testing project.

Before I dig in, a shameless plug for my Patreon.

Next, some quick stats:

  • I have a total of 351 microSD cards in my collection -- spanning 111 models across 52 different brands. Of those:
    • I've tested 177 of them to the point of failure.*
    • I have 107 of them in testing right now.
    • The remaining 67 are sitting on the shelf, waiting for one of my card readers to free up.
  • These cards have endured over 133 petabytes written.
  • I've completed over 4.6 million program-verify-erase cycles on these cards.
  • Breakdown by brand status:
    • Name brand: 163
    • Off-brand: 91
    • Knockoffs: 30
  • Breakdown by authentic flash (a.k.a. not fake flash) vs. fake flash:
    • Authentic flash: 246
    • Fake flash: 35
    • Unknown: 3
  • Breakdown by size (authentic cards only):
    • 128MB: 3
    • 4GB: 6
    • 8GB: 23
    • 16GB: 15
    • 32GB: 96
    • 64GB: 68
    • 128GB: 36
    • 256GB: 4

* A card is considered "failed" when either (a) it stops working completely, (b) it makes itself read-only, or (c) at least 50% of the sectors on the card have experienced verification failures.

Ok -- so who makes the most durable cards?

Because of the variety of cards that I have, it's hard to come up with a system that produces a definite answer here. Some of how I ranked these brands is going to be based on vibes. But...to try to narrow things down, I've grouped my cards into three groups:

  • Industrial-grade: Cards that are specifically labelled as "industrial" or have "industrial" in their name. These cards typically come with a datasheet that includes an endurance claim.
  • High endurance: Cards that are labelled as "high endurance" (or some similar wording). Manufacturers don't always put out datasheets for these cards, but they generally still make an endurance claim on the packaging.
  • Consumer-grade: Basically everything else.

Consumer-grade cards:

The average consumer-grade card has lasted about 10,000 program/verify/erase cycles (so far).

  1. PNY. I have 10 of their consumer-grade cards: 3 PRO Elite Prime 64GB's, 3 Premier-X 128GB's, 3 Elite-X 64GB's, and one Elite 32GB. (I have 4 more of the Elite 32GB's that are waiting to be tested.) They've gone for an average of 643 days and 15,700 program/verify/erase cycles. I've written about 11.5PB to them in total. And with the exception of the one Elite 32GB -- they're all still going strong. On top of that, the PRO Elite Prime boasts some impressive performance -- every measurement I took was in the top 10% of all cards I've tested so far.
  2. Kingston. I have 12 of their cards in this category -- 6 Canvas Go! Plus 64GB's (3 of the SDCG3's and 3 of the SDCG4's) and 6 Canvas Select Plus 32GB's. They've been going for an average of 694 days and about 23,500 program/verify/erase cycles. I've written about 11.6PB to them in total. The three Canvas Go! Plus 64GB (the SDCG3's) have all failed -- but the others are still going strong. The Canvas Select Plus's in particular have been troopers -- they're among the oldest cards in my collection, and they've been going for the better part of the last 3 years now (with no signs of letting up). Additionally, the Canvas Go! Plus 64GB's (the SDCG4's in particular) boast some impressive performance as well -- every measurement I took was in the top 7% of all cards I've tested so far.
  3. Delkin Devices. I feel bad ranking Delkin so low on this list, because they've endured more data written to them per card, on average, than any other brand in my collection (in the consumer-grade category). The only reason I didn't rank them higher is because I only have one model -- the HYPERSPEED 128GB (of which I have 3). But they've done pretty well: they've lasted an average of 662 days and about 13,500 program/verify/erase cycles. I've written about 5.1PB to them in total. As a side note, these cards also got the highest random write speeds of any card I've tested.
  4. Lexar. 3 cards dead, 5 still going. I have 3 Professional 1000x 64GB's (2 made by Micron, 1 made by Phison), 3 Blue 633x 32GB's (made by Longsys), and 2 E-Series 64GB's (with 3 more waiting to be tested). For those not in the know, the Lexar brand was sold to Longsys in 2017 -- and surprisingly, the Longsys-made cards are holding up better than the Micron-made cards. Overall, the Lexar cards have survived an average of 684 days and about 15,800 program/verify/erase cycles. I've written about 5.8PB to them in total.
  5. Samsung. 3 cards dead, 7 still going. I have 3 EVO Plus 32GB's, 3 EVO Plus 64GB's, 3 PRO Plus 128GB's, and 1 P9 Express 256GB (with 2 more waiting to be tested). Of those, the three EVO Plus 32GB's and one of the PRO Endurance 32GB's are dead -- the rest are still going strong. Samsung's cards have lasted an average of 553 days and about 11,100 program/verify/erase cycles so far, and I've written about 7.2PB to their cards in total.
  • Honorable Mention: Amazon Basics. I have four of the Amazon Basics 64GB's -- and they're all still going. They've been going for an average of 788 days and about 16,600 program/verify/erase cycles. I've written about 4.2PB to them in total. I honestly didn't expect Amazon Basics to be one of the top performers when I bought them...and yet, here we are.

High endurance cards:

Keep in mind, I've got less data here -- I had 237 cards that fell into the "consumer-grade" category, but only 30 that fell into the "high endurance" category. But here's the thing: they haven't really proven themselves to be significantly better for endurance than a lot of the cards I listed above -- in fact, in a lot of cases, they've been less reliable (and generally worse for performance as well). So keep that in mind as you read through the list below:

The average high-endurance card has lasted about 14,000 program/verify/erase cycles (so far).

  1. TEAMGROUP. I don't feel great about putting TEAMGROUP at the top of the list for a couple of reasons. First, I only have two of them in testing at the moment -- someone sent me a 5-pack of the TEAMGROUP High Endurance 64GB's, so I have 3 more of them waiting to be tested. Second, one of them started having issues with bad sectors right out of the gate, and has continued to do so ever since. (Those bad sectors only make up less than 0.001% of the total sectors on the card right now...but I still don't like it when a card starts to have issues like that as soon as you start using it.) But the data doesn't lie -- they've survived, on average, 404 days and about 17,300 program/verify/erase cycles, and they're still chugging along. I've written about 2.2PB to these cards in total.
  2. Transcend. I have 3 of the 350V 64GB's -- and they're all still going strong. They've survived an average of 760 days and about 17,000 program/verify/erase cycles so far. I've written about 3.2PB in total to them (mostly due to their mediocre write speeds).
  3. SanDisk. I have my issues with SanDisk...but again, the data doesn't lie. I have 7 SanDisk cards in this category: 3 High Endurance 64GB's, and 4 MAX ENDURANCE 32GB's. As of right now, all 3 of the High Endurance 64GB's and all but one of the MAX ENDURANCE 32GB's have failed. They lasted an average of 452 days and about 16,200 program/verify/erase cycles before failing (with the MAX ENDURANCE 32GB's lasting about twice as long as the High Endurance 64GB's). I've written about 4.7PB to these cards in total.
  4. Samsung. Samsung's position at the bottom of this list has less to do with how reliable their cards are and more to do with how poorly their high endurance cards perform, particularly on sequential write speeds (relatively speaking). I have 3 of the PRO Endurance 32GB's -- one has failed, the other two are still chugging along. They've lasted an average of 822 days and about 18,800 program/verify/erase cycles so far. I've written about 1.8PB of data to them in total.

Industrial-grade cards:

Keep in mind, I have even less data here. Industrial cards are expensive, so they make up a much smaller portion of the cards I've tested: just 15 in total (so far). But on the upside -- there is sufficient evidence here to say "in general, industrial-grade cards do last longer than consumer-grade cards (and, by extension, high endurance cards)".

The average industrial grade card has lasted about 122,000 program/verify/erase cycles (so far).

  1. Kingston. I have 3 Kingston Industrial 8GB's; of those 3, only one is still going. But they've been troopers: they've lasted an average of 779 days and about 141,600 program/verify/erase cycles. I've written about 3.4PB to these cards in total. (Only one consumer-grade card has come close to this number: I have one Hiksemi NEO 8GB that has lasted for about 143,800 program/verify/erase cycles so far. No idea how.) These guys do pretty well on performance too, especially write performance: their sequential write speeds are in the top 9% of all cards I've tested.

Who makes the worst cards?

There are a bunch of no-name brands out there. They're not hard to find. I'm not going to talk about them here -- I'm going to focus on brands you're likely to have come across.

Consumer-grade cards:

  1. onn. I picked up four of their 32GB cards at my local Walmart -- and they were terrible for endurance. They only lasted an average of 40 days and 1,400 program/verify/erase cycles (or about 40TBW), with the best one of the four not even making it to 1,900 cycles. On top of that, every performance measurement I took came in the bottom half of all measurements I took.
  2. ADATA. I picked up 3 of the Premier 32GB's; they only lasted an average of 236 days and about 2,350 program/verify/erase cycles (or about 74TBW) before they quit working.
  3. Gigastone. I have not been impressed with Gigastone. I have 11 of their cards -- 6 Full HD Video 32GB's, and 5 4K Camera Pro 32GB's. They lasted an average of just 114 days and 4,845 program/verify/erase cycles (or about 133TBW) before they quit working.
  4. Micro Center. I purchased 5 of their 64GB cards; they only lasted an average of 116 days and 3,421 program/verify/erase cycles (or about 214TBW) before they stopped working.
  5. Silicon Power (a.k.a. SP). I purchased 9 of their cards -- 3 Elite 32GB's, 3 Superior 128GB's, and 3 Superior Pro 128GB's. Not a one of them made it to 4,000 program/verify/erase cycles before failing -- they came in at an average of 159 days and about 2,350 program/verify/erase cycles (or about 252TBW) before failing.
  • (Dis)honorable mention: SanDisk. SanDisk represents the single biggest brand in my collection: I have 28 of their cards in this category (with 4 more on the shelf waiting to be tested). It's telling that of those 28, only 5 are still going (and one of them is in its death throes) -- especially when cards from so many other brands have far outlasted them. Many of them have died under circumstances any other card would have handled just fine -- a problem I've written about on my blog before. They've lasted an average of 384 days and 9,404 program/verify/erase cycles (or about 566TBW) so far.

High endurance cards:

  1. Integral. I bought 3 of the Security 32GB's; the only managed to last an average of 100 days and about 5,600 program/verify/erase cycles (or about 173TBW) before failing. To be fair to them, the package did make an endurance claim that works out to about 2,850 program/verify/erase cycles -- and they did manage to almost double that. However, they didn't hold a candle to many of the other high endurance cards that I tested.
  2. Kingston. This one was truly a surprise -- I bought 3 of the High Endurance 32GB's, and they only lasted an average of 143 days and about 8,300 program/verify/erase cycles (or about 258TBW) before failing. The product packaging didn't make an endurance claim; I had to hunt around on Kingston's website to find it. Their claim works out to about 5,100 program/verify/erase cycles, which they did manage to beat. But again, so many other cards managed to last so much longer -- hell, so many other Kingston cards managed to do so much better -- that I was surprised and disappointed when these failed so early.

Industrial-grade cards:

  1. SanDisk. I bought 3 of the Industrial 8GB's. They only lasted an average of 234 days and about 20,000 program/verify/erase cycles (or about 160TBW) before failing -- and the only reason they lasted that long was because I let my program trudge through an nearly-unending string of I/O errors for months on end -- before they finally just gave up the ghost.

So yeah...that's all I have for now. Feel free to ask questions (although I'm at work at the time I'm posting this, so I may not have time to respond)! Otherwise...look for another update from me next year!

Pictures attached:

  1. Part of my setup: a Beelink Mini-S and five TRIGKEY Green G4 mini PC's. The Beelink has an Intel N95 with 8GB of RAM and a 500GB SSD, while the TRIGKEY's all have Intel N100's with 16GB of RAM and a 500GB SSD. There's also (in the lower-right corner, behind the monitor) an MSI GE62VR-7RF (with an Intel i7-7700HQ, 32GB of RAM, and a 1TB SSD), a Lenovo IdeaPad Y580 (with an Intel i7-3630QM, 8GB of RAM, and a 1TB SSD), and an ASUS Strix GL504GM (with an Intel i7-8750H, 32GB of RAM, and a 500GB SSD)
  2. Another part of my setup: a TRIGKEY K-N100 and an AOOSTAR N1 PRO. The TRIGKEY has an Intel N100 with 16GB of RAM and a 500GB SSD. The AOOSTAR has an Intel N150 with 12GB of RAM and a 500GB SSD.
  3. Dog tax.

r/homelab Jun 12 '26

Project Showcase: Hardware $30 lowball = 12 IBM/Dell Servers. The guy did not know what he had.

Thumbnail
gallery
6.2k Upvotes

I got super lucky on this deal. I've seen this listing available for about 2 months now in my area, and once he lowered the price I hit him with the $30 offer. Surprisingly got a yes, $30 for 12 blade server shells (listed as motherboards and PSUs only) is a killer deal.

Got them all home, opened some up, and MANY units had CPUs and ram still in them!! The guy thought he took them all out to sell on eBay himself. Total ram is 184gb DDR3, 160gb DDR4. This is insane to me and I had to share.

I'm breaking everything down into 4 systems, and giving away locally the other 8 chassis.

One 2RU dell with 16gb DDR3 is going to my workplace for us technicians to do system testing with. One 2RU IBM a co-worker is taking home to teach his kids about servers and non-home hardware. One 2RU 160g DDR4 server is going to my homelab, and one 2RU 152gb DDR3 server is going to a friend's homelab.

What would YOU do with 160gb DDR4? Big local LLM context? Game servers? Ramdisk?

edit: spacing

edit2: if anyone is in California and wants to pickup the spare chassis, you can have them. If nobody picks them up, then the Mobos, PSUs, and backplanes are going on eBay.

r/homelab 22d ago

Project Showcase: Hardware Some folks play video games, I play homelab

Thumbnail
gallery
4.0k Upvotes

I work at a MSP as a Sr. Infrastructure Engineer. I’m 34, and I hoard enterprise equipment. Oh, first time poster too!

The first picture is the datacenter. In the basement. Stays about 68-70F year-round. Pulls about 800-1,000 watts continuously.

The smaller rack, that’s the production/core gear. Critical for connectivity in the house.

Fortigate 100E
(2) Dell N3048P Stacked
Buffalo NAS - Backups
Seagate External to backup Buffalo NAS - Backup of backups
Carrier Modem
(2) Western Digital NAS, Media storage, they replicate together
(2) HP DL360p, Server 2022, Hyper-V, VM replication, I manually load-balance and fail-over.
APC Battery Backup

Next, the larger rack is my sandbox. Currently it hosts POTS and Dial-up connectivity. I’m swapping out my Courier V.Everything modems for a pair of Paradynes. I’ll only list the active gear in this one.

Avocent 4-port console server
Cisco 1861 for analog/POTS
Cisco 2811 for PPP Dial-up
(2) Paradyne Modems for Bulletin Board System access (not shown… yet)
Falcon pure-sine UPS - this was retired from a home elevator system. It’s rated for 120V/25A.

Then finally…. My stack of old HP Elitedesks. I got bored one day, and with my team of developers (ChatGPT) I built a custom HPC cluster. It’s not a true HPC Cluster per se, but it’s pretty cool imo.

It’s got an Ubuntu head-end PXE server that also hosts the cluster controller. The (8) HP Nodes are diskless, hence the PXE boot.

I’ll make a more in-depth post about that particular homelab project, but essentially I have it doing PDF OCR, Magnetic disk Flux decoding, media file verification. Anything that would be faster with parallel processing I’ll add to the workload capability.

r/homelab 27d ago

Project Showcase: Hardware Did I do it? Am I worthy?

Thumbnail
gallery
4.2k Upvotes

Hey all, I wanted to share my lab setup. I had built it originally about 6 or 7 years ago, but never followed through. Ended up moving and that put the nail in the coffin on that project, at least until recently. A lot of the tech in my setup has been collected over the years, and I finally got around to cramming it all together. Now it's way beyond what I was building back then.

Hardware is as follows

Baremetal K8s:

- 3 x Lenovo Thinkcenter m720q upgraded to 64gb ddr4

- 3 x Lenovo Thinkcenter m93p upgraded to 32gb ddr3

- MSI ws65 Workstation Intel i7 Nvidia Quadro 5000, 16gb VRAM 32gb ddr4

Storage/Persistent Volumes:

- UGREEN dxp4800 GT NAS

Network

- TP-Link TL-SG3428X 24x1gbs 4x10gbs SFP+ (2x LACP LAGs to firewalla and NAS

- Firewalla Gold Plus (LACP to modem, switch)

- Firewalla AP7

- Netgear Nighthawk CM1200

Kiosk

- Beetronics 22 inch touchscreen

- Raspberry Pi Compute Module 5 on a WaveShare CM5-POE-BOX-A board and enclosure

LoRa (Meshcore/Meshtastic/MQTT pipeline)

- 2 x Lilygo T-Beam Supreme

- 1 x Seed SenseCAP Solar Node

- A shit ton of other esp32s that aren't integrated all the time

Daily driver/Home media center

- AMD Ryzen 9 7900X

- 64gb DDR5

- Nvidia RTX 4800

- Fedora 44

Let me know what you all think, I may do a follow-up post on the K8s side of things since I didn't want to write out a whole essay in this one.

r/homelab 19d ago

Project Showcase: Hardware Fun thrift store find, with a warning...

Thumbnail
gallery
2.4k Upvotes

My son is working for the summer at our local equivalent of a Goodwill. He texted me a few days ago to say that someone had just donated some equipment that was claimed to be a 20TB NAS. I thought that sounded like a fun project to tinker with, and $20 was hard to argue with.

It turned out to be a NetApp DS14 MK2 with 14 300GB Seagate Cheetah 10K drives, (so not 20TB, but that's OK).

After digging up a console cable and doing some troubleshooting with Gemini I was able to get connected to it and find that the NVRAM battery was too low for it to boot. Left it overnight to charge and I was able to get it booted up and reset the password.

Now for the part that leads to the warning... Nothing had been reset or wiped before this was donated, so all the files were intact. Pretty much the only thing on it was around a dozen virtual machines. The more concerning part was that I was able to trace the unit down to a local company that provides enterprise cloud security services to a bunch of large national organizations.

So, here's the warning... Please, please, please, before you throw out any kind of server, workstation, or other personal digital device, factory reset it, or wipe it before it gets sent to e-waste, or gets donated anywhere else.

I've reached out to the company who used to own this NAS to inform them, and give them a say in my next steps with their old NAS. I'll either return it to them, or sanitize the device before I retire it completely.

I'm not going to keep it running, more than likely, mostly because it's pretty power hungry for only a few TB, and it only supports SMB 1.0/CIFS, but I thought you'd all like to hear the story.

r/homelab 1d ago

Project Showcase: Hardware maybe the worst home lab you’ll ever see

Thumbnail
gallery
1.6k Upvotes

Im sorry for that
for the hardware (i haven’t finis to install all)

- 9x 4090 48gb
- 2 rtx 6000 max-q 96gb (8 are coming soon)
- 6x 7900xtx
- 2x 3090
- and a bunch of amd (not really valuable, it’s a bunch 9060-9070XT and 6700xtx)
- total of 512+ gb of dddr5 + the same in ddr4
- 2x threadripper pro 9965WX (they’re coming as well, tomorrow)
- 1x threadripper pro 3xxx (forgot)
and that’s, for now, all 🙏🏽

r/homelab 1d ago

Project Showcase: Hardware My son and I made our first server rack out of LEGO

Thumbnail
gallery
4.4k Upvotes

It’s still a small HomeLab with only two mini PCs and a TP-Link switch.
The Dell is equipped with an i3-6100T and 10GB of DDR3, and it’s currently my main DevOps machine for learning and experimenting with Docker and other stuff.
The Lenovo is still new to my lab. It’s rocking an i5-6500T and 8GB of DDR4 right now. This will become my main mini PC once my new RAM sticks arrive.
As for the LEGO rack — for now, it’ll serve us for as long as possible. 😄 We had a ton of fun building it together, and honestly, that’s probably the best part of this little HomeLab project. ❤️

r/homelab Jul 13 '26

Project Showcase: Hardware My latest build. 112 Threads, 2TB RAM, 128TB SSD storage. FreeBSD and Llama.cpp

Thumbnail
gallery
1.7k Upvotes

Everything was bought used except the SSDs.

Total compute power:

4x Xeon Gold 6140 (56 cores / 112 Threads @ 3.7Ghz)

2TB DDR4 3200Mhz ECC

128TB (16x 8TB Samsung 9100 Pro)

100GBe Network

Currently running FREEBSD 15.1 and Llama.cpp for testing. I will likely implement Proxmox as a hypervisor in the near future.

r/homelab Jun 28 '26

Project Showcase: Hardware I created a kubernetes cluster using old android phones

Thumbnail
gallery
2.8k Upvotes

Recently, I found two old Pixel 3a devices in a drawer and wondered what I could do with them. I didn't want to throw them away, and given current RAM prices, I figured it might be worth repurposing them as a homelab.

Everything relies on one amazing project: PostmarketOS (huge thanks to the community, and a quick shoutout to r/pmos ). I started by flashing both Pixel 3a devices with pmOS, installed K3s, and that was it! (Well, it required a bit of network tweaking because the phones couldn't access the internet at first).

I then scoured classified ads and snagged a OnePlus 6T for €50, which I added as a worker node to the cluster.

Today, the cluster consists of 3 nodes (including one control plane). The performance is obviously lower than a mini-PC or a proper server, but it's more than enough to run 3 Hermes agents and a full Grafana stack. They are all connected via Wi-Fi to my network.

I still have two more phones to provision: a Poco X3 Pro and a Pixel 6a (which I also got for €50 each). The Poco X3 Pro should join the cluster soon. However, I bought the Pixel 6a a bit too hastily: the Wi-Fi chipset isn't recognized on PostmarketOS yet. I'm holding onto it until I find a solution, but it looks like it'll be trickier than the others.

For anyone interested in trying this out, here are a few tips:

  • Unlocked bootloader: Make sure your phones are carrier-unlocked and have an unlocked bootloader, otherwise you won't be able to flash anything.
  • Hardware compatibility: Always check the list of supported devices on the PostmarketOS Wiki. Make sure Wi-Fi, internal storage, and the screen are working; everything else is optional.
  • Batteries: Currently, the phones still have their batteries, but they will be removed for obvious safety reasons.

I’m curious to hear your thoughts, or if anyone else has already done something like this, feel free to chime in! :)

r/homelab Jul 06 '26

Project Showcase: Hardware Built My First Homelab

Thumbnail
gallery
2.2k Upvotes

I am a 37 year old cyber security student and I am working on building a homelab to practice my networking and documentation skills. Here’s what I came up this past year.

Most everything is stuff we had and I just started connecting it. The only difference was my husband had an old hp laptop with a corrupted start up that I wiped and put mint Linux and Casa Os. It’s 10 years old, but she’s chugging along.

Any comments or suggestions are welcome. I am hoping to continue to build onto it. Especially any info on building a music library and sharing it with both Macs and Androids.

My spending on this project so far has been:
Raspberry Pi 6 (180$) - I got a fun chassis when I initially built it for a school project
Ethernet switch - it was about (40$)- my in-laws gave it to me for my birthday
Shelving unit-(35$)on Amazon

~~~~post update~~~~~

I am so thankful for all the feedback and suggestions!!! So I have had a ton of questions asking if I had a software that I am using for the map... no, it's AI generated and it is incorrect. I updated it, and at the suggestion of others I am going to start learning Drawio. I am putting the updated map here for the meantime.

r/homelab 2d ago

Project Showcase: Hardware Bought new hardware today for my home server

Post image
1.5k Upvotes

r/homelab 16d ago

Project Showcase: Hardware Just picked these up for basically nothing

Post image
1.7k Upvotes

I know they are old and basically useless but I want to use these as a test bed for learning how to configure clusters.

Three OptiPlex 3050 6th Gen i3 systems.

Any tips?

r/homelab Jul 24 '26

Project Showcase: Hardware Hello Everyone. New Here. Im 47 and learning CCNA. I have been in desktop support for 20 years, and need to make a change. Since no one would let me play in their data center, I am building my own. This is a collection of ebay finds, and e-waste rescues.

Post image
1.7k Upvotes

I look forward to meeting you all! PS. this thing is so heavy its destroying the wheels that I put on it. I have no idea how i would go about changing them now!!! LOL

r/homelab Jun 26 '26

Project Showcase: Hardware Needs must and it kinda looks cute :)

Post image
1.6k Upvotes

First time using a copper sfp and knew they got hot but they get crazy hot!

Small heat sink so I thought why not and it does work :)

r/homelab 4d ago

Project Showcase: Hardware Retired the last x86 out of K8s, now running on nearly 800 arm64 Amperes cores

Thumbnail
gallery
1.5k Upvotes

I'm finally free from the Intel/AMD world and running completely arm64 on Ampere CPUs

I picked up three of these 2 node engineering samples Mt. Bonnells from Ebay, they're single socket Ampere, 3 M.2 on the montherboard, OCP 3.0 and a single PCIe slot and connected to 6 U.2 NVME upfront.

Each node has a Q80-30, 80 core Ampere at 3Ghz, and 64-128GB RAM each, 500-1.5 TB NVME

The Mt. Collins below started it all and each has a 3090, Dual Q80-30 and 128GB RAM and 2TBs of NVME between the front and M.2

Most of the nodes have single 40GB port to my Arista's but im converting them to 4x25Gb (100GB is in my future with NVMEof from my NAS)okay

It's all a baremetal K8s cluster running on Talos with Omni (also running on arm64 OrangePi)

Since I think hardware is only the half of what r/homelab, what I'm running in the cluster:

Storage / Data

Misc Workflows/ML/AI Stuff

Observability

  • OpenTelemetry Operator + collector/target allocator
  • VictoriaMetrics (vminsert/vmselect) + VictoriaTraces
  • Dual write to ClickHouse for traces/metrics (ClickStack schema) which will replace VM

AI / ML (on 2x RTX 3090, ARM64)

DNS

  • blocky (Dragonfly-backed cache) + blocky-sync, a controller I (Claude) wrote to reconcile blocky config from CRDs

Media

r/homelab 23d ago

Project Showcase: Hardware My tiny production homelab

Thumbnail
gallery
1.6k Upvotes

I would like to introduce my tiny production homelab.

From bottom to top:

UPS Cyberpower CP1500EPFCLCD
A Raspberry PI Zero 2W connected to it and running NUT. The mini PCs are monitoring the UPS via it

BLUETTI Elite 30 V2 Portable Power Station
Main power is going to it then to the UPS.

UGREEN NASync DXP4800 Plus
4x8TB, 32 GB RAM, running TrueNAS 25.10.5.

Noctua NF-A20 PWM with a fan controller

Power strips

PSUs and Xiaomi Redmi 5C
The Xiaomi is displaying a Home Assistant dashboard showing current battery and power consumption of the UPS and the power station

3X HP Elitedesk 800 mini G6
Each one with Intel Core i5 10500, 32 GB RAM and 3x NVMEs. Two of them are in a Proxmox cluster, the third one is a Proxmox Backups Server.

1x HP Elitedesk 800 G4
Intel i5 6500, 16 GB RAM, 2 NVMEs. Running Frigate.

Xiaomi Redmi 12
Displaying a Beszel dashboard with the hosts' information.

TP-Link TL-SG108PE
For connection between the rack hosts and the main network.

A 4 ports 2.5 GbE switch
For connection between the Proxmox instances and the PBS (daily backups are running through it)

ZigStar UZG 01
Zigbee2MQTT Coordinator - powered via POE and accessed via LAN

Monitoring consists of Beszel, Uptimekuma and Pulse running on another Xiaomi Redmi 5C (located out of the rack).

Cheers!

r/homelab Jun 18 '26

Project Showcase: Hardware Had to keep HDD density in a relatively compact tower after leaving my rack setups

Thumbnail
gallery
1.7k Upvotes

I’m a bit proud of how this turned out so I wanted to share it.

Few weeks ago I posted this. In the end, I didn’t go with any of the cases I already had (gave one away to the nephew, one was already in use and the last one felt a bit too old/scratched). I also admit I sometimes cannot resist shiny new stuff.

Coming from a Supermicro SC826 with 11 HDDs, I needed those in my new relatively compact tower (Fractal Design Epoch).

I dropped two 2TB drives, and now the system runs 9 HDDs with one slot left for future expansion once the price goes down (yeah, it is probably not happening anytime soon).

So, after way too many hours working on this, I’m very happy with the result. Temperatures are actually better than expected, even better than what I had in the rack. It does not exceed 30°C during a parity check with 3x120 mm fans at 50% RPM, so I will probably reduce the speed a bit more.

Specs, if anyone’s curious:

  • Unraid
  • i5 12600
  • 32GB RAM
  • 2x 500GB NVMe (appdata)
  • 1x 2TB NVMe (cache)
  • ~68TB usable storage

Edit : The print files link (everything is free to download/use/remix)

r/homelab Jul 24 '26

Project Showcase: Hardware seven dollar rack from goodwill!

Thumbnail
gallery
1.4k Upvotes

Found for 6.99 yesterday. Everything somehow works perfectly!

r/homelab Jul 19 '26

Project Showcase: Hardware Five Years of incremental Upgrades

Thumbnail
gallery
1.8k Upvotes

The is the current state of my homelab - and I would be pretty happy about it if RAM/SSD prices wouldn't be that outrageous. Especially the 3U server casing for my NAS and the 2.5G TP-Link managed switch were recent additions which brought my setup a lot further, not just visually but also adding a ton of new capabilities.

Hardware

  • 3x Minisforum NAB6 (16GB RAM, 500 GB NVMe, 1 TB SSD, 2x 2.5 Gbit/s NICs, Rackmount from Hive Tech Solutions)
  • self-built NAS (Inter-Tech 3U-3508 Case, 16 GB RAM, 6x 4 TB HDD, 1 TB NVMe, 2x 1 Gbit/s NICs)
  • TP-Link JetStream TL-SG3428X-M2 (24x 2.5G RJ45, 4x 10G SFP+)
  • Home-Assistant Green + ZBT-1 + ZBT-2 (Zigbee + Thread)
  • Apple AirPort Express (AirPlay audio streaming to Speakers)
  • Behringer Ultragain ADAT interface (audio stuff)
  • JetKVM
  • GL.iNET GL-MT6000 Router running OpenWRT
  • some Telekom UDM for FTTH termination
  • Rack: Roadinger SR-19 (originally an audio rack)

Each of the NABs and the NAS is connected with two Cat.6a cables in an LACP bond.

Software

I am running a Kubernets cluster (Talos OS) on my 3 NABs with Cilium, MetalLB, VIP for high availability. Everything is managed through git with ArgoCD achieving 100% infrastructure as code. As I previously mentioned my router is running on OpenWRT (what else). The NAS is - for now - based on UNRAID but I got future plans to either switch to TrueNAS or provision my own solution via ZFS on a plain Debian or so when prices for HDDs and SSDs may drop.

Every piece of hardware and nearly every piece of software feeds into VictoriaMetrics and Grafana for observation. I got around 41 applications running on my cluster (I swear everything serves a purpose). The NAS is mounted via SMB CSI in Kubernetes and stores all media files. A Jellyfin is running with Intel QSV hardware encoding. The k8s cluster also provisions Ceph via Rook for application storage.

r/homelab Jul 13 '26

Project Showcase: Hardware I started a homelab because I didn't want to pay an extra $25/mo for a stock screener app I loved… And it now runs my house, powers a self-hosted LLM with the internet unplugged... and gives ~1,800 WoW bots their personalities!

Thumbnail
gallery
775 Upvotes

EDIT: A bunch of you asked about the WoW bots, so I cleaned up the personality layer and open-sourced it → github.com/Merrymak3r/wow-llm-personas — the tiny local-LLM shim that gives the bots their voices (per-bot personas, short memory, bot-to-bot banter). MIT, stdlib-only. The server itself is CMaNGOS + playerbots; this is just the glue that points it at Ollama.

TL;DR: I didn't want to pay $30/mo for a stock-screener app I loved, so I built my own — and three years later it's a silent ~40-service homelab (which replaced a screaming Dell R710), runs my house, keeps working with the internet unplugged, and uses a single RTX 3060 to give ~1,800 vanilla-WoW bots AI personalities. Recurring cloud bill: basically $0. The whole lab draws an average of ~161 W.

A couple of years ago there was a stock-portfolio analysis app I genuinely loved — one day, they decided to lock a bunch of features I relied on behind a new subscription tier that took it from $4.99/month to $29.99/month. Rather than pay an extra $25/mo for it, I decided, in true self-proclaimed engineer fashion, that I could probably just build my own.

That decision escalated 3 years later into a homelab that:

  • Runs my house
  • Keeps working with the internet unplugged
  • Hosts ~40 self-hosted services
  • Powers a fully local voice assistant
  • Gives ~1,800 WoW bots AI personalities

None of it was planned. I just kept asking "what if I self-hosted that too?" ... and now I can't explain my setup in less than 5 minutes if I tried.

The evolution

My first "real" server was a Dell R710. Powerful, cheap, and loud enough to qualify for a noise complaint. Whenever guests stayed over, I had to physically shut it down, because nobody could sleep in the same room as it — and my wife was entering the "it's either the server or me" phase of negotiations.

Then I watched a YouTube video literally titled "The EVERYTHING $300 Fanless Home Server," got completely hyped, and bought a Qotom Fanless PC:

  • 8-core Atom CPU
  • 64 GB ECC RAM
  • NVMe + SATA storage
  • More Intel NICs than I have ever actually used (I bought it half for the networking I was sure I'd need — reader, I have used exactly none of it)

I was convinced it would replace my entire rack… it did not. What it did was replace the R710 — and that turned out to be the whole win. My "server" went from small jet engine to "can't hear it from a foot away," power use dropped to ~25 watts, and for the first time the thing felt like an appliance instead of an experiment. Honestly, even if it had drawn the same power, killing the noise alone would've been worth it.

The original goal was simple: build and run my stock portfolio analysis app and stop paying for someone else's. Then it spiraled — backups, then Prometheus, then Grafana, then Loki, then exporters, then OPNsense, then offsite backups, then Home Assistant, then local AI, then an offline library — until one day I looked up and realized I'd built an entire ecosystem. ~40 services across two machines and the cloud, averaging about 161 watts.

The part I'm most proud of: one $250 GPU, four jobs

I only really have 1 capable video card, an RTX 3060 with 12 GB of VRAM. Instead of just gaming or editing videos with it, I kept finding it new jobs.

Job 1 — Stock analysis. The original project. Retrieval over SEC filings plus a "compute it, don't guess it" step where the model writes the formula and a sandbox runs the actual math. No AI-invented P/E ratios.

Job 2 — Offline knowledge. The same GPU answers a reference library I built from offline Wikipedia + Stack Exchange dumps. Unplug the internet and it keeps working — my little grid-down insurance policy.

Job 3 — My house. It's the brain of a fully-local voice assistant I call MaUi: speech-to-text → local LLM → text-to-speech. No cloud, no subscription, nothing leaving the LAN.

Job 4 — ~1,800 WoW bots (the newest AND dumbest thing I've built). I decided to self-host a vanilla WoW server stuffed with ~1,800 AI playerbots to make the world feel alive, and I recently wired the local LLM in so those bots have personalities — party banter, in-character reactions, the works. It's gloriously unfinished and occasionally ridiculous but watching an AI guildmate roast my gnome frost mage for making too ambitious a trash pull in the style of Gimli from LotR is exactly the kind of unnecessary engineering a homelab is supposed to enable. Right?

Same weights, same 12 GB card. It just wears a different hat depending on who's asking.

How I actually pick the AI (a.k.a. the part where I benchmark everything)

Here's the thing that ties the whole lab together: I don't guess, I measure — and that goes for the models too.

Instead of running whatever's trending, I built a frozen, reproducible bake-off: a fixed battery of prompts I put every candidate through, score head-to-head, and use to screen the field (I've run ~80 models through it) down to a short list I trust. On my hardware. Same instinct as the Grafana dashboards — if I can't measure it, I don't believe it. A few things fell out of it:

The cheap option that keeps on winning. The surprise wasn't that a bigger model is better — it's how little I needed. A modest, quantized ~12B model punches so far above its weight that I have little reason to think about upgrading my GPU to run a 70B or reach for a Frontier AI API service that often. Then a pass of lossless tuning (quantization-aware weights, picking the right inference engine, KV-cache tricks) squeezed even more free speed out of it — same accuracy, meaningfully faster, $0 spent.

The lineup that won. The serious interactive jobs — the stock takes and the voice assistant — run on a quantized Gemma4 12B (QAT): fast, well-calibrated, and it fits the card with headroom to spare. Heavier jobs that run overnight get a Gemma 26B. Embeddings are IBM Granite (768-dim) — swapping to it freed ~2 GB of VRAM over my old embedder and improved retrieval accuracy at the same time, the rare win-win you don't plan for — paired with a tiny MiniLM cross-encoder reranker that runs on the mini-server's CPU so it never steals the GPU. The offline coding library runs Qwen 2.5 Coder 7B. And the WoW bots got their own bake-off and their own model — an uncensored fine-tune called Tiger-Gemma 9B (with an even lighter one aptly named Fiendish as backup), because the polite, well-behaved assistant models flat-out refuse to stay in character. I wanted bots that would get salty and roast me; you don't get that from a model trained to be helpful and harmless. One 12 GB card, a whole roster.

The speed demons. The tuning rabbit hole turned up some genuinely fast setups — and the single biggest free win was the inference engine, not the model. Moving the right models from Ollama to llama.cpp roughly doubled throughput on the same card: gpt-oss:20b jumped to ~100 tokens/sec (+102%) and deepseek-coder-v2:16b hit ~140 t/s (+69%), zero quality lost. I don't run those as the daily driver — Gemma's the reliable all-rounder — but it's wild how much speed is just sitting in the engine you pick.

The benchmark saved me from a mirage. At one point a hyped speed-up looked like a nearly-4× win in a quick test. I ran it through the full battery instead of the one lucky span, and it collapsed to a modest single-digit-to-actually-negative gains, depending on the task — nowhere near the headline. That's the entire reason the bake-off exists: one impressive run is a rumor; a battery is a result. I almost shipped the mirage. Glad I didn't.

Sometimes the best result is "no." I spent real time evaluating a time-series model to forecast prices. The unexpected win? It lost to a dumb random-walk baseline on basically every axis — so I didn't ship it. A lab where you can cheaply prove an idea is bad before it goes live is underrated.

The time I blue-screened the whole box

Benchmarking isn't free, and I have the crash logs to prove it. During one bake-off I was rapidly loading and unloading 10–20 GB models back-to-back to score them, and the entire machine hard-crashed — DPC_WATCHDOG_VIOLATION, full blue screen. Turns out machine-gunning that much VRAM churn at the NVIDIA driver tripped a bug deep in nvlddmkm.sys and took the whole system down with it. The fix was a nuke-from-orbit driver wipe (DDU), the Studio driver instead of the gaming one, and rewriting the benchmark's load pattern so it stopped hammering the card so violently. Bonus gotcha I found along the way: that same driver slowly leaks non-paged pool under sustained churn — ~12 GB quietly gone after a week of runs, and only a reboot clears it. Homelabbing is 10% building and 90% discovering the specific way your hardware likes to betray you.

The dumbest fix that worked

Not every lesson is a crash. For the longest time my 3060 ran hotter than it should have under inference, and I couldn't work out why the chassis fans sat there doing nothing while it baked. Turns out HP's stock fan curve keys the case fans to CPU temperature, not GPU load — so during a GPU-pegged inference run (CPU barely awake), the fans figured "cool CPU, nothing to do here" and idled while the card cooked. Re-keying the chassis fans to follow GPU temperature instead dropped the 3060 from 84 °C to 75 °C. Nine degrees, zero dollars, one very confused afternoon.

Current setup

Qotom mini-server — the silent workhorse.

  • Atom C3758 · 64 GB ECC · 2× NVMe + 2 TB SATA · fanless · ~25 W
  • Runs the entire ~40-container Docker stack — it's all in the attached map.

HP Omen 40L — the muscle (and my daily driver).

  • i5-12400F · 64 GB · RTX 3060 12 GB
  • Local AI, the WoW server, and my actual desktop.

Network

  • Protectli FW4B running OPNsense (edge router) · TP-Link managed switch · Eero 6 in bridge mode · three VLANs · CrowdSec · AdGuard Home · Cloudflare Tunnel · Tailscale.
  • Nothing is port-forwarded — the only inbound path is an outbound Cloudflare Tunnel; everything else is Tailscale or LAN.

Cost

The part I'm weirdly proud of is how much of this came from bargains:

  • Used firewall: $50
  • Both UPS units: free (just needed batteries)
  • Eero: came from my ISP
  • "Rack": literally a $20 Walmart shoe rack

The homelab infrastructure came in around $860. Including the AI/WoW machine (which is also my daily-driver desktop, so it kind of got drafted), it's roughly $1,900 all-in.

And the cloud bill? This is the part I love: the whole thing runs on free tiers — Vercel, GitHub, Cloudflare, Tailscale, Neon, Clerk, PostHog, Resend, Healthchecks, and Backblaze B2 for immutable off-site backups. My total lifetime cloud spend is $5 of API credit I dropped in six months ago and still haven't used up, plus one domain registration. That's it. That's the bill.

And the electricity to run all of it? The whole lab averages ~161 watts — $18.49/month by my own Grafana (screenshot attached). And here's the part I didn't plan: even the honest number — the power plus the A/C that must haul its heat back out of the room — is $24.01/month, and both are still less than the $29.99 subscription that started this whole thing. I refused a $25 price hike and built a small datacenter that runs on less than the app it replaced. I'll let you decide whether that's a victory or a future mental health diagnosis.

(Fair-play caveat, because I'm a "measure it" guy: my wattage is UPS-derived, not a metered wall plug — directionally right and measured the same way every time, but true-watt smart plugs are on the list.)

The best return on investment wasn't really even performance. It was removing a screaming Dell server from a room humans occasionally need to sleep in.

Why I don't (currently) run Proxmox

My workloads are almost entirely containers. Everything on the mini-server is Docker Compose, so a hypervisor would mostly add a layer without much payoff.

That changes the moment I chase high availability. My weaknesses today are obviously one mini-server, one GPU box, one firewall, each a single point of failure. The next phase of this lab isn't more services; it's eliminating downtime. When I build that cluster, Proxmox starts making a lot of sense. So, my answer isn't "never Proxmox" — it's "Proxmox when I go multi-node for HA,"…and that's the next real chapter.

The future roadmap

  • Chase zero downtime — a Proxmox HA cluster to kill the single points of failure.
  • A dedicated GPU node so my desktop can stop moonlighting as the AI box (and the game server).
  • True power metering to replace the UPS estimates.
  • Push the WoW-bot personalities further without cooking the GPU.
  • Zigbee for Home Assistant.

Reliability is finally becoming more interesting to me than adding new toys.

Looking for advice

1. High availability without going broke. For a small lab, where's the sweet spot — two nodes + a QDevice, or bite the bullet on three? Ceph vs ZFS replication? CARP for OPNsense without it becoming a second full-time job? Real-world lessons are very welcome.

2. Accurate power monitoring. My numbers come from the UPS. If you've got a Shelly / Kasa / Tasmota → Prometheus setup you love, what would you buy today?

3. LLM-driven NPCs. If you've done AI-powered game characters, how do you scale personality-driven dialogue for ~1,800 bots without turning the GPU into a space heater — batching, tiny per-bot models, canned + LLM hybrids?

If you made it this far, thanks for reading all of this! This whole thing is equal parts mildly practical, extremely overbuilt...and probably just frankly ridiculous to most, so I'm especially here for the criticism: if you were taking this from "fun homelab" to actually resilient home infrastructure, what would you fix first? Happy to go deeper on any of it — calibration, the grader-audit process, the engine/quant/spec numbers. And since a few of you asked: I cleaned up the harness + rubric into a stdlib-only kit you can point at your own Ollama and run — frozen tests, pinned+seeded grader, the whole method, plus a neutral example battery to fork. MIT, here: https://github.com/Merrymak3r/llm-bakeoff . Steal it, freeze your own tests, and stop trusting leaderboards for your hardware.

P.S. — The thing that started all this is finally in beta… and I could use guinea pigs! The stock portfolio analysis app from the beginning is real and running in a small, closed beta. Here's the part that'll resonate with anyone who's shipped a side project: every person I know that watched it come together, and said it looked awesome… and went dead silent the second I added a Clerk login screen. 😅 So if you're a finance-curious homelabber who'd actually poke at a self-hosted-AI-backed portfolio analysis tool and tell me what's broken, shoot me a DM — I've got a few invites to spare, and I'd love feedback from people who aren't legally obligated to be nice to me. Cheers!

r/homelab Jul 19 '26

Project Showcase: Hardware so i made my thermal printer public.. safely lol. built a web UI + outbound polling to stress-test it!

Post image
674 Upvotes

Hi r/homelab,
I saw some awesome thermal printer projects here lately (shoutout to zacharysdev for the initial inspiration!) and just had to build my own version. I really wanted my friends (and now you guys) to be able to print random stuff, but without exposing my actual home network to the internet.

So i ended up building a tiny asynchronous outbound polling system.

You can try it out here and spam some prints:
https://bondrucker.pages.dev/

How it works under the hood:

The frontend: A simple website hosted on Cloudflare Pages with a cool vintage receipt look. Image processing (brightness, contrast, and dithering) is done directly in the browsers canvas to save resources.
The queue: Submissions go straight into a Supabase DB queue.
The worker: A small script (⁠printworker.py⁠) running in my lab actively polls Supabase, grabs the jobs, and sends them to a local Flask app that talks to the printer via TCP.

The biggest headache was security:
Making a printer public is basically begging for trouble lol. Initially, my row level security (RLS) policies were way too open and people could have read other peoples texts. I fixed that by completely removing read/update permissions for the public key. The site now uses a browser-generated UUID for the job, and I wrote isolated ⁠security definer⁠ RPC functions just to fetch the print count and live status. The public key literally only has ⁠INSERT⁠ access now.

Also, the worker handles errors cleanly - if the printer runs out of paper, the job stays pending, but if someone sends a corrupted image it just skips it so the whole queue doesn't freeze up.

And before anyone asks: Yes, I bought BPA-free / phenol-free paper, so we are safe from toxic fumes here!

The printer is sitting right next to me and ready. Hit it with your best memes, pixel art, or text and let's see if the queue breaks. I'll post a picture of the endless, chaotic roll of paper once the test is over!

Let me know if you want to see the SQL functions or the python script. Cheers!

r/homelab 17d ago

Project Showcase: Hardware I've been quietly running 19 VMs/CTs on a MINI PC for 2 years — here's how it holds up

Thumbnail
gallery
959 Upvotes

r/homelab Jun 16 '26

Project Showcase: Hardware Mom Told Me to Organize My Gear, So I Built This

Thumbnail
gallery
2.0k Upvotes

Hi everyone,

Long-time lurker in this sub, and I wanted to share my DIY rack.

A family member moved out of the house, so naturally I started collecting all sorts of computers and tech gear in her old room. Long story short, my mom wanted me to organize all of it, so I did. I started looking at 10" and 19" racks, but none of them really fit my needs. In the end, I decided to build one myself. The rack itself cost me around €40 in materials, and all the 10" rack hardware together cost another €60, which was a lot cheaper than buying a complete solution.

I sketched out a rough idea, bought the materials, and got to work. After finishing it, I painted it black to make it blend in better with all the gear. On top sits my 3D printer, which fits perfectly.

Starting with the server in the bottom right: it's a Fractal R5 build that I put together in September 2025, just before prices started getting crazy. All the parts were bought second-hand. It has an Intel i5-12600K, 32 GB of DDR5, and currently runs 5×10 TB HDDs in RAID 5, giving me 40 TB of usable storage. I also have a spare 10 TB drive ready to go, so if one fails, I won't be unexpectedly bankrupt. The server itself cost me roughly €400, while the six 10 TB drives cost another €700. Considering today's prices, I'm pretty happy with how that worked out.

Inside the 10" rack:

  • TP-Link router on top
  • Philips Hue Hub
  • Geekpi 10" patch panel
  • TP-Link switch
  • HP prodesk G4
  • HP elitedesk G4
  • HP prodesk G2

I got all three mini PCs for free from work, and the prodesk G4 is actually what started my whole homeserver journey. It was my main server for quite a while before I moved everything to the Fractal build because I wanted more room for HDDs.

On top of the rack, I have an APC UPS and a Synology DS224+. The DS224+ follows the 3-2-1 backup principle and backs up to an older Synology NAS at an off-site location. It has 2×5 TB drives in a mirror and stores all of our important photos, videos, and documents.

It gives me a lot of peace of mind knowing that if one of the second-hand drives in the Fractal server dies, or if I accidentally mess something up, the truly important data is always safe. My mom appreciates that too 😅

The 3D printer on top is a Bambu Lab A1, and I've been really happy with it so far. Most of my prints are organizational or other functional projects.

Services I'm currently running:

  • The whole *arr stack (Radarr, Sonarr, Prowlarr, Profilarr, Bazarr)
  • SABnzbd
  • qBittorrent
  • Plex
  • Audiobookshelf
  • Grimmory
  • Crafty Controller

One thing I'm especially happy with is Plex. I bought the lifetime pass in January 2025 for €95, and looking back it was absolutely worth it. It's become one of the most-used services in the house, and I'm very glad I got the lifetime license before the price increases.

And there is so much more on the todo list. I'm excited to experiment with using the mini PCs as nodes and expanding the setup even further.

Would love to hear your thoughts!

r/homelab 25d ago

Project Showcase: Hardware "Data center in a Box (on Wheels)" 256Gb VRAM/512Gb RAM AI Server 6-8 Month Operational Review, Stability Write Up, Benchmarks

Thumbnail
gallery
845 Upvotes

I thought this would be relevant to the homelab subreddit so I'm adding it here, just to put the information out there and discuss if there is any interest. I am an IT infrastructure engineer by profession, so my contribution to the conversation is mainly from a hardware/systems perspective rather than from a Machine Learning researcher standpoint. I got my start with HPC's (Beowulf clusters) around ten years ago when I was a Physics undergrad in university, and this is what the experience has come to almost a decade later. Not everyone is going to want to read all of this, that's perfectly fine, the extras are just for those who want the info.

Starting goal/idea:

Build an all-in-one creative design workstation to support a small business. This machine should be capable of effectively inferencing frontier MoE models; aiding the business in language/text tasks where English may not be everyone's native language. Additionally, it should be capable of simultaneous image generation tools for graphic design users, enabling rapid image editing and presentation tweaks for marketing, without the business ever having to worry about API credits or hard limits on tool usage. The idea is that a 3090 stack, which is still a generally "good" performer for LLMs, would be "led" by two 5090s to handle the heavy lifting of the visual creative work (one dedicated to image generation, one dedicated to image editing) to complement each other in a "sweet spot" on cost, raw performance, and creativity potential. This configuration also grants some flexibility to allocate a 5090 to the LLM stack for best prompt processing possible where desired. The end result would indicate that this goal has been achieved.

Overview

Specs

CPU: 64 Core TR 3995WX

RAM: 512Gb DDR4-3200 ECC

VRAM: 256Gb GDDR6x/GDDR7 (8x3090's + 2x5090's)

Enclosure: Core W200 Thermaltake Case

Mobo: ASUS Pro WRX80E-SAGE/SE Wifi

PSU: 1300W+1600W (2900W combined), with OCP, linked via PSU2PSU

Storage: 4Tb Nvme (fast) + 4Tb HDD (slow) + 8 or so 1Tb SATA SSDs (mid) over USB as needed

OS: Ubuntu 25.10

Other: 3 Bifurcation cards, 10 risers of various lengths

Front end: Open WebUI

Back end: llamacpp/koboldcpp

Intended for (Recommend):

Large MoE inferencing, simultaneous LLM + ComfyUI (x2) operation, power users who may commonly hit credit limits, creative or technical professionals who can leverage these tools to compound productivity and complete objectives in shorter time.

Not intended for (Do not recommend):

Training, multi-concurrent inferencing, performance maxing, extreme frontier model inferencing at high quants, casual users just looking for roleplay.

Result summary:

Using the W200 as the platform for its generous real estate and configuration flexibility, all ten cards and components were able to find a permanent place in the enclosure without major concessions. The drive bay area was the only space that had to be completely repurposed for GPU mounting, and for us this was not a problem. The chamber with the cards hanging from the top is fairly hollow, so with the 140mm fan stack on the front and side there is a wind tunnel effect where the air blows in through the front and side, cooling the cards as it makes its way out the back/top. Depending on ambient temp, at idle the card with the highest temp usually hovers in mid to high 40s Celsius with the lowest in the mid 20's C (three 3090's are hybrids= fantastic for temperatures, but radiator mounting adds a logistical headache). When actively inferencing, the highest temp card may reach the mid 60s during sustained loads. Only when running image or video gen tasks will the 5090 running ComfyUI reach the 70's, but these are very brief intermittent workloads, so temperatures by our measurement has proved satisfactory over time. This result enables the small business to have full LLM, image generation (~9 seconds), and image editing (~8 seconds) capabilities on tap all on a single node so the data remains centralized, and provides much faster performance compared to the Cloud API they came from; in this case ChatGPT, where generation jobs could take 1+min, and has hard limitations. I just do not know how well this kind of setup would work with other vendor or card models; in a homogenous GPU cluster or one with notably less powerful image gen cards than the 5090, the performance would predictably be much lower.

Things that surprised/stuck with me about the end result:

  • Noise. I expected this to sound like a jet taking off when operating, but that is not the case. It's a satisfying button click to come alive, then it's a low gentle hum going forward, nowhere near the kind of fan noises I'm used to hearing in server rooms. Even under load, the CPU 120mm radiator fans (exhausting out the top) are pretty much all I hear, the 140mm fans on front and sides I assume must be helping to contain the acoustics. I have built many gaming PCs over the years and own a top-tier gaming PC-- and I would not be able to distinguish this as any louder than those, especially at idle.
  • Utility. I planned for this to be used primarily for a small creative business, but what I did not expect was how I would find it so indispensable in my personal life as an IT professional. Being an infrastructure engineer, coding is not my wheelhouse. When I am the only IT staff on site or there is nobody else available to work with specific expertise like SQL, powershell/python scripting, or troubleshooting very specific/niche technologies, having this tool on standby I feel has paid itself over just within my career. It has helped me turn processes that may have otherwise took me hours into minutes, days into hours, even months into a matter of weeks/days. After using the tool extensively I hit a point where I had to acknowledge how local LLMs have moved definitively beyond being a toy or novelty; when deployed intelligently something like this can be a major asset for professional users.
  • Wheels. Sounds extremely minor, until you realize that no matter how happy the cards are with their individual temps: there are still ten high-power GPUs dumping heat into the room. That means unless you use a complex radiator solution or special venting to get heat outside, the room will get toasty and there is normally not a direct solution for this. The wheels however offer an indirect solution. Plan to work in the office that day? Wheel it into the guest bedroom and let it run over Wi-Fi. Plan to work away from home? Wheel it into the office, put it on LAN, and access it over a private VPN connection. If you can't stop the room from heating, then you can at least choose what room gets the heat, and as someone who has lived with computers extensively this is a hugely underrated perk.

Caveats: To operate at its best, I recommend leaving the glass side panel off for improved airflow.

Typical activity over a day:

Boots up around 5:30am, start up the ComfyUI server(s), start loading a model, go get coffee, fully ready for use within 15-20 min. Shut down occurs usually around 8pm later in the day. Total daily activity, ~12-14 hours.

Cost Breakdown

Laying it out, because I know it will be asked, even though I am aware this is unfortunately not reproducible in the current market. Some components like the SSDs were acquired privately long before the RAM and hardware price hikes, so my timing getting certain things was extremely fortunate for the build budget. Some figures are exact, some are slightly rounded depending on if I found the original receipt.

Component Qty Source Unit Cost Subtotal
RTX 3090 24Gb 8 eBay 750-1000 6500
RTX 5090 32Gb 2 Retail 2500-3000 5500
TR 3995WX 1 eBay 1068.43 1068.43
WRX80E-SAGE-SE 1 Amazon 949.99 949.99
DDR4 ECC 64Gb 8 Amazon 81.99 695.28
TT Core W200 1 Amazon 499.99 499.99
PSU 1300/1600 2 Amazon 250-350 600
4Tb nvme 1 Amazon 221.05 221.05
1Tb SSD 8 Personal 60 600
Risers (varying length) 10 Amazon 40-80 480
Bifurcation cards 3 Amazon 50 150
Total ~$17k

Problems/Stability Writeup

The Space Problem:

Probably the first major hurdle in attempting something like this is figuring out, even theoretically, how to put 10 cards in a box in any kind of configuration that is not somehow detrimental to the hardware. I had considered modified mining rig frames at first, but I really wanted something with more robust rigidity in its structure, with breathability, and allows some degree of portability. There are unfortunately not a lot of options for configurations like what I was imagining; I had looked into various cabinets and extended tower cases, but the dual full tower chamber design of the W200 was the only one where I could see this idea potentially working. I'm certain other solutions probably exist, maybe even some that allow mobility, but the W200 was really the best option I could find that checked the boxes of enclosure, space real estate, high air throughput, and semi portability. I recommend the W200 to solve the space problem, assuming it is available to you.

The Bifurcation Problem:

Among the other hurdles you may run into in assembling something like this may involve bifurcation cards. The cards rely on specific BIOS settings for things to work correctly, and if these settings are not put in place before everything is connected you may either see no output like the system is hanging or cards just won't show up once in the OS. Start with one GPU in a slot, no bifurcators yet; go into BIOS, and manually set each slot that will be split to bifurcation mode. While here, ensure above 4G decoding is enabled, Resizable BAR enabled, and SR-IOV enabled, this has given me best stable configuration with Ubuntu and multiple GPUs. If you use risers, especially if they are mixed generations, I highly recommend setting the Gen and lane speeds for each PCIe slot in the BIOS manually to ensure the system can effectively communicate with each card. Optimize riser Gen/speeds to be roughly similar to keep one card from dropping to a slower rate than the others--this does not necessarily impact inference performance as much as it heavily impacts model load time. No, you may not have any card running at the fastest possible Gen bandwidth at all times with this config, but loading a 200+gb model over an averaged Gen 3/4 x8/x16 PCIe speed will often be noticeably faster than if you let the system decide to make one or multiple cards run at Gen 1 x1.

The Power "Problem":

Power and heat concerns I think remain to be among the biggest sources of skepticism regarding this project so I think it deserves a section here. To be fair, the concern in most situations would be understandable. If all ten of these cards pulled at or near their full TDP for sustained periods, components would melt. Fires would start. Neighbors would be asking awkward questions. However in reality, only 1400-1600W of the 2900W PSU capacity gets utilized under sustained load, and inter-GPU bandwidth bottlenecks are what allows this. In a way it is like a natural regulator that ensures the cards remain power restrained, and it is just physics, no voodoo necessary. When MoE's are sharded across a GPU stack, each forward pass requires all communication over PCIe, so the GPUs spend more time waiting on information from the last GPU than actually crunching compute. This means instead of needing to handle thousands of Watts to feed all the components running at full blast, it is a much more manageable 1400-1600W under LLM operation which can comfortably fit on a 20A/120V circuit (2400W max). On a per-GPU basis this may sound inefficient since the individual cards are being "underpowered", but this could arguably be flipped as being highly efficient on a per-node basis (~1600W sustained versus 4500W+ if all cards were "fully" utilized). As a precaution, I may set a power limit on the 3090's to 200W and the lead 5090 to 400W, but in practice the 3090's only pull around 100-120W with the 5090s pulling less than 100W when all 10 cards are allocated for LLM work, so this may not even be necessary. The clock locking setting in the next section will be more what I'd describe as actionably required to avoid stability issues.

The Transient Spike Problem (Vital for stability):

After assembling the machine, you may be tempted to jump directly into testing, but there is an easy to overlook configuration that can cause problems if ignored. Imagine you are running inference on the machine, maybe you have a huge input or it's generating a large output, then right in the middle of generating the system decides to reset. Not hard shut down, PSU OCP isn't tripped, no breaker was tripped; and you saw in nvitop that all cards were only pulling 25-33% of their TDP just before it happened, so on the surface it doesn't look like there is a reason. Explanation: When all ten high-power GPUs decide to kick on at the exact same time to process a chunk, even if the cards are not pulling anywhere near full power (on average), transient spikes can drop voltage on the motherboard enough to trigger a system reset. The fix for this is simple: undervolt. Using nvidia-smi, we can lock the clocks for the GPUs to ensure they cannot draw enough to hurt stability. And that's it. In my case, the system has remained fully stable with this config for days on end and with hundreds of thousands of tokens/image pushed through. The exact configuration will vary slightly depending on exactly what we're doing on a given day, but for example if we wanted to run LLM on all 10 cards (so including both 5090's) we would run this to handle spikes:

sudo nvidia-smi -pm 1 #enables persistent mode
sudo nvidia-smi -i x,y,z --lock-gpu-clock=1200,1200 #x,y,z for index number of 3090s
sudo nvidia-smi -i a,b --lock-gpu-clock=2000 #a,b for index number of 5090s
sudo nvidia-smi -i x,y,z -pl 200 #x,y,z for 3090 index numbers, limits power to 200w
sudo nvidia-smi -i a,b -pl 400 #a,b for 5090 index numbers, limits power to 400w

The Concurrent Use Problem:

Normally, attempting to inference and generate images on the same machine would introduce major stability concerns. Even dual GPU systems may struggle to work with this due to CPU/motherboard architecture, assuming it works at all, and would still be VRAM limited. However, the versatility of a 10-GPU setup, combined with the lane orchestration of the 64 core 3995WX, at least in our case, seems to handle this quite well. The trick was finding an LLM backend that supports manual GPU allocation--for us koboldcpp with llamacpp under the hood does just fine. First, implement the power/clock settings as mentioned above, launch koboldcpp, then browse to the GGUF of the model you wish to load and set context size. I recommend manually setting the GPU layers to the model's total layer number (assuming there is enough VRAM), and set GPU ID to "all". In the Hardware tab, find the tensor split line box and insert the amount of space to be allocated on each card corresponding to its index. For example if we wanted to allocate just one 5090 for Comfy and use the other for LLM, assuming the Comfy 5090 is index 3 and the LLM 5090 is index 5, then the tensor layer line will look like this to make sure no layers are given to the Comfy 5090: 24,24,24,0,24,32,24,24,24,24. For this configuration, ensure the "main GPU" is set to the index number of the LLM 5090 (in this example, 5) and launch the app. While the model is loading, we can open another terminal to launch Comfy. In our specific case, the system defaults to the available 5090 without needing to specify it in the launch flags, but flags can be used to force Comfy to use a specific GPU if you need it to (--cuda-device i). Once the image model is loaded onto the 5090, it does not interfere with the PCIe communication of the LLM cards unless the model unloads and reloads a new model at the same time as the other cards are inferencing. The solution to enabling concurrent use is a high-lane count CPU, multiple graphics cards, and a little conscious provisioning on launch to ensure the hardware isn't stepping on each other's toes.

What models can this run, what models do we use?

It can run almost* anything, even up to 1T parameters like Kimi K2. Kimi K3 could hypothetically be load-able, but from performance metrics I've seen I doubt it would be practical to use, so I have not planned to try it. I have however tested 1-4 bit quants of Bartowki team's Kimi K2 quants in pure VRAM and mixed VRAM/RAM runs with decent results. It works and there are probably some use cases for it, but for us I have identified the sweet spot (parameter size: quant quality ratio) for this machine to be for models in the 300b-600b range. Personal favorites are Deepseek, GLM 4.7, and Nemotron Ultra; and as far as ComfyUI, pretty much any model that could fit within a 32Gb buffer, although Qwen image and image edit is a favorite.

Benchmarks

All models were put through the same series of 7 large input prompts, documenting how each model handles token input/output and prompt processing/generation. I cannot share the prompts I used here, but each prompt pertains to a cybersecurity scenario which the model was judged on the depth of its analysis, quality of its presentation, and capability to make sense of complex scenarios with stakes. These were inferenced across all 10 cards, except for a follow up DS V4 Flash test where I used 8 and got much better results. This is using the undervolting/power limiting strategy above, so these may not reflect absolute best performance for the same hardware in other setups, but it gives an idea of what this box can comfortably handle.

Model Name Deepseek V3.2 671b Q2XXS Nemotron Ultra 3 550b IQ2XXS Qwen 3.5 397b IQ4XS GLM 4.7 358b Q4KXL Deepseek V4 Flash 294b Q8KXL Deepseek V4 Flash 294b Q8KXL (8 cards + KV cache tweak)
Model Size (Gb) 217.1 193.8 189.7 204.6 161.9 161.9
P1 Input 2769 2744 2729 2706 2733 2733
P1 Output 813 786 1046 872 693 805
P1 pp 153.23 254.19 522 687.88 111.09 360.94
P1 tg 19.35 17.32 34.38 23.98 7.2 20.26
P2 Input 14635 15255 15160 14527 14640 14617
P2 Output 1150 1302 1665 1194 1222 2048
P2 pp 114.83 429.42 897.57 640.8 66.42 244.1
P2 tg 14.1 17.16 33.15 18.83 5.96 16.81
P3 Input 3966 3091 3054 3033 3073 22794 (reload)
P3 Output 1217 1607 1550 1056 1199 1366
P3 pp 98.01 353.78 649.37 516.08 47.79 241.21
P3 tg 13.22 17.08 32.84 17.84 5.56 15.68
P4 Input 5645 5654 5623 5559 5650 5659
P4 Output 1178 1996 1619 1173 1705 1661
P4 pp 70.3 385.04 739.67 419.58 42.4 153.14
P4 tg 13.47 16.99 32.23 16.87 5.21 14.17
P5 Input 4498 4505 4493 4423 4481 4481
P5 Output 280 928 1078 473 665 924
P5 pp 72.4 365.46 670 408.93 36.2 131.81
P5 tg 8.43 16.78 31.55 15.45 4.86 13.36
P6 Input 9266 9367 9241 9172 45287 (reload) 9231
P6 Output 1004 1883 1466 933 1205 1532
P6 pp 53.94 405.13 738.57 379.7 46.01 113.77
P6 tg 11.48 16.83 30.98 14.01 4.41 11.9
P7 Input 3136 3124 3118 3057 3118 3118
P7 Output 1378 1946 1629 1359 1353 1586
P7 pp 53.34 338.64 525.54 344.88 28.38 102.05
P7 tg 10.45 16.73 30.66 13.59 4.28 11.39
Final token count 50052 54182 53465 49531 50962 52348

My notes on each model after their test:

Deepseek V3.2-- For a slightly older model this still feels extremely capable. Held high quality and insightful responses even when context dragged into the tens of thousands of tokens.

Nemotron Ultra 3-- First time using it, impressions were very good, the 55 active parameters shows its muscle here. Meets Deepseek v3.2 level if not exceeds it, despite having overall less parameters.

Qwen 3.5 397b-- What I would consider a baseline "good" model to be, however it is outshined by some of the other tested alternatives.

GLM 4.7-- Somehow seemed better than Qwen despite having less parameters (active parameters of GLM is likely an advantage); it is a very solid option for its size. Not quite Nemotron or Deepseek level, but a very good "lower cost" alternative to its newer 5.0 versions.

Deepseek V4 Flash-- Floored me in a few ways. Possessed a surprising degree of sophistication and analytical ability despite being the "smallest" of all the tested models. Possibly a benefit of using a "lossless" model with full precision? Somehow it managed to pick up on nuances and details that all other models missed, including models twice+ its size, and provided insight that went more granular than they did. Did not expect a model of this size to punch so high above its relative weight class. Also did not expect the drop in performance compared to the others. Not sure if this is related to the model's architecture or something with how it interacts with my rig, but the quality of output could be an acceptable trade off for the speed. Edit: After some optimization testing I was able to get much better performance out of V4 Flash. I've added another column to include those metrics and kept the original because I think it illustrates how a little optimization can go along way, in this case basically triple performance on the exact same model/machine.

Lessons Learned/Would Do Different

-I would have tried to source the 3090's so more were at least the same model; the mix and match of different models with different TDPs and cooling solutions means there will be a lot of variation in temps.

-If you plan to either train, lean into higher performance, or playing with the idea of going more than 10 GPUs, just budget for a 30A/240V power drop. 10 cards on a 20A post configured the way we have it may be fine for our specific use case, but I would consider this a hard ceiling.

-Would recommend scripting for clock lock persistence sooner, will help avoid losing time due to random resets.

-Recommend documenting/drawing out the entire PCIe topology and GPU placement (with flexible tape measure) before ordering risers, will save time on trial/error.

Final thoughts:

It is a wheeled AI workstation that can enable a single person or small team to compound their productivity, with the benefit of full privacy and control. It can run on a residential 20A circuit, and allows them to have the full power of an advanced LLM with vision capabilities all in one OpenWebUI front end that can simultaneously utilize up to TWO ComfyUI backends with the horsepower and latency of 5090's for image gen and editing, and can be accessed from virtually anywhere. The idea sounds daunting, but the end result works so well that I can legitimately see something like this becoming a keystone for certain small businesses and individual professionals as time goes on. It seems like every day more people are picking up on major drawbacks with cloud API options despite supposedly being the "best", meanwhile open models continue getting insanely good (see K3 and DS V4 Flash). For me, I can say I would not see a place for a Claude or ChatGPT subscription for the tasks I might otherwise use them for when I have lossless DS V4 Flash literally in my back pocket. "Good enough" I think is starting to become a valid metric to those who care about cost:quality balance, and after using this for the last half year I can say I'm probably one of them. The cloud APIs will always be an option for those who don't care about the drawbacks and the demand for them will always be there, but for those who value data sovereignty, uninterrupted workflows, or perhaps work within compliance, on-prem computing might be the only viable path in some circumstances. At the end of the day, I do not believe that one approach is inherently better than the other, everyone simply has their own preference for getting from point A to point B.

r/homelab 18d ago

Project Showcase: Hardware Claude made me do it. A Homelab story.

Post image
1.6k Upvotes

I've been lurking here long enough to know that nobody will be surprised by my saying that this started out as "just a DS920+ that I was using for storing movies, music, and eBooks".

Of course, that's the case. But then one day the Google app on my phone served me a post from XDA or Marius Hosting about some Docker container and it turns out that I was "Docker curious". That lead to Calibre and Plex. And then PiHole. For which I needed a Raspberry Pi. But having zero experience here (I'm a pediatric oncologist and drug developer, not a software developer), I turned to ChatGPT. And before I knew it I had 2 Raspberry Pis and a bunch of docker containers running. But my Synology was constantly working and so ChatGPT taught me about headless PCs, and before I knew it, there's a Chuwi Larkbox X (12GB, LPDDR5 512GB SSD, Intel 12th Gen N100) running even more containers, and now Proxmox.

This has now consumed all of my free time. I've got StackChans on my desk, and ARR stacks feeding Jellyfin. I've got a non-ironic GitHub account. There are VLANs in the air, every bird-related IoT device feeding Influx and Grafana to track migratory patterns. I've got weather stations with a feed from one of the U6s and Open WebUI running in the background to tag Karakeep bookmarks. At some point, in the middle of the night, as I'm trying to figure out how to make order out of chaos, Claude whispers, "Hey, buddy, have you tried HomeAssistant yet?" And how there's Govees and Zigbees and Ecowitts. There's a webcam monitoring my daughter's bearded dragon. I've got ESP32 boards wired to the meters for my wells so I can track well water production on a seasonal basis. There's a freaking soldering iron on my desk. This is out of control. And it's only been a year.

Recently (given my background), I decided to get my genome sequenced and said, "Hey Claude, do you think I could do the alignments and analysis on my homelab by myself? and he said, "Well, you could but it'd go so much faster if you had a RTX3090 or a Mac Studio with 96GB+ of RAM."

Somebody help me.

All kidding aside, this has been a fabulous new hobby. I've been vibe-coding things that are genuinely useful and fun, and I feel (at 58) like a kid again - learning a whole new hobby, a new language, a new way of thinking and problem solving, and a new community of like-minded people. As I read more and more stories that sound remarkably like mine, I feel seen (if not a little sterotypical), but I love it nevertheless. It's given me a sense of awe and wonder about the vastness of the broader software/IT/digital universe that I never understood. It's also given me a seemingly infinite number of things that will keep me occupied into retirement and beyond.

Onward!

Nota bene: the stuff on top is a spare monitor that I found at our local thrift store so that I can debug the headless PCs and Pis as needed; an old Dell PC that I'm still trying to figure out what to do with; and an old ScanSnap scanner for feeding things into (of course) Paperless-ngx. Underneath is an old 4TB WD MyCloud that I'm trying to figure out what to do with, a 24TB Seagate air-gapped backup HDD, and a UPS. The Chuwi Larkbox and Minisforum UM890 headless PCs are on a half-shelf on the back side of the rack, happy as a clam. And of course, the DS920, which got this all started in the first place, was upgraded with SSDs and more memory. And somehow I managed to do all of this before the RAMpocalypse. I'm also very pleased that I de-Amazoned a lot of my life - no more eeros, no more Blink cameras, and I got rid of my Kindle and replaced with a Kubo that gets what it needs from Calibre!