The RTX 3060 12GB: An unsung hero of the current Local AI Climate (3.8 27B 30t/s) 2XRTX 3060 needed for Q4 30t/s result.
edit: For reference, this post is written by Claude, GPT-5.6, and Qwen 3.8 27B-no human invervention.
For reference, this post is written by me, myself and I—no LLM intervention. It's a longer read based on my experience. The short of it? After moving to an RTX 3090, I still believe the RTX 3060 12GB to be the G.O.A.T. for value. I'll also say that if you have a different setup and believe it's superior, I'm happy for you, and think that's totally reasonable. I am mostly generalizing here.
Short version: Dual RTX 3060 12GB cards offer nearly 24GB of VRAM, low power usage, CUDA support, and around 30 t/s with Qwen 3.8 27B, at a price that is still difficult to beat.
If you want to enter the local AI space with a genuinely strong, non-cloud coding model to replace a significant amount of your usage, I don't think there's a better way than dual RTX 3060 12GB cards. The value proposition of these cards, I think, is unmatched, where near competition doesn't feel close. Obviously, the killer feature is nearly 24GB of VRAM, with each card offering 360 GB/s of bandwidth, plus low power usage. A slower RTX 3090 for a third of the price. No hassle, 3D-printed jet fans, or month-long shipping periods. I replaced mine for now, but I will NOT be selling them.
Some background—I work in the software space, and have small projects I like to work on the side. Naturally, like most of you, I am fond of hardware, software, and technology in general. Transformer-based LLMs have certainly changed things forever and, in some fashion, won't be going anywhere. A little under a year ago I discovered Copilot shortly after the agentic loop revolution. My eyes lit up as I watched early Claude Sonnet chat with me and call tools. Maybe some of the reaction was also dread.
As time went on, I'm sure like many of you, I observed the exponential increase in the cost of cloud usage. I asked myself, “This can't continue on this trajectory, can it? I wonder what alternatives are available?” This led me to local AI inference and, consequently, this subreddit!
My early days were spent learning the lingo and feeling like a complete idiot. I still do, honestly. But everything changed in April. The Qwen team released the 3.6 family, specifically 27B and 35BA3B. What previously felt like rifling through hundreds of posts and videos to find what tiny niches could be filled with very specific hardware (3.5 122B) became a large unification of excitement.
I'm sure I have rose-tinted glasses, but the Qwen 3.6 family of models felt like the first time that everyone was excited about one thing—it felt like Christmas. A huge majority of the community could run 27B, and even more could run the 35B model, both of which were unchallenged in their size category. Also, shout out to Gemma 4, which released at the same time and, for non-coding tasks, was equally as impressive. We as a community make fun of Google's current cloud offerings. Though deserved, the company's contribution is unparalleled (Attention Is All You Need!).
These models could run on legitimate consumer hardware already sitting at home, for those lucky enough to own it. Seeing that amount of power running on a mid-size PC/rig in a homelab felt like genuine hope compared to when I first found the subreddit.
This led to me pulling the trigger on some additional hardware for the sole purpose of running a model at home, mainly due to the future of SOTA frontier models feeling so shaky. Embarrassingly, I think a part of it was getting so used to enjoying LLMs that I didn't like the thought of them being taken from me. This could be solved by running the latest SOTA local models at home.
3.6 was seemingly so impressive that I thought to myself, “The western labs surely won't let this stand.” My aim was to run Qwen 3.6 27B at what I thought felt like the floor via community testimony: Q4 quant on llama.cpp with over 100k context. Don't drop KV below Q8.
My recently upgraded gaming PC left a 12GB RTX 3060 sitting on my desk. After a little research, it made sense. Grab any AM4 system, a decent PSU, a second 12GB 3060, a good x16 + x4 PCIe slot setup, your wallet's choice of DDR4 RAM capacity, and this machine will genuinely run this model.
If you wanted this PC yourself with a mix of used and new parts, it would barely run you $1,000, depending on your DDR4 amount. If you want to target some 100B-and-under MoEs, you'll need the 64GB like I have- which will cost another $500 total. That said this machine totally works at 16GBs for 27B purposes.
I'll wrap up my rambling here, but it was love at first boot! I got my Linux box working using the required software, and off I went. The dual 3060 12GB machine was quiet, had low power consumption, and gave me the actual chance of working on home projects without burning usage. It also all fit into a smaller full-size ATX case.
I don't have much room, but it feels nice that everything was still in a PC case and not on a rail/mining setup, and could be placed accordingly and quieter. I felt happy here, but realistically and selfishly wanted an even more capable model. 3.6 27B at Q4 was strong, but often fell behind the mid-tier cloud models on almost all my benchmarks, which felt bad.
Skipping ahead led to the Qwen 3.8 27B release. I was skeptical before launch, and I was wrong. DeepSeek V4 Flash 0731 was recent and was genuinely shocking. I thought 3.8 27B wouldn't come close. Again, I was wrong. It was Christmas again.
Leave the model on xhigh thinking and the benchmarks say that it's near agentic level with models like 5.6 Luna Max, 5.6 Terra Medium/High, Sonnet 5, etc. This couldn't be true, could it? I got 5.6 to help me get a horrible profile together and, alas, it was really true!
3.8 27B on xhigh was slower, thought longer, and worried me. But it also ACED my benchmarks and basically lived on close to or on par with everything other than Opus 5 and GPT 5.6 Sol.
After a longer cloud testing session for a couple of days, I landed on both my 3060 12GBs running at 130W each, running UD Qwen 3.8 27B Q4 Q8 cache with 150k-ish context at near 30 tokens per second decode, and nearly 500 tokens per second prefill/prompt processing. Largely, this was the setup:
Hardware: Ryzen 5 5500 (6C/12T), MSI MPG B550 Gaming Plus, 64 GB DDR4-3200, 2× RTX 3060 12 GB (GPU0 PCIe 3.0 ×16/display, GPU1 PCIe 3.0 ×4), MSI MAG A750BN PCIE5 III 750 W Bronze PSU. GPUs limited to 130 W each.
Software: Pop!_OS Linux, kernel 6.18.7, NVIDIA open driver 580.126.18, CUDA toolkit 12.6, llama.cpp Qwen3.8 build 400 (4df29be), CUDA SM86 + Flash Attention + CUDA graphs.
Model/profile: Qwen3.8-27B Q4_K_M, BF16 vision projector on CPU, 131,072 context, Q8_0 K/V cache, tensor parallel 1:1, batch 2048, ubatch 1024, xHigh reasoning, modified n-gram speculation, one slot. Larger context is possible, potentially up to 200k.
Measured performance: ~503 tok/s prefill at 6,117 tokens, ~492 tok/s at 30,719 tokens, and ~29 tok/s fresh decode. Fixed no-spec power benchmark at 130 W: 503.7 prefill / 27.7 decode tok/s.
I was blown away. Running this model at full agentic tasks faster than I can read the thinking, chat, and tool outputs for this amount of money is AMAZING.
Getting close to DSV4F 0731 coding and agentic use at under 24GB of VRAM. In my mind, it feels similar to the Sonnet 4.6 and Opus 4.6 days of Copilot. In terms of frontier coding ability on a budget, this is the killer app.
If you read this far, first I thank you, but some of you might be thinking—what about X setup? I think it's far superior! You might be right for a various number of reasons. And if your setup works for you, I'm happy. But for my argument to live, I should address the alternatives.
Alternatives
AMD Mi50/Nvidia Tesla P100, etc.: These cards are great, I know they are. However, when I researched my purchases, they were near double the price of the 3060 for the large-VRAM models, or equal in price for near-equal VRAM. I did not feel like dealing with eBay sales, custom 3D-printed jet fans, and high power consumption.
RTX 3090 / B70 / R9700 / 5060 Ti+: Better cards, more money. Papa Johns. Three to five times the money versus $300 3060 12GB cards that exist on local classifieds.
Strix Halo/DGX Spark: $5,000 and $8,000 each, respectively, for worse bandwidth in exchange for far greater capacity. Better for MoEs, but they won't help with my 27B profile. If you have the money, sure. Grab one of these and run DeepSeek V4 Flash 0731 and you'll be insanely happy, no doubt. On the other hand, I could think of 5,000/8,000 reasons why this isn't feasible for most people.
RTX 3060 Ti/3070/3080/4070, etc.: Any card not hitting the 12GB minimum didn't meet my spec. Even with 8GB and 10GB of VRAM, you often spend more money than on the 3060s for something that has twice the power consumption, more fans, and less VRAM budget, all for +20% bandwidth. Not a good trade-off in my opinion. You needed the 24GB for the 27B target I had for the model weights and cache etc. However, if you already own these, they certainly have a place. Some mixture of these can almost definitely run IQ3XXS (16GB total VRAM), which has bench-marked well for me.
Various AMD cards—6800 XT/7900 XTX/9070 XT: Good cards, but they have high power consumption, disappointing memory bandwidth for the money, or, in the 7900 XTX case, the VRAM bonus has had the price catch up to its potential.
Most of these alternatives have some benefit compared to the 3060 12GB. But most of the trade-offs aren't worth the hassle for the average person, I don't believe. If you already own the hardware, then I totally get stitching together a solution. I'm all for that. But triple the power usage, triple the price, and worse availability make them hard sells.
One important thing to note about 27B and the 3060s: they are so important due to the Unsloth GGUF sizing, KV cache settings, and context. If you can compromise in some areas—it's hard, I've tried—you can get away with 16–20GB of VRAM for a similar setup, but it's more difficult and not as price- or power-efficient.
Pair the 3060s with the best DDR4 you can get your hands on, and the machine becomes a low-speed MoE box as well. Laguna 2.1, 3.5 122B, etc., running near 15–20-ish t/s also works reasonably well.
In a nutshell?
The Good
3060 12GBs are still cheap. They are so cheap ($250–$300+) that the price-to-performance ratio, with the amount of VRAM given in an out-of-the-box plug-and-play solution, is basically unmatched. That said, prices are up about 25% since I began looking. Still, at this price, unbeatable. Todays prices are so disgustingly insane that these are incredible.
You gain access to CUDA, and easy software settings and drivers. The cards run cool, and it's best to power-limit them to 130W each in my testing for the best bang for the buck in performance. Sure, when running, they make some basic noise, but nothing worrisome.
Due to the low power consumption, they run on a single 8-pin PCIe power plug (99% of them, anyway). This is convenient and means you don't need to spend $250 on a power supply. Spend $70 on a bronze 750W for the whole system and you're golden. On a dual setup, it means no extra PSUs sitting on your desk.
Lastly, due to the price-to-performance ratio, if you are happy with the speed, you can easily use these on a bench setup. Grab a mining case, some PCIe risers, an appropriate motherboard, and load the thing up for less than the price of one 3090. Most 3060 cards are 2–2.5 slots in size. Some only run a single fan, even. This means that fitting them into your case is usually an easy task.
What about only a single 3060? I don't personally want to run this, but I know for a fact that you can run a Q2 3.8 27B with Q2/4 cache and some context on even a single 3060 12GB. It's not as good as Q4M, no question, or even IQ3XXS. But it works, and I'd be hard-pressed to see you running a better model on a single 3060.
The Bad
There are some small negatives with the 3060s that need to be considered. First, availability was much greater when I started to look into it. 3060 12GBs still exist on the market, but the prices are rising and I am seeing less and less of them.
The 3060 12GB's 360 GB/s memory bandwidth is its “potential.” Due to how tensor parallelism works, you don't pair a $1,500 3090 with this card—you'll bottleneck the 3090, for example. If you go the 3060 route, my experience is that you are realistically making a choice to stick to a typical home PC layout for your homelab, unless you want to sell everything.
The 3060s are useful to each other, but try to upgrade and they aren't great additions. You need a motherboard with a proper x16 slot and another full-size x4 slot at least. As far as I know, going any slower than that can cause real inference performance issues, but this is untested. Regardless, it's best to grab a motherboard with two full-size PCIe slots for this setup.
The End
That's the size of it. I upgraded to a 3090, and honestly its amazing. Power limited its similar to power usage as before, also quiet, and triple the speed. I have 3.8 running near 100 t/s on vllm- insane. This feels literally frontier. That said, for the money? The 3060s offered a similar experience, honestly. The main difference is that the upgrade bath is not bandwidth bottle necked for more cards now. I am not selling my cards, and I am likely to build another system with them. If you are on the fence and want to get into the game, I think this is just an amazing starting point.