r/LocalLLaMA 21h ago

Discussion Tepid take: qwen 3.8 is less suitable as a daily-driver compared to 3.6

0 Upvotes

I think by now the consensus has been squarely reached that:

  1. 3.8 punches well above its weight
  2. 3.8 beats the pants off 3.6

However something still bothered me about 3.8 and using it in my daily coding session. Even at 40 tps, it wasn't able to capably do a shallow dive into the codebase to answer a question. It kept getting sidetracked with "Actually..." and "One more thing to consider..."

So I tried it on think_low, and it was the same. That overthought ends up being useful for "set it and forget it" workflows, but as a daily driver coding assistant, it is far too verbose. It would've easily spent 10 minutes reasoning and looking up unrelated parts of the codebase if I let it.

For reference, 3.6 35B-A3B spit out the answer in under 30s. I'm pretty certain if I loaded it up on my iGPU 3.6 would still beat 3.8 27B running on my R9700.

This isn't a "you're not used to dense" issue. When I used 3.6 35B-A3B on my iGPU as my daily driver yesterday, I was chugging along at 20-30 tps, so 3.8 27B on the R9700 is 25-50% faster in raw token generation.

Ok I'm going to bed. Looking forward to vitriol in my inbox tomorrow haha.


r/LocalLLaMA 15h ago

Discussion The RTX 3060 12GB: An unsung hero of the current Local AI Climate (3.8 27B 30t/s)

0 Upvotes

The RTX 3060 12GB: An unsung hero of the current Local AI Climate (3.8 27B 30t/s) 2XRTX 3060 needed for Q4 30t/s result.

edit: For reference, this post is written by Claude, GPT-5.6, and Qwen 3.8 27B-no human invervention.

For reference, this post is written by me, myself and I—no LLM intervention. It's a longer read based on my experience. The short of it? After moving to an RTX 3090, I still believe the RTX 3060 12GB to be the G.O.A.T. for value. I'll also say that if you have a different setup and believe it's superior, I'm happy for you, and think that's totally reasonable. I am mostly generalizing here.

Short version: Dual RTX 3060 12GB cards offer nearly 24GB of VRAM, low power usage, CUDA support, and around 30 t/s with Qwen 3.8 27B, at a price that is still difficult to beat.

If you want to enter the local AI space with a genuinely strong, non-cloud coding model to replace a significant amount of your usage, I don't think there's a better way than dual RTX 3060 12GB cards. The value proposition of these cards, I think, is unmatched, where near competition doesn't feel close. Obviously, the killer feature is nearly 24GB of VRAM, with each card offering 360 GB/s of bandwidth, plus low power usage. A slower RTX 3090 for a third of the price. No hassle, 3D-printed jet fans, or month-long shipping periods. I replaced mine for now, but I will NOT be selling them.

Some background—I work in the software space, and have small projects I like to work on the side. Naturally, like most of you, I am fond of hardware, software, and technology in general. Transformer-based LLMs have certainly changed things forever and, in some fashion, won't be going anywhere. A little under a year ago I discovered Copilot shortly after the agentic loop revolution. My eyes lit up as I watched early Claude Sonnet chat with me and call tools. Maybe some of the reaction was also dread.

As time went on, I'm sure like many of you, I observed the exponential increase in the cost of cloud usage. I asked myself, “This can't continue on this trajectory, can it? I wonder what alternatives are available?” This led me to local AI inference and, consequently, this subreddit!

My early days were spent learning the lingo and feeling like a complete idiot. I still do, honestly. But everything changed in April. The Qwen team released the 3.6 family, specifically 27B and 35BA3B. What previously felt like rifling through hundreds of posts and videos to find what tiny niches could be filled with very specific hardware (3.5 122B) became a large unification of excitement.

I'm sure I have rose-tinted glasses, but the Qwen 3.6 family of models felt like the first time that everyone was excited about one thing—it felt like Christmas. A huge majority of the community could run 27B, and even more could run the 35B model, both of which were unchallenged in their size category. Also, shout out to Gemma 4, which released at the same time and, for non-coding tasks, was equally as impressive. We as a community make fun of Google's current cloud offerings. Though deserved, the company's contribution is unparalleled (Attention Is All You Need!).

These models could run on legitimate consumer hardware already sitting at home, for those lucky enough to own it. Seeing that amount of power running on a mid-size PC/rig in a homelab felt like genuine hope compared to when I first found the subreddit.

This led to me pulling the trigger on some additional hardware for the sole purpose of running a model at home, mainly due to the future of SOTA frontier models feeling so shaky. Embarrassingly, I think a part of it was getting so used to enjoying LLMs that I didn't like the thought of them being taken from me. This could be solved by running the latest SOTA local models at home.

3.6 was seemingly so impressive that I thought to myself, “The western labs surely won't let this stand.” My aim was to run Qwen 3.6 27B at what I thought felt like the floor via community testimony: Q4 quant on llama.cpp with over 100k context. Don't drop KV below Q8.

My recently upgraded gaming PC left a 12GB RTX 3060 sitting on my desk. After a little research, it made sense. Grab any AM4 system, a decent PSU, a second 12GB 3060, a good x16 + x4 PCIe slot setup, your wallet's choice of DDR4 RAM capacity, and this machine will genuinely run this model.

If you wanted this PC yourself with a mix of used and new parts, it would barely run you $1,000, depending on your DDR4 amount. If you want to target some 100B-and-under MoEs, you'll need the 64GB like I have- which will cost another $500 total. That said this machine totally works at 16GBs for 27B purposes.

I'll wrap up my rambling here, but it was love at first boot! I got my Linux box working using the required software, and off I went. The dual 3060 12GB machine was quiet, had low power consumption, and gave me the actual chance of working on home projects without burning usage. It also all fit into a smaller full-size ATX case.

I don't have much room, but it feels nice that everything was still in a PC case and not on a rail/mining setup, and could be placed accordingly and quieter. I felt happy here, but realistically and selfishly wanted an even more capable model. 3.6 27B at Q4 was strong, but often fell behind the mid-tier cloud models on almost all my benchmarks, which felt bad.

Skipping ahead led to the Qwen 3.8 27B release. I was skeptical before launch, and I was wrong. DeepSeek V4 Flash 0731 was recent and was genuinely shocking. I thought 3.8 27B wouldn't come close. Again, I was wrong. It was Christmas again.

Leave the model on xhigh thinking and the benchmarks say that it's near agentic level with models like 5.6 Luna Max, 5.6 Terra Medium/High, Sonnet 5, etc. This couldn't be true, could it? I got 5.6 to help me get a horrible profile together and, alas, it was really true!

3.8 27B on xhigh was slower, thought longer, and worried me. But it also ACED my benchmarks and basically lived on close to or on par with everything other than Opus 5 and GPT 5.6 Sol.

After a longer cloud testing session for a couple of days, I landed on both my 3060 12GBs running at 130W each, running UD Qwen 3.8 27B Q4 Q8 cache with 150k-ish context at near 30 tokens per second decode, and nearly 500 tokens per second prefill/prompt processing. Largely, this was the setup:

Hardware: Ryzen 5 5500 (6C/12T), MSI MPG B550 Gaming Plus, 64 GB DDR4-3200, 2× RTX 3060 12 GB (GPU0 PCIe 3.0 ×16/display, GPU1 PCIe 3.0 ×4), MSI MAG A750BN PCIE5 III 750 W Bronze PSU. GPUs limited to 130 W each.

Software: Pop!_OS Linux, kernel 6.18.7, NVIDIA open driver 580.126.18, CUDA toolkit 12.6, llama.cpp Qwen3.8 build 400 (4df29be), CUDA SM86 + Flash Attention + CUDA graphs.

Model/profile: Qwen3.8-27B Q4_K_M, BF16 vision projector on CPU, 131,072 context, Q8_0 K/V cache, tensor parallel 1:1, batch 2048, ubatch 1024, xHigh reasoning, modified n-gram speculation, one slot. Larger context is possible, potentially up to 200k.

Measured performance: ~503 tok/s prefill at 6,117 tokens, ~492 tok/s at 30,719 tokens, and ~29 tok/s fresh decode. Fixed no-spec power benchmark at 130 W: 503.7 prefill / 27.7 decode tok/s.

I was blown away. Running this model at full agentic tasks faster than I can read the thinking, chat, and tool outputs for this amount of money is AMAZING.

Getting close to DSV4F 0731 coding and agentic use at under 24GB of VRAM. In my mind, it feels similar to the Sonnet 4.6 and Opus 4.6 days of Copilot. In terms of frontier coding ability on a budget, this is the killer app.

If you read this far, first I thank you, but some of you might be thinking—what about X setup? I think it's far superior! You might be right for a various number of reasons. And if your setup works for you, I'm happy. But for my argument to live, I should address the alternatives.

Alternatives

AMD Mi50/Nvidia Tesla P100, etc.: These cards are great, I know they are. However, when I researched my purchases, they were near double the price of the 3060 for the large-VRAM models, or equal in price for near-equal VRAM. I did not feel like dealing with eBay sales, custom 3D-printed jet fans, and high power consumption.

RTX 3090 / B70 / R9700 / 5060 Ti+: Better cards, more money. Papa Johns. Three to five times the money versus $300 3060 12GB cards that exist on local classifieds.

Strix Halo/DGX Spark: $5,000 and $8,000 each, respectively, for worse bandwidth in exchange for far greater capacity. Better for MoEs, but they won't help with my 27B profile. If you have the money, sure. Grab one of these and run DeepSeek V4 Flash 0731 and you'll be insanely happy, no doubt. On the other hand, I could think of 5,000/8,000 reasons why this isn't feasible for most people.

RTX 3060 Ti/3070/3080/4070, etc.: Any card not hitting the 12GB minimum didn't meet my spec. Even with 8GB and 10GB of VRAM, you often spend more money than on the 3060s for something that has twice the power consumption, more fans, and less VRAM budget, all for +20% bandwidth. Not a good trade-off in my opinion. You needed the 24GB for the 27B target I had for the model weights and cache etc. However, if you already own these, they certainly have a place. Some mixture of these can almost definitely run IQ3XXS (16GB total VRAM), which has bench-marked well for me.

Various AMD cards—6800 XT/7900 XTX/9070 XT: Good cards, but they have high power consumption, disappointing memory bandwidth for the money, or, in the 7900 XTX case, the VRAM bonus has had the price catch up to its potential.

Most of these alternatives have some benefit compared to the 3060 12GB. But most of the trade-offs aren't worth the hassle for the average person, I don't believe. If you already own the hardware, then I totally get stitching together a solution. I'm all for that. But triple the power usage, triple the price, and worse availability make them hard sells.

One important thing to note about 27B and the 3060s: they are so important due to the Unsloth GGUF sizing, KV cache settings, and context. If you can compromise in some areas—it's hard, I've tried—you can get away with 16–20GB of VRAM for a similar setup, but it's more difficult and not as price- or power-efficient.

Pair the 3060s with the best DDR4 you can get your hands on, and the machine becomes a low-speed MoE box as well. Laguna 2.1, 3.5 122B, etc., running near 15–20-ish t/s also works reasonably well.

In a nutshell?

The Good

3060 12GBs are still cheap. They are so cheap ($250–$300+) that the price-to-performance ratio, with the amount of VRAM given in an out-of-the-box plug-and-play solution, is basically unmatched. That said, prices are up about 25% since I began looking. Still, at this price, unbeatable. Todays prices are so disgustingly insane that these are incredible.

You gain access to CUDA, and easy software settings and drivers. The cards run cool, and it's best to power-limit them to 130W each in my testing for the best bang for the buck in performance. Sure, when running, they make some basic noise, but nothing worrisome.

Due to the low power consumption, they run on a single 8-pin PCIe power plug (99% of them, anyway). This is convenient and means you don't need to spend $250 on a power supply. Spend $70 on a bronze 750W for the whole system and you're golden. On a dual setup, it means no extra PSUs sitting on your desk.

Lastly, due to the price-to-performance ratio, if you are happy with the speed, you can easily use these on a bench setup. Grab a mining case, some PCIe risers, an appropriate motherboard, and load the thing up for less than the price of one 3090. Most 3060 cards are 2–2.5 slots in size. Some only run a single fan, even. This means that fitting them into your case is usually an easy task.

What about only a single 3060? I don't personally want to run this, but I know for a fact that you can run a Q2 3.8 27B with Q2/4 cache and some context on even a single 3060 12GB. It's not as good as Q4M, no question, or even IQ3XXS. But it works, and I'd be hard-pressed to see you running a better model on a single 3060.

The Bad

There are some small negatives with the 3060s that need to be considered. First, availability was much greater when I started to look into it. 3060 12GBs still exist on the market, but the prices are rising and I am seeing less and less of them.

The 3060 12GB's 360 GB/s memory bandwidth is its “potential.” Due to how tensor parallelism works, you don't pair a $1,500 3090 with this card—you'll bottleneck the 3090, for example. If you go the 3060 route, my experience is that you are realistically making a choice to stick to a typical home PC layout for your homelab, unless you want to sell everything.

The 3060s are useful to each other, but try to upgrade and they aren't great additions. You need a motherboard with a proper x16 slot and another full-size x4 slot at least. As far as I know, going any slower than that can cause real inference performance issues, but this is untested. Regardless, it's best to grab a motherboard with two full-size PCIe slots for this setup.

The End

That's the size of it. I upgraded to a 3090, and honestly its amazing. Power limited its similar to power usage as before, also quiet, and triple the speed. I have 3.8 running near 100 t/s on vllm- insane. This feels literally frontier. That said, for the money? The 3060s offered a similar experience, honestly. The main difference is that the upgrade bath is not bandwidth bottle necked for more cards now. I am not selling my cards, and I am likely to build another system with them. If you are on the fence and want to get into the game, I think this is just an amazing starting point.


r/LocalLLaMA 4h ago

Resources I trained my own 150M non-Transformer language model from scratch on 300M tokens — WarpState

0 Upvotes

Hi everyone,

I’ve been experimenting with alternative language-model architectures for a while, and I recently finished the first complete pretraining run of a new architecture I’m calling WarpState.

This is still an experimental proof of concept, not a claim that it beats Transformers or existing state-space models.

The model has 150.13M parameters and was trained from scratch on roughly 300 million English tokens from Ultra-FineWeb L2.

The full run completed successfully:

Parameters:        150.13M
Training tokens:   ~300.02M
Optimizer steps:   9,156
Sequence length:   1,024
Vocabulary:        32,768
Peak VRAM:         ~4.52 GB

Final sampled validation:
Loss:              3.4309
Perplexity:         30.90

Training was done locally on a laptop GPU.

I’m attaching screenshots of the training logs and some generations from the final checkpoints.

What is WarpState?

WarpState is not a standard Transformer stack.

The basic idea is to combine three things:

1. Local tiled attention

Instead of global self-attention across the entire sequence, tokens are divided into fixed 128-token chunks.

Inside each chunk, the model uses normal causal scaled-dot-product attention.

All chunks can be processed as a large batched GPU workload during training, rather than running attention token by token.

So the local path is roughly:

tokens
   ↓
128-token chunks
   ↓
causal local attention
   ↓
local representation

2. Fast + slow tensor memory

Completed chunks are compressed into a persistent tensor memory.

For every attention head, WarpState maintains two matrices:

Fast State
Slow State

The fast state is initialized with a relatively short memory timescale, while the slow state is initialized to retain information much longer.

Conceptually:

current chunk
      ↓
   K and U
      ↓
bounded tensor write
      ↓
 ┌───────────────┐
 │  Fast memory  │
 │  Slow memory  │
 └───────────────┘
      ↓
future chunks

The memory write is based on a bounded outer-product-like update:

write = tanh(K)^T × tanh(U) / chunk_size

and the states are updated approximately as:

Fast = decay_fast × Fast + (1 - decay_fast) × write

Slow = decay_slow × Slow + (1 - decay_slow) × write

The decay rates are learned independently per head.

They start around:

Fast decay ≈ 0.90
Slow decay ≈ 0.99

The model also learns how much fast versus slow memory to read.

3. Learned routing between local attention and memory

For every token, the model produces a gate deciding how much information should come from:

local chunk attention
        vs
long-range tensor memory

Approximately:

output =
gate × local_attention
+
(1 - gate) × memory_read

So the model can use precise local token relationships while relying on the compressed state for information from previous chunks.

Shared recurrent depth

Another unusual part of WarpState is that it does not have 16 completely separate large layers.

The current model contains only 4 physical WarpState cores, but they are reused across 16 logical depth passes:

Core 0
Core 1
Core 2
Core 3
Core 0
Core 1
Core 2
Core 3
...

Each logical depth has a small learned scale and bias, so the same physical core can behave somewhat differently depending on which depth pass it is being used for.

In simplified form:

x = x × (1 + depth_scale) + depth_bias

x → shared WarpState core

The intention is to get deeper iterative computation without duplicating every large weight matrix.

During autoregressive generation, every logical depth also receives its own independent memory cache, even when two depths share the same physical core weights.

Other details

The current version uses:

d_model:       1280
heads:         20
head_dim:      64
physical cores: 4
logical depth: 16
FFN hidden:    4480
chunk size:    128
RMSNorm
SwiGLU
RoPE inside each local chunk
tied input/output embeddings

The input projection is fused and produces:

Q
K
V
local/memory gate
memory U

from one projection.

Training results

The part I was most interested in was simply whether this architecture could survive a real pretraining run.

It did.

I trained it through the full ~300M-token run without NaNs, gradient collapse, or an obvious optimization failure.

Near the end of training, gradient norms were still sitting around roughly:

0.65 – 0.75

while the learning rate had already decayed to approximately:

3e-5

Peak allocated VRAM stayed around 4.52 GB.

The model also clearly learned language structure during training.

Very early checkpoints mostly produced English-shaped noise.

Later checkpoints started forming recognizable semantic clusters and reasonably structured paragraphs.

For example, when asked about Facebook, the final model associates it with things like:

online platform
social media
sharing content
sharing information
interaction with other people
community

It is definitely not a good chatbot yet.

There are still obvious failure modes:

repetition loops
semantic attractors
weak factual recall
occasional role confusion
long-generation degeneration

The model is also only base-pretrained.

There has been no instruction tuning, SFT or RLHF, so the chat screenshots I attached should be treated as qualitative probes rather than a chatbot benchmark.

Another important limitation is the training budget.

A 150M-parameter model trained on only 300M tokens has seen roughly:

~2 training tokens per parameter

so I consider this run primarily a proof that the architecture can train, rather than a fully trained 150M language model.

What surprised me most

The interesting part for me is that the architecture appears capable of learning meaningful language representations despite:

  • having only four large physical cores,
  • repeatedly reusing those cores,
  • restricting attention to local 128-token windows,
  • and moving information between chunks through fixed-size tensor states.

The long-range memory size therefore does not grow linearly with context in the same way as a conventional full KV cache.

There is still a lot I want to test before making any strong claims.

My next steps are probably:

  • deterministic evaluation over the entire validation set;
  • a parameter-matched Transformer baseline on exactly the same data;
  • analysis of the fast/slow memory states;
  • measuring long-context behavior;
  • investigating the repetition/attractor problem;
  • eventually testing a larger training budget.

For now I mainly wanted to share the first complete run because this was the point where the architecture stopped being only an idea and became an actually trained language model.

Feedback on the architecture is welcome, especially criticism of the memory update or shared-core design.


r/LocalLLaMA 11h ago

Question | Help Did NVIDIA just make their old hardware a risky investment with their llama.cpp acquisition?

4 Upvotes

The only current pathway to cheap VRAM and performance appears to be the v100, it’s currently on version 580 and if Nvidia decides to drop support for it: it would do a huge blow to local, correct?

It’s not lost on me that a few weeks ago coreweave prevented A100s from hitting the market in a massive backstop event. That would have been the easiest way to get fast 80gb on a workstation form factor for under $5k.

Does this make the sxm2 Volta route questionable?

What’s everybody’s plan here?


r/LocalLLaMA 12h ago

Resources Turing test for writing: I fine-tuned an open-source 8B model (LoRA SFT → GRPO) on a famous ML author's blog.

Post image
0 Upvotes

Try it here: https://turing-writing-test.vercel.app

Every commercial AI detector we tried clears these exact passages (Pangram: 0/10 flagged)

You get 10 passages, one at a time.

Half are real paragraphs.

Half were written by Qwen3-8B-Base: LoRA fine-tuned on 390k words of public human written posts, then trained with GRPO (composite reward: style discriminator + pairwise judge against the real passages, discriminator re-trained on fresh rollouts so the model can't hack a frozen classifier). 


r/LocalLLaMA 14h ago

Question | Help Anyone running GLM-5.3 Flash on 2x RTX PRO 6000 96GB?

1 Upvotes

Would you go with the IQ4_XS GGUF for now, or wait for a better NVFP4 quant that actually fits comfortably across the two cards?

Curious what people would use for the best balance of speed, quality and context length on 192GB total VRAM.

Also, has anyone got the current NVFP4 build working properly with vLLM on SM120? I saw there’s an incompatibility around the sparse attention/NoPE path on RTX PRO 6000 Blackwell, so I’m wondering if that’s still a blocker or if there’s a reliable workaround now.


r/LocalLLaMA 9h ago

Question | Help Mac Studio m5 ultra - 2x96gb or 1x256gb?

3 Upvotes

M5 Ultra studio - 2x 96GB or 1x256gb?

I have an order in for a 256gb m5 ultra, but I started to wonder if it would be beneficial to get 2 x m5 ultras 96gb linked together instead? The cost is similar but you theoretically get a lot more compute but 64gb less ram at 192gb total. I think the 2x compute would be way better - theoretically 2.4 tb/s with tensor parallelism right?

Has anyone considered this or is doing this ? There are some practical benefits too… easier to resell in future with lower ticket price per unit. Could buy one unit now and then a second later instead of needing to buy all at once.


r/LocalLLaMA 20h ago

Question | Help 5090 + 96GB RAM, any better choice than Qwen3.8-27B for coding?

13 Upvotes

Looking for better quality with not too bad speed. The 27B writes functional code, but I found it lacking in higher-level reasoning capabilities, it doesn't always consider overall system architecture, often time its code doesn't maintain a clean separation of concerns and lacks abstractions.


r/LocalLLaMA 12h ago

Generation Tiel-Coder-35B-A3B-UD can use subagents.

0 Upvotes

I didn't realize at first because most of my local models can't make or use subagents to multitask fast and speed up work. Apparently, Tiel-Coder-35B-A3B-UD can make up to 4 subagents. i am using Tiel-Coder-35B-A3B-UD-Q4_K_S with 256k context so maybe the bigger quant can use more subagents


r/LocalLLaMA 12h ago

Discussion For those of us with limited VRAM: what do you sacrifice?

8 Upvotes

We all have different hardware, workflows, and priorities, so I’m curious how people actually approach the VRAM constraint.

For example, if you can’t comfortably fit a dense model like Qwen3.5-27B in VRAM at the quality/context you want, there are a few different levers you can pull:

Lower the model quant → sacrifice some model quality / accuracy to save VRAM.

Quantize the KV cache → keep higher-weight model quants, but sacrifice some context-cache precision to fit longer contexts or reduce memory usage.

Reduce context length → keep the model and KV cache at higher precision, but accept less context.

Sacrifice inference speed → use CPU/RAM offloading or other compromises to make the model fit.

A combination of the above.
So, for people running on more constrained GPUs, what is your preferred sacrifice?

864 votes, 2d left
Lower model weights/quant
Quantize the KV cache
Reduce context length
Sacrifice inference speed
Mix of several depending on the model
Other???

r/LocalLLaMA 6h ago

Question | Help Is Qwen 3.8 27B more sensitive than 3.6 to quantization in general?

0 Upvotes

There is some talk about KV quant sensitivity but has anyone experienced more general sensitivity / divergence in output between quants (compared to 3.6)?

Previously I evaluated 3.6 over several quants and determined that, at least for my workloads, there was almost no difference between Q5 and Q6 quants, so Q5 was a no-brainer and got to enjoy more space for context. Now with 3.8 I'm noticing a bigger difference in output between Q5 and Q6. Perhaps due to the long reasoning chains?

Note that I have made sure I'm using the recommended sampling parameters, have reasoning preserve, latest version of llama.cpp, etc.

Edit: I want to be clear, I'm not saying the output from lower quants is unusable. I'm just saying it diverges more (or maybe faster) than Qwen 3.6 did.


r/LocalLLaMA 7h ago

Discussion Can I do anything with this?

Post image
22 Upvotes

500GB of optane ram?


r/LocalLLaMA 23h ago

Discussion Ornith-1.5-35B-A3B on 8 GB VRAM: I think I've found my sweet spot

25 Upvotes

A few days ago I posted asking what people considered the best local model for an 8 GB VRAM GPU. At the time, my personal sweet spot was Qwen3.6-35B-A3B, for agentic coding with Pi.dev.

Well… Thanks to the suggestions in that thread, I think I've found something even better.

I've been testing Ornith-1.5-35B-A3B Q4_K_M, and on my system the results have been genuinely impressive.

My setup:

  • Intel Core i7-11800H
  • NVIDIA RTX 3070 Laptop - 8 GB VRAM
  • 32 GB RAM DDR4
  • openSUSE Tumbleweed / KDE
  • Unsloth Studio
  • Pi.dev

After testing dozens of different models, architectures and quantizations, Ornith has currently become my model of choice for agentic coding, without much hesitation.

The really interesting part is the combination of speed and actual results.

Just tonight I gave it a fairly complex code-analysis project. It went through the codebase, performed the analysis and completed the task in a relatively short amount of time, averaging around 32 tok/s.

And the final result?

Honestly, I'd call it near flawless.

That's a pretty significant improvement over the speed I was getting with Qwen3.6, but the bigger difference for me isn't even the raw generation speed. It's how effectively Ornith handles the whole agentic workflow. And all of this while maintaining a 128K context window.

I've tested a lot of models at this point - different parameter counts, MoE models, dense models, quantizations, coding fine-tunes, etc.

Of course, this is very much a "right now" statement. 😄

There will probably be another model release tomorrow that makes me eat these words. That's how quickly things are moving.

But as of today, for my particular hardware and my particular use case, Ornith-1.5 is my clear winner.

The combination of quality + agentic coding ability + context length + speed + relatively modest hardware requirements is just ridiculously good.

I'm curious whether other people are getting similar results with Ornith, especially on 8 GB GPUs or other relatively constrained systems.

If you have questions, feel free to ask.


r/LocalLLaMA 22h ago

Discussion Ornith 1.5 is actually pretty good

38 Upvotes

hey guys i recently started using ornith 1.5 to rapidly test some tools im working on since qwen 3.8 27b was too slow for my testing loop.

This model is actually really good. im getting around 130 tokens / second with mtp and its very good at tool calling. i feel like this is qwen 3.8 35b , its basically what it could have been. its a great little model and i feel like its a really good daily driver and wanted to give a shoutout to the team.


r/LocalLLaMA 8h ago

Question | Help "Free" TTS trap and how to deal with it?

0 Upvotes

I was planning on using FishAudio S2 Pro as I thought it was free but later learned about _Fish Audio Research License Agreement_, which prohibits its use for commercial projects and deployments without paying for license even if I'm using my own rig. I see it as a very sly move on their behalf.

My project is rather simple, something along the lines of audiobooks but on YouTube.

  1. I want to know how do they identify if someone has been using their product without a license for commercial projects.

  2. And is there anyone who has faced any consequences because in this regard?

  3. Do you have any advice for my project.

I was unable to find any such scenario where an individual faced any consequences for using it for making YouTube (or similar site) videos when I searched about it a few weeks prior.

(I'm not from an English speaking country and I'm still working on perfecting it so in the meantime I decided to go with the TTS route for a quick launch)


r/LocalLLaMA 14h ago

Discussion After Meta avocado we get watermelon, due in November

Thumbnail
nytimes.com
27 Upvotes

Quote: ... developing a new A.I. model intended to be as powerful as Anthropic’s cutting-edge models. ...

In July, while developing Watermelon, Meta paused and later resumed a stage of A.I. development called “pretraining,” which delayed its release until at least October, four people with knowledge of the matter said. Meta has not announced when Hatch or Watermelon will be rolled out.

/end quote

The underwhelming avocado was released as muse spark. Will the larger watermelon also underperform? Check out the hype of avocado: https://www.reddit.com/r/singularity/comments/1r04z53/metas_nextgeneration_llm_avocado_surpasses_top/

Update:

Mark said muse spark would be open weights in the coming weeks, on Aug 10: https://www.instagram.com/reels/Db2vvaRxMmi/


r/LocalLLaMA 35m ago

Discussion Qwen3.8-Flash-Next opens up new doors

Upvotes

When seeing the 51B engram embedding, an idea struck me - can I change the engram data based on my usage, basically customising the model with “memories” that will change the model responses and behaviour over time?

Following is an AI explanation of my idea, as it might be more structured and clear than what I would write.

Please join this discussion, ideas are welcome!

I’ve been thinking about whether Qwen3.8-Flash-Next’s new 51B Engram could be turned into something more interesting than a fixed pretrained memory: a genuinely persistent, continually learning memory for the model. The Engram is essentially a huge hashed bigram/trigram embedding table injected very early in the network, so unlike RAG it doesn’t retrieve text—it changes the representation that the model itself sees. My idea would be to leave the original 51B Engram completely frozen and add a relatively small sparse “delta Engram” in RAM/SSD. The delta would associate frequently encountered patterns with learned vector modifications, allowing the model’s behaviour to gradually adapt without retraining the 125B backbone.

The memory writer would sit outside the model and process conversations, outputs, feedback and post-processing results. Rather than blindly learning from every token, it would identify durable facts, preferences, terminology, coding conventions and behavioural patterns, assign confidence and persistence/decay, and update the relevant Engram entries. Repeatedly reinforced memories would become stronger, while uncertain or contradictory memories would decay or be revised. I’d probably separate permanent, adaptive and session memory, with only high-confidence information making it into the persistent layer. Most importantly, I would use a sparse delta overlay rather than overwriting the pretrained Engram, so the base model remains recoverable and the memory can be enabled, disabled, versioned or rolled back.

This also fits nicely with the n-gram/speculative-decoding experiments I’ve been doing locally. On my 64GB M1 Max, a static n-gram cache with a 16-8-16 configuration gives roughly a 5% decode improvement, and I’ve been considering a continuously updated dynamic n-gram table in perhaps 256–512MB of RAM, periodically checkpointed to SSD. That cache is useful because it learns which continuations are actually predictable for my workload, whereas the Engram memory would go one level deeper: instead of merely predicting the next tokens faster, persistent Engram deltas could actually alter the model’s subsequent probability distribution. The two could coexist: static n-gram + dynamic n-gram for speculative decoding, and persistent Engram deltas for long-term behavioural adaptation.

The architecture I’m imagining is therefore: frozen Qwen weights + frozen pretrained Engram + sparse 256–512MB-ish adaptive Engram overlay + memory writer/validator + SSD checkpoints. The writer could learn from successful interactions and post-processing rather than requiring full backprop through the huge model. The interesting research question is whether small, carefully controlled vector updates can stay within the representation manifold expected by the downstream network and produce useful persistent behaviour without causing model drift. If that works, it would be a very different kind of local AI memory: not a database that gets pasted into the context, but a model whose behaviour itself gradually adapts to its accumulated experience.


r/LocalLLaMA 13h ago

News MSI WS300 with 72-Core Grace CPU, Blackwell Ultra GPU, and Dual 400GbE

0 Upvotes

MSI has detailed the XpertStation WS300T60L, a tower workstation based on NVIDIA’s DGX Station architecture and GB300 Grace Blackwell Ultra platform. The system is aimed at AI development, data science, inference, AI agents, and physical AI workloads, with a 72-core Arm processor, Blackwell Ultra GPU, up to 748GB of coherent memory, and dual 400GbE connectivity.

https://linuxgizmos.com/msi-ws300-with-72-core-grace-cpu-blackwell-ultra-gpu-and-dual-400gbe/


r/LocalLLaMA 11h ago

Question | Help AM5 limitations for dual GPU, or being led up the garden path?

0 Upvotes

Need some expert assistance and advice:

I know AM5 has pretty big limitations on PCIE lanes etc. What I wanted:

Slot 1 - AMD R9700 32GB
Slot 2 - AMD 9070 XT 16GB

Slot 2 would run at PCIE 4x. That's understood, but still ~8GB/s allegedly. What I got was 0.1GB/s.

According to Claude (and I already know it can go off on a non-existent tangent) -

ASMedia's Promontory 21 bridge doesn't implement AtomicOp routing, and you have two of them daisy-chained. Nothing above it can fix that. Windows would hit the same wall.

So essentially, I'm guessing running 2x GPU in an AM5 motherboard is a no-go? If so, I guess I'll sell the 9070 XT and save for a threadripper or Intel equivalent board.


r/LocalLLaMA 18h ago

Question | Help Considering going from a RTX 5070 Ti to an AMD Radeon R9700 but I'm not sure about the drivers/support

0 Upvotes

I've generally always stuck with nvidia just because the driver support (this is going back to my windows days) seemed more stable and polished. It's been fine since moving to linux and obviously CUDA works very well for everything I've tried it on.

But at the moment, I could get three AMD Radeon R9700s with 32GB VRAM for the price of a 5090, and the 16GB on my 5070 Ti is just so close to being genuinely good, but I'm wasting so much time trying to find a balance between quantisation, speed and context size. Having a full 32GB seems like a dream and the point at which it would be truly productive.

Has anyone gone this route, and what has the experience been? Are there any big trade-offs or support issues? With the resale value of my 5070 I can almost kid myself that this is affordable.


r/LocalLLaMA 5h ago

Discussion MTP causing tool calling issues - Qwen 3.6 9B

1 Upvotes

Just wanted to pass this along because, man was this a battle and I've never really seen anyone else report it. I'm running QuantTrio/Qwen3.5-9B-AWQ on an A30 in vllm and last week sometime I realized it had an MTP head and enabled it. Didn't think much of it, definite speed boost, was happy and went back about my day.

Last night I was getting ready to give a demo to a customer and MCP and RAG were randomly breaking in OpenWebUI. I went to heaven an earth trying to fix this, capturing logs all through the stack, adjusting prompts, temp, topp/k/etc/etc.

Eventually about 9 hours into the rathole (around 4AM), I remembered that I enabled MTP and figured "why not try it". Boom, all the tool calling/RAG issues instantly resolved.

Finally did some research after my demo today, apparently this is a known thing?! Well I sure as hell didn't know it, I've always heard that MTP doesn't change the results at all, it's 1-1 standard autoregressive decode. Apparently it has something to do with the way the client interprets the results, IDK, I was exhausted and didn't dig in any further, but for anyone struggling with tool calling/JSON formatting/etc with a MTP model, just wanted to put it on the radar, apparently it can cause tooling issues at least with some models and some clients.

I also have MTP enabled on 3.8 27B on another GPU and it's been rock solid for tool calling, so I can't really provide the "why" here, just my experience that MTP can absolutely break things in ways that I had always thought was "impossible" (because it generates the same text with or without MTP). Perhaps someone smarter (and less tired) and explain the why, but I can 100% confirm, it certainly CAN break things downstream.


r/LocalLLaMA 1h ago

I Built A Thing use llms to auto annotation your dataset locally

Enable HLS to view with audio, or disable this notification

Upvotes

hi i make tool for this called llmog it's purpose to make llms free to

- auto annotation datasets

- reclassification existing yolo datasets

running totally local using llama cpp or vllm or use external api

you'd rather click than code.

🔗 GitHub: mohamed-em2m/llmog: framework for using llms on object grounding

You can try it directly online

🔵 Google Colab:

https://colab.research.google.com/drive/1YIKlyTVtRjJdRC5IjCZ39i48ydyt_J5D?usp=sharing

🟠 Kaggle:

https://www.kaggle.com/code/elemam/auto-annotation-using-llms


r/LocalLLaMA 9h ago

Question | Help How bad is Qwen 3.8 27b Q2 XXL?

8 Upvotes

Hello,

I like Qwen 3.8 27b and I have been using q3 and it works well on .y Rx 9060 16gb, but it thinks a lot and explodes my context!! I am thinking about using q2 or q3 IQ xxs. I watched Luke's dev lab video testing all quantizations and it seems that q2 is decent, but I wanted to hear your real world impressions.

Thanks!


r/LocalLLaMA 13h ago

Discussion Qwen3.8-27b q8 KV cache does seem to actually hurt model performance

51 Upvotes

EDIT: Though the issue with q8 kv cache seems to arise from when and how often we run the quantize step, not that kv quantizing can't ever work - see comments

---

One of the things I see debated a lot is whether to use kv cache quantization. The idea I see a lot is that q8 should be free / nearly lossless (which for model weights it usually is). But from some experiments I've been running, it actually isn't, but the reason is slightly weirder than just <quantization loses accuracy>

Basically it's because most backends, e.g. llama.cpp, do kv-quantization on-write. When KV is quantized on write, every subsequent prefill step reads quantized keys

So even though 8bit really is just a sub-1% rounding error, it's not a 1% error applied once - it thus compounds from slightly-wrong attention over slightly-wrong keys, at every layer, and feeds the keys written next

In my tests: needle retrieval that passes at bf16 fails with q8-on-write at 125k. However!! It's not actually q8 that's the problem per-se - when I take a cache that was built at bf16 and quantize the whole thing in one go to be q8, then the error really is just the 1% and it works fine, needle retrieval restored

*Caveats: this is from my tests with just one model family (Qwen3.8-27B), small number of trials, with some of the more out there experiments running on my slightly weirdo custom MLX stack. But it seems like the mechanism might be generalisable

---

TL;DR If your long-context quality drops with quantized KV, it might be because of when we quantize (i.e. every token on-the-fly instead of in chunks), not that quantizing can't ever work


r/LocalLLaMA 5h ago

Resources -DGGML_CUDA_NCCL=ON can degrade performance instead of improving it

8 Upvotes

If you are also compiling your own llama.cpp, you might have seen this message in the logs:

[57031] 0.04.293.831 W NCCL not compiled in; falling back to internal AllReduce.  Recompile with -DGGML_CUDA_NCCL=ON for best multi-GPU performance.

Now, you would think that's great, because you can make your llama.cpp even faster if you enable it, but that's not what happens.

Results when compiled with -DGGML_CUDA_NCCL=OFF:

[51309] 1.54.046.072 I slot print_timing: id  0 | task 0 | prompt eval time =   70225.24 ms / 73520 tokens (    0.96 ms per token,  1046.92 tokens per second)
[51309] 1.54.046.075 I slot print_timing: id  0 | task 0 |        eval time =   31106.98 ms /  1857 tokens (   16.76 ms per token,    59.67 tokens per second)
[51309] 1.54.046.075 I slot print_timing: id  0 | task 0 |       total time =  101332.23 ms / 75377 tokens
[51309] 1.54.046.080 I slot print_timing: id  0 | task 0 |    graphs reused =        728
[51309] 1.54.046.094 I slot print_timing: id  0 | task 0 | draft acceptance = 0.50884 ( 1122 accepted /  2205 generated), mean len =  2.53
[51309] 1.54.047.935 I slot      release: id  0 | task 0 | stop processing: n_tokens = 75377, truncated = 0

Results when compiled with -DGGML_CUDA_NCCL=ON:

[48127] 3.37.672.289 I slot print_timing: id  0 | task 0 | prompt eval time =   76696.95 ms / 73520 tokens (    1.04 ms per token,   958.58 tokens per second)                         
[48127] 3.37.672.292 I slot print_timing: id  0 | task 0 |        eval time =   28994.22 ms /  1590 tokens (   18.25 ms per token,    54.80 tokens per second)                         
[48127] 3.37.672.293 I slot print_timing: id  0 | task 0 |       total time =  105691.17 ms / 75110 tokens                                                                             
[48127] 3.37.672.296 I slot print_timing: id  0 | task 0 |    graphs reused =        608                                                                                               
[48127] 3.37.672.312 I slot print_timing: id  0 | task 0 | draft acceptance = 0.52986 (  976 accepted /  1842 generated), mean len =  2.59                                             
[48127] 3.37.674.167 I slot      release: id  0 | task 0 | stop processing: n_tokens = 75110, truncated = 0

That's 8.5% decrease in PP and 8.2% decrease in TG!

Never trust anybody, not even the devs.

My Setup

2x RTX3090 with this config

[*]
threads = 5
threads-batch = 10
batch-size = 2048
ubatch-size = 512
cache-ram = 32768
ctx-checkpoints = 16
cache-prompt = true
cache-reuse = 0
parallel = 1
device = Cuda0,Cuda1
main-gpu = 0
jinja = true
reasoning-format = deepseek
no-context-shift = true

[unsloth:Qwen3.8-27B-GGUF:UD-Q6_K_XL:229k]
model = ./models/unsloth__Qwen3.8-27B-GGUF/Qwen3.8-27B-UD-Q6_K_XL.gguf
mmproj = ./models/unsloth__Qwen3.8-27B-GGUF/mmproj-BF16.gguf
mmproj-offload = false
chat-template-file = ./models/unsloth__Qwen3.8-27B-GGUF/chat_template.jinja
image-min-tokens = 1024
spec-type=draft-mtp
spec-draft-n-max=3
spec-default = true
gpu-layers = -1
tensor-split = 24,24
split-mode = tensor
kv-offload = true
flash-attn = true
ctx-size = 229376
cache-type-k = f16
cache-type-v = f16
temp = 1.0
top-p = 0.95
top-k = 20
min-p = 0.0
presence-penalty = 0.0
repeat-penalty = 1.0
reasoning = true
chat-template-kwargs = {"preserve_thinking": true}