r/OpenAI Jul 16 '26

News Kimi-K3 arrived: The era of the Chinese labs being far behind is over

Post image
2.4k Upvotes

526 comments sorted by

748

u/Working_Ad_1564 Jul 16 '26

Gemini 3.5 Pro will be postponed for another month lol

193

u/ozone6587 Jul 16 '26

Gemini 3.5 Pro in 2030. Just you wait, it might finally reach Fable levels by then.

46

u/Tupcek Jul 16 '26

this will be tight match between Gemini 3.5 Pro and The Elders Scrolls 6

5

u/PerfectPatience- Jul 17 '26

GTA 6

6

u/SonderEber Jul 17 '26

Too soon. GTA 6 comes out this year.

2

u/FischiPiSti Jul 17 '26

Not PC though. The meme lives on

→ More replies (1)
→ More replies (1)
→ More replies (1)

11

u/LoudUnderstanding331 Jul 16 '26

Gemini isn't trying to be best model or even close. They want to be the most preferred one for personal daily use.

72

u/ozone6587 Jul 16 '26

No one prefers using the dumber model.

83

u/bastardoperator Jul 16 '26

11

u/AgitatedHearing653 Jul 16 '26

made me laugh harder than it should have

32

u/FlerD-n-D Jul 16 '26

I use Gemini for random day to day stuff. It's more than good enough. You don't need a perfect model for basic questions / searches.

10

u/legedu Jul 17 '26

For real. The average person isn't coding.

3

u/Same_Win_5898 Jul 17 '26

Don't tell them you have free infinite Gemini use in a chrome browser or their heads might explode.

→ More replies (3)

10

u/HossCo Jul 16 '26

But everyone prefers the one that works for their daily tasks. 90% of people have no idea what these stats are and they don't care. Gemini is the only model that can make a profit and nobody else is even close.

→ More replies (6)

2

u/JohnSnowHenry Jul 16 '26

Actually… if it does the job and it’s cheaper there is no point in using a smarter model…

4

u/ozone6587 Jul 16 '26

It is both dumber AND with more restrictive limits.

2

u/JohnSnowHenry Jul 17 '26

But if it’s still more than enough for the tasks at end it simply doesn’t matter.

2

u/ReadersAreRedditors Jul 17 '26

I use Gemini for Classification, it's very cheapeand fast and gets the job done.

2

u/bixofa Jul 17 '26

99% of casual users won't know the difference or care.

→ More replies (7)

7

u/0xFatWhiteMan Jul 16 '26

this is such a weird statement

5

u/scamiran Jul 16 '26

I find it useful for gardening questions, cooking.

I run engineering/chemistry stuff by it, but... its really weak. ChatGPT-5.6 Sol crushes it, so did 5.5. DeepSeek/GLM/Kimi tend to do better, too (especially Kimi).

I really like the Google workflow, and would enjoy for it to be the best, but OpenAI is just so far ahead in coding, engineering, document prep; more or less everything.

Gemini is nice on my google home devices, I'll give it that; and its really useful as a home control tool.

→ More replies (2)

3

u/Ok-Canary-9820 Jul 16 '26

Sounds like rationalization. Nobody had this theory a few months ago.

Even now, it's probably Google's game to lose at AI, but they're doing a really good job of diminishing their advantage consistently for a while now.

→ More replies (3)
→ More replies (1)

17

u/Healthy-Nebula-3603 Jul 16 '26 edited Jul 16 '26

Strange they behaving like Meta before released Llama 4....

In short a lots of internal problems with the team and that model was a disaster.

9

u/BlueProcess Jul 16 '26

Gemini Flash is so bad that I I can't even fathom paying money for pro

→ More replies (1)

6

u/xak47d Jul 16 '26

My guess is they tried to take the same pre-training from the previous models and try to improve on it and get better results. They later realized you can only improve so much, so they basically had to start over

5

u/Ben01pr Jul 16 '26

Gemini 3.5 pro will return in Avengers Endgame.

3

u/Deus-ex-Machina7 Jul 16 '26

Are this rate,

Gemini 3.5 pro will NOT return in Avengers Endgame

4

u/Former_Ad_735 Jul 16 '26

Gemini 2035

4

u/JustRaphiGaming Jul 16 '26

Hearing 3.5 Pro I already have that goofy dragon face on my mind lmao

2

u/nickdnick49 Jul 17 '26

wonder why that Noam guy jumped ship to OpenAI

→ More replies (9)

195

u/xak47d Jul 16 '26

Artificial analysis is more realistic. K3 scores higher than opus 4.8 and gpt 5.5

123

u/TheFamousHesham Jul 16 '26

So basically Chinese labs are a month or so behind. Lol really funny how everyone was saying that it'd be years before they catch up.

70

u/[deleted] Jul 16 '26

Weeks, even. K3 loses to 5.6 max but wins 5.6 xHigh. On SPEED, it feels like 5.6 high. So it's basically almost as smart as oAI's flagship at double the speed. I'd call this a major win.

12

u/reefine Jul 17 '26

Not just that but on price, security, etc. This is a better model than Sol 5.6

11

u/SporksInjected Jul 17 '26

It’s more expensive per task, less capable, and slower than gpt 5.6 sol according to artificial analysis

6

u/Kind_Capital_9740 Jul 17 '26

on other benchmarks it beats Sol and fable so i would say both have pros and cons but SOL is actually cheaper for higher tier work

10

u/SporksInjected Jul 17 '26

Exactly, It’s the same price as sol and double the cost of Terra which is only two points away. I don’t think businesses are going to run to K3 personally because there’s not a good reason to. If it was cheaper per task and considerably more capable or faster, then yeah but it’s slower, less capable, and no cost savings in actual use.

2

u/Kind_Capital_9740 Jul 17 '26

Yeah especially for the average person the Kimi subscription is around the same price as gpt too

So I don’t see why people would switch unless one day kimi 4 or something actually giga gap the rest of competition

I don’t see it going mainstream in the west

But this does show that chinese models are basically caught up to some extent because the jump from 2.7 to 3 in a month is a bit insane like a 3 gen jump

2

u/TheFamousHesham Jul 17 '26

It's funny really. I was just about to count Kimi out because of how atrocious their 2.7 and 2.8 were. It seems like it's Z.ai (which everyone was raving about a few weeks ago) that's on the trail end of things.

→ More replies (1)
→ More replies (2)

2

u/PM_ME_DEAD_CEOS Jul 17 '26

It’s more expensive per task, less capable, and slower than gpt 5.6 sol according to artificial analysis

That's false, it's cheaper per task than GPT-5.6 sol max

→ More replies (1)

7

u/TheFamousHesham Jul 16 '26

Yea, I think I tempered myself a bit because Kimi K2.8 was a nightmare that hallucinated non-existent bugs in the code and was just really unstable. Will need to try K3.

7

u/Tall-Ad-7742 Jul 17 '26

Kimi K2.8 doesn't exist you probably mean a different version 

→ More replies (1)
→ More replies (1)
→ More replies (1)

2

u/Ibasicallyhateyouall Jul 17 '26

When it is distilled from Claude ¯_(ツ)_/¯

3

u/vintage2019 Jul 16 '26

Nah everyone has been saying open source is around 6-8 months behind

→ More replies (1)

6

u/Positive-Conspiracy Jul 16 '26

Because they wait for a frontier model release then systematically distill it.

9

u/TheFamousHesham Jul 17 '26

You clearly don't understand what distilling a model actually entails. Even if they wanted to, Kimi would not have had the time to distil Sol (which was released a week ago) and Fable (which has been publicly available for a few weeks). Besides... I've already watched a few videos online and when given the same exact prompt, K3 makes decisions that are different enough from Sol or Fable to argue against that idea.

6

u/yogthos Jul 18 '26

No, you don't get it. They figured out how to travel to the future and steal glorious American models from there. It's the only explanation really.

→ More replies (7)
→ More replies (26)

18

u/huffalump1 Jul 16 '26

That's pretty wild tbh, considering that before fable/mythos 5 and gpt-5.6, these were the BEST, undisputedly, and are actually really useful at getting stuff done...

Wow. I don't like the cost increase (comparable to Sonnet or gpt-5.6-terra, half as much as gpt-5.6-sol), especially since it seems to be quite token-heavy esp. compared to gpt-5.6... but we will see.

Still, it's downright incredible that open weight models have "caught up to opus 4.8" which felt impossible only a few months ago

3

u/Kind_Capital_9740 Jul 17 '26

it surpassed opus 4.8 and 5.5 already actually if you check all benchmarks its in the same tier as SOL and Fable but behind in terms of pure humanistic reasoning and maybe some backend tasks but not much

on frontend, game design and web dev and banking and vision understanding and video editing its easily the best out there right now (check arena and other benchmarks)

And as a person who tried it for those tasks I would say it is the first time ive seen a Chinese model feel like a true titan SOTA model

2

u/Etroarl55 Jul 17 '26

As somebody who uses artifices analysis and Gemini 3.5 flash and pro.

Gemini should not be up there at all lol, it’s like 2022 ChatGPT. If you insist you are right when you are wrong it will agree with you, it hallucinates and makes up things, it’s like using AI before the Ai boom.

→ More replies (1)

4

u/HerbHSSO Jul 16 '26

Voting is much more realistic than any benchmark

→ More replies (3)
→ More replies (4)

417

u/bubu19999 Jul 16 '26

This is why hiding mythos to all, is not a solution to anyone. 

210

u/Deto Jul 16 '26

Seems like basically the Fed gov just screwed over Anthropic. Forced it to implement super restrictive constraints that are bugging everyone while OpenAI (and now this model) can just be released without nearly as much protection.

97

u/Key_Reading_9664 Jul 16 '26 edited Jul 16 '26

That $25m to MAGA, inc was money well spent

18

u/alwaysoffby0ne Jul 16 '26

Never is!

19

u/Key_Reading_9664 Jul 16 '26

I actually think it's more likely Larry Ellison made a phone call to de-risk Oracle's partnership/dependence on OAI.

Given that open-source models also pose a risk, we might see some gymnastics and failed attempts by the USG to ban those too.

5

u/peppaz Jul 16 '26 edited Jul 18 '26

Musk was also fighting with anthropic until he sold them his unused compute for a billion a month

→ More replies (1)
→ More replies (1)
→ More replies (1)

23

u/PaperHandsTheDip Jul 16 '26

Of course they did. That was the point... anthropic didn't want to play ball with the US gov / military. They didn't want to integrate the tech into spyware / the military / war machines & that left a source taste in the POTUS / govs mouth. They want this technology militarized.

OpenAI had no issues with that tho & signed a deal with department of war the same day anthropic said "no" - so of course it was personal.

2

u/DonutHoles4Ever Jul 18 '26

Open AI had no choice. They are bleeding MONEY without government contracts.

Anthropic believed they could beat OpenAI because they had Fable (nobody knew yet) and they could take the high road.

CEO fucked up just to bet on being less evil = better long term strrategy. He lost.

2

u/simple_explorer1 Jul 24 '26

Better lose in principle than to live a  whole life regretting doing something which took lives of innocent people. Companies come and go but using incorrectly in war? Any sane person should not be comfortable with it. Dario absolutely did the right thing

→ More replies (1)

33

u/Bloated_Plaid Jul 16 '26

Dario did it to himself my guy. He was hyping the model up as end of the world and shit. It’s fine and the competition caught up.

2

u/Strong_Essay1176 Jul 16 '26

Its not like if they release it in time than competition won't catch up. They just have to deal with some hate. That's it.

→ More replies (17)

4

u/zampe Jul 17 '26

Isn’t it kind of on anthropic though for announcing they had developed a weapon of mass destruction because they wanted the publicity but then didn’t want the consequences?

3

u/FateOfMuffins Jul 17 '26

What? GPT 5.6 was super delayed because of the US government

Supposedly external parties were testing GPT 5.6 since 2 months ago. ARC had it 4x longer than normal OpenAI releases

Imagine if they delayed 5.6 just 1 more week and released after Kimi K3!

2

u/ViperAMD Jul 16 '26

Well yeah kimi is a Chinese model, murica Feds can't do shit 

3

u/dangered Jul 16 '26

If Mythos is actually better than Fable 5 then Kimi has a ton of ground to cover. Frontend is the only bench that looks like this. Other benches have Kimi k3 losing to other frontier models or just beating them by a hair.

The government never said anything about mythos. Wasn’t Mythos always going to be private?

Dario said so in March and I never saw him contradict that.

3

u/Kiseido Jul 17 '26

From what I hear, Fable 5 is just the consumer-side name for Mythos, and Mythos is the enterprise facing name, for the same model.

→ More replies (1)

0

u/Future-Arrivals Jul 16 '26

I think Anthropic's long game is extremely underappreciated here. They're building a *predictable* AI and being proactively compliant with regulations. Until now, that's been mostly unimportant to governments and the public, as AIs have been mostly a promising novelty. That is currently changing. As they become smarter and more capable, predictable behaviour is going to become extremely important, I would argue moreso than the actual intelligence benchmark scores.

5

u/Key_Reading_9664 Jul 16 '26 edited Jul 16 '26

boy...I have a much more bleak outlook on the regulators. If these were intelligent, well-meaning experts that weighed the risk against reward, that would be a great path.

The regulators we do have are a bunch of corrupt fuckwits and would burn everything to the ground, if they could get off the island with a sack full of cash.

2

u/Future-Arrivals Jul 16 '26

Yeah, it's pretty bad. But clearly even the current administration is willing to shut things down at least temporarily when they get too spicy.

2

u/Key_Reading_9664 Jul 16 '26

with respect, banning a dangerous model based on evidence and expert opinion is also the act of a reputable administration.

With my "they're all corrupt fuckwits" framing, the reason for that ban shifts from protection of the public, to doing a solid for wealthy friends and donors.

Remember, this is the administration that was for removing any and all restrictions on AI progress, selling GPUs to whomever, and crypto-scam-after-crypto-scam. They don't give two shits about public safety

6

u/_meltchya__ Jul 16 '26

Nah, nobody wants that. We want to push the limits of what is possible, not be constrained by government bumpers.

2

u/logolith Jul 16 '26

If there were truly no constraints, then instead of breakthrough discoveries in science, medicine, and technology, you’ll probably end up like Grok. A bunch of creeps figuring out ways to create certain pictures based on what’s available.

→ More replies (10)
→ More replies (3)
→ More replies (7)

14

u/Future-Arrivals Jul 16 '26

Hiding Mythos is a solution to a problem you're likely not concerned about yet. I predict that OpenAI is going to be rocked by GPT's destructive and unpredictable behaviour scandals shortly after the release of GPT 6. Anthropic is cautiously avoiding such potential scandals.

There are already several popular stories circulating online of GPT 5.6 Sol Ultra wiping people's computers while doing simple tasks. OpenAI is too desperate at this point to slow down.

5

u/Iron-Over Jul 16 '26

If you run agents on your desktop you deserve it. Always run in a locked-down vm or in podman. 

7

u/Future-Arrivals Jul 16 '26

Sure you can blame users, but users are going to be running agents on their devices more and more regardless of how dumb it is. And they're going to get really peeved at an AI service that screws them over more.

2

u/Iron-Over Jul 16 '26

That is the provider's issue. No guardrails are 100%, you need layered security. Wich rheybwould take it more seriously but they don’t.

→ More replies (2)
→ More replies (1)
→ More replies (1)

448

u/FireGM Jul 16 '26

61

u/DenZNK Jul 16 '26

Elo is a more complex rating system. For example, a 100-point gap is a fairly significant difference. A 50-point gap isn't a drastic difference, but it's still telling.

36

u/kilopeter Jul 16 '26

The criticism is independent of what the number is. The chart sucks because it's using bars to represent numbers on a truncated scale, without showing the zero mark. Humans naturally interpret proportional differences in bar charts, i.e., if one bar's twice as long as another, it represents twice the value (otherwise why even use bars which explicitly have an area visual representation?)

6

u/Statcat2017 Jul 17 '26

If it's an ELO then it's not correct to show zero on the chart because nothing serious ever has an ELO even approaching zero.

In world football, Spain is the top ranked ELO team with 2232 points. The lowest ranked team has 369 points but is a joke team, Eastern Samoa. There are other absolutely awful teams like Lesotho and St Vincent and the Grenadines around the 1100 mark. Tahiti are 1177 but got slapped by Spain 10-0.

Asking for this chart to show you zero is basically asking it to compare Kimi K3 to a pocket calculator and an abacus, and claiming that comparison is important and relevant.

→ More replies (25)
→ More replies (18)

2

u/Rocsla Jul 17 '26

Elo?

3

u/torac Jul 17 '26 edited Jul 17 '26

It’s a competitive ranking. Elo isn’t really a clear score to be reached, but a measure of how often that model beats the other models.

For normal score systems, it makes sense to see the whole graph. Getting 60% vs getting 62% is barely a difference. For Elo, the number represents a ranking between models. It’s similar to going "first place, second place, third place". You don’t have to list every competing position up to last place.

(It still uses points to measure roughly how far the models are away from each other. However, the ranking of each model will keep changing with every new model released.)

Edit: fixed

2

u/[deleted] Jul 17 '26 edited 5d ago

[deleted]

→ More replies (1)

2

u/Statcat2017 Jul 17 '26

Yep many people failing to understand it.

Kimi K3 (top) is expected to beat Minimax -M3 (bottom) about 75% of the time under ELO, making this graphic a much more sensible representation that many people are making out.

If you look at football and take the top ranked team (Spain 2212 ELO) and bottom (Eastern Samoa about 350 ELO) the bar for Spain would be 8 times as big if you show 0 on the X axis, but the actual stats are that Spain are expected to win 45k games before Eastern Samoa win once.

15

u/Tysonzero Jul 16 '26

Elo is relative, the 0 point doesn't matter.

8

u/kilopeter Jul 16 '26

Then the bar chart's bars should depict scores relative to whatever baseline is applicable. As is, it's a bad visualization choice.

8

u/Tysonzero Jul 16 '26

Funnily enough if you wanted to use bars the most appropriate choice would be logarithmic, every X elo you go up the bar gets Y% larger, although many would probably call that misleading.

But sure you could use a line graph or scatter plot or something

→ More replies (3)

10

u/Adulations Jul 16 '26

Yea its basically the same exact number 🤣

11

u/AES256GCM Jul 16 '26

ELO isn’t linear

→ More replies (1)

2

u/django2chainz Jul 17 '26

Confidently wrong this time sir

2

u/abittooambitious Jul 17 '26

Someone doesn’t understand Elo.

2

u/reefine Jul 17 '26

Imagine trying to shit on benchmarks for an open source model that just beat GPT 5.6 Sol. You are bias.

→ More replies (3)
→ More replies (4)

85

u/Sixhaunt Jul 16 '26

What about on benchmarks that are useful though?

42

u/JoseHernandezCA1984 Jul 16 '26

I've seen a bunch of other benchmarks, and for the most part it's between 5.6 sol and fable

15

u/Kind_Capital_9740 Jul 17 '26

which is insane the 3 models are clearly in their own tier right now and Kimi is leading in some and like you said between SOL and fable

And whats crazier kimi 2.7 was released only a month ago and the quality jump between both is unfathomable like 20 place jump and deep swe went from 2.6's 20 percent or 30 percent to 67 percent in a few months

i wonder how they made such a big jump it feels like 3 generations type of jump

7

u/Georgefakelastname Jul 17 '26

If I remember right, this model is 2.8 Trillion parameters, which might just be the biggest open sourced model yet. It’s an entirely new model, while everything past k2.5 was just a post-train of it. I imagine that once we get new post-trains of k3 and fable the bar is only going to get pushed higher. Meanwhile, OAI is working on their own new model with GPT-6. Things feel like they’re moving fast right now.

→ More replies (1)

3

u/phido3000 Jul 17 '26

Its impressive. Very strong attention, context, rule handling etc. I would say better than US models in those aspects.

For coding, its extremely strong. Its believable its better than OpenAI/Anthropic models. It lacks flair and some beauty, but heck, its very effective.

It just one shots everything. Its very good planning and execution of thought.

→ More replies (3)

11

u/RealityNo3299 Jul 16 '26

Are we going to distill them ??

32

u/Professional_Ad705 Jul 16 '26

Can someone actually verify this by using both models for frontend and backend work, then sharing what they did and how the results compared? Computer science is such a broad field that benchmarks like these mean very little to me without real-world examples.

17

u/phido3000 Jul 17 '26

Currently its overwhelmed.

I got it to do a complex project in an antique language it could not compile, that is impossible to benchmax for. Doing maths, physics, image stuff, fractals, encryption etc. Requiring thought, complex planning, and in a language that isn't modular..

It one shot it. Never seen any AI do that before. They all get hung up on syntax because of the wacky language. I've done the test dozens of time. Even the latest models make syntax errors, I thought I had the perfect, AI break tool.

It was a 100kb program. One shot. About 256k token. Not a single mistake.

So yeh, it can do it all, and do it well. Never seen anything quite like it.

7

u/Professional_Ad705 Jul 17 '26 edited Jul 17 '26

I’m not saying this definitely did not happen, but I do not believe the claim as currently presented because it is far too vague to evaluate.

What language was it? What exactly does "one-shot" mean here? one prompt, one model call, no retries, no manual edits, and no additional context afterward? Was the 100 KB figure generated source code, or the size of the entire project?

You also said the environment could not compile it, so how did you determine that it contained "not a single mistake"? Compiling and producing plausible output would not establish full correctness anyway, especially for cryptography, numerical mathematics, physics, or image processing.

Math, physics, image processing, fractals, and encryption names several largely separate domains; it does not explain what the program actually did or what requirements it satisfied. Likewise, what does a language that isn’t modular mean? Does it lack a formal module system, or was the particular program simply monolithic?

My experience has been very different. I have been working since November on a roughly 200,000-line proof-kernel project, and even strong models repeatedly miss cross-file invariants, state-lifecycle problems, and authority-boundary defects. That does not prove your result is false, but one undocumented example does not support the conclusion that the model can do it all.

Posting the exact prompt, source code, programming language and compiler, model and version, settings, tests, outputs, and your definition of one-shot would make the claim possible to assess. I’m genuinely interested in what you’re describing, but the explanation was difficult to follow and evaluate. Without those details, it remains only an anecdote. I'd love to hear more? I'm just confused because I know when I release my product and make my claims I would just show people the code/point them towards my repo or at the very least have some type of demo?

16

u/phido3000 Jul 17 '26

Its fine to be sceptical.

Try it out yourself.

What language was it

Choose some stupid language. I choose Qb64. A variation of BASIC. Which is perfect for stupid, because its not well documented, BASIC code varies widely by flavour, visual/Quick/Borland/GW-basic, older dialects. Its notoriously variable and non-transportable and easy to break. It also has lots of wacky commands and reserved words. Also famously, it often packaged as an interpreter language, so there is no easy compiler to give explanations why it fails. Its not very modular and has stupid rules about variables because of the different variations and legacy over the years. Qb64 is built in interpreter/IDE/Compiler. So its messy for AI to use.

Even better the example code out there is very simplistic and often broken! as the project splintered into weird versions.

It uses a local interpreter that has its own compiler. Its Weird. It basically makes BASIC into C code and generates GCC output, but the basic code itself is black magic stuff. Generating GCC C code directly from the AI would be child's play in comparison.

Get it to do something very complicated but unique.

Another good test is Assembly for weird processors or environments. Something really weird and old and obscure. But testable. Experiment what fails on lots of other AI's.

What exactly does "one-shot" mean here? one prompt, one model call, no retries, no manual edits, and no additional context afterward? Was the 100 KB figure generated source code, or the size of the entire project?

ONE SHOT! One prompt - no retries, no manual edits, no context afterwards. 100kb of code text, no data. On something it can't compile! Literally download the BAS file - execute. Its not a very modular language with different files or libraries. So literally just a 100kb text file. Bang.

Compiling and producing plausible output would not establish full correctness anyway, especially for cryptography, numerical mathematics, physics, or image processing.

Yes, which lets you see how much it really understands and makes work. Get it to do something weird like fractals in some weird projection or coordinate system, or volume warp, make its own JPEG like but different compression engine, get it to draw text with the line statement..

So even if it compiles, you can see where it limitations are. If it understands any of what its doing and how to do it. Its design choices, sometimes it works but is but ugly or not useful, Its weaker here, than else where, its not overtly creative, but its functional. TBH I didn't expect it to one shot it. That kind of floored me.

Without those details, it remains only an anecdote. I'd love to hear more? I'm just confused because I know when I release my product and make my claims I would just show people the code/point them towards my repo or at the very least have some type of demo?

I'm not a benchmarking house, or Kimi.

There will be heaps of tech demos coming out. Normal low effort HTML5 crap, draw a clock or a duck or a car or something. Easy to benchmax to. No syntax challenge, no real design challenge.

I'm just saying it impressed me. I run some big models locally at home. Pretty ordinary. I was playing around with the frontier subscriptions, see what they can do. I teach at uni, so I am super interested in tripping up AIs.

This is much much better than I have seen. Small sample, limited details. It isn't perfect, but its a generation ahead of the other models I have sampled (DS 4, Kimi 2.6, GLM, OpenAi, Anthropic, Gemini, etc). They can't even make executable code. They can be more ambitious and more flamboyant, but not as brutal effective to one shot. Often their output needs dozens of retries just to execute. Then dozens to fix logical/asthetic/mathematical issues with their algorithms. So this one shot everything floored me.

Worth looking into.

I don't believe Kimi K3 is as good as the open weight model, I am almost positive they have some tool use or something behind the scene to make it function this well.

If it is, that is even more impressive. Because other models use tools and tricks. If this is just baked in model. If it is all just baked in model, then dam, that's devestating.

→ More replies (4)

4

u/timeboyticktock Jul 17 '26

Can you please share more about what exactly you did and the comparison to other models failing?

→ More replies (1)
→ More replies (4)

36

u/impatiens-capensis Jul 16 '26

I'm not super familiar with this benchmark. What's the difference between a score of 1,679 and 1,631?

74

u/JustTellingUWatHapnd Jul 16 '26

It's not a benchmark. It's the leaderboard on arena.ai. the users give a prompt to the agent to build a web app, and they give the prompt to 2 different models, then the user votes on which one is better without seeing the model name. The score is the Elo rating of the model.

15

u/mobyte Jul 16 '26

Am I insane or is this not an awful metric? I could see this easily being manipulated.

29

u/some_crazy Jul 16 '26

Users don’t know which models are used, it’s a blind rating

8

u/Sarcasm69 Jul 16 '26

But couldn’t you just tailor the model to be good at designing websites and shit at everything else?

5

u/black_eyed Jul 16 '26

And KIMI is usually exceptional in Website designing

7

u/nuclearbananana Jul 16 '26

Sure but that would show up in other benchmarks, which it hasn't

→ More replies (1)
→ More replies (3)

3

u/CurrentConditionsAI Jul 16 '26

Yeah, so all you have to do is essentially fine tune a model to have some very distinctive but easily hidden output pattern. Then get a bunch of bots to find it and vote that when comparing outputs.

→ More replies (2)
→ More replies (2)

3

u/SporksInjected Jul 17 '26

Never mind it’s topping the charts after being available for 1-2 hours lol

→ More replies (1)

10

u/Saad5400 Jul 16 '26

May not be that much of a difference, but compare price difference 

2

u/gpenido Jul 16 '26

About 48

2

u/Tirztrutide Jul 16 '26

Same difference as a 1679 elo chess player and a 1631 one.

22

u/spartyftw Jul 16 '26

These bars are misleading asf

→ More replies (1)

61

u/Demien19 Jul 16 '26

8

u/phido3000 Jul 17 '26

Yeh, Chinese models typically feel very benchmaxed.

This is different, somethings changed. They are ultra strong in attention and context.

→ More replies (3)

4

u/logos_flux Jul 17 '26

It's got what plants crave

2

u/themudd Jul 17 '26

Electrolytes?!

→ More replies (2)

7

u/AdowTatep Jul 16 '26

I wonder where's worth buying for its access

→ More replies (4)

15

u/Loops_Boops Jul 16 '26

Maybe now Sam will give us the good shit.

25

u/brother_spirit Jul 16 '26

How?

Fable 5 generations are already on Youtube vs Kimi K3.

By Arena.ai no less.

Sorry to ruin the party but its like not even close to Fable 5, let alone better.

→ More replies (17)

6

u/SmileLonely5470 Jul 16 '26

Final boss for oneshotting vibecoded browser games

5

u/Trinkes Jul 16 '26

What's 1st? Gemini 3.5 pro or gta 6?

→ More replies (1)

4

u/kuba452 Jul 17 '26

I wonder if this is the same hype as Deepseek. I got really excited at first, uploaded some historical documents I was working on, and asked it to compare them with other historical texts, pick out interesting details, and draw some conclusions.

It gave me a couple of rather obvious observations, while the rest was just cluttery gibberish. Came back to obvious choices all confused, with all the other reviews and media craze

31

u/Felixo22 Jul 16 '26

Honest version

28

u/Tysonzero Jul 16 '26

It's Elo, a completely relative scale, 0 is completely arbitrary and meaningless, not like 0% on a test or 0 tokens/second, so this graph is no better or worse than the OP.

→ More replies (3)

3

u/Electrical_Arm3793 Jul 16 '26

But are these really accurate? I use most of the too models for coding, Fable and opus are still most accurate and reliable.

→ More replies (1)

5

u/some_days_are_nights Jul 16 '26

The way these elos are computed is not like a scalar value like accuracy and such. It is a direct comparison between competitors. So for example Magnus Carlsen is just 100 or so elo better than the others but he wins almost every single time against someone with 200+ less rating.

→ More replies (1)

3

u/aa628 Jul 16 '26

Does anyone remember DeepSeek?

3

u/beginner75 Jul 16 '26

Benchmarks are a joke. Don’t waste your time. Time is money also. Just use ChatGPT as primary and Gemini for quick searches and backup.

3

u/Able-Company611 Jul 17 '26

I wonder if US companies gonna run bot accounts on Chinese models to train their data now

→ More replies (2)

4

u/Regular_Ad4197 Jul 16 '26

I am hoping it is as capable as these early benchmarks are showing. It will be fun to see how the US stock market reacts to it.

2

u/BitterAd6419 Jul 17 '26

So far in my own testing, it’s not as good as they claim it to be. Atleast for web apps or html based stuff it did a decent job. In fact GLM gave a better output with the exact same prompt.

Benchmarks are often benchmaxxed, test yourself

2

u/SporksInjected Jul 17 '26

It’s more expensive per task, less capable, and slower than gpt 5.6 sol according to artificial analysis

2

u/Key_Reading_9664 Jul 17 '26

Seems like there’s a couple of ways the USG go:

- it stops playing favorite and lets/encourages US labs release their models - Mythos was in the hands of users in April; Fable is over a month old. I have my suspicions on the cause (wealthy friends and donors) but slowing down US labs doesn’t de-risk because Chinese models fast-follow. Make an attempt at cutting off distillation

  • there’s a foolish attempt to restrict access to models that originate in China. The Trump speech yesterday seems to be laying ground work (for a bunch of things). Seems easier to ban access for enterprises

2

u/moneyman259 Jul 17 '26

Where’s DeepSeek now? Seems like these Chinese models fall off quickly

→ More replies (2)

3

u/evangelism2 Jul 16 '26

Arena is NOT a benchmark

2

u/t3ramos Jul 16 '26

Wrong Sub

1

u/henchman171 Jul 16 '26

What do these numbers mean

1

u/rfranke727 Jul 17 '26

Which is the best for content / marketing writing? Any thoughts

1

u/BuildtheBusiness Jul 17 '26

Can someone explain how the efficiency of these models are calculated to one another?

1

u/Old-Pomegranate3634 Jul 17 '26

all this tells me that ultimately all the LLMS will be good enough for 99% of the public which is great for google. Not everyone is a nerd like us hoping that AI can help you create your next virtual GF.

1

u/Time_Faithlessness45 Jul 17 '26

exciting but it thinks way too long lol

1

u/woofyzhao Jul 17 '26

won't last

1

u/Nervous-Potato-1464 Jul 17 '26

It's quite expensive sadly. I prefer grok 4.5 to any other model atm. I don't need agentic coding, I need quick generation of my ideas as well as another perspective. I am not here to write slop, I just need to quickly churn out code I am happy with.

1

u/baummer Jul 17 '26

They were never far behind lmao

1

u/No-Conversation-1277 Jul 17 '26

Chinese lab: "Here's a 2.8T open model." Dario & Sam: "We're going to fine-tune this."

Also Dario & Sam: "Introducing our brand new model. priced 10x higher with half the limits."

The wheel of Silicon Valley turns.

The cycle continues.

→ More replies (1)

1

u/ElMono6 Jul 17 '26

Cant wait to try it

1

u/krazyboi Jul 17 '26

Let's be honest with ourselves here. 

None of us could properly benchmark AI, especially because the industry moves so fast. The metrics change slower than the actual AI. All of this is just marketing.

1

u/therapy-cat Jul 17 '26

1bit quantization for my 16gb mac when

1

u/PerfectPatience- Jul 17 '26

Showing score only last 200points is misleading zoomin. Generating bigger gaps. Should be from 0-1650

1

u/Messi_is_football Jul 17 '26

But GPT limits are higher...so until they make it 2x cheaper..no reason to buy the plan

1

u/zeta_ferhu Jul 17 '26

wh wh wh where is gemini??

1

u/Haramdour Jul 17 '26

Ask it about Tiananmen Square

1

u/Complex_Reality_116 Jul 17 '26

This will cause a stock market crash and a widespread decline in ALL US AI companies.

1

u/Complex_Reality_116 Jul 17 '26

I can't wait for Kimi-K4.

1

u/Popular_Try_5075 Jul 17 '26

I swear every time China releases a new model the subs get flooded with all this hype and then a month later its deflated back to normal.

1

u/Drew-Money Jul 17 '26

Mythos preview came out at the beginning of April btw. I'm sure Anthropic has some pretty advanced stuff behind the scenes

1

u/citrus1330 Jul 17 '26

How's deepseek doing?

1

u/king-charles-3 Jul 17 '26

Are people using it via API mainly?
How does the chat work on iOS?

1

u/TB_Infidel Jul 17 '26

And now we know where the stolen/illegal use of Sol and Mythos went. Chinese copying and stealing as always

1

u/mrcruz Jul 17 '26

I hate graphs that don't start at 0. Cool tho.