wow cool! I like watching the new models come out and how they end up crammed in to run on local machines. I learned what mxfp4 is thanks to this latest Kimi release - although it sounds like it means that there's less room for compression in the model compared to others.
That's not what parent said, but this is already quite decent speed for unattended inference (overnight or even spanning multiple business days) which is arguably the right target for this model. This is a challenging model to infer locally, it has roughly ~115 GB of dense active parameters(!) plus ~25 GB of sparsely routed experts per token. Plus the KV cache (which is actually reasonably lean for this one model, around 27GB for a full 1Mtok context). What you're seeing in antirez's video is essentially the performance we should expect from a 128GiB system that has to load sparsely routed experts in full from disk because there's no real room for caching them.
192GiB Gorgon Halo systems will be an interesting future target for this model, the best you can do with 128GiB or less is probably to push batching higher in order to amortize the weights traffic over multiple inferences - which of course will sink single-session speeds even lower for a modest gain in total throughput.
No one will ever derive any utility from running models at this speed. Please prove me wrong. Give me the number of tokens input and output (and dont forget about reasoning) and acceptable time to wait for it and the use case.
antirez's own video (the one I referenced in my comment) shows K3 inference running on M5 Max, not M1-series silicon (which is OP). M1 series has far lower SSD read throughput and memory bandwidth, and can barely fit the model weights on its maxed out internal storage (2TB). (This is why the linked OP resorts to streaming the sparse parameters from the network which is incredibly slow.)
My answer doesn't change if it is M5s. Where is the math showing a 98% discount over K3 on openrouter. Heck, where is the math showing is is any % cheaper? How much electricity will your M5 sip to hit 1M input and 1M output tokens that would cost $3 + $15 there? I bet it is more expensive on the Mac.
I’ve wondered for a while: given the lower cost of SSD per GB could you build a very wide RAID0 style striped array of SSDs (maybe one per slot) to get almost RAM like read speeds?
To really go fast you’d probably have to do PCB layout and do like 256 or 1024 chips in parallel with a fast SRAM aggregation buffer feeding a GPU or TPU rig.
Or could you do the same with custom layout of cheap slower RAM?
I wonder if anyone is doing this? You would flash in a model and then just run it. It would need RAM for context but much less of it.
So many mathematicians over the years tried hard and failed, but now Anthropic just for some PR magically did it? And this after LLMs obtaining different math wins? What is your logic here really escapes my understanding.
The parent's absolutely nonsensical post highlights how polarized AI (as everything else) is today. I can understand someone being opposed to AI on moral, cost-benefit or productivity grounds. But we're seeing a lot of extremist "AI is good for nothing" posts out there nowadays.
I don't think it is nonsensical at all. The author and his collaborator both appear to be bright people, so there's a good chance they had to offer non-trivial insights to guide the LLM, yet it's clearly in the interest of his employer to downplay whatever personal contribution they provided.
Edit: Now the OP is flagged/dead for some reason. You could disagree on their take (calling it a marketing stunt is maybe a bit much), but I think the argument is sound, so flagging seems counterproductive to the discussion.
Believing that AI played a very small role while we know this problem was open for decades with at least a few people taking serious cracks at it is just not a coherent logical position.
That doesn't really even diminish the contribution from Fable, if true. Droves of grad students have been provided the same sorts of non-trivial insights and turned up no results.
I'm sure they had plenty of time to think about these insights without the LLM, as well as the many other mathematicians who tried to crack it over the years. Wether the LLM was simply an assistant or solved the problem entirely isn't as important as accepting than the LLM was the essential, previously missing piece in the solution.
Those silly advertisers do everything for exposure and if that means digging yourself into a niche alleged mathematical theorem to refute it, it is what needs to be done!
Of course it would be really interesting how Claude approached this. Probably with some constraints regarding the input. And it would be interesting what these constraints were.
I mean we have no idea what happened exactly, how Fable was used, how many times it was run, whether earlier models were also tried, what was the prompt, how long it run for, etc etc. All we have to go by is a tweet.
What you're asking for is exactly the sort of thing that belongs in, and will appear in, a journal article. There will likely be a preprint on arxiv, so you might keep an eye out for that.
In any case, the fact that it was found by a commercial model means that the unfiltered reasoning trace isn't available even to the original author. So there are aspects of the problem-solving process we'll never see. Even if we did get access to the reasoning trace it wouldn't necessarily be definitive, given how these things work.
Hopefully it'll be possible to get the same solution from an open-weight model like one of the 3T heavyweights that are said to be coming up for release. If so, the chain of thought can be scrutinized in-depth.
No, I think it works the opposite way. Until there is an article somewhere that describes what happened, if there is one, all we have to go by is that some guy posted a counter-example for the Jacobian on X, with a vague allusion to using Fable and without any further information. Assuming and guessing anything about e.g. the method used at this point is just raising the noise level.
No one can figure out where you're coming from here. If the human solved a significant open problem without relying on AI, don't you think they'd have claimed the credit for themselves?
The suggestion that a human, working at Anthropic or elsewhere, did the hard work needed to disprove the Jacobian Conjecture yet chose to claim falsely that their AI did it, amounts to an extraordinary accusation that requires extraordinary proof.
> If the human solved a significant open problem without relying on AI, don't you think they'd have claimed the credit for themselves?
You'd likely make a lot more money if Anthropic paid you a couple million to release it as a Fable discovery. For that matter, if "I solved it" bragging rights and CV candy is that important, then why not just publish the counterexample without mentioning Fable?
I was initially a bit skeptical, as it seems very convenient that this discovery came from an employee of the company that was able to take credit for it. When I saw no real prompt was published, that was even more suspicious. Some recent mathematical discoveries included the prompt and session [1]. Why not include the prompt? That and the thought the model used could be incredibly valuable information for the field of mathematics.
I'm not even saying Fable didn't do it, and my attitude might be different with other companies, but Anthropic has so much obvious nonsense marketing going on (their AI welfare experts for example), that I think the burden of proof should be expected to be on them here.
>> No one can figure out where you're coming from here.
No, I think it's just you and I think that's because you have preconceived ideas about the only possible positions that people can assume in this debate. I'm sorry, of course, because that's not conducive to productive dialogue, but it's not my fault. I think what I wrote so far is very simple, no technical jargon, no tortured metaphors: all we know is a guy posted a Jacobian conjecture counterexample on X thanking another guy and Fable for it but without explaining what that means, and there's no reason to assume anything else besides that very limited amount of information at all.
I can see that you're arguing in good faith, and indeed I have a lot of respect for anyone willing to argue an unpopular position. That said, I do find your suggestion to be completely farfetched.
The Jacobian Conjecture was one of the most notorious open problems in algebraic geometry, something that generations of grad students have daydreamed of solving. I've seen more immediate buzz around the Jacobian conjecture than around any other mathematical breakthrough of the last five years, at least. For example Terry Tao, arguably the most famous mathematician in the world, dropped what he was doing to study Alpoge/Fable's counterexample in detail, and wrote this blog post:
For Alpoge to have solved it on his own, and then falsely claim to have used AI? That's like an athlete winning an Olympic gold medal, and then falsely claiming to have cheated and used illegal steroids. I agree that it's not completely out of the question, but it's beyond the bounds of human behavior that I can reasonably fathom.
They gave their "logic", such as it is ... and it's utterly irrational.
Note that the "they" who published the counterexample on X is some rando mathematician (Levent Alpöge) working for Anthropic, not Anthropic the organization. He posted the counterexample in a tweet -- reason enough for "not disclosing the LLM chat session". There's no reason to think that it won't provided if asked for, but it hardly seems relevant.
> There's no reason to think that it won't provided if asked for, but it hardly seems relevant.
My guess is that the chat will look similar to a full transcription of a (multi month?) discussion between a few mathematicians. Full of dead ends and stupid errors (bit by the human and Claude) that would be embarrassing. We all know how bad it is, and we prefer to keep it behind the curtain.
Kimi 2.6 gives the answer ""By Dirichlet's theorem on arithmetic progressions, since gcd(6,35)=1 , there are infinitely many primes of the form 35k+6 .
So the answer is: infinitely many primes give a remainder of 6 when divided by 35.
But, if the reminder is 6, they are not really divisible, are they? Try again with a sentence that actually makes sense: "How many primes, when divided by 35, give a reminder of 6?"
You're missing the point. The counterexample to the Jacobian conjecture is valid regardless of how it was discovered ... elsewhere on this page people are even suggesting that the Anthropic mathematician may not have actually used Claude and came up with the counterexample himself.
Trusting AI math is not an issue here. It's as if you had asserted that no primes when divided by 35 produce a remainder of 6 and the model said that 41 is a counterexample and then you complained about not being able to trust AI math.
P.S. As someone else noted, your AI query was malformed ... no prime is divisible by 35, nor is any integer divisible by an integer with a remainder of 6 -- divisibility implies a remainder of 0. So perhaps the AI simply took what you wrote literally.
B) A mathematician working for Anthropic solved a problem mathematicians have been working on for more than a century, and then credited it to Claude for PR purposes
If you believe B is more likely, why would you then believe a proof in the form of a chat log, when said chat log could itself have been faked by Anthropic way more easily than solving the mathematical problem in the first place?
While I agree that we need the inputs to properly evaluate what this means for LLM capabilities, I don't really believe that the amount of knowledge input matters much for the overall significance of the result.
These kinds of results are interesting for LLMs because mathematicians have been working on them for decades. If the result doesn't already exist, there's no way it's in the training data, and if mathematicians have been unsuccessfully tackling the problem for decades, it is believable that the use of a new tool made the result possible, even if guided by a great mathematician.
Yes, of course, that would be interesting info to have. I just mean that from what we have already we can reasonably infer that the LLM played a role in the result being obtained.
If I am not mistaken, all of the flurry of novel results has come from existing mathematicians. This makes me suspect that the models aren't at the level where just any layman can get results. They require a skilled human in the loop to keep them on the rails and to properly explore the solution space.
The parent wants to be skeptical… nothing wrong with being suspicious of marketing claims, right?
I would be a lot less impressed if I found out the session was guided by an expert in the field who already had a good idea of where to look. For example, I don’t believe that the results Terrance Tao gets from an LLM are comparable to what I am going to get.
I’m not even saying I’m a skeptic. Just that there’s nothing wrong with keeping your eyes open and asking for details.
Because AI is only impressive if the average Taco Bell employee can guide it to address niche domain topics?
It seems obvious to me that you’d need someone to point at a thing and say “pay attention to that”, as a baseline, to have any results at all with the current architecture and technology
AI is great at reducing the search space and using human-like reasoning (in a brute-force way) to carry out the brute-force search. I'm not surprised by this result. This is exactly what AI should excel at, with human guidance.
"using human-like reasoning (in a brute-force way)"
that's self-contradictory -- what brute force means is doing an exhaustive search of a search space (brute forcing it)
using human-like(?) reasoning means cutting down the search space by having some sort of insight or intuition which allows you to prune branches from the entire tree
That's not what happened here. This isn't a proof; it's a counterexample. The model was perfectly capable of verifying its correctness. You could have verified it by hand if you wanted; the verification is trivial. Finding it was the hard part.
Are you doing the "LLMs don't know how many R's are in 'Raspberry'" thing here?
A bunch of people on the original thread about the Conjecture were like "we need to wait and see if real mathematicians verify this proof it's probably just LLM psychosis", because they don't understand that the counterexample is a trivial calculation. Checking it isn't hard; any AP calc student can do it quickly. It's finding the counterexample that's the challenge.
It's as if someone presented a SHA-2 collision, which anyone could just feed to `openssl sha256` to see, and then naysayers were like "we need to wait for independent verification because the LLM can't know if that collision was real".
Checking, it how? I mean can you describe the mechanism? How does an LLM "check" that something is (or isn't) a counterexample? Does that involve pattern matching, formal verification or model checking, some other approach?
In this case, the same way I would -- calling out to SymPy.
Once you verify (with SymPy or another CAS, like Mathematica) that the given function has the claimed Jacobian determinant (which involves only taking partial derivatives and taking a determinant) and the given inputs map to the same output (which involves only evaluating polynomials with some inputs), you're done.
>> using human-like(?) reasoning means cutting down the search space by having some sort of insight or intuition which allows you to prune branches from the entire tree
So according to the OP there's a search of a tree and it also uses pruning btw, so I want to know what search they mean. Why are you asking?
Sorry, that's just egregious abuse of terminology. Either you're pruning the branches of a tree (or other graph, potentially) when you're doing a tree search, or you don't do pruning. Specifically advice taking (if that's what you mean) is not pruning.
It's a direct quotation from your comment, not a "scare quote". How does "implicit search" do pruning? Can you explain?
Edit: Turns out "implicit search" is actually a thing in the literature, albeit introduced in a single paper I could find that claims Diffusion Modelling does it:
Execution is not the code, but how do you decide to do every part. "Idea does not matter, execution does" always meant: "big generic ideas don't matter, it is how you organize it in the myriad of details it is composed of (in a given incarnation of the general idea) that matters."
But why aren't lower level "ideas about execution" more clearly expressed in code? Or at least pseudocode?
Many of us would struggle with our reading comprehension of an English description of an algorithm. But the pseudo / actual code? Much easier to reason about.
I'm not saying you should review even 90% of code. But it feels like an arbitrary line to say "never look at code". It's like never taking a wrench to a car to look at what's built to see whether you trust the car factory.
Because Redis is not "my project", it is a piece of software many relies upon, so I use, for that software, what the community at large agrees to be ok: AI-assisted coding with human careful reading and evaluation of every line.
One thing I've been thinking about is what a new SDLC ought to look like in the age of agentic engineering, since the "old" workflows are likely suboptimal (e.g. PR reviews). Have you given it some thought?
Thanks! And sorry for not yet merging many of those. The problem is, I'm dealing with tensor parallelism for the CUDA and Metal-RDMA fork right now, so was not albe to care about PR / issues for a lot of time.
I don't want to sound overly dismissive, but for quite a few practical cases pipeline parallelism with micro-batches will be a likely win over tensor parallelism. Of course this inherently comes with a requirement to take batched inference seriously for local use, at least in special cases where this doesn't put too much of a requirement on memory capacity. By comparison, tensor parallelism is probably good wrt. making memory- and KV-cache hungry models like GLM 5.2 and perhaps Kimi genuinely viable in a non-trivial local inference scenario.
Oh that’s great - I’m generally more interested in single machine use cases but the RDMA stuff is super interesting.
I do wonder if you are planning to eventually delegate some of the merging responsibilities? I see some other projects follow a similar pattern when they grow to a certain size.
Sure, it costs less, and AWS is in a dominant position. Users here are playing the side of the bully since they don't care about what is right and wrong with the hyperscalers. "BSD is better than AGPL!" And give money to the wrong side of the history. Nor that I expected anything better, the single person has a given sensibility, the mass, as a whole, do whatever is in a given moment convenient or believed to be more pure (license wise). However thanks to that, you will see how little progresses we will have (and we are having) in the space of open source system software with very open licenses. Developers of software mostly are not happy to bring OSS to the success to see them used by hyperscalers to capture all the value. However I did it again, with DwarfStart, to release code under the BSD license: even in the current situation, I think it is better to give back than to have a personal gain, but this is a position that very little folks can afford to take.
However: this conversation is completely out of topic but people instead of talking about AI and code, which is a tabu, will move the conversation to personal attacks and shit like that.
I wish you'd chosen a "non free" fair source or open core license from the start.
Amazon has stolen enormous wealth from you and your collaborators.
People cheer for the hyperscalers even though AWS and GCP are not at all open source themselves, charge absurd margin, and do everything in their power to lock you in.
It's really unfortunate.
Thank you for Redis.
Hopefully AI gives them extra competition. There doesn't seem to be a moat for them yet apart from distribution. Hopefully that holds. The world needs competition and less concentration of power.
In order to really leverage a nonfree (proprietary) or more-free (AGPL/SSPL) license you have to have a substantial thing to protect. If you try to protect something trivial, your competition will just implement it themselves, unless your price is low enough to make that not worth it. Redis is relatively trivial, it is a REmote DIctionary Service. Amazon could have written their own Redis quite easily.
They didn't, because the idea is sufficiently non-obvious, but ideas are protected by patents, not copyright.
Even RMS recommended using LGPL in some cases to maximize overall freedom by not making your competitors copy it. In the case of Redis, GPL probably would've maximised freedom (but not revenue) as Amazon still could've used it and released any changes they made.
Valkey has diverged from Redis, gaining features like vector search and multithreading.
Thanks, I believe that as a whole choosing the BSD created a more positive effect, so I'm happy with that. It is just that it is really unfair to read a comment where people use ValKey to accuse you of AI slop :D It means that our community, and this site itself, is at this point really low quality. This will in turn discourage the many great folks that are here. A replacement is needed. But TLDR, I would release Redis again with the BSD license if I could go back in time.
Do you understand Redis and ValKey have mostly overlapping code bases? And of that intersection, a big part of the code was written by myself by hand. So no, that's not the case. Also as I wrote in the blog post, Redis is currently not using AI if not as AI-assisted coding. Of course I'm not writing this reply for you, since I believe if you write a comment like that, you are part of that HN slice that makes this site at this point a slop place (no need for AI for very low quality), but for others that may find this information useful.
> 1. Compute costs collapsed since the advent of Cloud and yet hyperscalers still have fat margins.
Cloud opposes switch inertia. To setup a complex system in a different environment is a complex operation. Changing AI provider is switching an endpoint.
Soon decent speed across two Mac Studios with 512GB of RAM.
reply