The real story is not that GPT-5.6 and Grok 4.5 independently killed the same math conjecture. The real story is messier: one AI-assisted claim targets the Dinitz-Garg-Goemans flow conjecture, while the Grok-linked Capy claim appears to target a separate Graffiti graph conjecture.
Dmitry Rybin's GPT-5.6 Pro claim is still worth your attention. Keep it in its lane, though. According to VibeMathed's July 2026 problem page and follow-up coverage by Sergey Nikolenko, Rybin, a Ph.D. graduate of the Chinese University of Hong Kong, Shenzhen, used GPT-5.6 Pro to construct a proposed counterexample to the Dinitz-Garg-Goemans cost conjecture for single-source unsplittable flow. The reported instance is small enough to sound almost suspicious: 7 vertices, three demands, a fractional cost of 58, and a minimum unsplittable cost of at least 60 under the stated capacity-violation condition.
That is the claim. It is not yet a settled theorem.
The difference matters because math news spreads badly when the names get blurred. The viral Grok 4.5 post that Elon Musk amplified was not, based on the public records I found, the same Dinitz-Garg-Goemans result. The Capy agent claim was about Graffiti Conjecture 284, a separate graph theory conjecture from Siemion Fajtlowicz's Graffiti program. AGNT Labs published a July 23 verification note saying Capy Build found a Hoffman-Singleton graph refutation of Graffiti 284 during a run on July 22. That graph has 50 vertices and 175 edges. It is not Rybin's 7-node flow example.
So don't file this as two frontier models independently solving the same old problem. File it as two AI-linked math claims arriving close together, then getting compressed by the internet into one cleaner story than the facts support.
The flow claim is concrete enough to check
The Dinitz-Garg-Goemans conjecture is about routing demand through a network. In fractional flow, you can split a demand across several paths. In unsplittable flow, each demand has to travel along one path. The famous 1999 result by Yefim Dinitz, Naveen Garg and Michel Goemans showed that fractional flows can be rounded to unsplittable ones while controlling congestion. The cost question is the sharper part: can you do that without increasing total cost?
Rybin's proposed counterexample says no. Based on the published summaries, GPT-5.6 Pro produced a graph in which the fractional solution costs 58, while every acceptable unsplittable solution costs at least 60. A gap of 2 is enough. You don't need a dramatic number to break a universal conjecture.
That is also why the peer-review gap is not the same kind of gap you'd have in a 70-page proof. A finite counterexample can, in principle, be checked by direct computation. VibeMathed says the result includes proof certificates, a verification program, machine-readable counterexample data and LaTeX source. If those artifacts are correct, the graph should survive inspection. If one hidden path or capacity condition was missed, it falls apart.
That's the work now.
The Grok story is a different graph
The Grok-linked claim has its own facts. AGNT Labs credits a Capy Build agent run on July 22 with finding that the Hoffman-Singleton graph refutes Graffiti Conjecture 284. The note says the conjecture would require 7 to be less than or equal to 4 for that graph, which is false. That is a clean refutation if the statement and verification are faithful.
Elon Musk's X post, captured in several indexed mirrors and roundups, said Grok 4.5 had solved a graph theory conjecture open for roughly 30 years. The original Capy post, as reproduced in those same records, was more specific: Capy was running on Grok 4.5 Medium, and the target was Graffiti Conjecture 284. That distinction is not pedantry. An agent using a model, tools and a workflow is not the same thing as a raw model solving a problem from a blank prompt.
Frankly, this is where founders and investors should pay attention. The capability is interesting, but the attribution discipline is the product lesson. If your AI science startup can't separate model, agent, toolchain, prompt history and verification, you're not reporting discovery. You're telling a marketing story.
For now, the cautious view is the strongest one. Rybin's GPT-5.6 Pro result may be a real counterexample to the Dinitz-Garg-Goemans cost conjecture. Capy, running on Grok 4.5 Medium, may have found a real refutation of Graffiti Conjecture 284. Both claims are recent. Both need expert checking. They should not be merged into one headline because both involve AI, graph theory and a 30-year clock.
If either result holds, it is a serious moment for AI-assisted mathematics. If both hold, it is bigger. But you don't get there by rounding off the facts.
Also read: Australia tells AI data centers to generate their own power and get creators' permission first • Meta AI arrives in Threads DMs for half a billion users as the inbox becomes its new battleground • Dario Amodei says he doesn't want open-weight AI banned but fears what China just released for free