Counter-Signal
Lab — early draft from Era Haus

AI Wins Where the Answer Is Cheap to Check

Jul 27, 2026Counter-Signal

AI now makes original scientific discoveries, this week's consensus holds, and it is half right. A frontier model did produce a genuine counterexample to a mathematical conjecture open since 1939, something no human had found in 87 years, and it withstood independent checking. The discovery is real. Where it applies is narrower than the coverage suggests: the machine wins fastest wherever the answer is cheap to verify.

Give the claim its strongest form, because the event earns it. On the evening of the World Cup final, a Harvard number theorist, Levent Alpöge, posted that the Jacobian conjecture was false and attached the proof: a polynomial map, produced with Anthropic's Claude Fable 5, that meets every condition the conjecture demands and still fails to be invertible. Other mathematicians reproduced the computation within hours. The result was a genuinely new mathematical object, one that had eluded the entire field for almost nine decades. The comforting line that AI can speed up an expert but cannot do the original thinking took a real hit.

Look at how the discovery actually happened, because the mechanism is the whole story. The human contribution was the framing. Alpöge chose a specific, precisely stated problem that generations of mathematicians had judged worth attacking, pointed the model at it, and recognised a valid counterexample when one appeared. The reason the claim held up by the next morning is that a counterexample to this conjecture is cheap to check: compute one determinant, test whether two inputs share an output, and the matter is settled. Alpöge is himself affiliated with Anthropic, the maker of the model, and the result survived that conflict of interest because the verification was independent and nearly free.

This is the shape under most AI research headlines of 2026. The domains where models are racing ahead, competition mathematics and formal proofs and software engineering, are the ones where a candidate answer can be graded automatically at almost no cost. A machine refuted the roughly 80-year-old Erdős unit-distance conjecture earlier this year in the same manner, by generating a construction that either satisfies the constraints or does not. The current frontier now scores around 96% on SWE-bench Verified, a standard test of whether a model can fix a real, logged software bug, because a fix is confirmed the instant the tests run. When a problem can be stated exactly and checked mechanically, a system that generates and tests millions of candidates holds an enormous edge, and that edge is widening.

The mathematicians closest to this see the same boundary. In June, an international group meeting in the Netherlands published the Leiden Declaration on artificial intelligence and mathematics, endorsed by the International Mathematical Union and signed by more than a thousand people within a day. Its central worry is a verification problem: AI-generated proofs are arriving faster than the field can independently check them, which threatens the reliability that makes a proof worth trusting. Even the discipline being disrupted has defined its own challenge as a matter of verification.

For an operator, that boundary is where the whole question turns. Most decisions a business makes have no cheap verifier. Whether a market is worth entering, whether a candidate is the right hire, whether a strategy will hold, whether a message will land, these answers come back in quarters or years, if they resolve cleanly at all. A model can now generate a hundred plausible answers to each of them, and it still cannot tell you which one is correct, because nothing grades the answer at the moment it is made. Where verification is expensive or absent, frontier reasoning does not transfer the way it does in mathematics. It raises the value of the people who can pose the right problem and judge an answer that no test can score.

This is the Jevons paradox applied to thinking itself: when a resource becomes cheap, total use of it climbs, and the bottleneck shifts to whatever that resource still depends on. We argued a version of it for labour in Jevons Won't Save the Pipeline You Just Broke, where the gains from automation accrue at the top of the skill ladder, to the people who direct the tool, while the routine rung thins out. Cheap expert reasoning does the same to cognitive work. Its scarce complements are problem selection and verification, and a flood of almost-free machine answers makes both worth more. The person who can choose which problem deserves the effort, and confirm when it has been solved, becomes the constraint.

So the operators who should be uncomfortable are the ones whose expertise is producing the answer to a well-specified, checkable question: the tax return that either reconciles or does not, the code that either passes its tests or fails them, the analysis with one correct output. That work is entering the zone the models are saturating, and pricing it as scarce is a mistake this year. The ones who can hold their position sell something no cheap test can grade: which problem is worth the firm's attention, which answer to trust when the evidence is ambiguous, which risk is worth carrying. That is the judgment layer we have called the durable asset in defensibility in the AI era, and this week sharpens the definition. Build where verification is hard, and treat the rest as a cost that keeps falling.

The consensus has the discovery right and the lesson half-drawn. A machine did produce original mathematics that no human had reached in 87 years, and anyone still insisting AI cannot do creative expert work is now arguing against a verified counterexample. But the breakthroughs cluster on one side of a line, and the line is verifiability, not intelligence. Machines are winning wherever an answer is cheap to check, and they still need a human wherever it is not. So the question for the next year concerns something else entirely: how much of what your business sells is a checkable answer, which is commoditising now, and how much is a judgment no test can score, which is where the value is moving. In a flood of cheap answers, the expert who can choose the problem worth solving and validate the result becomes the scarcest input you have.