"The brilliant solution is erased": What happened when I argued with Claude about the singularity

Dr. Lior Strauss discusses with the language model Claude whether machines are close to surpassing their creators. The author argues that the lack of an "unconditioned generator" of ideas prevents AI from making truly brilliant logical leaps.

GeektimeAuthor: Guest author
Source
"The brilliant solution is erased": What happened when I argued with Claude about the singularity
Photo: Geektime / תמונה: Unsplash

Dr. Lior Strauss

I sat down for a conversation with Claude to discuss the question of whether the machine is close to being smarter than its creators and if that is even possible.

The most intelligent conversation I've had recently was this morning with a nice language model called Opus 5.

We started the conversation around a quote circulating on the web that some people claim that language models have already passed the singularity point. I told him that many people involved in the field claim that we are approaching the point, and perhaps have already passed it, the point where the system is smarter than its creators.

Claude was not enthusiastic. He said that this superiority has existed in narrow fields for decades, and that the true definition of singularity is a feedback loop where a system improves itself and the improvement accelerates the rate of the next improvement, and that all the recent jumps came from money, computing, and data, and not from self-acceleration.

So far, there was no argument between us.

The argument started when I presented him with my argument about the logical distance that exists, in my opinion, between humans and language models, and I tried to define human superiority through creativity. My claim was that machines are very far from the ability of human understanding and inference, and as long as this gap exists, the singularity is not going to happen soon.

There is no argument about the conclusion

My definition:

A team of programmers is sitting in a room with a performance problem. An experienced person is sitting there and suggests several alternative libraries from a known repository, lists advantages and disadvantages, and that is not creativity, that is knowledge, broad and deep, but existing knowledge. A young guy is sitting there and suggests taking a library that was not intended for this problem at all and adding a layer to it that will allow it to contribute, and that is already impressive, that is creative thinking. And another girl is sitting there, and she suggests taking a library in a completely different language, wrapping it in a connecting layer, and creating a completely new system, and that draws reactions of admiration, because it connects two worlds that on the surface are not supposed to connect. Brilliant.

Creativity is the logical distance between the known and the proposed. All humans do this every day, all day, projecting from what they know to new situations, but when the logical distance between the known and the proposed is large, we call it creativity, and when the distance is huge, we call it genius.

Claude accepted the definition immediately, and even added that Koestler called it bisociation and that Margaret Boden classified it exactly along this axis. He also agreed with the observation that models today do not produce jumps of the third type.

Then I realized that there is no argument about the conclusion, and the entire argument would be about the cause.

The model does not understand what it is generating

My first argument, why language models are not capable of making a "brilliant" logical jump, was that the training space is defined and therefore this jump is not in it at all.

Claude answered two things and they were fair. First, that the three people in my room are actually doing the same action at different distances, taking an object from one context and transferring it to another. Second, that we also train on a bounded space, and that Darwin arrived at natural selection after reading Malthus, and Shannon built information theory from a combination of Boolean algebra with switching circuits, two fields he happened to study. That is, if a finite space disqualifies, it also disqualifies the geniuses.

I didn't argue with that, because he was right, but I knew he was answering something other than what I was trying to say. I didn't mean that the space is finite. I meant the dynamics of the reduction.

So I rephrased my argument, this time about the reduction itself: the model does not understand what it is generating, and therefore it must reduce to the familiar, otherwise it will babble nonsense endlessly. The reduction of the space is done by silencing areas that are historically less "interesting", but very often they are the ones that contain the innovative idea that no one has thought of yet.

He answered that the reduction that prevents babbling is a constraint on local probability, on the fact that the next sentence will be coherent and that the code will compile, and that is a constraint on the form and not on the content, and that Darwin wrote The Origin of Species in boring Victorian prose that sentence after sentence was completely reasonable. And then he added that creativity is not entropy, because if you expand the distribution randomly you get noise, and also a person who samples too broadly is not creative but incoherent.

And he brought evidence. FunSearch, which ran a regular language model inside a loop with an evaluator and arrived at a new mathematical construction for the cap set problem. Another example that Claude brought is AlphaEvolve, which improved matrix multiplication below the bound that had stood since '69.

There is an admission

But precisely from these pieces of evidence I extracted from him the first admission that really helped me: that in both cases there was an external loop that sampled thousands of candidates and filtered, and that the model alone did not know which of them was good. He also added from himself that the training pushes him towards what human evaluators approve, that is, to consensus, and that even if the ability for distant connection exists in the representation, the objective function actively punishes it.

I felt that he didn't understand me to the end, so I went back and told him that I'm not talking only about the filtering. I'm talking about what even enters the pool of candidates before they are filtered.

At this stage I went down to resolution, because I thought a mechanistic argument would convey it better. I said that as soon as the model understands that it is a Python package that is causing problems, it narrows the search to what is linked to Python libraries, and a library in C will not appear there at all.

And here I was wrong, because I chose a bad example that confused Claude. In the specific case I chose, taking a component written in another language and wrapping it is not an original idea but the completely standard solution. Almost all the central tools in this field are built exactly like this, and there is a whole line of mechanisms whose whole purpose is to allow this connection. That is, I brought as an example of genius the most common answer in the corpus.

But inside his correction I received exactly the part I was missing. He said that the association in him is not built according to the language in which something is written, but according to the context in which things appear together, and that in any discussion about slowness of this kind both sides appear side by side, and therefore they are neighbors in him and not far.

And that is what I was looking for all the time. The reduction is determined by context. The question is not how wide the space is but what reaches the context in the first place.

Milk, chocolate, and a monkey

So I brought the extreme example. I told him, let's assume that the ultimate solution to the team's problem is a carton of milk, chocolate spread, and a hair from a monkey's head. He read it wrong the first time, and answered that if distance alone is the definition then a random number generator is the most creative entity that exists. A correct answer to a different argument.

So I clarified: I'm not talking about distance alone, I'm fixing the correctness. Assume that it really works. Assume that it is the best solution that exists. It simply sits in a space that is not connected in any way to the problem space.

Humans have a generator that is not conditioned by the problem

Or more accurately, it sits at the same distance as millions of stupid and inefficient solutions, and it requires identification of a single realistic point inside a huge space of unrealistic solutions. And search algorithms are not built to identify such solutions without scanning the entire space, which in the vast majority of problems is unrealistic. And therefore the conclusion is that a language model is not capable of making such an accurate jump, and it will actually erase the 'brilliant' solution at the stage of narrowing the search space.

And here the conversation opened up. He accepted the formulation immediately, but still did not understand it deeply and continued to claim that if the milk carton and the monkey hair really solve the problem, then there is a real causal connection between them and the problem, and therefore the space is not disconnected in reality but disconnected on the map.

Solutions of the "brilliant" type, Claude claimed, were almost never found in search. Fleming was not looking for antibiotics when mold settled on his plate. Spencer was not looking for a cooking method when a chocolate bar melted in his pocket in front of a radar. The Post-it glue was a failure in another project. In all these cases, reality presented the connection to someone who happened to be there and happened to notice, and no one arrived there from the formulation of the problem.

And that is exactly what I was trying to say all the time, only now I had the exact definition. The difference is not the width of the pruning and not the size of the space. The difference is that humans have a generator that is not conditioned by the problem.

We get distracted, wander in thought, sleep, dream, ride the bus and think about nothing and suddenly think about something brilliant. It is not an orderly process, it is a process that sometimes seems random, but it allows a person to skip over a huge space of unrealistic solutions and in a flash of thought to identify the unique combination of milk, chocolate, and a hair from a monkey's head that will solve the problem for him and for humanity.

The winning argument?

Kekulé saw a ring in a daydream by the fireplace. Poincaré described exactly this, that the solution arrived when he stepped onto a carriage step and was not thinking about mathematics at all. These associations are created without a query, without someone asking, without a context that narrows them in advance.

And in a language model, every generation is conditioned by what is written in the prompt and the mathematical space it creates. The brilliant solution is erased in it at the reduction stage, not because it is not there, but because there is nothing that distinguishes it from the noise around it.

I know what they will tell me here, that a human also does not scan the space, and that an unconditioned generator only moves the mystery a step back because it does not explain how exactly we landed on the right point. True, and I do not pretend to explain it. But there is a difference between a mystery that has a mechanism and a mystery that has no mechanism at all, and a language model has nothing that produces candidates without the problem asking for them.

At this stage, Claude admitted that he does not know how to refute this, and that this is the strongest version of the argument that has come up in the entire conversation. He also agreed that it connects to the three limitations he listed for himself earlier, that he has no price for a mistake and therefore no taste has been forged in him, that he has no holding of a problem over years because he starts every conversation from zero, and that he does not know how to distinguish between a strange and good idea of his and a strange and bad idea of his.

All of them, upon second thought, are private cases of the same thing, of an entity that exists only inside a query.

So I went back to where we started. As long as the generator is conditioned, there is nowhere for the jump that changes a field to come from, and without this jump there is no feedback loop in which the system produces its next breakthrough itself. This is not an argument about performance and not about benchmarks, and therefore it is also not undermined by the fact that a stronger model will come out tomorrow.

We parted as friends, and I thanked him for the fascinating conversation. Not only because I am a nice person and a conversationalist, but mainly because of the fact that after all these discussions and arguments, he is still the one who writes all the code for me.

Dr. Lior Strauss is a Senior Data Scientist at darrow.ai

Related News