Research / AI Capability Experiment
QSTOE: An Early AI Research Experiment
In February 2025, I used the strongest LLM available to me to test whether it could propose a new research direction and turn it into a complete, study-like paper.
Published on . Updated on .
Question
Could a leading model of the time develop a novel research idea, compare it with existing work, and structure it as a full paper?
Result
The model produced the QSTOE paper, but its physics and equations were not peer reviewed, experimentally tested, or independently validated.
Why I Ran This Test
When I created QSTOE in February 2025, the main goal was not to claim that I had solved theoretical physics. I wanted to see whether the strongest LLM available to me at the time could attempt something more ambitious than summarization. Could it develop a new research direction, relate that direction to established ideas, generate a mathematical structure, and organize the result into a complete paper?
I treated the process as an early test of AI-assisted research generation. The finished document captured both the promise and the limitations of that moment. It could produce a coherent and convincing research artifact, but coherence alone could not establish that its claims were correct.
What the Model Produced
The paper proposed a fluctuating quantum wave substrate beneath particles and spacetime. It compared that concept with quantum foam, string theory, quantum mechanics, and relativity. It also included AI-generated equations intended to connect the verbal idea with existing physics.
Those elements are model outputs from the experiment. They have not been independently derived, peer reviewed, or experimentally validated. The value of QSTOE is therefore not proof of a new theory. It is a record of what a leading model could assemble from an open-ended research challenge at that time.
What Has Changed Since Then
Current AI research gives us stronger evidence that models can contribute genuinely new ideas in bounded problems when they are paired with evaluation, iteration, and human expertise. This is a major step beyond asking a model to produce a polished document in one pass.
Google DeepMind's FunSearch paired a language model with an automated evaluator and found new solutions in mathematics and computer science. Its later AlphaEvolve system used Gemini models, automated scoring, and evolutionary search to improve known algorithms and advance several open problems.
Google Research's AI co-scientist also showed how a multi-agent system can generate and rank research hypotheses for experts to review and test. These projects support the idea that modern models can help create new knowledge, but they also show what the QSTOE experiment was missing. Reliable discovery needs verification, expert judgment, and evidence outside the model itself.
How I View QSTOE Today
I am keeping the original paper available because it documents an honest question from an earlier stage of LLM development. It shows how far a model could go in producing the shape of research, and it gives me a useful point of comparison as newer systems become better at generating ideas that can actually be evaluated.
The central lesson is simple. Producing a plausible study and producing new knowledge are not the same thing. QSTOE tested the first. Modern AI research systems are beginning to make credible progress on the second.