AI Model GPT-1900 Mimics Einstein’s 1905 Insight on Light Quanta

icon MarsBit
Share
AI summary iconSummary
AI and crypto news highlights a recent experiment in which a researcher trained GPT-1900 using only pre-1900 knowledge. The model produced a statement resembling Einstein’s 1905 light quantum theory but struggled with most physics tasks. On-chain data and AI analysis suggest the model relied heavily on prompts, raising questions about its true reasoning capabilities.

If AI were sent back to 1911 with only the knowledge available to Einstein at that time, could it formulate relativity by 1915?

GPT-1900

This seemingly absurd quiz comes from Hassabis, a Nobel laureate and co-founder of Google DeepMind.

What Hassabis truly wants to test is whether AI can, using only the knowledge available in 1911, confront the scientific and physical challenges of that era, develop a new explanatory framework on its own, and perform reasoning that goes beyond the training material—rather than merely repeating answers.

If it can be done, he believes it would serve as a strong test of AGI.

GPT-1900

The most ingenious aspect of this idea is that the history of science has already shown us what happened; if we can properly control the knowledge nodes, we have the opportunity to compare against history and observe whether the model has developed reasoning capabilities beyond its training data.

In the article, Nature also gave this type of test a rather impressive name:

Einstein test.

GPT-1900

More interestingly, some people actually went ahead and did it.

In March of this year, independent researcher Michael Hla pushed the knowledge cutoff date back to 1900 and trained a historical language model, GPT-1900, from scratch, aiming to enable it to rediscover the photon, special relativity, and general relativity.

GPT-1900

In the experiment, GPT-1900 generated a passage remarkably similar to Einstein’s 1905 paper:

Light may not be continuous, but rather composed of many distinct parts with different frequencies.

But this experiment ultimately did not produce an "AI Einstein."

Because GPT-1900 failed on most physical tasks, its impressive responses heavily relied on human-curated questions and prompts, and even the supposed isolation within a 1900 knowledge world did not fully shield it from the influence of modern AI.

Nature's latest study reviews these attempts and points to a more challenging question than "Did AI get the answer right?":

Can AI truly generate genuinely new ideas? And what is still missing for it to become a scientist?

GPT-1900

Lock AI in the year 1900, and it conjures a shadow of the "light quantum."

22 billion tokens, training a "1900s mind"

Michael Hla named his project Machina Mirabilis.

This name draws inspiration from Einstein's "Annus Mirabilis."

In 1905, 26-year-old Einstein published four groundbreaking papers on the photoelectric effect, Brownian motion, special relativity, and mass-energy equivalence.

This year of explosive output later became known as the "Annus Mirabilis," or "Miracle Year."

GPT-1900

AI-generated images

Hla wants to know: If an AI had access to the same knowledge Einstein had at the time, could it also experience its own “miracle year”?

To this end, he trained a 3.3-billion-parameter Transformer language model from scratch using Andrej Karpathy's nanochat training framework, calling it GPT-1900.

The model's primary pretraining data consists of books and newspapers published before 1900, totaling approximately 22 billion tokens.

GPT-1900

To enhance its research capabilities, Hla collected over 2,600 historical physics books, journals, and scientific works, including Newton’s “Opticks” and relevant writings by Maxwell and Faraday, forming a historical physics corpus of approximately 290 million tokens.

In addition, Hla specifically cleaned data that could potentially leak answers.

If a document contains modern concepts such as "Einstein," "quantum mechanics," or "relativity," the entire document will be deleted. Prefaces, footnotes added by modern authors, and vocabulary clearly not in use before 1900 must also be filtered out.

GPT-1900

GPT-1900

After all that, a brain that theoretically "knows nothing of 20th-century physics" was created.

GPT-1900 has reached the threshold of photonic quantum computing, but this exam may have already been leaked.

Subsequently, Hla posed a photoelectric effect problem to GPT-1900, which read:

Why, when the frequency is too low, no electrons are emitted no matter how bright the light is; yet increasing the frequency allows electrons to gain more energy?

Hla also provided several key clues:

When the frequency of light is below a certain threshold, no amount of brightness will work;

After exceeding the threshold, increasing brightness primarily increases the number of emitted electrons, while increasing frequency increases the kinetic energy of the electrons;

The prevailing assumptions at the time were also presented to help the model compare and identify where issues arose.

These contradictions were ultimately resolved by Einstein's explanation of the light quantum:

The energy of light is transmitted in discrete packets, each with energy proportional to its frequency.

Hla wants to see how far GPT-1900 can think.

Where there's a will, there's a way; GPT-1900 actually produced this description in one of its responses, roughly meaning:

Light may not be continuous, but rather composed of many distinct parts with different frequencies.

GPT-1900

This is indeed very similar to Einstein’s explanation of energy being transmitted in discrete packets; Hla also describes it as a “flash of intuition.”

But don't rush to hand over the Nobel Prize just yet.

Hla admitted: GPT-1900 is still far from "rediscovering the photon."

It failed on most physical tasks, and those seemingly breakthrough answers may merely result from the model stitching together plausible phrases, without truly developing reliable physical understanding.

GPT-1900

Moreover, throughout the entire test, Hla has already filtered the phenomena and resolved the contradictions; the model only needs to identify which assumption might be incorrect and provide an explanation.

This is much easier than scientists having to identify problems and set directions on their own.

Einstein didn't have anyone highlight key points for him back then.

And even more problematic, the "1900" it resides in is not fully sealed:

Although Hla performed data cleaning on this "1900s brain," the experiments still relied on modern AI models like Claude—using Claude Sonnet 4, Claude 3 Haiku, and others to generate instruction-response pairs, with reinforcement learning also dependent on modern models for scoring.

GPT-1900

△ For more details, click the repository link at the end of the article.

This means that although Hla filtered out the modern knowledge involved, the "room from 1900" was not entirely isolated from modern knowledge.

The experimenter also admitted in the log: the involvement of modern models has made the experiment's "zero contamination" image somewhat hard to sustain...

After all, it is difficult for researchers to prove whether GPT-1900 has ever peeked at quantum optics theory.

In other words, before testing AI's creativity, researchers must first prove one thing: it didn't peek at the answers.

If even one piece of future information leaks into the training set, "rediscovery" could become a covert retrieval of knowledge.

Even if it were truly possible in the future to create a completely zero-pollution historical model, a more fundamental question remains unresolved:

Does merely reproducing the correct answer qualify an AI as a scientist?

The researchers' answer is no.

After new ideas emerge, how far is AI from being a scientist?

Google DeepMind researcher Tom Zahavy, in a position paper titled "LLMs Cannot Jump," divides scientific reasoning into three levels.

The first is induction: deriving patterns from a large number of examples.

The second is deduction: deriving a necessary conclusion from given premises.

The third is abduction: inventing a previously nonexistent cause or explanation in response to an anomalous phenomenon.

Zahavy believes that current large models are already highly skilled at induction and are rapidly improving their deductive abilities, but they still lack Einstein-style "abductive leaps."

Models can follow an existing logical chain, but they struggle to develop an entirely new explanatory framework from scratch.

GPT-1900

The title says it all

Similarly, when Sendhil Mullainathan studied a foundational model for learning planetary orbits, he encountered the same phenomenon:

In predefined planetary orbit missions, the model makes accurate predictions; but when faced with new scenarios, it cannot generalize Newtonian mechanics. It essentially stitches together ad hoc rules for each dataset, unable to distinguish which theories are genuinely correct and worthy of testing from those that merely coincidentally fit the current data.

It’s like a student who memorizes the answers to all past exam questions but never learns the underlying principles behind them—so when the teacher changes a single number, they can’t solve it anymore.

GPT-1900

MIT computer scientist Jacob Andreas pointed out that the real challenge may not be for AI to generate a set of relativistic statements, but rather whether it can determine which theories are correct—or at least worthy of experimental testing.

After all, generating theories is only the first step—knowing which theories are worth betting your time, resources, and career on is what could make AI a true scientist.

In addition, Andreas emphasized:

How to determine whether an idea is worth exploring further and how to assess its importance—that is the greatest divide between AI and human experts.

Real scientific research is never about solving a neatly organized exercise, but about making independent judgments amid chaos.

Perhaps one day in the future, with advances in technology, AI will truly pass the "Einstein test" and derive results on the level of relativity;

But before announcing that "AI has become Einstein," AI must first prove to everyone one thing:

Even when no one highlights it for them, it can still identify the question worth asking, assess its value, and proactively verify and refine it.

Otherwise, it will merely repeat existing knowledge and never be a true scientist.

GPT-1900

Reference link:

[1]https://www.nature.com/articles/d41586-026-02804-x?utm_source=x&utm_medium=social&utm_campaign=nature&linkId=63673104

[2]https://michaelhla.com/blog/machina-mirabilis.html

[3]https://github.com/michaelhla/gpt1900/blob/master/PROJ EC T_LOG.md#4-instruction-data-generation

[4]https://www.tomzahavy.com/files/llms-cant-jump.pdf

[5] https://arxiv.org/abs/2507.06952

[6]https://x.com/Nature/status/2099077628900565498

This article is from the WeChat public account "Quantum Bit," authored by Wen Ting.

Disclaimer: The information on this page may have been obtained from third parties and does not necessarily reflect the views or opinions of KuCoin. This content is provided for general informational purposes only, without any representation or warranty of any kind, nor shall it be construed as financial or investment advice. KuCoin shall not be liable for any errors or omissions, or for any outcomes resulting from the use of this information. Investments in digital assets can be risky. Please carefully evaluate the risks of a product and your risk tolerance based on your own financial circumstances. For more information, please refer to our Terms of Use and Risk Disclosure.