Periodic Neon Outperforms GPT-6 Astra in XRD Analysis Using 1,300 H200 GPUs

icon MarsBit
Share
AI summary iconSummary
Periodic Neon outperformed GPT-6 Astra in XRD analysis with a 55.3% success rate. The model, trained on 1,300 H200 GPUs, surpassed Astra, which utilized over 100,000 Grace Blackwell GPUs. FrontierXRD evaluated 134 complex problems. Periodic Labs, co-founded by Liam Fedus, is building a materials laboratory for AI training. On-chain analysis reveals growing interest in AI-driven material discovery. Market analysis indicates strong potential for AI in scientific research.

Using only 1,300 H200s, outperformed GPT-6 Astra on a demanding scientific benchmark.

Recently, Liam Fedus, former Vice President of Research at OpenAI, posted on X to unveil Periodic Labs' first model: Periodic Neon.

X-ray diffraction

The most compelling part of the post is this sentence:

Using only 1,300 H200s and several months of experimental data, Neon outperformed GPT-6 Astra on their internal benchmark.

X-ray diffraction

Real-life footage from the periodic Menlo Park High Magnetic Field Materials Laboratory, with a robotic arm operating above a sample.

In the image above, rows of sample trays hold newly synthesized candidate materials, which will be placed into an X-ray diffractometer to capture their "fingerprints," then converted into spectra that Neon can read.

The lab runs 24/7, synthesizing, measuring, and feeding data.

Fedus said their first targets would be the toughest materials in superconductors, magnets, and semiconductors.

Why would someone with a background in large models go into materials science?

Fedus was one of the core creators of ChatGPT and led post-training at OpenAI, making him the most knowledgeable about what makes the model stronger.

When he founded Periodic last year, he made it clear:

The total amount of text on the internet is approximately 10 trillion tokens, and the best state-of-the-art models have already read through it all. For AI to advance further, it needs new data not available on the internet.

Where to find it? The lab.

So, he brought on Ekin Dogus Cubuk, head of chemistry and physics research at Google DeepMind, and from day one, built his own lab and generated his own data—with one goal: to create an AI scientist.

A year later, the first answer has arrived.

Neon, with approximately 1 trillion parameters, achieved a success rate of 55.3% on a challenging set of X-ray diffraction (XRD) analysis tests, outperforming both GPT-6 Astra and Claude Fable 5.1.

In terms of parameters, according to industry rumors, Astra has parameters in the trillions, while Neon has only 1T;

In terms of computing power, according to Jensen Huang’s public statements, Astra is backed by over 100,000 Grace Blackwell GPUs; the peak scale of Neon’s final training amounted to only 1,300 H200 GPUs.

This contrast quickly sparked a frenzy within the community.

Jeff Dean immediately commented under Fedus's post to offer congratulations.

X-ray diffraction

1,300 to 100,000—behind these numbers lies another path to making AI even stronger.

A test paper that raises the score from 2.7% to 55.3%

The problem Neon solves is called X-ray diffraction analysis (XRD).

What is X-ray diffraction?

For example, when a beam of X-rays strikes a crystal, it produces bright peaks at specific angles. The positions of these peaks vary depending on the material, much like a fingerprint.

Whether the synthesized material is a superconductor, magnetic material, or semiconductor depends first on whether this "fingerprint" matches.

X-ray diffraction

The X-ray diffractometer in Periodic Lab. The XRD data that Neon learns is generated by this type of instrument.

The problem is that laboratory-synthesized samples are often highly impure, with multiple substances overlapping, like hundreds of fingerprints smudged randomly on the same piece of paper.

Materials scientists typically need to open software, search databases, and compare papers, going through a lengthy process to identify what’s inside.

Periodic selected 134 of these "even experts find challenging" difficult problems to form the benchmark FrontierXRD.

Neon earned 55.3% on these 134 questions.

All models use Periodic’s own scientific tool environment, Periodic Harness.

Compared to the solution based on Claude Code with standard XRD tools, Periodic Harness achieves a success rate 3.8 times higher under similar per-item costs.

X-ray diffraction

AI research is half about intellect and half about the tools at hand.

To prevent the model from memorizing answers, the team also prepared an additional 198 questions, whose chemical systems were deliberately excluded from the training data. Neon still outperforms a host of state-of-the-art models.

However, the official team also acknowledged that this set of questions is simpler than FrontierXRD.

Neon is also more cost-effective.

At the highest tier, Neon achieves a 55.3% success rate with a cost of approximately $4 per question; GPT-6 Astra costs over $7 per question but performs worse than Neon; Claude Fable 5.1 has a cost similar to Astra but a success rate of only about 40%.

The internet is about to be fully consumed by AI; they want to build an endless "well of data."

Setting aside the benchmarking, what Periodic truly wants people to see is the high-throughput materials laboratory it has built in Menlo Park.

X-ray diffraction

Periodic Lab footage: Robotic arm picking up a sample and placing it into the equipment.

From the earliest experiments to today’s continuous 24/7 operation, Periodic has broken down materials discovery into a cycle:

First guess what needs to be done, based on predictions of material stability and properties;

How to synthesize it again;

Finally, clarify exactly what was achieved and whether the nature is correct. Obtain the results, revise the hypothesis, and go through another round.

X-ray diffraction

The periodic material discovery cycle: what to do, how to do it, what was done—three sections seamlessly connected.

With each cycle, the lab produces a batch of fresh, private data with physical feedback—this data is the raw material for Neon.

The periodic training corpus includes academic literature, code, and experimental data, and its size is currently doubling every month.

They also found that pre-training with supervised learning to instill scientific knowledge leads to better subsequent RL performance.

Moreover, during the experiment, the GPU doesn't sit idle—the model can use existing experimental data to continue analysis and improve predictions.

The title of that research blog post by Periodic is "Nature Is Our Learning Environment."

It targets an unavoidable hurdle in the AI community:

Over the past few years, scaling has relied on accumulating ever more internet data. However, text on the internet is nearly exhausted by state-of-the-art models.

Next, whoever has an environment capable of continuously generating new data gains another path to scaling.

The internet is a data mine that depletes with every bit extracted; a laboratory is a well that continuously yields water.

Nature doesn't provide a standard answer; Periodic hired an AI referee.

Moving RL into the physical world presents a completely different level of difficulty.

In the digital world, RL has enabled agents to write nearly all code and solve long-standing mathematical conjectures—by deploying vast numbers of agents to tackle problems that are fast to solve and automatically scorable.

In the physical world, all three are stuck:

Agents can't be added on a whim; each additional experiment requires more electricity, equipment, and engineering effort; a single experiment may take several days, and the results are often ambiguous.

The last point is the most critical. Periodic brilliantly stated in their research blog: Science can be falsified, but it is not easily verified.

An XRD pattern cannot be reliably assessed for success at low cost by simply checking the fit; expert evaluation is still required to determine whether each phase is supported by the pattern and whether the results are chemically plausible.

The periodic approach involves distilling experts' judgments into arbiters.

They enlisted materials experts to label thousands of XRD patterns, then used these labels to calibrate an LLM judging panel composed of Opus 5 and GPT-5.6 Sol; three sets of numbers are particularly telling:

Among experts, the agreement rate is 77.2%;

LLM judge and individual expert, agreement rate of 74.6%;

LLM judge and expert consensus, agreement rate of 84%.

X-ray diffraction

In other words, the reliability of this AI referee is now approaching the level of agreement among human experts.

Turning the ambiguous into trainable rewards is a core technical strength of Neon.

1,300 to 100,000—how exactly were these cards saved?

1,300 H200 GPUs represent the peak number of GPUs running simultaneously during the final training phase, which includes both intermediate training and reinforcement learning.

Astra's more than 100,000 Blackwell GPUs correspond to large-scale training of a cutting-edge general-purpose model.

So, 1,300 compared to 100,000 doesn't yield an efficiency multiplier, but it at least shows that Periodic has maximized the use of these 1,300 cards.

Reinforcement learning on scientific tasks has a peculiar quirk: when the model solves a problem, it may spend over an hour thinking, consulting references, and running tools—but updating the model based on the result takes only a few minutes.

The periodic approach involves separating problem-solving and learning, assigning each to its own GPU so they do not wait for one another.

Here's a real story of a failure.

Their earliest RL loop bundled together the code for training, inference, and the model, with no isolation.

Until one time, the model-generated code requested 80 GB of memory, causing the entire task to crash immediately.

Thus, we developed our own sandbox, pbox.

It directly uses idle CPUs on the GPU node where the RL task is running to execute scientific tools, with all data remaining within the cluster.

Compared to the custodial sandbox service they tested, data transfer is 4.5 times faster, and throughput is 3.3 times higher.

Train idle computing power and allocate it to scientific computing tasks such as physical simulations, maintaining the cluster's utilization above 95% year-round.

The concept behind Periodic is straightforward: since you can't outmatch the giants in total computing power, focus on maximizing scientific output per GPU hour.

The flywheel is still the same, but the fuel has changed.

Neon has been deployed in Periodic’s laboratory to analyze real experiments and help the team discover better superconductors and magnets.

What step is still missing?

Returning to the three-stage cycle mentioned earlier, Neon has now demonstrated only the final stage: understanding what the laboratory actually produced.

Regarding having AI oversee entire research projects, design synthetic pathways, and determine which experiments to run next, Periodic's official stance is that they are expanding this approach in those directions.

Fedus’s wording is also restrained: the model currently “helps us decide,” rather than making decisions for scientists.

More precisely: The flywheel of experiments, data, and training has already started turning, but two key links remain unconnected to achieve full autonomous closure: deciding what to do, and how to synthesize it.

Periodic also doesn't intend to deny the value of hashing power.

The official stated that scaling training compute to the level of today’s most advanced models could unlock stronger scientific capabilities. They also mentioned that, in the future, a more powerful open-weight model could be used as the foundation.

GPUs are still important, but GPUs alone are no longer enough.

Looking back at Fedus’s resume: in the GPT-4 technical report’s list of contributors, his role was “Data Flywheel” lead.

Three years later, he moved this flywheel from internet text into a real physical laboratory.

The flywheel remains the same; what has changed is only what is fed into it: from text written by humans, to answers provided by nature.

The next phase of scaling may no longer be about who can buy a hundred thousand chips, but rather who can continuously feed real-world data into their models without pause.

Reference materials:

https://periodic.com/news/building-labs-that-learn

https://periodic.com/news/nature-is-our-learning-environment

https://periodic.com/news/ai-infrastructure-at-periodic

https://x.com/LiamFedus/status/1973055380193431965

https://a16z.com/announcement/investing-in-periodic-labs/

This article is from the WeChat public account "New Intelligence Yuan," authored by ASI Revelation.

Disclaimer: The information on this page may have been obtained from third parties and does not necessarily reflect the views or opinions of KuCoin. This content is provided for general informational purposes only, without any representation or warranty of any kind, nor shall it be construed as financial or investment advice. KuCoin shall not be liable for any errors or omissions, or for any outcomes resulting from the use of this information. Investments in digital assets can be risky. Please carefully evaluate the risks of a product and your risk tolerance based on your own financial circumstances. For more information, please refer to our Terms of Use and Risk Disclosure.