Tag Archives: theoretical physics

Amplitudes 2026, Part II

This is a continuation of my conference coverage from last week. The same warnings apply: this is much more technical than my usual posts, readers beware!

Last week I covered the talks from Monday through Wednesday, so I’ll jump right in here with Thursday morning, where the first speaker was Francesco Riva, who after a brief ad for his new board game Tutti Quantum gave a review of positivity constraints, first-principles restrictions on quantum field theories based on their behavior at high energies. After covering some of the research program’s successes like arguments against Galileons and massive gravity, he talked about how new methods allow one to take into account the possibility of loops of massless particles, essentially by invoking a formal version of the idea that real experiments have finite size. He was followed by Grant Remmen, who talked about his work deriving string theory-like amplitudes from increasingly minimal assumptions. While one can always quibble with the assumptions they impose, I do find it encouraging that they are now managing to do this game with both gravity and gauge theory, and with five-particle amplitudes, not just four.

Paul Heslop then covered applications of bootstrap techniques to non-supersymmetric theories, where analytic superspace still finds a way to be useful. Tomasz Taylor covered progress in calculating Yang-Mills amplitudes in de Sitter space. Andrea Puhm and Nima Arkani-Hamed don’t have slides online yet: Puhm’s title suggests her talk was part of the celestial holography field, while Nima’s was likely similar to his talk at Lancefest, where he talked about calculations of amplitudes in the limit of a very large number of loops (represented in the field with a capital L) or very large numbers of particles (represented in the field with a lower-case n). Given the different context, I’m guessing he left out the self-effacing jokes where he was “little n” and Lance was “big L”, though I’m hoping he at least mentioned his students had been checking their results against what they called the “lanswer”.

The evening ended with a gong show, which for the non-initiated is a series of short student talks with a strict time limit (hence the gong). I’m not quite so intrepid as to read all of the slides for these, commenters who attended are welcome to highlight special examples.

Friday began with a talk by Agnese Bissi, who reviewed the connection between holographic correlators and amplitudes in AdS space. I liked her emphasis on this as a lab to find nice new representations of amplitudes, and her summary at the end of the current frontiers. Axel Kleinschmidt followed with a talk on one-loop string amplitudes, where he and his collaborators have gotten gradually more proficient at manipulating the rich structure of elliptic functions that make an appearance. Piotr Tourkine talked about his work using the S-matrix bootstrap to find scattering amplitudes in higher dimensions, a context where accounting for thresholds presented a new challenge. This is a problem people are approaching with genuine supercomputers, he quoted one calculation at 100,000 CPU hours. Lauren Williams is one of a small community of mathematicians who have been intrigued by the amplituhedron, her talk was a walk through a series of conjectures, some made by physicists, some by mathematicians, most with counterexamples found in the last few years.

Finally, Zvi Bern closed the conference with a talk on the frontier of amplitudes-based gravitational wave calculations, referring to it, probably to David Kosower’s annoyance, as 5PM. Some parts of this frontier have been calculated, but a few have proved hard going, bottlenecked by immensely challenging integrals, which go beyond the capabilities of publicly available codes. The new integrals have a variety of strange functions, including the elliptics and Calabi-Yaus I spent time on in my own career, as well as Heun integrals, which I imagine I will have to learn more about from the slides of Elliptics & beyond ’26, as I don’t remember people talking about them when I was in the field. New integration strategies have led to improvements in the public codes and seem to be making good progress, but Zvi highlighted that ideally they want to not need supercomputers at all, as they’ll need to go higher in loops to see effects from, for example, the deformability of neutron stars.

Amplitudes 2027 will be in Munich, with the summer school in Mainz under Stefan Weinzierl’s capable hands. I’m looking forward to seeing what the state of the art looks like then!

Amplitudes 2026

This week was Amplitudes, my old subfield’s big yearly conference. This year, it’s at Queen Mary University of London.

I’m too busy to attend these days, now that nobody is paying me a salary to do that sort of thing. But it’s still a good chance to keep up with the field, which helps me find stories. And I know I have a few readers who are interested. So I read the slides when I can, and fill you all in.

As of writing this, I’ve read through slides from the first few days of talks. I’ll likely post on the rest next week. As usual when I’m conference-blogging, this post will be quite a bit more technical than my average post, so readers beware: I’ll be mentioning a lot of amplitude-ish ideas without much explanation. I’m happy to explain in the comments if you’re curious, though!

Before I launch into talking about the content, I should mention something I’ve heard about the venue. Apparently this year registration filled up surprisingly quickly. The rumor is that the folks at Queen Mary weren’t able to find a venue with enough space to host the full community, leading to a smaller conference than usual. As the Amplitudes subfield grows, I suspect this will be more and more of a challenge. I remember back when I was organizing in 2021, rooms that could hold enough people were shockingly expensive to book. I wouldn’t be surprised if this becomes more of an issue going forward.

David Kosower opened the conference with a review of the state of the art in amplitudes for gravitational waves. I enjoyed an early slide showing how LIGO has increased in precision over the last ten years, which really helped illustrate the value that good theoretical predictions can bring as the experiment gets better. I also appreciated his attempt to get people to stop saying “post-Minkowskian” and start using a more normal term like Relativistic Perturbation Theory, though based on the other talks it doesn’t look like it’s catching on. After covering the overall state of the art (currently around four loops) he talked about his own work looking for ways to more directly get waveforms for orbiting black holes out of an amplitudes-style calculation, a theme his collaborator Donal O’Connell covered in more detail later that day. (The trick, apparently, is background fields!) Other gravitational wave-related talks came from Gustav Jakobsen, who covered four-loop results using the worldline method, Graham Brown, who explained how to calculate the Magnusian Jakobsen mentioned (it’s the log of the amplitude, basically), Canxin Shi, who talked about how a concept called Stratonovich-Weyl quantization helps explain structures that keep showing up in classical observables, Giulia Isabella, who covered a way to get six-loop contributions to gravitational waves via a wave equation, Lara Bohnenblust, who showed how to get strong-field information from a shockwave limit, and Gang Chen, whose slides were brief enough that I didn’t really get a clear idea of what he was up to. Rounding out the day were two talks that didn’t involve gravitational waves: Zihan Zhou, reporting on work with Nima on a water wave polytope called the hydrotope that they found a general formula for with help from Claude (the AI, not Duhr, to steal a recurring joke from Lancefest), and Michael Ruf, who probably snuck in due to the fact that he has done a lot of work with gravitational waves, but was reporting on a QCD calculation using a tool called Scorpio.

Tuesday had more of a QCD theme. Thomas Gehrmann gave the day’s opening review talk, where he pointed out that to get 1% precision for predictions for the LHC, we’ll likely need three-loop calculations. He pointed out the important role IR real emission calculations and improvements in parton distribution functions need to play, and pointed out that collider physics isn’t just the LHC: between reanalyses of older electron-positron collider data and the upcoming electron-ion collider and FCC, there are even more contexts where high-precision QCD will matter. Simone Zoia talked about one aspect of the current state of the art, five-particle processes involving massive particles, where the functions get strange and elliptical and one has to think hard about the methods one uses (for example, do you want canonical differential equations, or almost-canonical? Bulirsch-Stoer or AMFlow for numerics?) He also included an excellent Laplace quote, “Nature Laughs at the Difficulties of Integration”. Yang Zhang talked about a different calculational frontier, covering progress in the planar limit and with no masses, but for two-loop six-particle and three-loop five-particle scattering. The talk included some nice branding, like an “epsilon collaboration” proposing an algorithm to find epsilon forms for any Feynman integral and a program called Effortless to find symbol letters. Yang Zhang is also a skilled amateur photographer, and he mentions taking the conference photo for Amplitudes five times before. Personally, I’m surprised it’s only been five times! Dmitry Chicherin presented results bootstrapping QCD amplitudes. This was a dream of mine back in my hexagon function days, and while they aren’t quite at the point of being useful (my understanding is they’re still only getting the leading transcendentality, and none of the amplitudes they’re finding are new) it’s still pretty cool that this is even possible now. Bo Feng proposed what he claims is a general algorithm to find generating functions for IBP reduction. It’s not clear to me whether his setup bypasses the computational difficulties in existing methods (mostly involving solving large systems of equations), or whether it shifts the issue elsewhere, though. Pierre Vanhove’s talk also involved QCD, though at an effective field theory mediated step removed, with chiral perturbation theory, a theory of low-energy QCD that he used to shore up the lower energy ranges of lattice calculations for contributions to the muon anomalous magnetic moment (a natural calculation to involve him due to the presence of novel elliptic integrals). Overall, the QCD talks impressed me with the wide range of software packages mentioned, some of which already existed when I was in the field but many of which are new. I do wonder if this is just a result of many research paths maturing at the same time, or if people are finding it easier to write packages with tools like Claude Code.

The remaining three Tuesday talks were more mathematical or theoretical in focus. Henrik Johansson talked about black hole Compton scattering in N=8 supergravity, where it’s possible to use the wave equation to extract all-loop results. Cristian Vergu reported on his progress with Landau analysis. It’s been really fun watching this grow from a small reading group at NBI to what appears to be at this point quite a deep understanding, including a picture of how to understand amplitude singularities with quite a lot of breadth and detail, and the persistent hope that this could allow one to manipulate singular quantities without needing dimensional regularization. Michael Borinsky gave an update on tropicalization, where he showcased a theorem he proved demonstrating a way to compute certain classes of amplitudes in polynomial time in the loop order, even when the number of Feynman diagrams increases factorially. The examples he talked through were impressive, but I’m still a bit skeptical this could work for the Standard Model. I’d want to talk to him about it to figure it out, anyway!

Wednesday was a short day, as is tradition, to give people some time for tourism in the afternoon. Alexander Zhiboedov began the day with a talk on energy correlators in N=4 super Yang-Mills at finite coupling, where he is now able to do a bootstrap with two-sided bounds to hone in on the actual quantity even outside of the planar limit. Arthur Lipstein and Daniel Baumann both talked about cosmological correlators, the former with amplitudeologists’ favorite toy model of the conformally coupled scalar and the latter with Yang-Mills and gravity in de Sitter space. Dave Dunbar closed the day with a historical talk, walking through the milestone amplitudes papers before the Amplitudes conference existed. It’s the kind of talk he could have given at Lancefest the week before if others hadn’t already covered the material!

I’ll cover the rest of the conference (Thursday and Friday) in next week’s post, so see you then!

At Lancefest

It’s been a while since I’ve said this: I’m at a conference this week!

Specifically, I’m at Lancefest, which is not just any old conference, but a birthday conference for Lance Dixon. When a renowned academic turns 60 or so, their students and collaborators hold a birthday party-flavored conference for them. The conferences are usually a mix of academic talks and reminiscences, with the occasional roast thrown in.

I went to my advisor’s birthday conference four years ago. Lance wasn’t my advisor, but in many ways he might as well have been. When my advisor took a sabbatical in the middle of my PhD, he sent me to work with Lance. It was my first real experience doing research in a team, not just puzzling away by myself with occasional feedback. And I was hooked: I spent the rest of my academic career in Lance’s field. We collaborated time after time, and even when I started to branch out he remained a frequent presence.

In part, that’s because Lance’s field was really Lance’s field. Amplitudeology has grown a lot since I started out, with several hundred people going to the field’s big yearly conference and subfields like Elliptics having their own yearly conferences. In such a world, it’s tough for anyone to feel like a truly central figure. But Lance tends to. He’s been able to keep up with that growing world, to keep finding important problems and keep understanding others’ ideas. While some have specialized, or stepped back, Lance seems to somehow manage to be a father figure for the whole field all at once. It’s a capability his advisor Jeff Harvey might have predicted, he mentioned in his talk that Lance’s interests were always broad. Over the years the young students I met who joined when the amplitudes field was already large saw Lance as a kind of mysterious titan, and were occasionally awed that I had worked with him. “What was that like?”

Well, it was like working with Lance. Lance isn’t a manager at heart, like some senior academics end up. He wants to understand everything he works on, and will happily dive in, Maple subscription in hand, and try to figure things out for himself. He wants you to keep up with him, understanding on your own terms, to keep him honest, to provide an independent check. But that often wasn’t possible, because the man is just so damn fast. He’d be miles ahead of me, with his Maple and his laptop, while I was churning overnight Mathematica runs on a twenty-machine cluster.

(Was the difference due to our taste in software? Partly. I did learn to use Maple later, it genuinely is faster at some things. But Lance is faster at almost everything.)

And he cared so damn much sometimes. About getting things right, like a scientist should. About making things nice, too: finding a pretty basis of functions, a better notation for the paper, something that might jostle out the next big insight. Working with him, you could feel like that one paper was the most important thing in the universe.

Others at the event have had similar stories. Fernando Febres Cordero remembers noticing a potential issue, emailing Lance about it, and in a few minutes hearing back with a potential explanation.

Lance is someone who became a leader without really being a politician. He doesn’t have the legions of students in tenured positions that some do. I trimmed that count by one, and it wasn’t huge to begin with. But for someone who isn’t “everywhere” in that sense, he manages to be “everywhere” all the same.

So Lance, happy birthday! You’re really the only person who could have had a birthday conference quite like this, a cross-section of the field, all with something kind to say. Thanks for putting up with any embarrassment associated with having this much attention for three days, and I wish you many Maple-fueled mysteries to come.

An AI Opinions Chart

You ever read something and suddenly a whole classification scheme lights up in your head?

A thread on X from “stringking42069” showed me a combination of opinions I hadn’t seen before. stringking42069 is a pro-string theory commentator with a macho gym bro memer gimmick. He’s openly contemptuous of many physicists who describe themselves as string theorists, arguing that only a smaller number really deserve the name.

To be clear, none of that is the new combination. Long-time readers of this blog will remember a frequent commenter with a very similar attitude, if much less tendency to use the word “bro”.

The new thing, from my perspective, is how he thinks about AI. As he explains in that thread, he sees AI as great at certain kinds of physics calculations, ones where the methods and goals are mostly known and the challenge is working out the math. He doesn’t expect it to be able to contribute real creativity or judgement, the messy decision-making that physicists use to decide what is worth building in the first place.

Others with that perspective tend to argue that this will be a boon for scientists, who AI will free up to do creative work, multiplying their output. The difference is, stringking42069 thinks a lot of scientists are not doing creative work in the first place, including most of the people making extensive use of AI. So if anything he’s happy to see them go, and only pissed that they’re sucking up resources and attention on the way out, and discouraging students who could be joining the parts of the field that do real creative work.

It made me realize that there are two axes to thinking about AI in physics.

On the one hand, there’s where you think AI capabilities are. Is AI going to lead to “a nation of geniuses in a data center”, an AI-powered super-(cyber-)Ed Witten for everything and everyone? Is AI great at routine work and coding, but will never be able to do anything really creative or novel? Or is AI total hype, almost always a waste of time?

On the other hand, there’s another axis: misanthropy about science. For some of the people arguing about AI online, most scientists are good people trying their best to do worthwhile things. For others, most scientists are complacent and cliquish, wasting time and money on ideas that are going nowhere and forcing the real geniuses out of the field.

Put those together, and you get the table below:

Thinks academia is mostly fineMisanthrope
AI geniuses are comingThe practice of science will change. We’ll play at science like chess, and have fun trying to read and understand amazing AI insights.Soon all scientists will be out of a job when the public notices AI can do it all better. Then the real breakthroughs will come.
AI can do routine workAI frees scientists to focus on what we do best: creativity. We should think carefully about how to train junior scientists now, though.AI is comparable to bad scientists who only do derivative work. If they leave, we real paradigm-changers could inherit the field.
AI is complete hypeMost scientists don’t use AI. AI is worrying because it misleads students and the public, who should listen to real scientists.Scientists are shilling for AI companies, as you should expect for people who waste the public’s money on reputation games.

This classification is missing a lot, of course. One important question is not just what AI can do in principle, but what it can do cost-effectively, and whether anyone is actually willing to pay for it. A point where I agree with stringking42069 is that companies get a lot of good PR out of building AI physicists right now, and that PR benefit won’t be relevant forever. I’m also leaving out the more general questions of AI’s effect on society, for example people who think AI geniuses will lead to the end of the world as we know it.

But I suspect if you look at this table, you can already start matching the scientists you see on social media. I’ve seen examples of all of these in the wild (though the bottom-left is somewhat rare, as far as I can tell). Where do you fall?

Breakthrough Prize 2026

Because of last week’s “bonus info” post, I’m only now getting around to commenting on this year’s Breakthrough Prizes in Fundamental Physics. While I don’t comment on them every year, I know enough about several of this year’s winners that I figured a post would be helpful.

For those who haven’t heard of it, the Breakthrough Prizes are a bit like the Nobel, if it was created by a 21st century rich person instead of a 19th century one. They give out more money, and instead of an organization like the Swedish Academy of Sciences they pick winners via a committee of past winners. They’re more flexible in structure than the Nobel, with extra prizes for early-career researchers and a tendency to reward accomplishments that are either entirely theoretical or solid experimental work that doesn’t show a new discovery, both of which are things the Nobel Prize is structured to avoid. They’ve also shown willingness to reward large collaborations, rather than following the Nobel’s informal rule to only give the award to three people at a time.

This last was on display this year in their main award in physics this year, for the muon g-2 collaborations. The award is going to collaborations of scientists and engineers at three different particle colliders, for work done over a span of over fifty years to measure the magnetic properties of the muon. These measurements have shown a tantalizing discrepancy with predictions that inspired many to conjecture new physics. However, in the last few years it’s looked more and more like the discrepancy was due to an imprecise prediction, and better methods seem to be converging to the experimental value. At this point, smart money is that there is no disagreement with the Standard Model here, but as always in science there’s a chance some mystery remains.

The Breakthrough Prize also offered a special, out-of-schedule prize to David Gross. Already a Nobel laureate, Gross had a crucial role in our understanding of the force of quantum chromodynamics that binds protons and neutrons together. He was also a major founding figure in string theory, and since the Breakthrough Prize is more comfortable recognizing theoretical contributions they get to mention this as well. Gross is also known in the community for his personality, which tends to fill up any room he’s in. I can only imagine the conversations that led to Breakthrough’s decision to add a special prize for him this year.

Breakthrough is also adding a new recurring prize, the Vera Rubin New Frontiers Prize, honoring women who make important contributions to physics within two years of their PhD. The prize is a bit smaller than the exiting early-career New Horizons in Physics Prizes, presumably because it goes to even younger researchers. This year’s winner is from my old field, scattering amplitudes. Carolina Figueiredo is part of the latest evolution of the research program behind the amplituhedron. The new framework of “surfaceology” seems like a promising geometry-flavored way to understand particle physics calculations in more realistic theories, and unlike its predecessors may have some practical value eventually as well. Congrats Carolina!

Finally, the New Horizons in Physics Prizes are for impressive early-career researchers. I don’t know much about the first recipient, Benjamin Safdi, who works on searches for axions and axion-like particles, today’s most trendy dark matter candidate. I know a bit more about the work done by Clay Córdova, Thomas Dumitrescu, Shu-Heng Shao, and Yifan Wang, having met several of them in my physics career. They work on what are called generalized symmetries, concepts which go beyond the usual idea of how symmetry is supposed to work by involving more complicated tensors. I saw these crop up a fair bit in talks, but they were distant enough from my area that I never had a particularly clear grasp of what people were doing with them. I know even less about the work of the last three, Dillon Brout, J. Colin Hill, Mathew Madhavacheril, Maria Vincenzi, Daniel Scolnic, and W. L. Kimmy Wu, on cosmological measurements, but I was friends with Mathew in grad school and am impressed that he’s now working on cosmology given how little cosmology research there was at Stony Brook at the time.

A Window on Absolutely Everything

It’s often said that in quantum physics, everything that can happen will happen.

One way this comes up is in something called a path integral, used to calculate the probabilities of quantum events. If you want to find what happens to a particle traveling from point A to point B, you have to add up a contribution for every path, no matter how windy, that goes between A and B. These contributions mostly cancel out, and matter less the further they are from a straight line, so the straight-line path is, for the most part, a good description of what happens. But in principle, all of the other paths matter too.

The same thing happens in quantum field theory, in more elaborate form. Instead of a path from one place to another, the paths are from one configuration of quantum fields to another, via all the different ways fields can in principle interact. We are almost never able to take account of all these possibilities mathematically, so we have to approximate, organizing the interactions into more and more complicated pictures called Feynman diagrams, each with a smaller and smaller effect.

In principle, these diagrams need to contain every single combination of interactions that might result in the end-state we’re interested in. These combinations can have a Rube Goldberg flavor, with one field activating another, which activates another, only to all cancel out in the end. Because of this, any field that exists, any particle no matter how rare, can matter, if only a little.

And from that, physicists can learn something.

Because absolutely everything matters, physicists get to reason about absolutely everything that exists.

The best example involves something called an anomaly. These aren’t the anomalies of experimental physics, unexpected results that have a tendency to go away with better measurements. Instead of something unexpected, a theorist’s anomaly is something impossible.

Anomalies are combinations of particles that, if they were to show up together in a sum of Feynman diagrams, would break the rules that the theory was made with in the first place. If they show up, they’re a sign of an inconsistent theory, one that doesn’t obey its own rules and thus doesn’t make sense.

In order to have a theory without anomalies, different calculations involving different particles need to cancel. For example, it might be that the charge of different particles has to add up to zero. This means that if you’ve only discovered a few particles, and their charges don’t add up to zero, then you know you’re missing one. There is an extra particle there, which you haven’t observed, that together makes charge add up to zero.

This logic actually works! It was used to predict the top quark. Before the top quark was discovered, the list of quarks, electrons, and neutrinos had electric charges that didn’t add up to zero. One particle was missing, with the same charge as the up quark and charm quark. It was found in 1995, after being proposed almost 20 years earlier.

What AI Physicists Are Missing and What They Aren’t

I’ve seen a couple more thoughtful takes on use of LLMs for physics lately. This blog post by Minas Karamis is particularly nice.

He points out something that I’ve said a version of: an AI that must be supervised like a student isn’t very useful, because the main point of student projects isn’t the paper at the end: it’s training the student. If students don’t struggle through all the mistakes of a project, they won’t get the expertise to one day do greater things.

Someone might object that not all suffering is educational. In the 1700’s, Leonhard Euler calculated digit after digit of transcendental numbers by hand. Nobody asks students to do that anymore, and they still seem to turn out alright. Why would using an LLM for science be worse than using a computer for numerical calculations?

In a word: different skills. Programming numerics teaches you some of the same skills as calculating the numbers by hand: skills at being specific about what you mean, aware of the consequences of the details and their implications. Prompting an AI still requires those skills, to check whether the AI’s output is correct. But it’s much worse at teaching them: unlike programming or calculating, when prompting AI, the consequences of your actions aren’t predictable.

For some, though, there is another objection. Sure, using AI reliably might require those skills now. But when it gets better, surely being careful will stop mattering. Surely the AI will end up doing science on its own, and all that training will be as useful as if we trained the students to play football.

I’m skeptical, but not as strongly as some. I think we’re still living in a time when it makes sense to hire scientists, and train people to think, and invest in your retirement.

I don’t think I have any knock-down arguments for that, though. Just some suggestive ones.

One I’ve talked about before is that a lot of the most important parts of thinking aren’t written down. An AI physicist is going to have a hard time replicating the kinds of methods and approaches that people use behind the scenes, but rarely describe or spell out. It will be easier to suss this out over time, as more data accumulates of people working with LLMs and correcting them. But ultimately there isn’t going to be a lot of documentation of this kind of thing.

Another limitation is memory. A mature scientist can draw from experiences across their entire career. For an LLM, any problem it’s solved in the past is by default lost in each new session. People build structures around this, taking notes and reminding the AI when it “wakes up”, or making documents the AI can be prompted to check. But nothing in this vein so far seems to get nearly as wide-scope or powerful as human memory. A scientist career is still the best way we have to build durable, functional expertise.

Finally, there is a question of costs, and efficiency. Here I’m not an expert, and I get the impression the actual experts disagree. I don’t know whether we should expect scaling to hit a wall, but I wouldn’t be that surprised if it did.

There are other common reasons for skepticism that seem more dubious to me. I don’t think AI is inherently worse at creativity just because they’re trained on existing work, though some of the skills we associate with creativity aren’t very well-documented, and thus are hard to train for. I don’t think AI’s randomness or unreliability is a deal-breaker, because human intuition is also random and unreliable: we solve that with tools, and that’s something AI can in principle do as well. I don’t think humans are “more agentic” or something, except in the sense that most AIs are made by companies who need to make them behave in a customer-friendly way. But an agent is just a game-theoretic construct, a way to figure out can win or lose in situations with defined stakes, and anything you can train or engineer to try to win can be modeled by that construct.

Coming from a place of uncertainty, my main appeal to you is to not get hung up on the bad reasons, either yourself, or from the people you’re arguing with. Focus on the best arguments, and see where they take you.

About the OpenAI Amplitudes Paper, but Not as Much as You’d Like

I’ve had a bit more time to dig in to the paper I mentioned last week, where OpenAI collaborated with amplitudes researchers, using one of their internal models to find and prove a simplified version of a particle physics formula. I figured I’d say a bit about my own impressions from reading the paper and OpenAI’s press release.

This won’t be a real “deep dive”, though it will be long nonetheless. As it turns out, most of the questions I’d like answers to aren’t answered in the paper or the press release. Getting them will involve actual journalistic work, i.e. blocking off time to interview people, and I haven’t done that yet. What I can do is talk about what I know so far, and what I’m still wondering.

Context:

Scattering amplitudes are formulas used by particle physicists to make predictions. For a while, people would just calculate these when they needed them, writing down pages of mess that you could plug in numbers to to get answers. However, forty years ago two physicists decided they wanted more, writing “we hope to obtain a simplified form for the answer, making our result not only an experimentalist’s, but a theorist’s delight.”

In their next paper, they managed to find that “theorist’s delight”: a simplified, intuitive-looking answer that worked for calculations involving any number of particles, summarizing many different calculations. Ten years later, a few people had started building on it, and ten years after that, the big shots started paying attention. A whole subfield, “amplitudeology”, grew from that seed, finding new forms of “theorists’s delight” in scattering amplitudes.

Each subfield has its own kind of “theory of victory”, its own concept for what kind of research is most likely to yield progress. In amplitudes, it’s these kinds of simplifications. When they work out well, they yield new, more efficient calculation techniques, yielding new messy results which can be simplified once more. To one extent or another, most of the field is chasing after those situations when simplification works out well.

That motivation shapes both the most ambitious projects of senior researchers, and the smallest student projects. Students often spend enormous amounts of time looking for a nice formula for something and figuring out how to generalize it, often on a question suggested by a senior researcher. These projects mostly serve as training, but occasionally manage to uncover something more impressive and useful, an idea others can build around.

I’m mentioning all of this, because as far as I can tell, what ChatGPT and the OpenAI internal model contributed here roughly lines up with the roles students have on amplitudes papers. In fact, it’s not that different from the role one of the authors, Alfredo Guevara, had when I helped mentor him during his Master’s.

Senior researchers noticed something unusual, suggested by prior literature. They decided to work out the implications, did some calculations, and got some messy results. It wasn’t immediately clear how to clean up the results, or generalize them. So they waited, and eventually were contacted by someone eager for a research project, who did the work to get the results into a nice, general form. Then everyone publishes together on a shared paper.

How impressed should you be?

I said, “as far as I can tell” above. What’s annoying is that this paper makes it hard to tell.

If you read through the paper, they mention AI briefly in the introduction, saying they used GPT-5.2 Pro to conjecture formula (39) in the paper, and an OpenAI internal model to prove it. The press release actually goes into more detail, saying that the humans found formulas (29)-(32), and GPT-5.2 Pro found a special case where it could simplify them to formulas (35)-(38), before conjecturing (39). You can get even more detail from an X thread by one of the authors, OpenAI Research Scientist Alex Lupsasca. Alex had done his PhD with another one of the authors, Andrew Strominger, and was excited to apply the tools he was developing at OpenAI to his old research field. So they looked for a problem, and tried out the one that ended up in the paper.

What is missing, from the paper, press release, and X thread, is any real detail about how the AI tools were used. We don’t have the prompts, or the output, or any real way to assess how much input came from humans and how much from the AI.

(We have more for their follow-up paper, where Lupsasca posted a transcript of the chat.)

Contra some commentators, I don’t think the authors are being intentionally vague here. They’re following business as usual. In a theoretical physics paper, you don’t list who did what, or take detailed account of how you came to the results. You clean things up, and create a nice narrative. This goes double if you’re aiming for one of the most prestigious journals, which tend to have length limits.

This business-as-usual approach is ok, if frustrating, for the average physics paper. It is, however, entirely inappropriate for a paper showcasing emerging technologies. For a paper that was going to be highlighted this highly by OpenAI, the question of how they reached their conclusion is much more interesting than the results themselves. And while I wouldn’t ask them to go to the standards of an actual AI paper, with ablation analysis and all that jazz, they could at least have aimed for the level of detail of my final research paper, which gave samples of the AI input and output used in its genetic algorithm.

For the moment, then, I have to guess what input the AI had, and what it actually accomplished.

Let’s focus on the work done by the internal OpenAI model. The descriptions I’ve seen suggest that it started where GPT-5.2 Pro did, with formulas (29)-(32), but with a more specific prompt that guided what it was looking for. It then ran for 12 hours with no additional input, and both conjectured (39) and proved it was correct, providing essentially the proof that follows formula (39) in the paper.

Given that, how impressed should we be?

First, the model needs to decide to go to a specialized region, instead of trying to simplify the formula in full generality. I don’t know whether they prompted their internal model explicitly to do this. It’s not something I’d expect a student to do, because students don’t know what types of results are interesting enough to get published, so they wouldn’t be confident in computing only a limited version of a result without an advisor telling them it was ok. On the other hand, it is actually something I’d expect an LLM to be unusually likely to do, as a result of not managing to consistently stick to the original request! What I don’t know is whether the LLM proposed this for the right reason: that if you have the formula for one region, you can usually find it for other regions.

Second, the model needs to take formulas (29)-(32), write them in the specialized region, and simplify them to formulas (35)-(38). I’ve seen a few people saying you can do this pretty easily with Mathematica. That’s true, though not every senior researcher is comfortable doing that kind of thing, as you need to be a bit smarter than just using the Simplify[] command. Most of the people on this paper strike me as pen-and-paper types who wouldn’t necessarily know how to do that. It’s definitely the kind of thing I’d expect most students to figure out, perhaps after a couple of weeks of flailing around if it’s their first crack at it. The LLM likely would not have used Mathematica, but would have used SymPy, since these “AI scientist” setups usually can write and execute Python code. You shouldn’t think of this as the AI reasoning through the calculation itself, but it at least sounds like it was reasonably quick at coding it up.

Then, the model needs to conjecture formula (39). This gets highlighted in the intro, but as many have pointed out, it’s pretty easy to do. If any non-physicists are still reading at this point, take a look:

Could you guess (39) from (35)-(38)?

After that, the paper goes over the proof that formula (39) is correct. Most of this proof isn’t terribly difficult, but the way it begins is actually unusual in an interesting way. The proof uses ideas from time-ordered perturbation theory, an old-fashioned way to do particle physics calculations. Time-ordered perturbation theory isn’t something any of the authors are known for using with regularity, but it has recently seen a resurgence in another area of amplitudes research, showing up for example in papers by Matthew Schwartz, a colleague of Strominger at Harvard.

If a student of Strominger came up with an idea drawn from time-ordered perturbation theory, that would actually be pretty impressive. It would mean that, rather than just learning from their official mentor, this student was talking to other people in the department and broadening their horizons, showing a kind of initiative that theoretical physicists value a lot.

From an LLM, though, this is not impressive in the same way. The LLM was not trained by Strominger, it did not learn specifically from Strominger’s papers. Its context suggested it was working on an amplitudes paper, and it produced an idea which would be at home in an amplitudes paper, just a different one than the one it was working on.

While not impressive, that capability may be quite useful. Academic subfields can often get very specialized and siloed. A tool that suggests ideas from elsewhere in the field could help some people broaden their horizons.

Overall, it appears that that twelve-hour OpenAI internal model run reproduced roughly what an unusually bright student would be able to contribute over the course of a several-month project. Like most student projects, you could find a senior researcher who could do the project much faster, maybe even faster than the LLM. But it’s unclear whether any of the authors could have: different senior researchers have different skillsets.

A stab at implications:

If we take all this at face-value, it looks like OpenAI’s internal model was able to do a reasonably competent student project with no serious mistakes in twelve hours. If they started selling that capability, what would happen?

If it’s cheap enough, you might wonder if professors would choose to use the OpenAI model instead of hiring students. I don’t think this would happen, though: I think it misunderstands why these kinds of student projects exist in a theoretical field. Professors sometimes use students to get results they care about, but more often, the student’s interest is itself the motivation, with the professor wanting to educate someone, to empire-build, or just to take on their share of the department’s responsibilities. AI is only useful for this insofar as AI companies continue reaching out to these people to generate press releases: once this is routinely possible, the motivation goes away.

More dangerously, if it’s even cheaper, you could imagine students being tempted to use it. The whole point of a student project is to train and acculturate the student, to get them to the point where they have affection for the field and the capability to do more impressive things. You can’t skip that, but people are going to be tempted to.

And of course, there is the broader question of how much farther this technology can go. That’s the hardest to estimate here, since we don’t know the prompts used. So I don’t know if seeing this result tells us anything more about the bigger picture than we knew going in.

Remaining questions:

At the end of the day, there are a lot of things I still want to know. And if I do end up covering this professionally, they’re things I’ll ask.

  1. What was the prompt given to the internal model, and how much did it do based on that prompt?
  2. Was it really done in one shot, no retries or feedback?
  3. How much did running the internal model cost?
  4. Is this result likely to be useful? Are there things people want to calculate that this could make easier? Recursion relations it could seed? Is it useful for SCET somehow?
  5. How easy would it have been for the authors to do what the LLM did? What about other experts in the community?

Hypothesis: If AI Is Bad at Originality, It’s a Documentation Problem

Recently, a few people have asked me about this paper.

A couple weeks back, OpenAI announced a collaboration with a group of amplitudes researchers, physicists who study the types of calculations people do to make predictions at particle colliders. The amplitudes folks had identified an interesting loophole, finding a calculation that many would have expected to be zero actually gave a nonzero answer. They did the calculation for different examples involving more and more particles, and got some fairly messy answers. They suspected, as amplitudes researchers always expect, that there was a simpler formula, one that worked for any number of particles. But they couldn’t find it.

Then a former amplitudes researcher at OpenAI suggested that they use AI to find it.

“Use AI” can mean a lot of different things, and most of them don’t look much like the way the average person talks to ChatGPT. This was closer than most. They were using “reasoning models”, loops that try to predict the next few phrases in a “chain of thought” again and again and again. Using that kind of tool, they were able to find that simpler formula, and mathematically prove that it was correct.

A few of you are hoping for an in-depth post about what they did, and its implications. This isn’t that. I’m still figuring out if I’ll be writing that for an actual news site, for money, rather than free, for you folks.

Instead, I want to talk about a specific idea I’ve seen crop up around the paper.

See, for some, the existence of a result like this isn’t all that surprising.

Mathematicians have been experimenting with reasoning models for a bit, now. Recently, a group published a systematic study, setting the AI loose on a database of minor open problems proposed by the famously amphetamine-fueled mathematician Paul Erdös. The AI managed to tackle a few of the problems, sometimes by identifying existing solutions that had not yet been linked to the problem database, but sometimes by proofs that appeared to be new.

The Erdös problems solved by the AI were not especially important. Neither was the problem solved by the amplitudes researchers, as far as I can tell at this point.

But I get the impression the amplitudes problem was a bit more interesting than the Erdös problems. The difference, so far, has mostly been attributed to human involvement. This amplitudes paper started because human amplitudes researchers found an interesting loophole, and only after that used the AI. Unlike the mathematicians, they weren’t just searching a database.

This lines up with a general point, one people tend to make much less carefully. It’s often said that, unlike humans, AI will never be truly creative. It can solve mechanical problems, do things people have done before, but it will never be good at having truly novel ideas.

To me, that line of thinking goes a bit too far. I suspect it’s right on one level, that it will be hard for any of these reasoning models to propose anything truly novel. But if so, I think it will be for a different reason.

The thing is, creativity is not as magical as we make it out to be. Our ideas, scientific or artistic, don’t just come from the gods. They recombine existing ideas, shuffling them in ways more akin to randomness than miracle. They’re then filtered through experience, deep heuristics honed over careers. Some people are good at ideas, and some are bad at them. Having ideas takes work, and there are things people do to improve their ideas. Nothing about creativity suggests it should be impossible to mechanize.

However, a machine trained on text won’t necessarily know how to do any of that.

That’s because in science, we don’t write down our inspirations. By the time a result gets into a scientific paper or textbook, it’s polished and refined into a pure argument, cutting out most of the twists and turns that were an essential part of the creative process. Mathematics is even worse, most math papers don’t even mention the motivation behind the work, let alone the path taken to the paper.

This lack of documentation makes it hard for students, making success much more a function of having the right mentors to model good practices, rather than being able to pick them up from literature everyone can access. I suspect it makes it even harder for language models. And if today’s language model-based reasoning tools are bad at that crucial, human-seeming step, of coming up with the right idea at the right time? I think that has more to do with this lack of documentation, than with the fact that they’re “statistical parrots”.

The Timeline for Replacing Theorists Is Not Technological

Quanta Magazine recently published a reflection by Natalie Wolchover on the state of fundamental particle physics. The discussion covers a lot of ground, but one particular paragraph has gotten the lion’s share of the attention. Wolchover talked to Jared Kaplan, the ex-theoretical physicist turned co-founder of Anthropic, one of the foremost AI companies today.

Kaplan was one of Nima Arkani-Hamed’s PhD students, which adds an extra little punch.

There’s a lot to contest here. Is AI technology anywhere close to generating papers as good as the top physicists, or is that relegated to the sci-fi future? Does Kaplan really believe this, or is he just hyping up his company?

I don’t have any special insight into those questions, about the technology and Kaplan’s motivations. But I think that, even if we trusted him on the claim that AI could be generating Witten- or Nima-level papers in three years, that doesn’t mean it will replace theoretical physicists. That part of the argument isn’t a claim about the technology, but about society.

So let’s take the technological claims as given, and make them a bit more specific. Since we don’t have any objective way of judging the quality of scientific papers, let’s stick to the subjective. Today, there are a lot of people who get excited when Witten posts a new paper. They enjoy reading them, they find the insights inspiring, they love the clarity of the writing and their tendency to clear up murky ideas. They also find them reliable: the papers very rarely have mistakes, and don’t leave important questions unanswered.

Let’s use that as our baseline, then. Suppose that Anthropic had an AI workflow that could reliably write papers that were just as appealing to physicists as Witten’s papers are, for the same reasons. What happens to physicists?

Witten himself is retired, which for an academic means you do pretty much the same thing you were doing before, but now paid out of things like retirement savings and pension funds, not an institute budget. Nobody is going to fire Witten, there’s no salary to fire him from. And unless he finds these developments intensely depressing and demoralizing (possible, but very much depends on how this is presented), he’s not going to stop writing papers. Witten isn’t getting replaced.

More generally, though, I don’t think this directly results in anyone getting fired, or in universities trimming positions. The people making funding decisions aren’t just sitting on a pot of money, trying to maximize research output. They’ve got money to be spent on hires, and different pools of money to be spent on equipment, and the hires get distributed based on what current researchers at the institutes think is promising. Universities want to hire people who can get grants, to help fund the university, and absent rules about AI personhood, the AIs won’t be applying for grants.

Funding cuts might be argued for based on AI, but that will happen long before AI is performing at the Witten level. We already see this happening in other industries or government agencies, where groups that already want to cut funding are getting think tanks and consultants to write estimates that justify cutting positions, without actually caring whether those estimates are performed carefully enough to justify their conclusions. That can happen now, and doesn’t depend on technological progress.

AI could also replace theoretical physicists in another sense: the physicists themselves might use AI to do most of their work. That’s more plausible, but here adoption still heavily depends on social factors. Will people feel like they are being assessed on whether they can produce these Witten-level papers, and that only those who make them get hired, or funded? Maybe. But it will propagate unevenly, from subfield to subfield. Some areas will make their own rules forbidding AI content, there will be battles and scandals and embarrassments aplenty. It won’t be a single switch, the technology alone setting the timeline.

Finally, AI could replace theoretical physicists in another way, by people outside of academia filling the field so much that theoretical physicists have nothing more that they want to do. Some non-physicists are very passionate about physics, and some of those people have a lot of money. I’ve done writing work for one such person, whose foundation is now attempting to build an AI Physicist. If these AI Physicists get to Witten-level quality, they might start writing compelling paper after compelling paper. Those papers, though, will due to their origins be specialized. Much as philanthropists mostly fund the subfields they’ve heard of, philanthropist-funded AI will mostly target topics the people running the AI have heard are important. Much like physicists themselves adopting the technology, there will be uneven progress from subfield to subfield, inch by socially-determined inch.

In a hard-to-quantify area like progress in science, that’s all you can hope for. I suspect Kaplan got a bit of a distorted picture of how progress and merit work in theoretical physics. He studied with Nima Arkani-Hamed, who is undeniably exceptionally brilliant but also undeniably exceptionally charismatic. It must feel to a student of Nima’s that academia simply hires the best people, that it does whatever it takes to accomplish the obviously best research. But the best research is not obvious.

I think some of these people imagine a more direct replacement process, not arranged by topic and tastes, but by goals. They picture AI sweeping in and doing what theoretical physics was always “meant to do”: solve quantum gravity, and proceed to shower us with teleporters and antigravity machines. I don’t think there’s any reason to expect that to happen. If you just asked a machine to come up with the most useful model of the universe for a near-term goal, then in all likelihood it wouldn’t consider theoretical high-energy physics at all. If you see your AI as a tool to navigate between utopia and dystopia, theoretical physics might matter at some point: when your AI has devoured the inner solar system, is about to spread beyond the few light-minutes when it can signal itself in real-time, and has to commit to a strategy. But as long as the inner solar system remains un-devoured, I don’t think you’ll see an obviously successful theory of fundamental physics.