It Got to My Field

For the folks who found this blog from Anthropic’s site, welcome!

For everyone else, I’d better give some context.

Last month, I had a blog post titled “It Only Counts When AI Gets to My Field”. The title was a joke, the content less so. I said that if some of the longstanding problems of my old research field got solved by AI, then I’d sit up and take notice.

The post turned out to be a bit of a self-fulfilling prophecy, after some folks at Anthropic read it and decided to tackle one of the outstanding problems I mentioned. They then invited me to write a post about it on their science blog.

The post is here. I recommend reading it, then coming back. The rest of this post will be a little Q&A.

Q: At the end of that post, it says Anthropic paid you for your time. Can we trust what you wrote?

A: When I agreed to write the post, I made it clear that I was going to give my own opinion, not write an ad for Anthropic. They could propose light edits, but that’s it. And I avoided signing anything with them, not even an NDA, so I could freely tell you if they pushed my boundaries.

They didn’t push my boundaries. They asked me to clarify a few things, and to give more detail on the science. I didn’t change the message or the takeaways.

The reason I asked them to pay me is that, as a freelance writer, I don’t have a salary to fall back on. Time I spend on a project like that post is time I’m not spending on journalistic projects, so if I worked on it for free I’d essentially be using my vacation time for it. The rate I’m charging them is roughly in the middle of what I’d have gotten paid if I used that time on journalism: a bit more than I would have made working for the lower-paying outlets, a bit less than I would have made with the higher-paying ones.

I’m also not expecting this to be the start of a longer-term business relationship or anything like that. So overall, I don’t think I’m incentivized to lie on their behalf. You can trust me.

Q: How confident are you that they did what they said they did?

A: I didn’t get the feeling they were lying. But I didn’t start out skeptical.

I haven’t seen the logs from the LLM, or anything like that. That’s the kind of thing I would dig into if I were more suspicious of their story. If there’s a reason to be suspicious, I’ll ask.

But so far, I don’t feel that I have much reason to be suspicious. They were able to look up details on the fly when I asked them, they didn’t seem to have a polished message they were trying to push past me. And more importantly, as I’ll mention in a bit, I don’t think anything they described is all that outlandish. It lines up, largely, with the capabilities I’d expect their Claude Science platform to have.

It’s also relevant that they didn’t use an internal model for this. It means that scientists will likely be trying out similar problems soon, so if it does turn out they exaggerated something, people are going to figure out quite quickly.

Q: So how big is this? We just saw AI solve a Millennium problem, after all.

A: This was a lot easier than a Millennium problem. But it was also a lot cheaper.

The problem they solved is one with a clear recipe, honed and explained over multiple papers. It’s something I expected to be hard to do without access to a lot of computer power, so I thought it would need to be approached with a novel technique, in order to avoid using that computer power.

In the end, it didn’t need that. As I mention in the post, another group got most of the result at around the same time, with a much smaller amount of AI assistance.

It’s not an easy recipe, to be clear. I think other people will be surprised that a reasonably affordable program like Claude Science can do this without a lot of guidance. I’m not surprised, mostly, because I’ve been paying enough attention to what people have been doing with these things, and carrying out this kind of recipe consistently is actually something the good science AIs can do right now. If you haven’t been following as closely, that’s going to be a lot more surprising.

So maybe the best way to answer this is: instead of paying $15 million to solve one of the most famous problems in mathematics, they paid $1,000 or so to make a significant next step in an ongoing research program in theoretical physics. This doesn’t tell you much about whether AI can achieve field-defining breakthroughs, but it tells you a lot about the kinds of things a theoretical physicist with $1000 spare budget can do right now.

(As an aside, it feels crazy to me that 90% or more of the compute cost was from the LLM, not the calculation itself. On the one hand, it feels nuts to essentially use ten times more computer power to do this than would have been needed if a human had done the coding. On the other hand, I’ve applied for grants that budgeted more than that per year for travel costs, and spending this kind of money on getting a result seems a lot more useful than spending it on airfare.)

Q: Isn’t it reckless to publicly post a problem for AI like that?

A: I posted my challenge before OpenAI announced their Navier-Stokes result. At that point, there had been a few awkward surprises, but for the most part AI companies had a pattern of talking to an academic first before trying to solve a problem. It’s what they did back in March.

I expected that was what they’d do this time, and got more than a little blindsided when they just solved the problem on their own and reached out to Lance afterwards. I’m lucky that Lance doesn’t seem mad at me about it, it would have been quite understandable if he was.

I did tell them that if they wanted to tackle any of the other problems in that post, they should reach out to one of the scientists involved first, before attempting it. In addition to being polite, it’s a way to make sure that they understand the problem correctly and have a plan to verify the result. They happened to pick a challenge that was particularly easy to verify, but it didn’t have to go that way.

Q: Any thoughts about AI’s future impact on science?

A: If there’s a problem you can’t solve, but you think someone with expertise in another field or better programming skills could, then it’s probably solvable with AI.

A lot of problems don’t fall into that category. I’ve seen people talking online about AI solving quantum gravity, but solving quantum gravity is a question of which bullets you’re willing to bite, not a question of technical skill. That doesn’t mean AI will never be able to address it, but if so it will come from some sort of superpersuader aspect, not merely scientific capabilities.

And of course, some fields do require experiments. People are increasingly building AI-powered labs. I’ll leave it to people from actual experimental fields to think about the potential there.

Q: Can you say something explicit about the bigger risks from AI, now?

A: As I mention in the post, I didn’t learn that much about the bigger questions from this. I’m still not an expert.

But there’s one thing I think is worth emphasizing:

This technology is clearly getting more effective over time. I don’t think they could have done this six months ago. If you’re trying to predict what will happen next, you shouldn’t just assume this is the most powerful it will get, or the cheapest. If you think there’s a limit, you need to argue for it.

Data Comes From Papers

There’s a lovely resource that I suspect most non-physicists don’t know about. It’s called the Particle Data Group, and its interactive site PDGLive.

If you want to know the most up-to-date information on a subatomic particle, PDGLive has your back. From their home page, you can click on familiar particles like photons (\gamma), electrons (e) or gluons and find the best info experiments have provided on properties like their mass and charge. It’s a great place to get an authoritative number for any particle physics property you’re interested in.

Of course, scientists don’t just accept one authoritative answer for anything, even whether Pluto is a planet. That’s why each measurement in PDGLive comes with an arrow you can expand to see a table of past measurements, which you can compare to.

Each measurement comes with a link to a source. And each source is not an experiment webpage, or some entry in a database. It’s a document, a publication, a paper.

In fact, PDGLive itself is just a website version of a paper, the Particle Data Group’s “Review of Particle Physics”. Click enough times, and you’ll find sections of a pdf on the site, explaining their reasoning for picking this experiment over that, emphasizing this or the other thing.

People talk about papers as how academics persuade one another, or how they show off and establish credit. But papers are also just a way to organize data. Each time an experiment figures something out, all of their reasoning and procedures are summarized together with the numbers they got. And anyone who reports those numbers to you will include a link, so you can go back, and check where the number came from.

It probably feels a bit weird, all of these numbers and technical details bottoming out in an archaic practice of writing down words for other human beings. But it means that all of the richness of the process is there somewhere, linked together and collated by the same social forces that keep track of credit, all in one navigable whole.

So when you run into a number, spare some thought for where it came from. You can probably find out.

Everybody Who Isn’t ”Viewers Like You”

Last week, I talked about how truthseekers get paid. But truth-tellers and truth-seekers are different things.

Consider educational kids’ shows on public television.

Nobody who works on Sesame Street is out there uncovering new letters and numbers. Bill Nye’s show wasn’t bringing analysis fresh from the lab.

The purpose of these shows is to educate. The purpose of education is to change minds.

So who pays for educational kids’ shows on public television?

If you’re from the US and watched PBS growing up, you remember one answer: “viewers like you!” US public television is supported by donations, ordinary people across the country who want it to keep on educating kids.

But you also might remember the lists of names that came before “viewers like you”. Some of those were things like “the Department of Education” or “a grant from the National Science Foundation”: government programs, in other words. Others were philanthropists and private foundations. Some were tied to companies, like the Intel Foundation, or Juicy Juice.

All of these groups, from government departments to donors, are trying to change kids’ minds. They support specific shows on specific topics, where they want kids to be better-informed. The same groups have the same kind of impact on schools. For example, I remember in elementary school we all learned to play a recorder, because a wealthy donor had given the school recorders out of the idea that music education was especially important.

For a truth-seeker like a journalist, accepting that kind of funding would be a problem. Grants for journalists tend to support things like travel, letting journalists learn more about specific topics, not pre-judging the conclusion. But children’s television is about truth-telling, not truth-seeking, so our standards are different. We trust the people making children’s television to care about whether they’re telling the truth. And because the topics aren’t new, we don’t usually worry about their judgement being biased.

All this is rather obvious. But now, consider science YouTube.

Some science YouTubers seem to have a mission much like children’s television. They’re there to teach, not to make independent judgements. They don’t search for truth on their own. And some of them are funded by educational grants, much like children’s television.

Others are a bit more like journalists, or even activists. People follow them for their opinions, to hear their assessment. They’re trying to be truth-seekers.

On YouTube, it’s not always obvious which is which.

There’s a particular group of philanthropists called Effective Altruists, and many of them are concerned about AI. So in between funding things like anti-malaria bed nets, some of them are giving grants to YouTubers to make educational content about AI-related risks.

Apparently, they reached out to Sabine Hossenfelder, which was a bad idea. Sabine Hossenfelder’s followers aren’t just looking for education on known facts. They’re looking for her judgements, her literal bullshit-rating on ideas. And so while she’s paid by “viewers like you”, she’s not really the type to get paid by that type of grant.

What I want to emphasize, and what looked like it was getting lost in the discussion, was that their pitch would have been totally reasonable for other YouTubers. Educators do occasionally get grants to educate on specific topics. This is in fact a totally normal thing. Some YouTubers are educators first and foremost, they aren’t there as truth-seekers, but truth-tellers, with a real difference in how careful they need to be about bias.

Some YouTubers are different from other YouTubers. News at 11.

Paying the Truthseekers

Academics and journalists have a lot in common, at least in principle.

Whether you’re a reporter or a professor, your job is to go out into the world and figure out the truth. You’re supposed to be careful, to check and correct for how you might be wrong. And at the end of the day you’re supposed to communicate what you found.

The differences mostly come in how you’re paid.

You could imagine some sort of pure truthseeker, paid purely by how well they tell the truth. People would ask them to find out the truth about something, and pay them for the service. And the truthseekers with the best track record would get the most clients. But neither profession really works like this.

Journalism comes closest. Once upon a time, people bought newspapers in order to be the first to know when something important happened. While there’s still a little bit of that going on (I guess this is what Bloomberg Terminals are for?), it’s a lot less central because of the internet. Now, there are hundreds of ways to find out about things, from a multitude of news sites to social media. More and more, people expect to be able to get information for free.

In that environment, the news has to compete not on the facts themselves, but on how it presents them. People pay for news that’s curated well to match their interests, or news that feels more respectable. And more than either of those, they pay for news that’s entertaining. So while truthseeking skills pay, writing skills often end up mattering more. In a sense, it’s why it’s possible for me to do journalism at all. I was trained in the academic truthseeking tradition, not the journalistic one. I got into journalism by impressing editors with my writing, not my ability to suss out the truth.

That academic truthseeking tradition is quite different, in part because the rewards for it are much more indirect. Academics pay comes from two main sources: research grants, and student tuition. Students are mostly there to learn old facts, not new ones, so that source of money supports research only in so far that students believe that a successful researcher with time for research will also be a better teacher.

Research grants, in principle, pay for truthseeking. But they’re typically paid by governments, which often don’t have a clear idea of what they’d like to learn, since the more practical questions are already being researched by private companies. So the decision gets delegated out to other academics, who have a vague shared sense of what’s worth knowing and what’s not. Accuracy should have an impact: that is, it should be easier to get grants if you’re better at finding the truth. But in practice, unless someone does so badly they trigger a scandal, academics don’t usually get things all that wrong. So grants are mostly based on other factors.

Paying someone purely to deliver the truth, not to entertain or match a culture, seems tricky. You could imagine sci-fi scenarios. What if we could track the logic people used to make decisions, and demand payment if those decisions were based on facts we uncovered, like a journalist getting a percentage of every short made in response to bad news they dug up about a company? What if governments paid in proportion to how valuable academic ideas turned out to be, centuries after they were discovered, and modern-day academics sold shares in that future payout to fund themselves? What if prediction markets something something?

For the moment, academics and journalists are both in a weird middle space. They’re truthseekers, still, by culture and inclination and desire. But they’re paid for something else.

Don’t Judge an Explanation by Its Cover

Dark matter bugs people.

I’ve talked before about why, and why it, and other beyond-the-standard-model proposals like those inspired by MOND, are nonetheless credible with physicists. But beyond the logic in that post, there’s a deeper reason people find dark matter strange. It’s that they don’t know what kind of an explanation dark matter is.

Dark matter sounds very lazy. If you can’t explain the movements of stars based on the matter you can see, then proposing invisible matter sounds like the easy way out. But it’s actually a lot less easy than it sounds, because matter is something quite specific. Matter gravitates and bends light. Matter moves. Matter can be described with a pressure, one like gas and dust and not like other things like light or the Higgs field. If you propose a new type of matter, you have to check and see that all of those consequences hold, with detailed implications for almost every observation every astronomer takes.

For the most part, those consequences have been checked, and they do hold. Sometimes they fail, and it’s those failures, and not the idea that dark matter is “lazy”, that drive dark matter’s critics in the physics profession. Physicists who oppose dark matter have other explanations with their own consequences, for example new types of quantum fields that often get described to the public as “modified gravity”. When they argue against dark matter, they do it by comparing those consequences in detail, working through the implications and seeing which phenomena hold.

Dark matter, as it turns out, is a very constraining explanation, one with strict consequences. There are other corners of physics where the explanations may seem less lazy, but actually have fewer consequences, and thereby less scientific heft.

For example, consider the debate about evidence for dark energy I wrote about last month. A key question there was how to interpret light from supernovae. Some groups argued that supernovae change in brightness with distance, others that they change based on how old their galaxies are. Sabine Hossenfelder glossed the debate by saying it comes down to how you model supernovae. And while that’s true, it can give the wrong impression.

You might think that these people are comparing detailed computer models of supernovae, and making different assumptions when they set their models up. But in reality, it’s much less detailed. The people on both sides of this debate are looking at correlations, trying to draw statistical lines through supernova datasets. The difference between one model and another isn’t a complicated physical setup you can put into a simulation, it’s just which lines on a graph you account for and which you ignore.

Because of that, while these models may sound much more sophisticated than dark matter, they actually have much less scientific weight. The different supernova models don’t have grand, widespread consequences, they’re not mucking with the laws of physics or proposing new classes of object that every astronomer needs to account for. They’re pretty much just proposing tweaks to how to interpret one very specific type of data. That makes their questions much harder to resolve, and their answers much less universally convincing.

If you’re not a scientist, if you read science news, it can be hard to tell the difference. Some ideas in science may sound simple, but have a whole raft of consequences that distinguish them from other ideas. Others may sound sophisticated, but are much more like “fudge factors”, only distinguished by statistical arguments, not by a rich trail of qualitative evidence.

For the most part, as an outsider, you’ll never know which is which. But as always, it’s best to be aware of your limits.

Newsworthiness Guide for Scientists

I had a recurring “elevator pitch” at Lancefest earlier this summer. After explaining that I’m a science journalist now, I’d end with “so if you run into a story, let me know!”

One person had a question that left me stumped: “What counts as a story?”

For those of us who don’t happen to be Einstein

I didn’t have a good response then. I’ve got a better one now, though I’m afraid it doesn’t fit in an elevator pitch. This is all based on my experience, so take it with a grain of salt. But here are the criteria that seem to matter:

First, a story usually needs a news hook. News is, in particular, supposed to be “new”. That doesn’t mean I can’t write about history, or established science. But editors like those stories a lot better if there is some recent development, within the past year or so, to tie it to. The new development doesn’t have to be all that important, the story can mostly focus on something else. But it needs to be somewhere in there.

Second, news stories are usually qualitative, not quantitative. I need to be able to tell a story about what happened, what actions people took and why they mattered. Quantitative developments usually only make the news if they’re so big that they shade into the qualitative: something doubling unexpectedly, for example.

Third, ideally a science news story is something that is getting the experts excited. Journalists aren’t supposed to judge the scientific merit of ideas on their own, they’re supposed to rely on experts. The most solid stories, the ones that are easiest to pitch, are ones where there’s a community of experts that largely think something is cool. That makes it easier to get good quotes, and easier to justify its relevance. If you accomplished something and you’re having trouble convincing anyone it matters, don’t start with me, start with your colleagues!

Fourth: less importantly, it helps when stories have a human angle. If you can tell a tale about how you came up with an idea, if you came from an unusual background, if something was hotly debated but now is deemed essential: these things sweeten a story, they capture readers’ interest, and editors see their value.

Finally, stories involve something changing. It can be something that just changed now, for a news piece, but it can also be something that changed over time, for a feature in a magazine. The key is change. “Old method still works” is not going to excite people, and it won’t count as news.

After writing all that out, I’m still not sure I answered the original question. But hopefully I’ve at least given some tips that can get you started. If you’re a scientist, and you see something that hits most of the boxes on this list but hasn’t been covered in the news yet, consider reaching out to me. You may have run into a story!

Better Bounds

I swear this isn’t turning into an AI blog. But did you see the one about the Riemann hypothesis?

Someone at Anthropic did something I’m sure they’re all tempted to do, and tried to use an internal version of their Claude AI system to prove the most famous open conjecture in mathematics. It didn’t work, to be clear, and I get the impression they didn’t expect it to. But out of six hundred or so fruitless tries, one attempt did prove a new bound. Previously, mathematicians had been able to prove that at least 41.6% of the zeroes of the Riemann zeta function satisfied the Riemann hypothesis. Now, the new proof shows that at least 67.2% satisfy it.

Anthropic’s press release is impressively careful. As someone who’s had to think about how to write content that both excites the public and doesn’t piss off experts too much, they do an admirable job walking that line. They even say, straight-out, “We don’t expect that the techniques Claude used will lead to proving the Riemann hypothesis.”

Bounds are like that, sometimes.

I should know. Physicists also find bounds.

Physics has its own conjectures with the fame of the Riemann hypothesis. Dark matter might be made of detectable particles. Protons could decay. There might be extra dimensions, or magnetic monopoles, or cosmic strings. General relativity might be subtly wrong.

It would be an amazing achievement to demonstrate any of these things. But most physicists won’t manage that. Instead, they bound them.

Physicists compete to get better bounds, excluding unusual possibilities with greater and greater care. They find evidence that dark matter can’t be of a specific mass with a specific charge, so the next experiment has to look somewhere else, or find evidence that general relativity holds to even greater precision, so any deviation must be even smaller. Some work to improve experiments with better and better bounds. Others analyze data from older experiments, or find under-appreciated consequences of known facts, and can get even better bounds.

Bounds aren’t typically newsworthy (though occasionally they make it through), so most people don’t hear about them. If you read the news, you hear about positive claims much more often than negative ones: evidence for something new, not evidence that our current knowledge holds. But the nature of physics is that most work supports the status quo. Most work improves bounds.

Do bounds lead, with time, to the positive claims? Sometimes, but not always. Often, bounds are just bounds. They’re attempts to use the methods physicists have to learn something new about the world. Even if the new fact is just “don’t look here”.

It Only Counts When AI Gets to My Field

It’s a meme at this point.

When Deep Blue beat Kasparov, Go players could say that their game, unlike Chess, was too complex to fall to a computer program. Then AlphaGo showed they were wrong. It just took a better approach.

When AlphaFold leaped ahead of human experts in predicting how proteins fold, it was due to a mountain of carefully labeled protein structure data. Other scientists and mathematicians could argue that nothing like that existed in their field, so a similar success was unlikely. But LLMs can now navigate scientific literature, and loosely imitate the reasoning process of a mathematical proof. And increasingly, the math and computer science results coming out of AI labs are ones that humans find impressive.

Now, experts argue about how far AI can really go. Will AI mathematics only be good at finding counterexamples and solving cute puzzles, not introducing new concepts and frameworks? Will AI only be meaningfully good at fields like mathematics with clear rules and carefully collected conjectures, not fuzzier fields like physics? Will AI-powered labs only manage to optimize specific procedures, and not carry out entire experimental programs? Each time, the pattern seems to be that scholars are skeptical, until AI gets to their field.

I’m aware of this pattern. But I’m willing to take the risk.

I think my old field, scattering amplitudes, is special. And when AI can do something meaningful there, I’ll really start to worry.

By “something meaningful”, I don’t mean the student-level results that have come out so far. I mean tackling some of the field’s big outstanding problems: determining whether N=8 supergravity diverges at seven loops, or finding the six-particle amplitude in N=4 super Yang-Mills to nine loops. Getting another loop past the state of the art for gravitational wave physics or collider physics would also count.

These problems are difficult not just because people haven’t had the right ideas, but because they’re hard in a computational sense. Each loop, a rough measure of the precision of the end result, represents an increase in complexity, in calculations that typically scale exponentially or even factorially in the number of loops. In principle, amplitudes researchers could do any of these with no new ideas, just using known methods. They’d just need access to a lot more computing power.

See, while everyone else is preoccupied with whether AI can come up with genuinely new ideas, I think the real measure is what those ideas accomplish. And the most important measure of accomplishment, if you’re worried about how scared to be about AI, is whether it can do things that seem like they would take too much computing power.

In the past, when people dreamed up the scariest hypothetical things AI could achieve, critics argued they were impossible due to a lack of computing power. Apocalypse scenarios often involve designing self-replicating nanobots based on computer models of molecules, or unstoppable social manipulation based on simulating the minds of the humans the AI interacts with. If AI is going to manage these things, or something like them, it will take an approach that somehow bypasses that need for more computers than we can build.

So if AI companies want to impress people like me (or scare us, for that matter), then they need to tackle my old field. Show that an AI can take the kinds of computer resources an academic has access to, and solve one of the scattering amplitudes field’s big outstanding problems. Show that a computational limit everyone expected to be a problem doesn’t actually matter. Give us N=8 supergravity to seven loops, or N=4 super Yang-Mills to nine loops.

Bonus Info on Dark Energy and Muons

I had two pieces up this month, one in New Scientist and one in Quanta. I figured I’d give a bit of “bonus info” for both here.

The New Scientist piece covered a paper arguing that, according to the supernova evidence, there is no need for dark energy, because the universe isn’t speeding up after all. That’s a pretty dramatic claim, and it’s one the majority of cosmologists don’t agree with. But a small group has been hammering at the consensus.

Long-time followers of this blog have already heard of these folks. I criticized news coverage around a paper with one of the authors, Subir Sarkar, back in 2016, when he was only arguing that the supernova evidence was insufficient to establish dark energy, not that it was literally the wrong way around. He and his co-authors wanted an opportunity to state their case, so I posted their response a few weeks later. Later, I got to know Rameez, another author on the new paper. He ended up doing a guest post in 2019, on the next step in the story.

That story went from a statistical critique to an active proposal for something outside the consensus, in the form of inhomogeneous cosmology, the idea that the universe may be lumpier than typically assumed, with flowing currents that explain things otherwise attributed to exotic physics. That’s another topic I’ve covered before, but in this article I didn’t have the space to go into it.

Sarkar and Rameez are claiming something much more radical than most proponents of inhomogeneous cosmology, arguing against dark energy as a whole, not just more subtle effects. They told me a lot about their reasoning, most of which also didn’t make it into the piece.

It’s largely not going to make it here, either. Not due to space limitations: I’ve been pushing the blog post lengths you folks tolerate for a long time, no reason to stop now. Rather, it’s because I genuinely don’t feel qualified to judge this.

The issue is, cosmology is messy.

How do you figure out how the universe is expanding? You can look at supernovae, and use how dim they are to estimate how far away they are. The issue is, supernovae vary, even famous “standard candles”: they move in different ways, they shine differently depending on different things. Simulations aren’t nearly advanced enough to tell you all this from first principles, and so there are a wash of different corrections based largely on empirical observations, taking this or that correlation seriously and factoring them out of the measurement, while disregarding other correlations as spurious. Which of these combinations of corrections you trust determines whether the universe seems to be accelerating or decelerating.

It makes me glad I went into particle physics instead of cosmology, where you can mostly model everything from first-principles, and compare to your models.

Of course, even that can get you into trouble.

Take that Quanta piece. I wrote an update of the muon g-2 measurement. In some ways, this is an issue that many think of as already solved, by precisely that kind of first-principles modeling. Muons seemed to be violating the Standard Model, according to a calculation based in part on empirical formulas. Lattice QCD researchers figured out how to do that part of the prediction via a first-principles simulation instead, with enough accuracy that it could substitute for the empirical formulas. And lo and behold, the new prediction agreed with the experiment. Eventually, the experts accepted it, and all was well.

Except, as I commented at the time, not really. Because the empirical formulas were based on other experiments, colliding electrons and positrons. And as I now understand in much more detail, the disagreement between those electron-positron experiments is worryingly large.

So nobody is resting on their laurels. The lattice folks improved their calculation, with a result published in Nature this year that prompted Quanta to ask me to report on it. The people using the empirical method are still trying to sort out what happened. And the experimentalists are busy scrutinizing how they analyze their data, and collecting and analyzing more.

A few details that didn’t make it into the piece:

First: it goes beyond the level I was aiming for in the piece, but I really do want to emphasize the role of error estimates. Ultimately, every step in the story was not just one where experiments and predictions disagree, but one where they disagree past their reported error. The problem with disagreement between the new experiments isn’t that they disagree, it’s that they disagree so badly that they’re straining the statistical methods the people using the empirical method use to combine results together, so badly that if taken seriously, those methods would have to throw away ten years of progress of increasing precision. I don’t expect it, but I really hope for a postmortem in which we learn how to better estimate the kinds of experimental errors that can cause this. I haven’t seen that kind of postmortem for other results, so I don’t expect one here. But I have to believe that in the background someone is learning something, and getting better at this.

Second: one interesting question my editors raised was whether it was possible to do a first-principles calculation to compare with the electron-positron experiments. The answer is yes, but with some caveats. One will be recognizable to physicists: lattice QCD can only compute energy-integrated cross-sections, not measurements at particular energies like experiments find. But they can compute integrals with a modulating function…such as one strongly peaked at a particular energy. It’s a familiar type of trick for a quantum field theorist, though in this case it’s trickier than it sounds. I actually misunderstood, and thought that the same people I was talking to about the muon g-2 calculation had done that, and found it agreed with one of the electron-positron experiments over another. That’s not true: their comparisons were more indirect, while other groups attempting more direct comparisons haven’t quite gotten something good enough to do more than gesture at an agreement. So while the direction was correct, there are indeed lattice-based suggestions that favor one experiment over the others, the piece phrased things much too definitely. There’ll be a correction fixing this.

Quantum Apologetics and Quantum Theology

As an atheist, I started out frustrated by how little interest religious people had in debating their beliefs. Much of that was probably to do with how obnoxious it was to be “debated” by a socially awkward ten-year-old. But as I appreciate now, defending religious beliefs is just not a core activity for most religious people, even the experts. While there are many theologians who study the doctrines of their various religions, only a few engage in apologetics: arguments designed to convince people on the outside. The rest work within a particular religion, working out its implications.

There’s a similar, less-often-noticed behavior when it comes to interpretations of quantum mechanics.

Much like religions, there are many different interpretations of quantum mechanics, from people who envision a fundamentally undetermined world to those who picture a vast multiverse of all possibilities, to people who think quantum mechanics needs to be supplemented with faster-than-light signals, deterministic rules, or even consciousness. And as there isn’t yet any broad consensus for any of these options, you’d be forgiven for assuming that these people are trying their hardest to convince others that their interpretation is right.

But most of the work these people do is “quantum theology”, not “quantum apologetics”. The average paper connected to a quantum interpretation isn’t designed to convince people with different interpretations. It starts out with an interpretation in place, and tackles a more detailed question of what the interpretation should actually mean in practice. This work can be quite valuable and impressive…provided you’re already convinced. But if you’re not, and you see someone glowingly praise a paper on say the many-worlds interpretation, you might mistakenly think they’re on the cusp of closing the question of which interpretation is right, and not merely solving a technical issue within many-worlds.

Are there other areas of science with this pattern?

Let’s talk about string theory.

Physicists really started getting excited about string theory around forty years ago. Some people hear that number and wonder what happened. Have string theorists been looking for evidence for forty years and not found any? Isn’t that a huge waste of time?

That would be “string apologetics”. And these days, string apologetics is not actually that common. It’s a priority for some, to be sure, but most string theorists aren’t working on proving string theory. It’s clearly not an easy thing to do, and as a result, most people don’t spend their time on it.

Realizing that, some people assume that string theorists are actually practicing “string theology”, and get mad all over again. If string theorists are just assuming string theory and spending their time figuring out the “string versions” of known facts, then many would also deride their work as pointless.

But actually, string theology is also not very common. There are certainly some people who work to figure out the “string version” of this or that, or do research that only makes sense assuming string theory. But most of the string theory community doesn’t do that either.

Instead, they do something that, to continue the analogy, you might call “string pastoral care”.

Most theologians aren’t apologists, but similarly, most priests don’t spend their time doing theology. They use their training to advise their congregations on how to live their lives, solving day-to-day problems with some religious inspiration.

Similarly, most people these days who call themselves string theorists are working on more general questions about the types of theories used in particle physics. They’re doing this making use of their background in string theory, as an inspiration for solutions, a source of mathematical tools to solve problems, and a motivation for which questions are the most interesting. But if string theory turns out to be false, most of these peoples’ work will still be useful. It’s “string pastoral care”, used to solve problems for the people around them, not “string theology”.

Do you know any other fields that this applies to? Let me know in the comments!