It Got to My Field

For the folks who found this blog from Anthropic’s site, welcome!

For everyone else, I’d better give some context.

Last month, I had a blog post titled “It Only Counts When AI Gets to My Field”. The title was a joke, the content less so. I said that if some of the longstanding problems of my old research field got solved by AI, then I’d sit up and take notice.

The post turned out to be a bit of a self-fulfilling prophecy, after some folks at Anthropic read it and decided to tackle one of the outstanding problems I mentioned. They then invited me to write a post about it on their science blog.

The post is here. I recommend reading it, then coming back. The rest of this post will be a little Q&A.

Q: At the end of that post, it says Anthropic paid you for your time. Can we trust what you wrote?

A: When I agreed to write the post, I made it clear that I was going to give my own opinion, not write an ad for Anthropic. They could propose light edits, but that’s it. And I avoided signing anything with them, not even an NDA, so I could freely tell you if they pushed my boundaries.

They didn’t push my boundaries. They asked me to clarify a few things, and to give more detail on the science. I didn’t change the message or the takeaways.

The reason I asked them to pay me is that, as a freelance writer, I don’t have a salary to fall back on. Time I spend on a project like that post is time I’m not spending on journalistic projects, so if I worked on it for free I’d essentially be using my vacation time for it. The rate I’m charging them is roughly in the middle of what I’d have gotten paid if I used that time on journalism: a bit more than I would have made working for the lower-paying outlets, a bit less than I would have made with the higher-paying ones.

I’m also not expecting this to be the start of a longer-term business relationship or anything like that. So overall, I don’t think I’m incentivized to lie on their behalf. You can trust me.

Q: How confident are you that they did what they said they did?

A: I didn’t get the feeling they were lying. But I didn’t start out skeptical.

I haven’t seen the logs from the LLM, or anything like that. That’s the kind of thing I would dig into if I were more suspicious of their story. If there’s a reason to be suspicious, I’ll ask.

But so far, I don’t feel that I have much reason to be suspicious. They were able to look up details on the fly when I asked them, they didn’t seem to have a polished message they were trying to push past me. And more importantly, as I’ll mention in a bit, I don’t think anything they described is all that outlandish. It lines up, largely, with the capabilities I’d expect their Claude Science platform to have.

It’s also relevant that they didn’t use an internal model for this. It means that scientists will likely be trying out similar problems soon, so if it does turn out they exaggerated something, people are going to figure out quite quickly.

Q: So how big is this? We just saw AI solve a Millennium problem, after all.

A: This was a lot easier than a Millennium problem. But it was also a lot cheaper.

The problem they solved is one with a clear recipe, honed and explained over multiple papers. It’s something I expected to be hard to do without access to a lot of computer power, so I thought it would need to be approached with a novel technique, in order to avoid using that computer power.

In the end, it didn’t need that. As I mention in the post, another group got most of the result at around the same time, with a much smaller amount of AI assistance.

It’s not an easy recipe, to be clear. I think other people will be surprised that a reasonably affordable program like Claude Science can do this without a lot of guidance. I’m not surprised, mostly, because I’ve been paying enough attention to what people have been doing with these things, and carrying out this kind of recipe consistently is actually something the good science AIs can do right now. If you haven’t been following as closely, that’s going to be a lot more surprising.

So maybe the best way to answer this is: instead of paying $15 million to solve one of the most famous problems in mathematics, they paid $1,000 or so to make a significant next step in an ongoing research program in theoretical physics. This doesn’t tell you much about whether AI can achieve field-defining breakthroughs, but it tells you a lot about the kinds of things a theoretical physicist with $1000 spare budget can do right now.

(As an aside, it feels crazy to me that 90% or more of the compute cost was from the LLM, not the calculation itself. On the one hand, it feels nuts to essentially use ten times more computer power to do this than would have been needed if a human had done the coding. On the other hand, I’ve applied for grants that budgeted more than that per year for travel costs, and spending this kind of money on getting a result seems a lot more useful than spending it on airfare.)

Q: Isn’t it reckless to publicly post a problem for AI like that?

A: I posted my challenge before OpenAI announced their Navier-Stokes result. At that point, there had been a few awkward surprises, but for the most part AI companies had a pattern of talking to an academic first before trying to solve a problem. It’s what they did back in March.

I expected that was what they’d do this time, and got more than a little blindsided when they just solved the problem on their own and reached out to Lance afterwards. I’m lucky that Lance doesn’t seem mad at me about it, it would have been quite understandable if he was.

I did tell them that if they wanted to tackle any of the other problems in that post, they should reach out to one of the scientists involved first, before attempting it. In addition to being polite, it’s a way to make sure that they understand the problem correctly and have a plan to verify the result. They happened to pick a challenge that was particularly easy to verify, but it didn’t have to go that way.

Q: Any thoughts about AI’s future impact on science?

A: If there’s a problem you can’t solve, but you think someone with expertise in another field or better programming skills could, then it’s probably solvable with AI.

A lot of problems don’t fall into that category. I’ve seen people talking online about AI solving quantum gravity, but solving quantum gravity is a question of which bullets you’re willing to bite, not a question of technical skill. That doesn’t mean AI will never be able to address it, but if so it will come from some sort of superpersuader aspect, not merely scientific capabilities.

And of course, some fields do require experiments. People are increasingly building AI-powered labs. I’ll leave it to people from actual experimental fields to think about the potential there.

Q: Can you say something explicit about the bigger risks from AI, now?

A: As I mention in the post, I didn’t learn that much about the bigger questions from this. I’m still not an expert.

But there’s one thing I think is worth emphasizing:

This technology is clearly getting more effective over time. I don’t think they could have done this six months ago. If you’re trying to predict what will happen next, you shouldn’t just assume this is the most powerful it will get, or the cheapest. If you think there’s a limit, you need to argue for it.

Leave a comment! If it's your first time, it will go into moderation.