It Got to My Field

For the folks who found this blog from Anthropic’s site, welcome!

For everyone else, I’d better give some context.

Last month, I had a blog post titled “It Only Counts When AI Gets to My Field”. The title was a joke, the content less so. I said that if some of the longstanding problems of my old research field got solved by AI, then I’d sit up and take notice.

The post turned out to be a bit of a self-fulfilling prophecy, after some folks at Anthropic read it and decided to tackle one of the outstanding problems I mentioned. They then invited me to write a post about it on their science blog.

The post is here. I recommend reading it, then coming back. The rest of this post will be a little Q&A.

Q: At the end of that post, it says Anthropic paid you for your time. Can we trust what you wrote?

A: When I agreed to write the post, I made it clear that I was going to give my own opinion, not write an ad for Anthropic. They could propose light edits, but that’s it. And I avoided signing anything with them, not even an NDA, so I could freely tell you if they pushed my boundaries.

They didn’t push my boundaries. They asked me to clarify a few things, and to give more detail on the science. I didn’t change the message or the takeaways.

The reason I asked them to pay me is that, as a freelance writer, I don’t have a salary to fall back on. Time I spend on a project like that post is time I’m not spending on journalistic projects, so if I worked on it for free I’d essentially be using my vacation time for it. The rate I’m charging them is roughly in the middle of what I’d have gotten paid if I used that time on journalism: a bit more than I would have made working for the lower-paying outlets, a bit less than I would have made with the higher-paying ones.

I’m also not expecting this to be the start of a longer-term business relationship or anything like that. So overall, I don’t think I’m incentivized to lie on their behalf. You can trust me.

Q: How confident are you that they did what they said they did?

A: I didn’t get the feeling they were lying. But I didn’t start out skeptical.

I haven’t seen the logs from the LLM, or anything like that. That’s the kind of thing I would dig into if I were more suspicious of their story. If there’s a reason to be suspicious, I’ll ask.

But so far, I don’t feel that I have much reason to be suspicious. They were able to look up details on the fly when I asked them, they didn’t seem to have a polished message they were trying to push past me. And more importantly, as I’ll mention in a bit, I don’t think anything they described is all that outlandish. It lines up, largely, with the capabilities I’d expect their Claude Science platform to have.

It’s also relevant that they didn’t use an internal model for this. It means that scientists will likely be trying out similar problems soon, so if it does turn out they exaggerated something, people are going to figure out quite quickly.

Q: So how big is this? We just saw AI solve a Millennium problem, after all.

A: This was a lot easier than a Millennium problem. But it was also a lot cheaper.

The problem they solved is one with a clear recipe, honed and explained over multiple papers. It’s something I expected to be hard to do without access to a lot of computer power, so I thought it would need to be approached with a novel technique, in order to avoid using that computer power.

In the end, it didn’t need that. As I mention in the post, another group got most of the result at around the same time, with a much smaller amount of AI assistance.

It’s not an easy recipe, to be clear. I think other people will be surprised that a reasonably affordable program like Claude Science can do this without a lot of guidance. I’m not surprised, mostly, because I’ve been paying enough attention to what people have been doing with these things, and carrying out this kind of recipe consistently is actually something the good science AIs can do right now. If you haven’t been following as closely, that’s going to be a lot more surprising.

So maybe the best way to answer this is: instead of paying $15 million to solve one of the most famous problems in mathematics, they paid $1,000 or so to make a significant next step in an ongoing research program in theoretical physics. This doesn’t tell you much about whether AI can achieve field-defining breakthroughs, but it tells you a lot about the kinds of things a theoretical physicist with $1000 spare budget can do right now.

(As an aside, it feels crazy to me that 90% or more of the compute cost was from the LLM, not the calculation itself. On the one hand, it feels nuts to essentially use ten times more computer power to do this than would have been needed if a human had done the coding. On the other hand, I’ve applied for grants that budgeted more than that per year for travel costs, and spending this kind of money on getting a result seems a lot more useful than spending it on airfare.)

Q: Isn’t it reckless to publicly post a problem for AI like that?

A: I posted my challenge before OpenAI announced their Navier-Stokes result. At that point, there had been a few awkward surprises, but for the most part AI companies had a pattern of talking to an academic first before trying to solve a problem. It’s what they did back in March.

I expected that was what they’d do this time, and got more than a little blindsided when they just solved the problem on their own and reached out to Lance afterwards. I’m lucky that Lance doesn’t seem mad at me about it, it would have been quite understandable if he was.

I did tell them that if they wanted to tackle any of the other problems in that post, they should reach out to one of the scientists involved first, before attempting it. In addition to being polite, it’s a way to make sure that they understand the problem correctly and have a plan to verify the result. They happened to pick a challenge that was particularly easy to verify, but it didn’t have to go that way.

Q: Any thoughts about AI’s future impact on science?

A: If there’s a problem you can’t solve, but you think someone with expertise in another field or better programming skills could, then it’s probably solvable with AI.

A lot of problems don’t fall into that category. I’ve seen people talking online about AI solving quantum gravity, but solving quantum gravity is a question of which bullets you’re willing to bite, not a question of technical skill. That doesn’t mean AI will never be able to address it, but if so it will come from some sort of superpersuader aspect, not merely scientific capabilities.

And of course, some fields do require experiments. People are increasingly building AI-powered labs. I’ll leave it to people from actual experimental fields to think about the potential there.

Q: Can you say something explicit about the bigger risks from AI, now?

A: As I mention in the post, I didn’t learn that much about the bigger questions from this. I’m still not an expert.

But there’s one thing I think is worth emphasizing:

This technology is clearly getting more effective over time. I don’t think they could have done this six months ago. If you’re trying to predict what will happen next, you shouldn’t just assume this is the most powerful it will get, or the cheapest. If you think there’s a limit, you need to argue for it.

Edit: One more thing I should mention, now that I’ve confirmed she’s ok with it: the idea for my challenge came from some discussions at Lancefest, a conference in honor of Lance back in June, and particularly from Anastasia Volovich, whose talk declared the nine-loop calculation “a benchmark – or a ‘challenge’ – against which any “AI takeover” should be measured”.

12 thoughts on “It Got to My Field”

  1. Karl Young's avatarKarl Young

    Hey Matt,

    I’ve been following and enjoying the blog (as well as your pieces in Quanta) for a while now.

    I thought this was a really interesting story re. Claude as the new Feynman (OK maybe he had a few ideas as well as being better at integrals than any other human…).

    But it was an offhand comment that got me wondering; i.e. that AI was pretty much capable at the student level, and probably not a stretch to figure it’s at the level of a grad student or will be soon. So maybe this could be a bit of a danger to the field ? E.g. if you’re a student hell bent on getting tenure at Harvard why would you bother with “trivial” calculations re. standard assignments, if avoiding that provided even a slight advantage. I suppose the first order gate keepers are things like tests and qualifying exams, but it’s a little concerning that the whole field might slowly slip into a state of relying on automated curation of fundamental knowledge, and that seems like it could go in better or worse directions for the field as a whole.

    Liked by 1 person

    Reply ↓
    1. 4gravitons's avatar4gravitons Post author

      I think one of the takeaways of this post is that it’s no longer “student level” anymore, since these are calculations that typically take an expert to do in a sufficiently reliable way.

      But yes, in general the question of how to approach training these days is a big problem! I think we definitely need to keep reminding students that they need to actually practice more basic skills in order to get more advanced skills. But that’s likely not enough. And that’s putting aside the possibility of the technology getting even better, and what the community ought to look like if that happens.

      Liked by 1 person

      Reply ↓
  2. boldly91f5a7d879's avatarboldly91f5a7d879

    It’s not totally clear what kind of challenge you were trying to raise. If you wanted to see if AI could produce a significant result, it did. If you wanted to see how much computer power would be required, you found out. If you wanted to see if AI could do a problem that was hard for humans, it could. But if you wanted to see if AI could find a better method, you didn’t, and I think it is because you confused the last two. The 9-loop problem was hard for humans because it involved a lot of calculations over a lot of stages that required a lot of checking. But it was not that hard for AI because it had a model to replicate and against which to test its code before proceeding to solve the same problem for one more loop. And you can see in the Claude output how it used those exemplars. Basically, there is nothing too complicated for AI to do as long as it can discover an adequate description of the calculation, an exemplar on which to test, and a demonstration of how to present the results.

    However, if you had a problem that was known to have a solution for N dimensions (or some other feature) and not to have a solution for N+2 dimensions, and you asked it either to prove there was no solution for N+1 dimensions or to provide the solution, you would create a totally different workflow and require a lot more searching and trying. Would AI come up with a new insight? It might come up with a method from a remote discipline that it could map to the problem, but as for something completely original, I doubt it. It might find a proof by contradiction; Claude used proof by construction over the last few months to disprove Paul Erdős’s Unit Distance Conjecture and The Jacobian Conjecture.

    Liked by 1 person

    Reply ↓
    1. 4gravitons's avatar4gravitons Post author

      I didn’t think the problem was hard for humans purely because it required a lot of checking. I thought it was hard for humans because doing it the normal well-understood way would take more computational resources than the humans wanted to spend.

      As it happens, I was wrong about that. But I think you’d have to admit there should be a threshold where that would kick in. At some point, if you’d ask Claude to get the L-loop amplitude while only using $100 of external compute, it would have to invent new techniques, or it would fail. I’d just mistakenly assumed that L was 9. It still could be 10 or 11, or 20. It’s almost certainly lower than 100, right?

      Liked by 1 person

      Reply ↓
  3. Pingback: Some HEP-TH News | Not Even Wrong

  4. EternalLurker's avatarEternalLurker

    I think you’re missing the actual insight here (though admittedly I know nothing about amplitudeology outside of your post’s overview). Your claim appears to be that “oh well my expectation of the computation limit was just too low”. But that IS the key insight: the fact that the expected limit is almost always too low.

    Programmers have been telling people this for decades. “Everything in reality is just math in the end” -> “practically every problem is solvable with more compute”. And Moore’s Law therefore means that everything is solvable eventually. Your disappointment that “AI finally got to my field” was actually “compute finally got to my field” is understandable, but I think it sidesteps the actual realization, which is that compute gets to everything eventually; the universe just isn’t as complicated or “beautiful” as people like to think.

    (The irony is that programmers also frequently think their own field is the one immune to this. Hence Sutton’s Bitter Lesson trying to remind them all that actually this is even more true than normal in machine learning, empirically.)

    I think viewing this as “AI didn’t accomplish that much, just some compute shenanigans” is completely backwards, when in truth this is yet more evidence that “compute shenanigans will get to every field yes even your field even creative stuff like art”, and a huge part of the intelligence of AI IS just the raw power it can bring to bear (including through agent swarms in parallelizable problems).

    Liked by 1 person

    Reply ↓
    1. 4gravitons's avatar4gravitons Post author

      I think you’re missing a key aspect here: it wasn’t using that much compute!

      I picked the problem specifically because I knew that, if you had enough compute, it was easy. The point was to do it with roughly the same amount of compute people had tried it with before. This isn’t a case of “compute came for the field”. It looks like it could be a case of “professionalism came for the field”, if the SymPy setup was more efficient than the CASes people had been using before. But you shouldn’t expect professionalism to come for every field, right? Some fields should already be working to a professional level.

      Liked by 1 person

      Reply ↓
      1. EternalLurker's avatarEternalLurker

        I do understand that you thought the problem was easy with “enough compute”, but I got the impression that your issue was in overestimating how much would be “enough”; is that not what you meant in the article by “If people thought it were possible to just run the usual bootstrap method for another loop they would’ve,” and by “I was too naïve about where the computation limit was”?

        You yourself say in the article that “Running 96 CPUs for a week might have felt like a lot when I was doing this kind of work ten years ago, but it’s pretty affordable now”. And 96 CPUs today are not 96 CPUs of ten years ago. You claim in this comment that it “wasn’t using that much compute”, but the amount of compute it used WOULD have been a lot by prior standards, as per your own admission! That’s exactly my point; that is a perfect example of compute coming for a field, because the costs of compute are constantly coming down. What would have been computationally intractable in the past becomes less and less so as Moore’s Law continues to (mostly) hold.

        Liked by 1 person

        Reply ↓
        1. EternalLurker's avatarEternalLurker

          To clarify (maybe I should make an account so I can edit comments lol), I guess there are two separate issues here: “is increasing compute making problems tractable that previously weren’t” and “are we underestimating what problems we can solve with current (affordable) compute”. I’m arguing that the answer to BOTH of those is “yes”, with high confidence in the former versus relatively uninformed speculation for the latter (your article being a data point in its favor IMO), but I probably shouldn’t have combined them into one point like that; it’s a bit of a conflation of distinct questions, my bad.

          Liked by 2 people

          Reply ↓
        2. 4gravitons's avatar4gravitons Post author

          Ok, I think I see the misunderstanding. I personally haven’t been doing these types of calculations recently, but Lance and others were. My last paper bootstrapping an amplitude in N=4 was published in 2019, and that stage of the calculation was largely complete before that. What I was trusting in was that the people who were still working on these things were using the currently available compute as well as could be expected (sort of an academic version of the efficient market hypothesis).

          Liked by 1 person

          Reply ↓
          1. JollyJoker's avatarJollyJoker

            Lance is presumably the only one who knows about the compute requirements at this point. It would be interesting to see if it can do 10 loops with the same methods. And of course if the result tells us something more than “we have a result”.

            Liked by 2 people

            Reply ↓
        3. 4gravitons's avatar4gravitons Post author

          Doubling back on this re: Moore’s Law, with a thought that should help explain some of where my expectations were coming from: scattering amplitudes generically grow in complexity factorially, so super-exponentially. So one does expect it to take longer and longer for Moore’s Law to catch up, just by itself at least.

          Liked by 1 person

          Reply ↓

Leave a comment! If it's your first time, it will go into moderation.