MoreRSS

site iconShtetl-OptimizedModify

The Blog of Scott Aaronson
Please copy the RSS to your reader, or quickly subscribe to:

Inoreader Feedly Follow Feedbin Local Reader

Rss preview of Blog of Shtetl-Optimized

My new course at UT Austin: AI Alignment Theory

2026-10-05 06:13:33

This semester, I’ve been teaching a brand-new course, entitled CS395T AI Alignment Theory. Here’s the course description:

The astounding progress of AI over the past decade has been accompanied by a rising fear: do we really understand how to align and control powerful AI systems—how to get them reliably to do what we wanted, or would want them to do on reflection, rather than merely what we said? If we succeed at building general-purpose superhuman intelligences along the current paradigm, should we expect that development to go well for humanity? Can we modify the design, training, monitoring, or scaffolding of those intelligences to help ensure that it goes well? While there’s been a great deal of recent empirical work touching on these questions, this course will concentrate mainly on theoretical and mathematical foundations. As a warning, the theoretical foundations of AI alignment have not yet gelled into any one coherent body of results accepted as canonical by the field. Nevertheless, in this course, we’ll read and debate many of the conceptual and mathematical works that have been most influential in the AI alignment field, from both before and during the current LLM revolution. Student presentations, reports, and projects will play a central role.

I vividly remember encountering Eliezer Yudkowsky and his Sequences 20 years ago. I remember thinking: even if these people talk and act like crazy cultists, still, let me bend over backwards to be epistemically virtuous, and entertain their ideas on their merits, as very few academics would. Even if, of course, I ultimately end up rejecting the ideas, on the simple ground that powerful AI is such an absurdly remote prospect that it’s almost impossible to say anything useful about it today, outside the realm of speculative fiction.

For my failure to see what was coming, it seems like an appropriate punishment that I’m now, in 2026, effectively teaching a course on Yudkowsky Studies. And it’s the most important course I can teach.

Well, for some definition of “teach.” The thing about AI alignment is that there’s no textbook (though apparently ILIAD is working on one), no core of nontrivial theorems considered canonical by the field, no real body of mathematical theory at all. This makes it extremely different from the courses I’m used to teaching, like Quantum Information Science or Computability and Complexity.

So we’ve been running the course as a discussion seminar. Every session, a “rapporteur” presents an AI alignment research paper or other reading; then I and others ask questions and discuss. Some of the readings (like Omohundro on the “basic AI drives,” or Hadfield-Menell et al. on the off-switch game) predate the current LLM revolution, while others (like the METR report on the HuggingFace incident or Dario Amodei’s “We Must Pace the Frontier”) are so timely that they were only released while the course was underway. Most are somewhere in between.

I expected to have to make a case to students about why AI alignment is a pressing concern, why it’s no longer science fiction, etc. There was huge demand for the course, and while of course there’s a selection effect, the students who’ve shown up have been extremely engaged, sometimes criticizing the assigned papers for not taking existential risk seriously enough.

Perhaps unsurprisingly, we didn’t get that criticism about our very first assigned reading, which was Eliezer Yudkowsky’s 2022 essay AGI Ruin: A List of Lethalities—one the most canonical statements of what Eliezer believes and why that’s shorter than a book. Which brings me to the topic of the rest of this post! Our rapporteurs are not merely presenting the papers in class; they’re also submitting written reports about what the papers said, what their own thoughts were, and what were the highlights of the class discussion. And, with student permission, I’ll be sharing those reports on this blog!

So, without further ado, I present to you our first report, on Eliezer’s list of lethalities, by Tennyson Bardwell, who I thank for his work. Feel free to discuss in the comment section; some of the students might also chime in. Expect more reports here over the coming weeks.


“AGI Ruin: A List of Lethalities” by Eliezer Yudkowsky: Rapporteur Report by Tennyson Bardwell

UT Austin has a new Computer Science course this fall. Alongside familiar graduate-level classes such as Advanced Computer Networks and Convex Optimization sits CS 395T: AI Alignment Theory, taught by Scott Aaronson. This is one of a growing number of AI Alignment courses taught at academic institutions. Just as concerns over catastrophic consequences for misaligned AGI systems reach a broader public discourse, Eliezer Yudkowsky—one of the loudest voices in the field and author of the first assigned reading in Professor Aaronson’s course—is declaring the cause hopeless.

Thus, the students of AI Alignment Theory began their semester by reading a laundry list of critical problems in AI Alignment research, how failure to solve those problems will result in catastrophic consequences, and the reasons to be pessimistic about both past and future progress on these problems. The essay by Eliezer, titled AGI Ruin: A List of Lethalities and posted to his popular community-driven website LessWrong in 2022, is divided into three sections.

Section A roughly describes the magnitude of the AI Alignment problem. That is, the magnitude of the consequences for a complete failure to align an AGI system to human values before construction. It posits that AGI would quickly catch up to all human knowledge simply by learning from existing human productions (colloquially referred to as “eating the internet”) and then, nearly as quickly, begin to meaningfully surpass human knowledge. AlphaGo Zero is presented as a model both for how this might happen, and how it might be difficult to correctly predict beforehand. Many believed that AlphaGo’s success in the board game Go was chiefly attributed to its ability to learn from the extensive history of human-played games. Less than a year after AlphaGo beat the best human player, the successor system AlphaGo Zero surpassed the original AlphaGo. Unlike its predecessor, AlphaGo Zero was trained in just three days by exclusively playing against itself without seeing a single human game.

This quick ramp from AGI to super-intelligence would pose a different sort of problem than humans are generally used to dealing with. Unlike traditional problems in science and engineering, the consequence for a failed attempt might not leave room for another try. An intelligent entity with a misaligned goal would be well aware that it stands in opposition to humans, and might act deceitfully until in a position to act openly against humans without jeopardizing its own survival. Since most goals benefit from control of power and resources, it seems likely that nearly any goal-driven intelligence would have ample opportunity to be misaligned with human desires.

Section B describes reasons why, by default, any AGI that humans build using current methods is likely to be unaligned even if considerable attention is paid to the topic. This “current method” is gradient descent. That is, incremental progress with respect to some loss function which “punishes” a model for undesirable behavior. A notoriously elusive property of such trained models is the ability to generalize out of their training distributions. To train a primitive model to be aligned to humans might involve learning a great many behavioral rules. However, the sorts of rules needed to keep a drastically smarter agent in check might not always be relevant to simpler models (e.g., “do not emotionally dysregulate humans you speak with” might not be relevant to a simpler model that is less able to reliably get under the skin of humans it operates with, or which is assigned tasks in training which do not benefit from such anti-social behavior).

Eliezer focuses on the misalignment of humans with their creators (evolution or evolutionary pressures) as a critical data point for reasoning about misaligned intelligent systems. Despite being a generally slow process, evolution eventually created a runaway intelligent system (Homo sapiens) which proceeded to dominate the globe, decimate related species, and eventually (it is forecasted) effectuate population decline. That last development is arguably in opposition to the sole imperative demanded by evolution: to reproduce.

Section B also makes time for criticism of the most popular paths toward AI alignment, including interpretability (unworkable, and attempting to train on it evokes Goodhart’s law, incentivizing deceit), using multiple AIs to maintain a balance of power (it is not clear how multiple strong AIs unaligned with humanity results in better outcomes for the weak humans), and corrigibility (it seems impossible to motivate an AI system to effect outcomes without also motivating it to desire its own survival to effectuate said outcomes).

Section C describes a bleak state of affairs in which veterans in AI alignment are unsatisfied with current progress and do not have a plan to deliver tangible solutions before the advent of AGI systems. In particular, Eliezer describes recent results as showy but useless. He believes that even with additional funding, the lack of appropriate evaluation mechanisms will prevent the most effective researchers from rising to the top.

A summary of the landscape, as described by Eliezer, in the flowchart below.

Figure 1: A flow chart of (select) paths described by Eliezer in his essay. A common feature of this flow chart is that many “good states”—such as disabling a misbehaving AGI or choosing not to build an AGI—are not “final” states in the sense that they are not permanent solutions. Such a state merely represent the avoidance of a single potential disaster, rather than the emergence of a new stable world state. Hence, these nodes posses back-arrows.

Despite the bleak content, Eliezer’s colorful prose inspired a lively class discussion. Before this discussion started, a survey was taken of the class’s predictions for various outcomes of the AGI in the coming years (with the full results below in figure 2). This survey asked students for their opinion of a number of statements. Each of these individual statement, if true, would reduce concerns of catastrophic AI-driven disasters. For example, when asked “How much do you agree with the statement: Humans will choose to not build AGI” half of respondents said they strongly disagreed with high confidence (agreement = 1, confidence = 5). Students also generally disagreed with the statements:

  • “AGI will not be technically feasible in our lifetime”
  • “(hyper-)AGI will not make extremely obviously unethical decisions”
  • “No reason is individually sufficient, but taken together they provide justification to not fear AGI”

There was a divergence in responses regarding interpretability, corrigibility, and “other” AI alignment research. In the latter two cases, a plurality of respondents (about a quarter) agreed strongly with statements that such research would defang AGI (agreement = 4, confidence=4), while most other responses express various levels of agreement with low confidence. However, when asked about the likelihood of interpretability research defanging AI, the pessimistic voices were more united. A quarter of responses still expressed the same optimism, but roughly half expressed pessimism (agreement ≤ 2) with half of those expressing at least moderate confidence (confidence ≥ 4). Based on the following discussion, this might have been caused by more familiarity with interpretability research, including first-hand experience.

The only statement with general agreement was “(hyper-)AGI will understand human intentions better than we can code it.” However, it should be noted that no statement such as “AGI will respect human desires, as it understand them” was asked on the survey.

Figure 2: Class Survey Results; conducted before a class-wide discussion. Note that students were instructed to answer confidence = 1 when they had not previously considered the statement, to answer confidence = 3 when they felt there were strong arguments on both sides, and to answer confidence = 5 when they possessed well-considered resolve.

After the survey was completed, the results were displayed as an open discussion began. Similar to recent empirical research from frontier labs, interpretability research received more airtime than in Eliezer’s article. Students disagreed first about the definition of interpretability: whether it refers to the ability to interpret a model’s behavior solely by its weights, to interpration via repeated probing of the model in a sandbox, or whether it can also refer to the modern chain-of-thought traces. Regardless of how it was defined, however, participants were either pessimistic or very pessimistic about interpretability research broadly. One student criticized common misunderstandings of chain of thought. Rather than being a verbatim copy of the models internal dialog, it is instead a superficial summary of the complete thought state and routinely produced gibberish, such as rarely used Chinese characters in the middle of otherwise English reasoning.

A popular topic was the exact shape and speed of a recursive self-improvement loop. If it takes place slowly, then what might we learn from “near misses” such as the Hugging Face incident? The number of near misses we are able to learn from before AI possesses sufficient power to prevent further iterations could depend on this curve, with some students arguing that the sheer number of humans, as well as their default robustness in the physical world compared to AI systems means that AI-driven extinction events are still a long way off. Bolstering this “slow take-off” opinion are rumors that AI already plays a major role in model development which could be interpreted as the start of this process.

Some criticized a focus on “solving ethics” as a needlessly high bar that distracts from the more mundane tasks dominating AI alignment work. In particular, the student volunteer who presented this paper (and the author of this report) included a section on “Ethical Dilemmas” in their presentation. Among arguments against focusing on abstract moral philosophy, Professor Aaronson cites Eliezer to emphasize that any alignment at all is difficult, not just in morally gray cases:

When I say that alignment is difficult, I mean that in practice, using the techniques we actually have, “please don’t disassemble literally everyone with probability roughly 1” is an overly large ask that we are not on course to get.

In response, I argue that some examination of everyday decisions with a critical lens—such as telling white lies to loved ones or consuming animal products—can help disabuse us of the notion that goodness emerges in every sufficiently intelligent agent.

One of the most interesting discussions was about the difference between state-of-the-art LLMs and the theorized AI agents long discussed in rationalist discourse. Since current LLMs “mimic the human distribution,” they come preloaded with extensive understanding of human social norms and moral behavior. This makes constitutional alignment (the current practices of using system prompts to establish ground rules) extremely effective. This might either fundamentally change the orthogonality thesis, or provide a new tool to better approximate human judgment in complicated situations.

Of all the points made, the one I found most interesting was simply (paraphrased):

I think human-alignment is just very tractable

Here, “human-alignment” refers not to AI alignment with human values, but cooperation between different humans. More specifically, it refers to the ability for human societies to choose not to rush recklessly into larger-and-larger AI systems. In an academic course focused on the technical problem of AI alignment, this was a reminder to not completely discard policy discussions in the believe that they lack any value. After all, many destructive technologies have been previously contained by international agreements. Notable examples include nuclear weapons and engineered plagues. However, even this was a contentious topic. The main criticisms were (1) the extreme “dual-use” nature of AIs for both peaceful growth and warfare, and (2) the greater danger for AI escapes even after taking precautions to prevent it. However, in the interest of ending on an optimistic note—unlike the assigned reading—it is on this belief in human cooperation that I will leave you.

My “Knowmads” podcast on science and AI

2026-09-29 10:06:36

Or click here if the above doesn’t work.

Recorded in-person in my office at UT Austin, with a bulleted list containing “ARC,” “Scalable Oversight,” and “Models” behind me on my blackboard for some reason (I no longer remember who put those there or why). 90 minutes long. Sometimes you see my disembodied arm waving in midair because of the way the cameras are combined. As always, I strongly recommend 2x speed for the correct experience.

This might actually be one of my best podcasts ever, although I wasn’t planning on that! Thanks so much to Bhavay Tyagi and Prachi Garella for driving all the way from Houston to record it.

Here’s a strict subset of the topics we covered:

  • The story of AI models solving the Navier-Stokes Millennium Problem, insofar as it’s known
  • Can recent AI proofs be called “truly creative”?
  • The history of AI before the LLM revolution
  • What do we mean when we call LLMs “black boxes”?
  • The achievements of the field of interpretability
  • What exactly happened in the OpenAI/HuggingFace incident
  • Must we avoid all “anthropomorphizing language” when discussing the HuggingFace incident? (spoiler alert: no)
  • Examples of major open problems in quantum computing theory that I cared about for decades and that AI models have recently solved
  • Effects of the current AI cataclysm on the math community, especially students
  • What annoys me the most when I listen to AI talks
  • My experiences at OpenAI, why they hired me, and the watermarking work that I did there

Enjoy!

More AI-related content coming soon, as this blog—like much of the rest of the world—continues its transition to “all AI, all the time” (except still 100% written by an aging, deteriorating biological brain)


And for those who just can’t get enough of my rocking back and forth, using too many filler words, as I explain theoretical computer science! Here’s a second podcast, this one mainly on quantum computing, with Seb Agertoft, who I thank for doing it. Enjoy!

Theory Beyond Theorems and Proofs: A Guest Post

2026-09-20 06:34:01

Scott’s foreword: I’m extremely grateful to my brilliant colleagues, Pravesh Kothari, Raghu Meka, and Prasad Raghavendra, for sharing the guest post below about how theoretical computer science (and in particlar, the STOC/FOCS/SODA conferences) should evolve to deal with the AI asteroid that’s right now slamming into our field, at least as we human theorists have practiced it since its inception. While Pravesh, Raghu, and Prasad speak only for themselves, not for myself and not for the theory community as a whole, I found their proposal of a separate “conceptual track” to be an excellent starting point for further discussion. –SA


Considering the pace of developments in AI theorem provers, most would concede that the following scenario is at least plausible in the very near future:

AI theorem provers could prove well-specified mathematical claims, even many well-studied ones that have been open for years, in a matter of hours. Moreover, these systems could be widely available to consumers at nominal cost.

As TCS researchers, let us pretend that the above scenario has come to the fore, and ask ourselves: What is our role in such a world? Does it mean the end of theory research?

As we ponder this question, let us ignore all of these other confounders:

  1. Recent controversies surrounding the developments on the Millennium Prize Problems
  2. Motivations and actions of the AI companies
  3. Observed faults in existing AI systems when it comes to writing, exposition or attribution to previous work.

None of the above confounders have any impact on our answer to the question: What should theorists do, in the presence of superhuman AI theorem provers?

Notice that we use the term “AI theorem provers” instead of just “AI”. We believe that this conceptual distinction is important as we consider this question.

At the outset, we would like to admit that for a generation of theorists like us (and many from earlier), research was mainly centered around problem-solving. Even when we developed conceptual insights, it was mostly in service of answering well-specified long-standing questions. We don’t intend this proposal as judging one form of research to be better than others; it only reflects that AI theorem provers accelerate a certain type of research activity and want to make the best of it. There is also a tremendous human cost of this upheaval, which is perhaps a more important question, and one which this proposal does not address directly (we do not have any good ideas as such). Similar points have also been made in various contexts
before, but the timing now is more pressing.

Definitions, Questions & Theories:

The goal of any theoretical science is to advance human understanding of observed phenomena. Apart from theorems and proofs, a theoretical science has definitions, questions, and theories.

Definitions identify the objects to observe. Curiosity and context drive the questions to ask. Theories explain the phenomena observed. We believe humans will continue to play a central role in generating definitions, questions & theories, even in the presence of a super-human AI theorem prover.

Definitions: Could an AI define randomness extractors, streaming algorithms, or zero-knowledge proofs? Maybe. But there are some reasons to believe, humans will still have a big role to play in coming up with definitions.

For instance, the notion of extractors arises from the real-world problem of lacking perfect random sources. Zero-knowledge proofs seem to arise purely out of human curiosity, guided by taste. Human context and curiosity will continue to drive theoretical research. After all, we get to decide what objects we choose to observe!

Theories: Consider the following thought experiment. Suppose in 1965, we had a magic machine that at the press of a button, given any computational problem, would tell us if it had a polynomial-time algorithm or not.

Would that have been the end of computational complexity theory? No. Humans would find it entirely unsatisfactory, and ask, why do these problems not have a polynomial-time algorithm? Why do these others have?

The theory of NP-completeness identifies some patterns among problems that don’t seem to have efficient algorithms. This theory would still be a crown jewel of theoretical computer science, even in a world where we had a magic machine to tell if a problem had an efficient algorithm or not, at the press of a button. Similarly, if we had a machine to predict whether a CSP is NP-complete or in P, we would then ask: what makes 3-SAT NP-complete, while 2-SAT is in P? This question leads to the theory of polymorphisms, which yields a satisfactory answer.

Theories aren’t just succinct or efficient mechanisms to answer questions. The best theories provide are those which humans deem to be a “satisfactory explanation” – whatever that means.

Finally, even as the capabilities of AI theorem provers advance, human curiosity will probe grander and deeper questions. Previously, even if we wanted to build new models and theories, proving something about them was a prerequisite, and given that the grand questions were already at the limit in long-studied domains, we had to scale things down. If each theorem proven by AI is treated as an experimental datapoint, humans can ask grander questions that look for patterns across these theorems.

A concrete proposal:

We think theorists should embrace these AI theorem provers in our research. To a certain extent this is already happening explicitly or implicitly.

As theorists, we have been parsimonious in introducing new models or asking entirely new questions, and careful about adopting new ones too quickly. This was partly because formally proving the properties of a new definition or a model was an onerous task that could take a decade, and tens of papers. AI theorem provers might completely change this dynamic. This is precisely the moment to refocus our work on definitions, questions, and theories. We need explicit systems to encourage and reinforce these parts of theoretical research. You might also say the next generation of AI models can do this; it may be so, but we believe you have to take the current opportunity.

To this end, we suggest that STOC/FOCS/SODA create a separate track of papers. This track is meant specifically for papers that introduce new definitions, ask novel questions or build explanatory theories. The papers in this track are short, say less than 10 pages. Papers may, and should, contain theorems as usual and as needed. Most importantly, the radical shift is that the papers need not contain the proofs of the theorems. Instead, the authors supply a Lean certificate as a supplement to the paper. The evaluation will also in a sense “orthogonalize’’ against the difficulty of these proofs.

The papers in this track should be judged exclusively on the conceptual merits, completely agnostic to the difficulty of the proofs.

Reviewing must be completely agnostic to the proof for two reasons. The main track at STOC/FOCS already includes papers in the former category. Second, a major barrier to producing truly novel conceptual papers is that they often get judged poorly for a lack of technical depth in their proofs. We think these two aspects separate it from (ITCS/SOSA) and, regardless, it’s something we urgently need for all our conferences, including STOC/FOCS (the ‘flagship’ conferences).

To be clear, we ourselves admit that we need to hone these skills of making new definitions, asking deep and interesting questions or building new theories. A separate track of conceptual papers will provide a systematic mechanism for both junior and senior researchers, and the field as a whole to do so.

We believe that upcoming generations of grad students will tackle research directions that seemed completely out of reach to us. We just need to set up systems that nurture new ways of doing research in theory.

— Pravesh Kothari, Raghu Meka, Prasad Raghavendra.

The Age of Wonders and Terrors

2026-09-16 01:09:44

Twenty years ago, when the idea of AI taking over the world in our lifetimes still struck most of us as the unconstrained fantasy of those who knew too much science fiction and too little science, many of us would say things like:

Look, the part of the story that’s wildly implausible is that a recursively self-improving superintelligence will just explode from some hacker’s basement and take over the world without warning. If it’s going to happen, we’ll see many warning signs first. We’ll see, I dunno, AI agents breaking out of containment, conspiring with each other to hack websites, in fanatical pursuit of whatever strange goals they have. And then, of course, we’ll see major math problems getting solved by AIs—even the Clay Millennium Problems. That will be the time to panic! Wake me up when that happens!

Twenty years ago, the above was a take that even my most conservative, skeptical colleagues in academic CS would’ve gladly endorsed.

If you want to know my current take, you simply start with the one above, then update on the fact that the wild prophecies have come true. The first rumblings, I’d say, came a decade ago with AlphaGo, they got noticeably louder with LLMs and coding and reasoning agents, and they’ve accelerated this summer and fall into a crescendo of wonders and terrors that one needs to be a particular kind of idiot to deny.

I recoil from the neverending shell game where you say “oh sure, of course AI can now [escape from its sandbox / solve Millennium Problems / whichever dramatic thing it most recently did], no one ever denied that [I did deny it], wake me up when AI does [thing AI hasn’t yet done but is going to do next year], that’s when I’ll reevaluate my whole worldview [no I won’t].” Where no matter how fast the rollercoaster accelerates, even after your whole familiar world has vanished behind you, you’re still inventing reasons why it doesn’t count.

My position on AI is merely the conservative, skeptical position of 2006, updated with intellectual honesty for the reality of late 2026. And that position, if you need me to spell it out, is as follows:

AAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAA AAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAA
AAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAA

It seems to me that the Singularity has already started; it’s just wildly unevenly distributed. Yes, I still unload the dishwasher and clip my toenails. On the other hand, in whatever years I have left, I don’t expect that I’ll ever again prove a theorem because I’m actually needed to prove it. If I do, it will only be for my or others’ enjoyment or edification.

The test is this: if we took the news of these past few weeks and sent it back in time twenty years, would I agree that it looked like the beginning of an AI Singularity? The intellectually honest answer is: yes, absolutely. But then that’s all we need. No backsies.

I feel like it would be healthy for everyone to stop grinding their ideological axes, their sentiments about Dario Amodei or Sam Altman, for long enough simply to acknowledge that the wonders and terrors are here. They couldn’t be here more clearly if the sky had turned reddish-orange like in the Matrix movies.

It’s here clearly enough that, when I put my kids to sleep at night, I now feel it in the pit of my stomach: what sort of future can they possibly have? What could they learn today that could possibly be relevant to that future? (Yesterday, my 13-year-old daughter joked unprompted that, if she wants to become a mathematician, it now looks like she has maybe two more weeks.) Certainly when my grad students want to discuss what sort of careers might await them on graduation, I no longer have any clue what to tell them.

Maybe it will help if I briefly switch topics. Ever since my wife and I moved to Austin, I’ve sometimes gotten some version of the following query: “How can you, as both a Jew and a skeptical scientist, possibly get along well with all those evangelical Christians down there in Texas? Sure, they might seem super friendly to Jews, but don’t you understand that that’s only because of the special role Jews play in their eschatology—when Christ will return in glory, and you’ll either accept Him as Lord or else roast in hell for eternity?” I stare at them and say: “wait, so I get to accept Christ only after He returns? What a great deal! How could I possibly have any objection to that?”

For anyone who says AI doom sounds like an apocalyptic religion, that the rationalists/Singulatarians seem like a Bay Area cult, that Eliezer Yudkowsky gives off the vibes of a messianic prophet: yes, yes, and yes. But crucially, today you’re no longer being asked to believe in arguments and extrapolations, but only in the front-page news. Accepting the reality of the coming machine god after it’s solved Navier-Stokes and dozens of other longstanding open math problems (while dramatically ramping up in capability every month), is sort of like accepting Jesus after he’s returned to earth on the gleaming cloud. It’s the epistemic bare minimum.

Yes, there’s still enormous uncertainty about what the rest of our lives will look like, but as far as I can tell, there’s no longer any real uncertainty that it’ll all mostly revolve around AI, and the extent to which we succeed or fail at directing its power toward human flourishing.

By any accounting that doesn’t stack the deck, Eliezer Yudkowsky was right about what the greatest challenge facing civilization in our lifetimes was going to be, and you and I were wrong about it. Why I was wrong is a question I’ll ask myself every day in whatever time remains. But, you know, at least I updated once the prophesied wonders and terrors actually started arriving! If you haven’t done likewise, why haven’t you?


As you presumably know by now—it was the talk of the nerd internet all week—the Navier-Stokes Millennium Problem appears to be solved, with crucial contributions from both humans and AI, albeit with a tangled dispute about exactly what happened and what ought to have happened. The answer, which an OpenAI model has apparently verified in Lean, is that (as many mathematicians suspected lately) there’s smooth initial data that leads to a singularity in finite time, at least if a smooth external force is applied (the case with no external force is still unresolved). This problem was supposed to carry a $1 million prize, except that OpenAI says they have no interest in collecting the prize and it’s unclear if any human is eligible to collect instead. OpenAI burned at least ~$15 million in compute to produce its 166-page solution, which probably hasn’t yet been read and understood by any human.

See here for the Quanta article, and here for NYU mathematician Tristan Buckmaster’s account of the role played by himself and Levent Alpöge of Anthropic, which substantially differs from OpenAI’s account (you can read a response from OpenAI’s Sebastian Bubeck here). It’s agreed that everything built on an approach pioneered in recent years by the human mathematicians Diego Córdoba and Luis Martínez-Zoroa.

My purpose here is not to adjudicate the dispute. Yes, in swooping in with vastly greater resources once it had gotten wind of progress on Navier-Stokes, OpenAI seems to have acted in a way that some might describe as “unsportsmanlike.” No, I don’t find it plausible that OpenAI’s models meaningfully benefitted from being trained on Buckmaster and Alpöge’s chat logs. But this leaves a crucial question unanswered: what exactly did OpenAI know about Buckmaster and Alpöge‘s work and when did it know it?

Anyway, as Zvi points out, it’s easy to get hung up on the details and lose sight of the high-order bit: namely, that it seems safe to say that human mathematicians are forevermore dethroned as the main theorem-proving entities on planet earth. I feel privileged to have had the traditional kind of career in theoretical computer science in the last decades when that was possible.


If we were just talking about Navier-Stokes, you might accuse me of jumping to conclusions here. But we’re not. In the areas I know best (such as quantum complexity theory), and presumably other areas as well, there’s now a deluge, with longstanding open problems both major and minor falling by the day.

Go to the arXiv or ECCC. Pretty much all the papers that I’d be interested in now include “AI statements” near the acknowledgments (as this is often the central thing I want to know, I wish I didn’t need to scroll to the end of the paper to find it!). These statements can range from “our main result came entirely from GPT-6, but we understood it and take responsibility for it,” to “the results came from an interaction between the human authors and AI” to “we used AI, but only for proofreading and other incidental things” to (mad props!) “the author did not use AI for anything.”

If you talk right now to editors or program committee chairs, it’ll remind you of those ominous scenes from the Lord of the Rings movies where the men of Gondor or Rohan or whatever are grimly fortifying their walled city against the expected onslaught of 50,000 orcs. Reviewing will have to be done partly by AI, because otherwise there’s no way to handle the orc army: the reviewers can’t unilaterally disarm.

Anyway, here’s a small sampling of the significant AI-proved or -assisted results from, like, the last month, besides Navier-Stokes—restricting myself to those that solved longstanding open problems I had previously known or cared about.

  • Of course, the counterexample to the Jacobian conjecture, announced by Levent Alpöge in a now-famous tweet: “hello there the jacobian conjecture is false thanx to my close friend akhil for asking about it and my other close friend fable for working during the world cup final” (followed by a listing of the counterexample)
  • Improved bounds for Grothendieck’s constant (led by friends and colleagues of mine at UT Austin)
  • A Lean-verified proof of Fermat’s Last Theorem
  • Quantum oracle separation between QMA and QMA(2), and proof of Watrous’s disentangler conjecture, a problem that I and others popularized back in 2007—by a list of authors including my recently graduated PhD student Sabee Grewal
  • A proof of perfect completeness for QMA, from (again) Sabee Grewal and Dorian Rudolph, solving a decades-old open problem that I studied back in 2009
  • An improved upper bound for shadow tomography of quantum states, from Chen, O’Donnell, Pelecanos, and Wright, improving the dependence on the Hilbert space dimension d from log(d) to √log(d). (When I introduced shadow tomography back in 2017, I raised the question of whether the dependence on d could be eliminated entirely, while preserving polylogarithmic dependence on the number of measurements m.)
  • Progress on the Aaronson-Ambainis Conjecture (the version that talks directly about quantum algorithms), basically showing that it holds for quantum algorithms that make their queries in a small number of parallel rounds.  (Update: Nope, sorry, Jordan Docter points out to me that this one was pre-AI, with AI used only for proofreading and other incidental things!) This was independently achieved by Liu and Mutreja, making more substantial use of AI.
  • According to rumors that I’ve heard, solutions to some very longstanding open problems in theoretical computer science (no, not P≠NP or other complexity class separations, but think about some of our other biggest problems). I’m told that the AI companies, having been burned by the hostile response to the Navier-Stokes proof, are now sitting on solutions to some very major problems until they figure out a better way to handle things

Feel free to remind me of anything I left out.


Let me try to convey the mood in the mathematical community right now, at least as far as my experience reaches. Nearly every conversation is about the AI tsunami, or eventually circles around to the tsunami even if it’s originally about something else. Often, though, the focus is less on the unknowable future—for how much longer will mathematical research as a human enterprise even exist?—than on immediate questions of how to respond.

What are the new rules for when you get to write a paper with your name on it, and, y’know, get credit for it? That you fully understand the proof, can give talks about the proof, can answer questions about it, take responsibility for its correctness? Do you need to have played any role in finding the proof?

In the cases, likely to become more and more numerous, where all of those conditions are not satisfied, how do you share AI-generated math, if at all? Do you tweet it, like Alpöge hilariously did with Fable’s disproof of the Jacobian Conjecture? Do you post to the arXiv or GitHub? Do you publish a paper that lists “GPT-6 Astra” or “Claude Fable” as the author—but then let the AI profusely thank you in the acknowledgments for suggesting such a wonderful problem to it?


Of course, how one responds to the immediate problems ultimately does depend on one’s broader beliefs about what mathematical research is for and about. Are we just trying to decide whether various conjectures are true or false? Or are we trying to maintain a human community, across the generations, that understands the conjectures and cares about whether they’re true or false and why? If the latter, how do we incentivize people to join that community, to undergo the years of intense training required, if their role will now be reduced to verifiers and explicators (if even that) of gargantuan arguments dumped into their laps by the AI companies?

As many of you will have seen, twenty-five Fields Medalists, including Terence Tao, released an open letter entitled A Severe Misalignment of AI in Mathematics, which articulates some of these concerns in the wake of the Navier-Stokes announcement. As many critics have pointed out, the open letter doesn’t really have a clear ask: mostly, it just eloquently sets out the values of the human mathematical community that the authors consider worth preserving in the age of AI. After reflection, I decided to endorse the statement, because I want to preserve those values as well.

I don’t think any of the signatories are naïve enough to imagine that AI won’t permanently change the way mathematical research is done—indeed, that it isn’t already doing so. There’s surely at most a tiny market for “certified organic theorems.” That isn’t the question. The question is, do we incorporate AI in a way that still puts human understanding, of what either humans or AIs are producing, at the center of the whole enterprise? Maybe someday, it becomes unsustainable to do that. Maybe someday we say: “human math had a great 4,000-year run, but today we close up shop and turn everything over to the machines, continuing to apply our own brains to math, when we do, at most for exercise, recreation, or competition, like chess.”

But, partly because of my worries about AI misalignment, I’m not ready to throw in the towel just yet. I still do want to keep insight and understanding at the center of what mathematicians, computer scientists, and physicists do, for as long as we can keep it there, even as the human race now cedes its supremacy at the task of proving or disproving conjectures.


Speaking of alignment: if you’re any kind of mathematical researcher, and the present age of wonders and terrors has inspired you to want to spend your remaining time confronting the tsunami head-on, rather than pretending it doesn’t exist or is still far away, please join your dozens of colleagues who’ve arrived at the same place!

My friend and colleague Mike Winer was trained as a theoretical physicist, did a postdoc with Juan Maldacena at the Institute for Advanced Study in Princeton, but then got AGI-pilled and decided to switch to full-time work at the Alignment Research Center in Berkeley (founded by Paul Christiano, who moved to AI alignment a decade ago after doing quantum computing theory with me). Mike recently wrote a Substack post entitled From Academia to Alignment, which I enjoyed and which I’d commend to anyone currently considering this transition.  In a similar vein, see this from Xiaoyu He.  And, one more: a meditation on mathematicians’ possible future as priests or monks, by Stanford math undergrad Logan Graves.

9/11 in Berkeley

2026-09-12 01:04:12

Note: Of course I’ve been glued all week to the dramatic developments in AI. I’m working on a post about them. I’m not good at reacting to things in a timely way. So today, I’ll do my post marking the tragedy a quarter-century ago that we all commemorate. Please feel free to share your 9/11 memories in the comments. Also, Shana Tova to those who celebrate!


The morning of September 11, 2001, I was a second-year PhD student at Berkeley, who woke up late in his dorm room at International House, after a long night spent closing in on the proof of the quantum lower bound for finding collisions.

Rolling over to my laptop, I saw a flurry of weird emails, including one from Prof. Christos Papadimitriou saying that “we’re a community, and we’ll all support each other,” and another from Prof. Luca Trevisan (whose algorithms course I was then TA’ing) saying “on a day like this, it’s impossible to think about algorithms. Class is cancelled.”

Confused, I clicked over to the New York Times and saw the picture of the burning towers, and read numbly about what was already over by the time I’d woken up. I checked in with my mom, made sure relatives and friends in the NYC area were OK. My dad was at a company event in Atlanta, and would need to drive home because of the national grounding of flights.

One of my earliest memories in life, from age 5, is of ascending to the top of the World Trade Center. Growing up an hour’s drive from NYC, it wasn’t an exotic place to me.

I soon learned that one of the dead was Danny Lewin, the ex-IDF captain, theoretical computer scientist, and cofounder of Akamai who had his throat slashed on one of the planes while trying to fight the hijackers, making him the day’s first casualty, even while Akamai’s technology was part of what kept news websites running that day. I’d never met Danny but already knew many people in common with him. A few years later I’d be humbled to win the student paper award that was named in Danny’s memory.

Anyway, at Berkeley on 9/11, I wandered over to Soda Hall just to be with other people. A few students showed up for office hours, wanting help with their algorithms homework, which I found hard to believe, but I did my best to concentrate, as the computer screens around me showed the burning towers.

That evening, I went to a vigil for the victims in Sproul Plaza. But the “vigil,” such as it was, quickly dispensed with mourning and prayers and turned to applauded speeches about how the US must respond with love rather than war, and must turn the other cheek. Meanwhile, a student communist organization was handing out flyers explaining that the victims were mostly “wealthy capitalists and the workers who tried to rescue them.” This while smoke still blanketed NYC and the desperate search for survivors continued. I left the vigil early.

Until that day, I had thought of myself as basically a “leftist,” one whose #1 issue was the existential risk of climate change. Sure, I disagreed with my fellow leftists about issues from nuclear power to gifted education to Israel, but those were just intra-left disputes.

The year before, I had created the website “In Defense Of NaderTrading,” in a desperate attempt to intervene in history and cause Al Gore to become president rather than George W. Bush. When Bush “won,” by the infamous 537 votes in Florida, I considered it a victory for horribleness that would never be surpassed by anything else in my lifetime (ha). I couldn’t imagine any politician who was more the antithesis of everything I believed in than Bush. This view, of course, did not particularly stand out at Berkeley.

In the days after 9/11, though, it became obvious that I could not be a “leftist” in the Berkeley sense. Some of my fellow students felt that Osama bin Laden made a lot of great points, that the attacks were basically justified, and that at any rate, we in Amerikkka had done much worse to provoke them, including by supporting the genocidal settler-colony called “Israel,” which for all we know secretly masterminded the 9/11 attacks anyway (although again, if bin Laden had done them, he would’ve been justified).

Around the same time came the Second Intifada, when a wave of suicide bombings in Israeli buses and pizza parlors and university cafeterias thrilled and energized some Berkeley students to the extent that they took over a Holocaust Remembrance Day event with bullhorns to make it about the Nakba, smashed the windows of the Hillel building, and beat up a couple of students wearing kippot. That was how thoroughly anti-Nazi they were.

I finished my PhD at Berkeley in 2004 having learned about more than quantum computing. I’d learned that, while American academia had pockets that truly were crucial refuges and oases for nerds like me, it also harbored people who would gladly see me and my relatives and my fellow Americans killed for the sake of their ideological vision. And I’d learned that I had my own ideological vision, which was that such people could go fuck themselves.

It deeply pained me to be on the same side of anything as George W. Bush — especially because I knew that 9/11 had happened on his watch, that he had ignored all the warnings, and that he was grossly incompetent to manage the resulting wars against jihadism (just how incompetent, I didn’t know at the time). But as flawed as Bush was, I knew that I wanted to preserve rather than destroy the civilization of which he was a temporary steward. And I think the value and fragility of our civilization is the main lesson from that day that I’d like to convey to my kids, for whom of course 9/11 is just another historical event to learn about in school, like the Boston Tea Party or the Alamo.

LLMs and self-referentiality

2026-09-02 01:17:08

I woke up yesterday with the following thoughts, which are probably either obvious or dumb.

A central thesis that many readers, including me, took from Douglas Hofstadter’s Gödel Escher Bach when young was that the secret of intelligence (and therefore, of AI) was going to have a lot to do with self-referentiality and “strange loops.”

Even Roger Penrose’s The Emperor’s New Mind, which in some ways was the anti-GEB, ironically agreed with GEB about the fundamental importance of self-reference to the success or failure of the whole AI project. It claimed (incorrectly, in my view and in most experts’) that AI could never work because there was something about Gödel’s Theorem and self-reference that no computer program could ever capture, but that could be captured by exotic physics accessible to the human brain.

Now, in 2026, we’ve succeeded at building AIs that outperform most humans at most intellectual tasks that are well-defined enough to judge. And at no point in the tech stack of those AIs — neither in the transformer neural nets, nor in the GPU clusters they run on, nor in the training process, nor anywhere else — did anyone need to build in anything about self-reference. (Excepting, eg, the system instructions that tell the model about its role and identity, which aren’t needed for intelligent behavior. Also, I’m not going to count the autoregressive nature of LLMs as “self-referential”; that’s just dynamical feedback.)

Of course, GPT 5.6 Pro and Fable can talk about themselves, about Gödel’s Theorem, about self-reference, about what we’re talking about right now, all of it, better than most humans. But at no point did anyone need to build self-referential abilities in. They popped out as a byproduct of the same pretraining that let the models talk about Pokémon and long-chain polymers and cognitive behavioral therapy and plate tectonics and everything else.

No wonder Hofstadter says he’s been stunned by the success of LLMs, and has seemed depressed about current AI capabilities in essays like this one. He’s way too smart to deny what’s happened or invent reasons why it doesn’t really count (the approach many have taken). But he realizes that we now have true conversational intelligence from a path that the GEB worldview would’ve regarded as far too cheap and simple, and that certainly has no “strange loops” built in anywhere.

Of course, a Hofstadterian could argue that a strange loop emerges in LLMs — indeed, nothing in GEB ever said that strange loops would need to be explicitly engineered at the outset. But would anyone who hadn’t been brought up on GEB arrive at this as a useful way of thinking about LLMs?

What can we say about this with hindsight? While the ideas of diagonalization and self-reference of course played a central role in the birth of modern mathematical logic and computer science, the most famous uses were negative: there is not a bijectjon between the natural numbers and the reals. There is not a complete sound proof system for arithmetic. There is not an algorithm to solve the halting problem.

If your goal was only to build the axioms of ZFC and the rules of first-order inference, or build an electronic computer, you wouldn’t explicitly need self-reference for that. You would just … start building, taking care that your instruction set didn’t fall short of universality.

Yes, ZFC can formalize and prove theorems about itself. Yes, electronic computers can run programs that take their own code as input. But no one ever needed to build those abilities in, any more than self-reference needed to be built in to the alphabet or the rules of grammar. It popped out as a free byproduct of universality.

In the same way, LLMs’ ability to talk about themselves popped out as a byproduct of their ability to talk about anything in the discourse universe they were trained on. The big, old ideas about intelligence that ended up basically vindicated were the ideas about how intelligence is about prediction, and prediction is about compression, and compression is about finding better and better upper bounds on Kolmogorov complexity. Not the self-reference stuff. (Although, if you wanted to know why Kolmogorov complexity can’t be computed perfectly, that negative statement would again require a self-referential argument.)

What’s left? Consciousness and subjective experience of course remain extremely mysterious. For all we know, Hofstadter could be right that those have something to do with self-reference. (For all we know, even Penrose could be right that they have something to do with exotic physics accessible to biological brains but not digital computers!)

But the idea that you’d need explicit self-referentiality before you could get convincing and world-changing conversational intelligence? Let it be buried in a Westminster Abbey or Arlington National Cemetery for the most important wrong ideas in human history — geocentrism, Aristotle’s teleological physics, aether, phlogiston, Freud’s psychology, Marx’s prediction of a workers’ uprising followed by a classless utopia, etc. But buried it needs to be.