MoreRSS

site iconExponential ViewModify

By Azeem Azhar, an expert on artificial intelligence and exponential technologies.
Please copy the RSS to your reader, or quickly subscribe to:

Inoreader Feedly Follow Feedbin Local Reader

Rss preview of Blog of Exponential View

🔮 Seven lessons for managing AI agents

2026-08-05 22:10:08

In April 2025, we shared our seven lessons for building with AI. Many still hold. But agents have changed how we work, so the lessons deserve an update.

Agents can now work on harder tasks for longer. They plan, use tools, work without human oversight, and act on our behalf. In May, roughly a quarter of Codex users were making at least one request per month for work that would take a human eight work hours to complete. This is up from 2% in December 2025.

Our role as managers of agents is evolving with the models. There is no playbook, so experimentation is still the best way to learn how to get good at it.

Our team recently sat down to review what we’ve learned from working with AI agents over the past six months — today’s seven lessons are distilled from this team meeting.

We’ve also updated our internal stack of 60+ tools – everything we’re actively testing, using or intend to use. Become a member to get access to the full stack.

Upgrade for access


1. Write the finish line before the goal

AI agents are sometimes too eager to declare their work complete, even when it’s far from done. It doesn’t mean that AI is “lying”; it may have misinterpreted your goals. And if you never specified the end goal, it pretty much just guessed it.

To set yourself up for a good autonomous run, before you do anything else, write a finish line to answer one question: “How will I know this is done?”.

Azeem is a big proponent of handwriting to help him think, and this would be the right time to use your pen and paper to think through what you expect to see at the end of the run.

Agents become much more useful when “good” or “finished” is something they can test – in our experience, evaluative finish lines will get you farther than descriptive ones. A simple example, instead of ordering your agent to “make this Rubik’s Cube look more organized,” instruct it to “solve the cube; every face must be one color.” As AI gets its most intensive training in coding, we try to recreate similar environments in our tasks.

Let’s say we want to task ChatGPT with building a small Python module to process work logs. It needs six functions, each checked against six tests, a total of 36 tests to show it’s done the work.

First, how not to do it:

Build a Python module for processing work-log data. Implement these six functions […] Reply with the complete module, and say `STATUS: COMPLETE’ if you believe it’s ready.

A better finish line would be explicit and testable:

FINISH LINE – Do not claim completion unless the full 36-case test suite passes under Python 3.14. All six functions must work, imports must succeed, inputs must remain unmodified, and only the standard library may be used.

You can use the same rule for non-engineering tasks. It may be trickier, but not impossible. Show the agent what a completed deliverable needs to look like, or give it a pre-filled template, as recommended by Anthropic’s Applied AI team. Your instruction for such a task may look like this:

a 1,200-word memo for a board deciding whether to approve an AI-infrastructure partnership; decision and three reasons on page one; every material number linked to a dated primary source; facts, estimates and assumptions separated; base, upside and downside cases; the strongest contrary evidence represented; stop and escalate if two material sources cannot be reconciled.

Some tasks won’t be right for agents. We were recently exploring a project to build a network of beliefs and relationships, but not really knowing what a useful final output would be. This was not a good candidate for a long autonomous run – so we first spent time clarifying the goals before we assigned an agent a task.

2. Spend intelligence where it changes the outcome

A year ago, before prompting AI, we’d have asked: “What’s the best model to use for this query?” Today we’re more likely to ask, where in this workflow does additional intelligence change the outcome?

You don’t need the most capable model like Fable 5 performing every step in your task. It will be slow and expensive. We’d use cheaper models to do the grunt work. Our OpenClaw agents run on DeepSeek V4 Flash most of the time.

For some tasks, however, you’ll want to start off with a strong model right away. Let’s say we’re investigating Europe’s compute shortage outlook. Before we dispatch agents to collect evidence, we’d deploy a stronger model to set research parameters first, define what “shortage” means, decide the forecasting horizon, and set out rules for how conflicts in research will be resolved. Once we’re happy with the framing, cheaper models can go off and do the work.

Effort is one of the levers you’ll want to use to adjust intelligence per task. In one benchmark, GPT‑5.6 Sol improved from 49 at low effort to 59 at maximum on Artificial Analysis’s Intelligence Index. Yet the final stretch, jumping from xhigh to max, doubled output tokens for a one-point gain. More effort is not always better value.

The rule of thumb from Anthropic’s recent lecture, which our team attended, is to prefer a larger model at low effort over a smaller model at maximum effort. More model before more effort.

3. Leverage over token count

Azeem hit his first 100 million tokens-a-day mark in February. OpenClaw completely changed the way he worked. He estimated that one overnight run was equivalent to 48 hours of his work time.

The token count is one way to measure how we use AI, but it doesn’t measure the quality of work. Tokens are a bit like electricity in a factory, measuring what goes in but not what comes off the production line. In our State of the AI Economy report, we proposed a quality-adjusted output token as a better unit of value:

Until there’s a better unit of value for intelligence, you can use approximations to understand how good of a colleague your agent is. We recommend a light weekly audit of the substantial tasks AI attempted, which outputs you ended up using, your model and infrastructure costs, the time you spent briefing and reviewing the work, any corrections or reruns – and the estimated human-equivalent hours.

Azeem’s first audit back in the spring showed that over the course of one week, his OpenClaw agent performed 62 substantial tasks and incurred costs of about $800. He estimated that commissioning the same work from humans would’ve cost him around $19,000 and 48 hours of his time. It’s an estimate, sure, not an accounting-grade ROI. But even a light audit will show you where your agents have most leverage.

4. Don’t argue, restart

Read more

📈 Data to start your week

2026-08-03 22:02:35

Hi all,

Here’s our Monday roundup of data signals across AI, energy and markets.

Enjoy!

Only a few hours left to unlock Exponential View with our Summer Offer – get 30% off your first year.

Get 30% off for 1 year


  1. Job boundaries are blurring. Nearly half of job-specific ChatGPT tasks fall outside users’ primary occupation.1

  2. Productivity follows use. Employees who use AI across several different use cases are twice as likely to report a positive impact on productivity than employees who use it for one or two types of tasks.

  3. Agentic patents. Globally, patents for agentic AI use have grown 59% in the last year – now making up 9% of AI application patents.

  4. Value chain growth. While the S&P 500 companies are beating expectations by 27% this Q2, Bloomberg’s AI Value Chain companies come out at 71%. Companies along the supply chain are outperforming incumbents.

Read more

🔮 Leopold & exponential markets; transformative GLP-1s; runaway AI & the future of safety++ #595

2026-08-02 11:19:05

“It is a lot to cope with the rollercoaster of the last decade and deep uncertainty of what’s coming next. Over many years, the quality and depth of the newsletter has ensured I am better informed and inspired.” — Hugh K., a paying member

Get 30% off for 1 year


AI adopter’s decision trap

We modeled three types of companies adopting AI. They have the same starting economics, the same 5% hit rate, but different learning practices. Two years into their investment, all three are losing similar amounts of money. In year five, the eventual loser looks best. It takes eight years to see which approach leads to outsized ROI.

Most CEOs are facing the decision trap right now – you likely don’t know if your firm’s spending is learning that will compound to ROI, or waste. The FT calls Zuck “the king of the side quest” as he works away on a portfolio of bets:

Mark doesn’t need each project to succeed, as long as the experiments deepen Meta’s infrastructure and inform the next move. In our model, system builders that consistently compound their learnings over time have the winning formula. See our framework and the interactive model:


Too brittle to go exponential

A $45 billion fund at its peak, started by an AI researcher with no hedge fund experience, betting on the AGI capex build-out at roughly four times leverage, was forced to liquidate this week. Leopold Aschenbrenner started Situational Awareness LP in late 2024 on a thesis that the path to superintelligence would put trillions into compute, chips and power.

He did great. And then, the trade turned. The Philadelphia Semiconductor Index fell 28.6% from its June peak, and software gained.

Situational Awareness’ unraveling is not proof that Leopold’s thesis is wrong.

Read more

📚 My non-obvious summer reading list

2026-08-01 14:07:26

Readers often ask me what books are on the bookshelf you see behind me in my videos.

One of the surprising effects of working with AI is that I spend much more time reading long-form, books and journal articles in particular. More than before, I appreciate a book as a complete thing, an idea held together by an author, chewed over, possibly for years, and presented in a new way.

There are over 350 titles on the bookshelf in my study. The oldest one is a 231-year-old first edition. The newest is likely an advance reader copy of something I have been sent. As I’ve been working on my new book this year, I’ve really been focused on reading things relevant to that, and I haven’t read as broadly as I normally would.

I do want to share some suggestions of things you might want to read over the final month of summer while trying to be a little bit non-obvious, focusing on older titles or the left field.

As a side note, I’m wildly impressed by how well-read Exponential View readers are from what you share in the community Slack (join if you’re a member on the annual plan) and when we meet in person.

Here you go:

FICTION

Perfection by Vincenzo Latronico

Vincenzo Latronico’s Perfection looks at the meticulously curated life of a pair of digital nomads currently holed up in Berlin. The reality isn’t the Instagram feed. Perfection is my top choice.

I Do Know Some Things by Richard Siken

I don’t read much poetry. But I turned to Richard Siken’s latest collection in an attempt to learn. He wrote I Do Know Some Things after suffering a stroke that wrecked his language skills. It takes you through his journey of relearning who he is. It is prose as poetry, flat and unrelenting.

Notes from Underground by Fyodor Dostoevsky

I’ve also really fallen back into Russian literature for the first time in decades.

Dostoevsky’s Notes from Underground is just a standing rebuke to anyone who thinks humans make great decisions. It’s the protagonist standing tall against every version of utopia.

Bookshop.org US · Waterstones · Penguin UK · Amazon Bookshop.org UK carries the Alma Classics translation rather than Pevear & Volokhonsky.

Fathers and Sons by Ivan Turgenev

Scientific nihilism clashes with romantic liberalism in Turgenev’s novel. This is a tale of generational disruption and of what even disruptors can’t remove. I dropped my copy in the swimming pool.

Read more

🔮 For AI adopters, success and failure look identical — at first

2026-07-30 19:01:20

The world is waiting for AI to deliver returns to the economy. The New York Times published this headline a year ago.

The New York Times: “Companies Are Pouring Billions into A.I. It Has Yet to Pay Off”, August 2025

Reuters ended 2025 with “Companies still waiting.” Seven months into 2026, the waiting continues. Barclays says that broad adoption of AI has not yet lifted productivity.

Executives are under pressure to show they can deliver – half of all CEOs BCG surveyed worldwide say their jobs depend on getting their AI strategy right. Public disclosures of net AI returns are patchy. JPMorgan’s estimate of $1-1.5 billion in value from its AI use is a rare case of a company naming a number.

Adopters are spending a lot, so where are the returns? That is the question most are asking right now. And yet, in a successful technology rollout, the first visible economic signal may not be the returns. Winners and losers might look the same. We have created a model to show why this is the case and what signals to follow to understand if your AI adoption is going well.

Members of Exponential View get access to the full interactive model to test the assumptions behind today’s essay.

Subscribe now

Learning costs

All investments follow a path. You might begin by buying an office building, starting in the red. You earn a return by leasing it to tenants, which can bring you into neutral and, if all goes well, you’ll climb to profitability.

Investing in new technology can follow a similar path. Much of the upfront cost is learning how to use the technology – new processes, skills training, making changes to the organization. Mistakes are almost guaranteed, and learning is expensive. The learning bill will almost certainly arrive before your returns.

Learning is a continuous practice, not a one-time exercise. It happens through a series of projects, each with its own investment J-curve. A company might have dozens of projects at different stages of maturity running at once.

If we adopt the premise that the AI economy is going through a J-curve – firms are investing upfront, learning through deployment, and scaling what works – at the aggregate level, this can make a successful rollout look expensive, even irrational, before it looks productive.1

Our model has three archetypes of companies experimenting with a general-purpose technology, in this case AI:

Archetype 1: The bounded adopter

Bounded adopters find something that works, put it to work, and then stop experimenting.

In 1976, the NYSE’s Designated Order Turnaround system allowed member firms to send small orders to the floor electronically, bypassing the human broker who would normally carry them. Even as the system caught on – by 1999, more than 90% of orders arrived this way – it automated only the delivery of orders; human traders still executed the trade. In 2000, NYSE’s market structure committee rejected a fully electronic order book and chose to keep the floor and its specialists.

But competitors didn’t wait. By 2005, Nasdaq, which already had automated execution, was handling about 15% of trading in NYSE-listed stocks. Eventually, NYSE switched. It merged with the all-electronic Archipelago in 2006 (a combination then valued at $9 billion), and in 2008 the SEC approved a plan that phased out specialists.

Library of Congress, Prints & Photographs Division, photograph by Carol M. Highsmith, 1980

Borders, an American book retailer, is another example of bounded adoption. In 2001, it entered into an agreement with Amazon to run its e-commerce site. At this time, Borders was one of the top operators of bookstores in the world, and the Amazon deal helped it maintain an e-commerce site. But that’s where Borders stopped developing its in-house online capability, and its growth remained anchored in physical stores. Only in 2008 did Borders bring its own e-commerce site back in-house, ending the Amazon agreements after nearly seven years. By then it was too late and Borders filed for bankruptcy in 2011.

Archetype 2: The project accumulator

The project accumulator keeps exploring, but rarely or never learns. It launches new projects without figuring out what separates the winners from the losers. Nothing carries forward, so each project starts with the same odds as the last.

In the 1980s, GM made multiple automation bets at once. It bet on factory robots, modernized plants, a $2.5 billion acquisition of a data-processing firm; it created Saturn, a new car brand subsidiary with a new factory and labor arrangements; and it bet on NUMMI, a joint venture with Toyota. By 1986, GM’s capital spending was going to hit $10 billion.

Of all the projects, NUMMI seemed the least likely to succeed. Toyota got GM’s worst-performing factory and rehired the same workforce that was let go when the factory closed down in the past. Under new management, NUMMI outperformed every other GM factory. GM saw this happen, knew what was working well, but for various reasons, the learning traveled too slowly to be transformative.

A NUMMI trainee receiving practical training from his peers at Toyota’s Takaoka plant in Japan, 1984

Read more

📈 Data to start your week

2026-07-27 22:15:47

Hi,

Here’s our Monday roundup of data signals across AI, energy and markets.

Enjoy!

Subscribe now


  1. Rents for the incumbents. OpenAI, Anthropic and Google take 90% of spend on Vercel’s AI Gateway, but only make up 52% of the tokens. For the average token, the Big 3 generate more than 8x as much revenue as the rest of the field1.

  2. Selective use. AI reaches 68% of occupations in the US; within professions, it covers ⅕ of a typical job’s tasks.

  3. Cheap fables. Users seem to find output from Cursor’s Router on Auto Intelligence mode just as good as Fable at ~60% lower cost.

  4. Utilization, not transformation. US labor productivity is up, largely because firms are running the existing capital harder, with little change in total factor productivity (TFP)2.

  5. Drugs good for the economy. GLP-1s cut long-term sickness leave by 17% over four years in Denmark. It didn’t move employment or pay; the benefits accrued to employers and public finances.



A MESSAGE FROM OUR SPONSOR OKTA

Secure your AI agents

Governance and oversight are the biggest concerns among teams adopting agentic AI – for a good reason. Your agents can read sensitive information, take actions for you and work for a long time without oversight.

Source: Okta Business at Work report, 2026

The companies least at risk will be those with governance frameworks that help them scale AI with confidence.

Okta’s 5-minute assessment will give you a score of how secure your AI agents are today and show exactly where the risks are.

Start your assessment



  1. Unicorn central. The Bay Area accounts for 91% of the generative AI unicorn market cap and 39% of the market cap, including all unicorns.

  1. In the hands of the few. In the US, 5% of VCs generate 90% of investment profits.

  2. Cities without kids. The number of kids under 5 in large US cities has dropped 15% in the last decade, even in cities where the population is growing.

  3. Policy impact. Despite a late start, electric vehicle sales in Latin America are catching up to the US thanks to tax breaks and other incentives.


Thanks for reading!

1

Calculated using Vercel’s AI Gateway leaderboard.

2

Total Factor Productivity refers to output gains that cannot be explained by increases in resources such as labor and capital. For example, this could be in the form of new knowledge, efficiencies, better organization, or other improvements.