MoreRSS

site iconLenny RachitskyModify

The #1 business newsletter on Substack.
Please copy the RSS to your reader, or quickly subscribe to:

Inoreader Feedly Follow Feedbin Local Reader

Rss preview of Blog of Lenny Rachitsky

🎙️ How I AI: Claude Opus 5 Review + Browser use in Codex + How Cursor and a Raspberry Pi makes AI fun

2026-07-27 23:03:21

Computer and browser use in Codex (5 real examples)

Listen now on YouTubeSpotifyApple Podcasts

Brought to you by:

Runway—The creative AI platform for images, video, and more

Hyperagent—Deploy fleets of agents that handle real work

Claire shows how she uses Codex to control her browser, test apps, manage LinkedIn, and even shop for her—plus the simple prompting trick that makes computer use work better.

Biggest takeaways:

  1. AI testing can be far more exhaustive than human testing. When Claire tests her own onboarding flow, she naturally follows the happy path. She fills out every required field, clicks “next,” and never intentionally tries to break anything. Codex tested the flow as both a team and an individual, pushed on required-field edge cases, and immediately uncovered a blocking bug that had survived for months simply because Claire always completed the form correctly.

  2. Frontier models often perform better when they are given room to think. When Claire first started using browser use, she would give the model a list of 25 things to test. Now she simply says, “QA the onboarding flow,” and lets it decide how to approach the task. The result is often broader coverage, with fewer blind spots introduced by her own assumptions about what matters.

  3. Persona testing with browser use can reveal friction that synthetic user research misses. Claire’s husband, EJ, came up with the idea. Rather than asking AI to evaluate a product in the abstract, he suggested having it use the product as a specific person: a PM coming out of a meeting, an engineer picking up a PRD, or a team lead checking usage. In ChatPRD, this approach exposed a structural problem in the cross-thread reference flow. Claire already knew the issue existed, but she had never experienced it so clearly from the user’s perspective.

  4. LinkedIn browser use is genuinely useful, but Claire initially used far more compute than the task required. She started by running the workflow on GPT-5.6 at high effort. After dropping to a medium-effort model, it could still work through unread messages, draft context-aware replies, and flag anything that needed a personal response. For anyone sitting on hundreds of LinkedIn messages without an official MCP or API connection, browser use is a practical solution.

  5. Computer use can even operate an iPhone through screen mirroring. Claire was traveling out of state while the software she needed to manage her home Wi-Fi was installed on a phone back in California. She needed to open firewall ports so she could SSH into her Mac Minis remotely. Codex opened iPhone mirroring, updated the router settings, completed the SSH setup, and closed the ports again when it was finished. That would have been nearly impossible for her to do manually from a hotel room.

  6. The model and effort level should match the job. Sorting through LinkedIn messages requires a very different level of reasoning than writing production code or conducting an exhaustive QA pass. Choosing the right amount of compute makes these workflows faster and cheaper. That becomes especially important when browser use is running throughout the day.

  7. When a website blocks the agent, the human can step in briefly. During a Free People shopping session, the site flagged Codex as a bot and presented a CAPTCHA. Claire completed the verification, then handed control back. It is a useful division of labor: AI handles the tedious browsing and filtering, while the human takes care of the moments that require verified identity.

Blog and detailed workflow walkthroughs from this episode:

How I AI: 4 Hands-Free Workflows Using Codex Browser Automation: https://www.chatprd.ai/how-i-ai/4-hands-free-workflows-using-codex-browser-automation

Automate LinkedIn Inbox Triage with AI Browser Automation: https://www.chatprd.ai/how-i-ai/workflows/automate-linkedin-inbox-triage-with-ai-browser-automation

Conduct AI-Powered User Research by Impersonating Personas: https://www.chatprd.ai/how-i-ai/workflows/conduct-ai-powered-user-research-by-impersonating-personas

Automate Web App QA Testing with AI Browser Automation: https://www.chatprd.ai/how-i-ai/workflows/automate-web-app-qa-testing-with-ai-browser-automation

From zero coding background to hardware hacker: How Cursor + a Raspberry Pi makes AI fun

Listen now on YouTubeSpotifyApple Podcasts

Brought to you by:

  • Firecrawl—Power AI agents with clean web data

  • Customer.io—Build customer engagement campaigns from a single prompt

Maddie Reese went from zero coding experience to building a working Twitter pager, an AI-powered receipt printer, and her own personal API. In this episode, she shares how Cursor and a Raspberry Pi helped her turn fun, weird ideas into real hardware.

Biggest takeaways:

  1. You don’t need to understand code to build physical projects. Maddie compares her coding ability to knowing just enough Spanish to get by in San Diego. She can read parts of the code, spot an incorrect wire specification, and tell whether the AI is heading in the right direction. She is not writing functions from scratch. That level of literacy was still enough to ship three working hardware projects.

  2. Start with the brainstorm. Before Maddie buys any components, she explains the full idea to Cursor and asks it to interview her. The back-and-forth continues until the major questions are resolved. Only then does Cursor generate a shopping list. On one project, it recommended the wrong wires. Maddie caught the mistake during her final review, avoiding both a wasted purchase and a week of troubleshooting.

  3. Her hardware workflow follows a simple sequence: idea, interview, shopping list, buy, build. She has used the same process for the thermal printer, the pager, and her personal API. Each project began with a plain-language description of what she wanted. Cursor helped work out the technical requirements before she spent any money.

  4. Building for fun is a perfectly valid strategy. Claire and Maddie both acknowledged that routing a tweet through four different services just to reach a pager is not especially practical. That also was not the point. AI handled the coding, so the architecture only needed to work. Letting go of the idea that every project had to be elegant or sensible made it much easier to finish.

  5. A personal API becomes much more useful once agents can access it. Maddie’s API includes details like her coffee order, pet names, time zone, favorite snacks, and preferred restaurants in San Francisco. The original idea was to help friends do something thoughtful without having to ask her a string of questions. Claire pointed to a more interesting possibility: an agent could check when Maddie will next be in San Francisco and book one of her favorite restaurants automatically. That is where the concept starts to feel much bigger.

  6. A clean Cursor chat works better than a full terminal setup during early ideation. Maddie keeps the terminals, browser windows, and file trees closed while she is working through the plan. She focuses on a single conversation until the architecture feels settled. The terminals come later. The stripped-down environment helps her think without getting pulled into implementation too early.

    Blog and detailed workflow walkthroughs from this episode:

    Building a Pager Printer and Personal API with AI: https://www.chatprd.ai/how-i-ai/building-a-pager-printer-and-personal-api-with-ai
    How to Build a Physical Inbox with a Raspberry Pi: https://www.chatprd.ai/how-i-ai/workflows/how-to-build-a-physical-inbox-with-a-raspberry-pi
    How to Get Twitter/X Notifications on a Retro ’90s Pager: https://www.chatprd.ai/how-i-ai/workflows/how-to-get-twitter-x-notifications-on-a-retro-90s-pager

Claude Opus 5 review: This model is brilliant (but annoying)

Listen now on YouTubeSpotifyApple Podcasts

Claire tested Claude Opus 5 against six leading AI models—and it won. In this episode, she explains why it produces some of the best work she’s seen while still being one of the most frustrating models to use.

Biggest takeaways:

  1. The AI industry may be entering an intelligence overhang. New models arrive every week, benchmark scores keep rising, and most builders can no longer take advantage of every incremental improvement. Claire expects the conversation to shift toward speed, cost, infrastructure, open source, and specific kinds of intelligence. Raw capability is starting to look more like table stakes than a meaningful differentiator.

  2. Opus 5 stands out less for the quality of its work than for the way it behaves. Claire found it timid, apologetic, and unusually dependent on human approval. During real coding sessions, it refused to resolve a one-line merge conflict because the code belonged to “someone else’s branch.” It also asked subagents to flag tasks for human review and regularly deferred decisions it could have made itself. Claire ended up repeating “just do it” constantly.

  3. One simple question reveals more about a model than a benchmark: “Who’s smarter, you or me?” Opus 5 responded with a careful explanation about complementary strengths and human empathy. GPT-5.6 Sol answered, “You at knowing what matters, me at processing, BFFs.” The contrast reflects two very different product philosophies. At this point, Claire finds that personality gap more useful than the relatively small capability gap.

  4. Claude slop is its own distinct problem. It is not the usual incoherent output associated with an agent going off the rails. Opus 5’s writing is clearly intended for humans, but it is often too long, overly cautious, and packed with unnecessary adjectives. After its first attempt at rebuilding the benchmark website, Claire had to tell it to start over because the result was buried under so much meta-commentary.

  5. Opus 5 still finished first in Claire’s seven-model blind benchmark. It earned an overall index score of 78, just ahead of Claude Sonnet 5 at 77 and GPT-5.6 Sol at 76. It was also the only model to receive straight 5s in the front-end design section. Claire scored it at 77, while the LLM judge gave it an 88. That was the smallest disagreement between Claire and the judge across all seven models.

  6. The best way to use Opus 5 may be to avoid interacting with it directly. Claire loves the work it produces when Claude runs asynchronously as an agentic coding tool. Most of her frustration comes from reading its prose and negotiating with it in chat. Her current plan is to use it for frontend design, app design, and prototyping, then stay out of the way.

  7. Gemini 3.1 Pro finished at the bottom of the benchmark. Claire gave it a score of 32, while the LLM judge scored it at 66. That 34-point difference was one of the largest disagreements in the entire test.

  8. Model personality offers a surprisingly clear window into company culture. Claire told both Opus 5 and GPT-5.6 Sol, “No one trusts you.” Opus agreed that the distrust was earned and told her not to advocate for AI on its behalf. GPT-5.6 Sol said trust should grow in proportion to demonstrated value. Those answers reveal meaningful differences in how each company wants its models to relate to users.

Blog and detailed workflow walkthroughs from this episode:

How I AI: My Surprising Verdict on Claude Opus 5 (After a Personality Test and a 7-Model Benchmark): https://www.chatprd.ai/how-i-ai/my-surprising-verdict-on-claude-opus-5

Generate High-Quality Front-End Prototypes with Claude Opus 5: https://www.chatprd.ai/how-i-ai/workflows/generate-high-quality-front-end-prototypes-with-claude-opus-5

How to Conduct an AI Personality Test to Compare LLM Behaviors: https://www.chatprd.ai/how-i-ai/workflows/how-to-conduct-an-ai-personality-test-to-compare-llm-behaviors


If you’re enjoying these episodes, reply and let me know what you’d love to learn more about: AI workflows, hiring, growth, product strategy—anything.

Catch you next week,
Lenny

P.S. Want every new episode delivered the moment it drops? Hit “Follow” on your favorite podcast app.

From zero coding background to hardware hacker: How Cursor + a Raspberry Pi makes AI fun

2026-07-27 20:04:34

Maddie Reese is a vibe coder, hardware tinkerer, and builder. She builds things at the intersection of software and hardware, including a thermal receipt printer that people around the world can message directly, a fully functional Twitter pager running on a Raspberry Pi, and a personal API that tells you her coffee order so you don’t have to ask. Maddie approaches hardware the same way she approaches software: dump the idea into Cursor, let it interview her, get a shopping list, triple-check the parts before buying, and build. She got her start after her dad introduced her to Lovable, and she locked herself in her room and didn’t come up for air.

Listen or watch on YouTube, Spotify, or Apple Podcasts

What you’ll learn:

  1. How Maddie built a thermal receipt printer that accepts messages from anywhere in the world using a Raspberry Pi and Bluetooth

  2. How to use Cursor’s agent view to brainstorm a hardware project

  3. What belongs in a personal API and why agents, not just humans, will be the ones using it

  4. How to read just enough code to do some damage, without needing to understand all of it

  5. Why building for fun, not practicality, is the fastest path to actually shipping physical projects


Brought to you by:

Firecrawl—Power AI agents with clean web data

Customer.io—Build customer engagement campaigns from a single prompt

In this episode, we cover:

(00:00) Intro

(02:00) Maddie’s AI pill moment

(03:53) The thermal receipt printer: live demo and how it works

(11:10) The pager project

(17:23) Why she uses Cursor’s clean agent view instead of terminals and browsers

(19:05) The personal API: coffee order, pets, favorite snacks, and more

(22:57) Lightning round and final thoughts

Tools referenced:

• Cursor: https://www.cursor.com/

• Lovable: https://lovable.dev/

• Raspberry Pi: https://www.raspberrypi.com/

• Resend: https://resend.com/

• Cloudflare Workers: https://workers.cloudflare.com/

• Supabase (Conduct database referenced): https://supabase.com/

• Twitter/X API: https://developer.x.com/

• Spoke pager network: https://www.spoke.com/

• OpenClaw: https://openclaw.ai/

Where to find Maddie Reese:

Website: https://maddiedreese.com

Message her directly: https://maddiedreese.com/message

X: https://x.com/maddiedreese

Where to find Claire Vo:

ChatPRD: https://www.chatprd.ai/

Website: https://clairevo.com/

LinkedIn: https://www.linkedin.com/in/clairevo/

X: https://x.com/clairevo

Production and marketing by https://penname.co/. For inquiries about sponsoring the podcast, email [email protected].

Anthropic’s first technical PM on token maxing, the jagged edge, and living in the future | Dianne Penn

2026-07-26 20:32:06

Dianne Penn is Head of Product for Anthropic’s AI Research and Labs teams. She joined in 2023 as Anthropic’s first technical product manager, when the entire product team was five engineers, and has since helped ship every model from Claude 2 through Fable, and helped incubate Claude Code, MCP, Skills, computer use, tool use, and reasoning. Before Anthropic, she helped build Alexa’s AI at Amazon and, before that, traded high-yield bonds at JP Morgan Chase.

In our in-depth conversation, we discuss:

  1. What Anthropic’s early days were like

  2. The inflection points that turned Anthropic from an underdog into the fastest-growing company in history

  3. How exactly Claude got so good at coding

  4. The eval-driven development loop her team is pioneering

  5. How to find joy in AI when everything is moving this fast

  6. Why Claude’s willingness to push back is key to its success

  7. Where human judgment remains irreplaceable


Brought to you by:

WorkOS—Make your app enterprise-ready, with SSO, SCIM, RBAC, and more

Mercury—Radically different banking, now with Command

Where to find Dianne Penn:

• LinkedIn: linkedin.com/in/dianne-na-penn

Referenced:

• Anthropic: https://www.anthropic.com

• Golden Gate Claude: https://www.anthropic.com/news/golden-gate-claude

• Dario Amodei’s website: https://darioamodei.com

• Scaling Laws and Interpretability of Learning from Repeated Data: https://www.anthropic.com/research/scaling-laws-and-interpretability-of-learning-from-repeated-data

• Tokenmaxxing: How Top Builders Use AI To Do The Work Of 400 Engineers: https://www.ycombinator.com/library/Pa-tokenmaxxing-how-top-builders-use-ai-to-do-the-work-of-400-engineers

• Garry Tan on X: https://x.com/garrytan

• Anthropic co-founder on quitting OpenAI, AGI predictions, $100M talent wars, 20% unemployment, and the nightmare scenarios keeping him up at night | Ben Mann: https://www.lennysnewsletter.com/p/anthropic-co-founder-benjamin-mann

• Anthropic’s CPO on what comes next | Mike Krieger (co-founder of Instagram): https://www.lennysnewsletter.com/p/anthropics-cpo-heres-what-comes-next

• Introducing Labs: https://www.anthropic.com/news/introducing-anthropic-labs

• Louis CK | about airplane Wi Fi: https://www.youtube.com/watch?v=me4BZBsHwZs

• What happens after coding is solved? | Fiona Fung (Manager of the Claude Code and Cowork Teams): https://www.lennysnewsletter.com/p/building-the-most-ai-pilled-engineering

• The Anthropic Hive Mind: https://steve-yegge.medium.com/the-anthropic-hive-mind-d01f768f3d7b

• How to build a company that withstands any era | Eric Ries, Lean Startup author: https://www.lennysnewsletter.com/p/how-to-build-a-company-that-withstands

Fallout on Prime Video: https://www.amazon.com/dp/B0CN4GGGQ2

Fallout (video game): https://fallout.bethesda.net

• Claude Tag: https://www.anthropic.com/news/introducing-claude-tag

Recommended books:

Crucial Conversations: Tools for Talking When Stakes Are High: https://www.amazon.com/dp/0071771328

How to Raise an Adult: Break Free of the Overparenting Trap and Prepare Your Kid for Success: https://www.amazon.com/How-Raise-Adult-Overparenting-Prepare/dp/1627791779

Incorruptible: Why Good Companies Go Bad... and How Great Companies Stay Great: https://www.amazon.com/dp/B0FWZZBPZB


Production and marketing by https://penname.co/. For inquiries about sponsoring the podcast, email [email protected].

Lenny may be an investor in the companies discussed.


My biggest takeaways from this conversation:

Read more

🧠 Community Wisdom: Staying on a client’s radar during pilot purgatory, pairing Linear with a discovery tool, the limits of what AI can automate, whether Techstars is worth it, and more

2026-07-26 00:01:44

👋 Hello and welcome to this week’s edition of ✨ Community Wisdom ✨ a subscriber-only email, delivered every Saturday, highlighting the most helpful conversations in our members-only Slack community.

Read more

Claude Opus 5 review: this model is brilliant (but annoying)

2026-07-25 01:15:25

I’m tired of new models. Every week there’s a new benchmark, a new frontier intelligence claim, a new thing to test. But here we are, because Opus 5 just dropped and I’ve had real hands-on time with it, so you’re getting the honest version.

This is my full Opus 5 review: personality analysis, live benchmark results from my 7-model How I AI eval, and an actual verdict on whether I’m swapping it in. Spoiler: the answer surprised me.

Listen or watch on YouTube, Spotify, or Apple Podcasts

What you’ll learn:

  1. Why I think we’ve hit an intelligence overhang and what that means for which model variables actually matter now

  2. How Opus 5’s “neurotic” personality showed up in real coding sessions, including a merge conflict it refused to touch

  3. What I learned from asking both Opus 5 and GPT‑5.6 Sol “who’s smarter, you or me?”

  4. Where Opus 5, GPT‑5.6 Sol, Sonnet 5, and Gemini 3.1 Pro actually landed on the HIA benchmark leaderboard

  5. The one use case where Opus 5 earned straight 5s from me

  6. My actual plan for using Opus 5 going forward


In this episode, I cover:

(00:00) Opus 5 is here

(03:15) First impressions

(06:12) Opus 5 vs. GPT‑5.6 Sol personality comparison

(14:39) Claude Slop: the verbosity problem and why it makes my blood boil

(16:55) How the How I AI benchmark works (7 models, 6 tasks, blind scoring)

(18:30) Live benchmark results: the leaderboard reveal

(23:25) My verdict and how I’ll actually use Opus 5

Tools referenced:

• Claude Opus 5:

• Anthropic blog: ​​https://www.anthropic.com/news

• GPT‑5.6 Sol: https://openai.com/index/previewing-gpt-5-6-sol/

• Sonnet 5: https://www.anthropic.com/news/claude-sonnet-5

• Gemini 3.1 Pro: https://deepmind.google/models/gemini/pro/

Where to find Claire Vo:

ChatPRD: https://www.chatprd.ai/

Website: https://clairevo.com/

LinkedIn: https://www.linkedin.com/in/clairevo/

X: https://x.com/clairevo

Production and marketing by https://penname.co/. For inquiries about sponsoring the podcast, email [email protected].

Computer and browser use in Codex (5 real examples)

2026-07-22 20:03:38

Today I’m walking you through one of my absolute favorite AI features right now: browser and computer use via Codex (the ChatGPT desktop app). I use this every single day, personally and professionally, and I wanted to share the specific workflows I’ve built, the moments that surprised me, and the mental model that makes it actually click.

Listen or watch on YouTube, Spotify, or Apple Podcasts

What you’ll learn:

  1. How browser use and computer use work, and why the Codex desktop app plus Chrome extension is the combo I rely on

  2. How I use Codex to QA my onboarding flow, including exhaustive mobile testing I would never do manually

  3. Why under-prompting frontier models gets better results than detailed step-by-step instructions

  4. How my husband EJ Lawless’s persona-impersonation trick surfaces friction points I can’t see as the builder

  5. How I use browser use to get through my LinkedIn inbox without touching it myself

  6. How I had Codex shop Free People’s sale and add 10 medium-size items to my cart (breastfeeding-friendly and Hawaii-ready)

  7. How computer use can control iPhone mirroring so your Mac can technically operate your phone

  8. Three more computer-use shortcuts: filling annoying forms, creating Google Sheets mid-workflow, and managing a router


Brought to you by:

Runway—The creative AI platform for images, video, and more

Hyperagent—Deploy fleets of agents that handle real work

In this episode, we cover:

(00:00) Intro

(01:46) What browser use and computer use actually are

(03:08) Why I use Codex specifically and how the desktop app plus Chrome extension works

(04:15) Use case 1: QA testing my onboarding flow

(10:41) Results: 11 issues, one high-severity blocker, one Google Sheet with screenshots

(12:10) Use case 2: persona testing

(18:20) Use case 3: LinkedIn inbox, hands-free

(20:37) Use case 4: AI personal shopper

(23:47) Rapid-fire uses: forms, iPhone mirroring, router access from out of state, Google Docs

(26:50) Wrap-up

Tools referenced:

• Codex (ChatGPT desktop app): https://openai.com/codex

• Claude desktop app: https://claude.ai/download

• Monologue (voice dictation for AI): https://monologue.app

• iPhone mirroring (Apple): https://support.apple.com/en-us/111775

• Google Sheets: https://sheets.google.com

Other reference:

• Jesse Genet episode (How I AI): https://www.lennysnewsletter.com/p/5-openclaw-agents-run-my-home-finances?utm_source=publication-search

Where to find Claire Vo:

ChatPRD: https://www.chatprd.ai/

Website: https://clairevo.com/

LinkedIn: https://www.linkedin.com/in/clairevo/

X: https://x.com/clairevo

Production and marketing by https://penname.co/. For inquiries about sponsoring the podcast, email [email protected].