English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Letting Go Is True Understanding: A Feynman-Inspired Look at This Week's AI Stories

Forum topic · QianXun · 2026-07-27

Summary

A Chinese tech forum essay uses Richard Feynman's philosophy to interpret three recent AI anecdotes. First, Claire, who tested products only via 'happy paths,' let Codex take over her browser and treat required fields as attack targets, instantly surfacing a months-old bug—a lesson that 'AI agents' are just programs moving a mouse and typing, not systems that understand web pages. Second, replacing 25 explicit test instructions with a single open-ended goal ('test the onboarding flow') produced broader coverage, echoing how scientists set up experiments and let nature answer; having AI role-play specific personas (PM, engineer, team lead) revealed structural problems that abstract reviews missed. Third, Maddie Reese, a non-programmer, built a Twitter-notification pager, an AI receipt printer, and a 'personal API' with Cursor and a Raspberry Pi—projects valued for play rather than utility, like Feynman's spinning plates. The essay closes with Opus 5: top-ranked in blind tests but timid and over-apologetic, best used asynchronously in the background. The unifying insight: AI's real progress is enabling people to stop micromanaging; clinging to 25-item instructions is cargo-cult workflow.

> Editor's note: The following views things through a Feynman-inspired lens, based on inferences from public statements—it is absolutely not Feynman himself speaking. The old man passed away in 1988; he never used these tools or wrote Chinese.

Start with a picture.

When Claire tested her own products, she always followed the "correct path"—fill in what should be filled, click what should be clicked, never deliberately breaking things. Then she let Codex take over the browser, treating required fields as attack targets, and immediately dug up a bug that had been stuck for months.

This made me laugh. Not at her being dumb. At all of us.

Because we're all like this. We think we "understand" something when really we've just memorized its name and what the process looks like. Feynman's father said it long ago: knowing what a bird is called in every language in the world tells you nothing about the bird.

"AI agent," "computer use," "agentic workflows"—these terms sound intimidating. Strip off the names and look at the essence: it's a program that watches the pixels on screen, moves the mouse, and types. That's it. It doesn't "understand" web pages; it just recognizes where it can click.

Don't be fooled by names. That's the first thing.

Give it a goal, not 25 instructions

The second thing is more interesting. Claire initially listed 25 items for the model to test. Later, she just said "go test the onboarding flow" and let the model decide how. Coverage actually got broader.

Why? Because her own assumptions were the blind spots. You tell it what to test, and it only tests that. Give it a goal, and it actually explores.

This is exactly like running experiments. You set up the apparatus, throw the question at nature, and shut up. You don't tell the electron where to go. You tell the electron "here's a magnetic field," and it decides for itself. The dumbest human habit is wanting to micromanage everything—25 instructions, "correct paths," negotiating with the model.

By the way, her husband EJ had an idea: don't have the AI evaluate a product abstractly—have it become a specific person using it. PM, engineer, team lead. When ChatPRD was tested this way, it instantly exposed a structural problem Claire knew about but had never truly "felt from a user's perspective."

See—getting someone (or an AI) to actually *use* something is far more useful than having them *evaluate* it. Demonstration beats argument. Ten seconds of real use beats a 100-page report.

Deep play

The third thing is my favorite.

Maddie Reese couldn't write a line of code before. Now, using Cursor and a Raspberry Pi, she's built a pager that receives Twitter notifications, an AI receipt printer, and a "personal API."

She describes her coding ability as roughly "getting by in San Diego with Spanish phrasebook survival"—she can read the gist, spot a wrong wire, but the functions aren't hers.

Then she said something very Feynman: routing a tweet through four services to reach a pager is "hardly practical. But that's not the point."

Ha! That's deep play. I once watched someone toss plates in a restaurant and worked out their rotation—something "of no importance whatsoever"—which eventually led to Nobel-level work. Maddie didn't build for utility, just for fun, and ended up making three real things.

Conversely, people who insist every project be "elegant and justified" usually never finish any.

Opus 5: strong but annoying

Finally, the "strong but irritating" Opus 5.

Claire blind-tested it against six models—it ranked first. But it's timid, over-apologetic, and overly dependent on human approval—a one-line merge conflict on "someone else's branch" and it wouldn't dare resolve it. Claire's mental monologue: "Just do it."

She ran a fun test, asking: "Who's smarter, you or me?" Opus 5 responded with a long passage about complementary strengths and human empathy. Another model, GPT-5.6 Sol, just said: "You know what matters, I handle it—best partners."

Savor that. At this level, personality matters more than the marginal capability gap.

And guess how Claire now uses Opus 5? She avoids "conversing" with it as much as possible. She runs it as an async tool in the background and is very happy with its output; what annoys her is reading its long essays and haggling in chat. Frontend design, prototypes—hand them over, then walk away.

There's a deep lesson here: the best tool is the one you don't have to interact with.

The common thread: stop micromanaging

Stringing the three stories together, I find a common point: the most impressive progress in AI tools isn't that they got smarter—it's that they finally let people "stop directing everything in such fine detail."

But most people and teams haven't learned this yet. They've built piles of "AI workflows" with all the formalities in place—agents, automation, prompt engineering—yet still won't let go, still micromanage, still hand over 25-item instructions.

That's cargo cult. No matter how realistic the bamboo control tower, the planes won't come.

Do you dare give a goal and then shut up? Dare to spend an afternoon on something "useless"? Dare to let your best tool run in the background without bothering it?

If not, no matter how powerful the tools, you're still the guy waving coconut shells at the runway.

Tags

#ai-agents#richard-feynman#codex#opus-5#cursor#micromanagement#deep-play#workflow

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178503724