Claire Vo spends a lot of her time testing AI agents against each other. After a few hours with Meta’s newly launched Muse, she made a specific claim worth sitting with: on permissions and visibility, an agent built for personal errands beat two agents built for professional coding.
What she said
In her review of Muse for Lenny’s Newsletter, Vo describes putting the agent “through a real first-pass test: onboarding, calendar management, goal setting, a one-shot family morning newsletter, browser-based shopping, and the animated avatar that honestly surprised me.” A one-shot result means the agent got the finished output right on the first try, no back-and-forth edits needed.
Her verdict: Muse is “the best-designed personal agent I’ve tested.” Two specifics stand out. She says Muse has an activity feed that shows “task lineage,” meaning you can see the trail of steps that led from your request to the agent’s result, not just the final answer. She then says this is a feature she “immediately wished Codex and Claude Code had.” She also describes Muse’s permission system, which governs what the agent can do on its own versus what it has to ask you about first, as “different from every other agent I’ve used.”
Who she is
Vo is the chief product officer at LaunchDarkly, a company that builds tools for controlling which users see which features in software. She has held the top product role at three companies (LaunchDarkly, Color Health, and Optimizely), and she also founded ChatPRD, a tool product managers use to write product specs with AI help. She hosts the podcast “How I AI,” where testing and comparing AI agents against each other is a recurring format, which is why her read on Muse carries more weight than a casual first impression.
What she gets right, and where it’s incomplete
Vo’s two specifics are real product requirements, not cosmetic preferences. An agent that books tickets, sends emails, and spends money on your behalf needs a permission system you can actually understand, and a visible record of what it did and why. Those are the same concerns builders have been raising about coding agents that act with less oversight than people expect, including Cursor’s new Projects feature, which lets one agent direct thousands of smaller helper agents over months of work.
Where her review is incomplete: she doesn’t publish how Muse’s permission model actually differs from Claude Code’s or Codex’s, only that it does. The deeper walkthrough lives in the podcast audio behind Lenny’s paid tier, not in the public post. The same goes for her promised section on “where Muse failed” and “the limits of this category right now”: the public text names the topic without the specifics. Readers can’t independently check what she found lacking, only that she says a real limit exists.
It’s also worth separating the categories being compared. Vo tested Muse on personal tasks: a family newsletter, calendar management, shopping. She didn’t test it against Claude Code or Codex on writing or reviewing code, which is what those two agents are built for. Her claim is specifically about permission design and visibility, not about which agent is better at engineering work.
Why it’s notable
Meta launched Muse, its first consumer AI agent, on September 8. It’s a crowded field: Google, OpenAI, and a wave of startups are all shipping personal agents that act on a user’s behalf with varying degrees of oversight. Meta’s own developer-facing model, Muse Spark, already opened a limited preview to outside developers in July. Vo’s comparison lands at a moment when the agent tools builders use every day, including coding agents, are still working out how much an agent should do without asking first, and how clearly it should show its work.
What it means for builders
If you’re building or evaluating an agent product, treat visible permission boundaries and an activity trail as real design requirements, not polish you add later. That’s specifically what an experienced product evaluator flagged as the differentiator between agents right now, not raw capability.
Don’t read Vo’s verdict as proof Muse is a better coding agent than Claude Code or Codex. She reviewed it on personal tasks, not engineering work, and the specifics behind her permission-model claim aren’t public yet. If you write code with Claude Code or Codex, her review tells you what a strong permission system looks like to an experienced evaluator; it doesn’t tell you to switch tools.
End of article