What happened

OpenAI published a developer blog post telling builders to revisit the instruction files they hand their coding agents now that GPT-6 Astra is out. Astra is OpenAI’s newest model built for long, difficult coding and computer-use tasks, which BuilderWithin covered last week.

The post targets anyone who has built up an AGENTS.md file (a plain-text instructions file that agents like Codex read at the start of a session to learn a project’s conventions), “skills” (named, reusable instruction bundles an agent pulls up for a specific kind of task, like handling a database migration), or detailed standing task prompts over the past year. OpenAI’s point: the handholding you added for earlier, weaker models can now actively work against a more capable one.

What OpenAI says to change

Drop the blanket instructions. OpenAI writes that requiring the agent to read a full architecture document before every edit is “excessive for a typo fix.” The same applies to instructions that push the model to run tests every single time. Astra already tests its own work without being told, so an instruction written for a model that needed that nudge can now make it over-test simple changes and waste time doing it.

Make skills narrow and specific. Instead of a trigger like “use when working with databases,” OpenAI recommends something closer to “use when adding or changing a migration, or reviewing its rollout.” A vague trigger forces the model to guess whether a skill applies. A precise one tells it exactly when to reach for it.

Structure skills as routers, not recipes. OpenAI calls this “progressive disclosure”: the main file for a skill should be a short index pointing to more detail, not a long document the model has to read in full every time it’s invoked. The model pulls in the deeper instructions only when the task actually needs them.

Point to reference files by task, not by default. Rather than telling the agent to read every project file up front, OpenAI suggests mapping specific files to specific situations, for example “architecture.md for service boundaries, database.md for schema changes, deployment.md when preparing a deployment.” The agent opens only what the current task calls for.

Loosen overly cautious stop conditions. If your instructions tell the agent to pause and ask before continuing in situations where you’d actually be fine letting it keep going, OpenAI says that guardrail now costs more than it protects. It recommends defining what “done” looks like for a task explicitly, then letting Astra work through to that point instead of stopping early out of excess caution.

Why it matters

None of this is Codex-specific advice in spirit, even though OpenAI wrote it for Codex and Astra. Anyone maintaining an instructions file for a coding agent, whether that’s AGENTS.md, Claude Code’s equivalent CLAUDE.md, or a similar setup file, is maintaining a document written against whatever model was current when they wrote it. Models get more capable every few months. Instructions rarely get revisited on that same schedule.

An instructions file that over-specifies stops being neutral once the underlying model improves. Extra required reading burns time and tokens (the units a model is billed and rate-limited on) doing work a task doesn’t need. Redundant “please double check” instructions can push a model that already double-checks into doing it twice. Cautious stop conditions written for a model that used to make mistakes can end up interrupting a model that would have finished the job correctly on its own.

Who should care

Anyone with a maintained AGENTS.md, a library of skills, or a set of standing task prompts they reuse across coding-agent sessions, especially if that file was written more than a few months ago or before your current model’s most recent upgrade.

What builders should do next

Pick your longest-standing skill or instruction file and read it as if seeing it for the first time. For each instruction, ask whether it’s telling the model to do something it would already do on its own, or gating a decision more tightly than the task actually needs.

Then run a bounded test before you commit to trimming it. Pick two comparable tasks the file covers. Run one with the instructions as they are today, and the other with your trimmed version. Compare four things: whether the output is correct, how many separate steps or actions it took to get there, roughly how many tokens the session used, and whether the trimmed version stopped early somewhere it shouldn’t have. If the trimmed version finishes in fewer steps with the same correctness, the extra instructions were dead weight. If it makes a mistake the original version would have caught, that instruction was still earning its place. Keep only what survives the test.


End of article