A shift already underway in the tools you use every day.
Vivienne Su shared a piece by Addy Osmani that names a shift you are already half living with Claude and Codex: away from typing one prompt at a time and toward designing a system that runs on its own. The idea is called loop engineering. Instead of writing an instruction, reading the answer, then writing the next instruction, you build a small loop that finds the work, hands it to the agent, checks the result, writes down what got done, and picks the next step by itself. You design the loop once. After that the loop is the one prompting the agent, not you.
Two people building these tools said it plainly. Peter Steinberger: "You shouldn't be prompting coding agents anymore. You should be designing loops that prompt your agents." Boris Cherny, who leads Claude Code at Anthropic, put it the same way: "I don't prompt Claude anymore. I have loops running that prompt Claude and figuring out what to do. My job is to write loops." Osmani calls the layer just below this idea agent harness engineering, building the single environment one agent runs inside. Loop engineering sits one floor above that: it runs on a timer, spawns small helper agents, and feeds itself the next task instead of waiting for you to feed it one. He is upfront that this is early. He is still skeptical, token costs (what each run of the model costs you to use) can swing wildly, and letting quality slip while a loop runs unreviewed is a real risk, not a hypothetical one.
He breaks a working loop into pieces that fit together. Automations run on a schedule and do a first pass of discovery and filtering without you watching. Worktrees (separate copies of the same code, so two agents working at once don't overwrite each other) let you run agents in parallel. Skills are written-down project knowledge the agent would otherwise have to guess at every time. Sub-agents (a second, independent AI process that checks or extends the first one's work) split the work so one agent does the task and a different one checks it, because a model grading its own work tends to be too easy on itself. The piece Osmani says is easy to underrate is memory: a markdown file, a project board, anything that lives outside a single conversation and survives after the agent forgets everything at the end of a session.
The concrete detail is that both tools you already use ship these building blocks. In the Codex app you build an automation in the Automations tab: pick the project, the prompt, how often it runs, and whether it runs on your live code or a separate background copy. Runs that find something land in a Triage inbox; runs that find nothing archive themselves, so you are never wading through empty results. Osmani says OpenAI uses this internally for daily bug triage, summarizing failed test runs, and writing commit summaries. Claude Code gets to the same place through /loop, which reruns a prompt on a set interval, scheduled cron jobs (tasks that fire automatically at set times), hooks (small scripts that trigger at specific points while the agent works), and GitHub Actions if you want the loop to keep running after you close the laptop. One piece works differently: /goal does not just repeat on a timer, it keeps going until a condition you wrote is actually true, and a separate, smaller model checks after every turn whether that condition is met, so the agent doing the work is never the one grading whether it is done.
None of this needs a new tool. A year ago building a loop meant writing your own pile of bash scripts and maintaining it forever. Now the pieces ship inside Codex and Claude Code themselves, and once you notice the pieces are the same shape in both products, you stop arguing about which tool to use and just design a loop that works in whichever one you happen to be sitting in. The next task you catch yourself running by hand twice is the one to turn into a loop.