So what would you say...you do here?
📝 postLast time, I shared how instead of coding, I spend my time building and tuning "the factory" in which agents, me, and my teammates collaborate to get a ton of stuff done quickly, with high quality, and without nagging people about unimportant approvals or details. The post was more about how I think of my job now compared to before. I talked about how I spend my time now on tuning a "software factory" that builds a bunch of things in parallel, working through the tedious parts on its own while involving me only where I want or need to be involved. This post gets into some of the specifics.
The day-to-day and week-to-week of shipping software has changed a lot. To kick things off, here's a sort of "day in the life" of how I used to get things done and how I get things done now.
In the days of yore
Let's start with project work. This meant things like designing the user experience and the systems to drive it, implementing it, testing it, fixing it, and getting it operationally ready. This is work the team did while marching toward a feature launch. To make sure we built the right things, we wrote design documents and held meetings to talk about them. We wrote code and iterated on it until we were ready for it to be reviewed by teammates. If teammates had comments, we reworked the code to address the feedback. We each brought our own experience and preferences about how the code should be structured, and deliberated on it to make sure it's factored well for future plans. We planned sprints and had standups to coordinate on things to make sure we're not stepping on each other's toes, and to help each other or adjust the plan when we're blocked.
While many already feel nostalgia toward those days, it was not always rosy. On designs, we spent sometimes many rounds of back-and-forth, deliberating to make sure they were future-proofed and avoided pitfalls. And if we found partway through implementation that we were wrong or learned something new, we'd have to backtrack and rework, or would wring our hands to try to figure out how to plow forward rather than throw away code. On code reviews, we quibbled over code style and design patterns, blocking each other from checking in until we either came to an agreement (or a settled disagreement) and moved forward. And we'd wait hours or days between rounds of feedback because we were each in the zone on our hands-on work, or in a meeting or just working slightly different hours and couldn't drop it all on a dime to give the next round of feedback. Sprint planning and standups were mind-numbingly boring and consumed a ton of time of estimating work and thinking further down the road into the plan than at least I had patience for. But on balance, this stuff was necessary. It's how we put our heads together, learned from each other, built high-quality systems, and moved things along.
The world of tomorrow (today!)
Now, project work looks different, from design, to coding, to review, to sprint planning and standups.
Today, designs have a few different starting points. They may start with shorter conversations with teammates about the "what" and "why". As we collaborate, we want to have a cohesive vision for how the thing is being built and what we're building. I might start with writing a shorter paper talking about the problem statement and a shape of how I want to solve it. I do a lot of brainstorming with an agent, having it do research on tools or designs, having it pick apart my ideas. But throughout this process, I'm writing, reading, researching, sitting back and thinking, communicating, weighing tradeoffs, bringing in my expertise, experience, and taste, and also learning plenty. But in the end, with all of this back and forth, I produce some kind of design or spec that I'm happy with. Depending on the task, this might be a structured approach like spec-driven development, or any kind of /plan mode or prompting.
I don't code anymore, and code review looks very different too. I have an understanding of the code and the project, and an opinion on how it should be shaped going forward. I use this to steer the design and implementation plan, but that lengthy time I would have spent coding is freed up. My review loop rarely is about the end code. I do most of my review and back-and-forth in that earlier design phase, where I iterate with the agent on sketches of the implementation, like what interfaces it will add or change, architectural changes it needs, and what instrumentation it will add to make it debuggable. With this amount of detail, I can decide whether the agent's approach meets the long-term structure that we need to move fast.
Team planning is different too. Now things move easily 4x faster, so a detailed sprint plan like before is hardly worth doing. Standups on my team are very different than before. Instead of a routine status dump, we cover whatever design ideas, future plans for the next few days, or explore the customer experience of mocks or more often of working implementation. It's a hodgepodge of discussion that we don't plan until we get there. If there aren't things to review, discuss, or plan, they can be 10 minutes. Our team's record for longest standup was 10 minutes short of 3 hours. When someone described what they were going to do next, we realized we had some design and architecture that we had been deferring but finally came due, so we did it as a group. We still need to coordinate so we don't step on each other's toes with big changes to the same parts of the system, duplicate work, or even contradictory work. Coordination and communication still feel like the constraints for how much we can get done, so we break down tasks in a more coarse-grained, independent way than before. And we still agree on our team milestones so we don't go off in a bunch of directions and end up spending our coordination tax on the things that are needed less urgently.
Loop engineering
In an earlier blog post, Agentic bumper bowling, I talked about agentic loops. In the software factory, there are also loops outside the agentic loops of the design agent, or coding agent or testing agent, or review agent, or whatever agent. If the loops aren't built into the factory, then you're the one nudging everything along between steps or sending things back for rework. This is super tedious.

Some 20 years ago, a coworker described having to run tedious commands himself as "typing it with his face" - not the most ergonomic, speedy, or fun way to do things. Or later, some coworkers were teasing me when I tried to show someone some log analysis tricks on their Mac laptop. But since I was used to Windows (mostly for the productivity tools - my coding was all on a remote Linux machine), I didn't know how to do simple stuff like jump words. I forget who drew the comparison, but the imagery we ended up with was me being a t-rex who was attempting to type, but since t-rex arms are short I'd be using two of those grabby extension things, but of course each with a t-rex head doing the chomping. And so I immediately commissioned an artist on Fiverr to render that for us.
So what do I do to make this ergonomic in the factory? I make loops. Here are a few of the concrete loops that I make sure are always working smoothly. But guess what? In a lot of ways, it's the same stuff I was doing before.
Agent onboarding
Whenever I've moved to another team or product, I start by going through the onboarding docs. These often point to design docs, outline where the code is, link to our operational dashboards, deployment pipelines, customer feedback channels, and even some helpful historical documents that shaped the architecture like a few select retrospective "root cause analysis" docs. After all, I learn a lot from when things go wrong and from the decisions about what we changed as a result. Invariably though, these onboarding docs are a little stale. Maybe the commands to get things running locally aren't quite right anymore, or they were meant for how to install things on an older operating system. Maybe links are broken, or the architecture docs are out of date. I fill in the gaps by talking to teammates, and I make sure to spend some time to improve and fix those onboarding docs for the next person who joins the team. Or if I know that a new person is joining the team, I'll give my team's onboarding docs a quick pass to see if there's anything glaringly out of date.
When I start up an agent, it knows zero about the system I'm working on, how we build, how we test, what standards we have, what our roadmap is, or how to tell how things are performing in production. Every time an agent starts up, it's like I'm onboarding a new team member. When those "onboarding docs" are wrong, instead of me being able to fill in the gaps occasionally by chatting about things with teammates at lunch, there are at least tens of agent sessions that I have to start up every day and have that same conversation with to nudge it to do the right thing. That's exhausting.
So the ROI of that work has changed. Instead of "onboarding instructions" being a thing I could triage away in favor of more important and urgent work, this "team and project context" becomes one of the most important things to continuously improve. Failing to do so creates an enormous amount of frustrating work where I'm repeating myself all the time.
Quality gates
Another thing I used to do was build automatic quality gates. We'd have "push hooks" to make sure you have a clean build that passes all of the tests, checkstyle, linters, findbugs, code coverage, and whatever other scanners we could piece together to look for specific mistakes that we could tell from static analysis. And you know what? I fucking hated this. I hated getting a commit ready with all the tests passing, coverage awesome, and code reviewed by a teammate with a back-and-forth or two, to have some linter block my push because I had 161 characters on a line instead of 160. I'd have to fix it and wait another 4-24 hours for a teammate to hit the thumbs up on the new change. Such bullshit. Of course we'd try to position those checkers so they'd happen as early as possible, and ideally get used by the IDE itself to autoformat, but tuning these rules always seemed like a distraction because they were just "ergonomics", and I could tolerate a lot of, well, bullshit. It's just part of the job, and we'd have the judgement on how to balance all of the priorities.
Now, those quality gates - those deterministic annoying checkers, coverage standards, you name it - those are awesome. It's their time to shine. Agents don't care at all if they need to rework something. They thrive on feedback. They need feedback. So we task them with writing exhaustive tests using testing practices that used to be a little tedious or overkill sometimes, like property-based testing for exhaustively exploring the state space. Or mutation testing, where you intentionally break the code to make sure the tests fail. Or pen testing. There are even some pretty amazing new techniques to catch bugs in the requirements doc before any code is written using automated reasoning. These testing techniques are things we didn't always have time for before. But now these are the exact tools that the agent needs to do its job right, and I find they lead to higher quality code than before when I was doing things by hand and deciding which technique I had time to do by hand given the task at hand.
Code review
One of the hottest topics out there is on human review of code changes. It's time-consuming, and I've never met someone who loved doing it as a part of their job. But it's essential, right? How else would you know whether the agent is doing a good job or if it's going to create a bunch of slop?
Well let's talk about what your job is in a code review. Is it to find bugs? Eh, not really. Sometimes I find a bug when reviewing code, but I'm not a compiler and I'm not a unit test framework. I'm not checking out the code and finding gaps in testing and writing them myself. What I am looking for though is whether this code is taking the right approach for where we're taking the product, doesn't miss some important design consideration, or blurs the encapsulation of some part of the code in a way that will be hard to maintain later. If we're not reviewing the code, how do we avoid vibecoded slop?
Well at first, you do review the code. But every comment in a code review is a defect. Not a defect in the code that was written, but a defect in the codified rules enforced by auto-reviewers. If we want code to be styled a certain way, add a linter. Add a team of review agents that look for problems or a failure to enforce best practices, like security, resiliency, usability, you name it. If we want to make it so you can't break encapsulation, add agent steering to your review agent to check for that. And then of course add a gate that prevents the agent from claiming that it's done until it has addressed all of the auto feedback, or if it raises its hand for a human because it thinks the feedback is unreasonable or points to a larger problem in the requirements.
There's still plenty of human specification and review that needs to happen. Designs need to be deliberated between you and the agent and even as a team. Sometimes for a task you do want to thoroughly review the code yourself. Great - have the agent wait for you for that task. If there's a part of the code that you always want detailed human review of, add some kind of agent or workflow rule that requires human approval when that part of the code is changed, like the authorization code or the replication algorithm code, or whatever your most thoughtful business logic is.
And if you still need to review code for some kind of policy reason, then 1) challenge the specific need and focus on the outcome that the requirement is actually getting at, and 2) consider when the review needs to happen. In my design loop where I'm going back-and-forth with the agent early in a task, I have it describe the interfaces that it's going to change and what they'll look like. I know people talk about "shift-left" for testing, but reviewing code before it merges to mainline is such a slowdown in throughput. Maybe some code review should "shift-right", and happen after integration testing instead. And 3), consider what you actually want to look at when you review. If in my feature definition or design feedback I said I wanted the UI to look a certain way or the implementation to be done a certain way, I might want to focus my attention on those areas. If the coding agent found new things that needed to happen that it didn't talk about in the design, then I need that called out so I can see if it made changes that I was happy with.
Independent, thorough, fast test cycles
Another thing I used to do when I'd join a team is set up a version of our service running entirely on my development machine. This was even my first task when I joined amazon.com 20 years ago: get the web server called "Gurupa" to run on your developer desktop under my desk. To do that I had to learn all the tools and how to troubleshoot it. It taught me a lot about how the system worked. But other than learning, why did I need to set this up? So that cycle time of coding to testing would be faster. Who wants to wait hours for their tests to run before they find out if their change breaks everything? And you especially want those tests to run before you push the changes and break the shared "devo stack". Instead of that, it was important to be able to run the whole system - often many microservices - together on my developer desktop. And for more complex testing, to have my own deployed stack in my own AWS account where I could do end-to-end testing without risking "breaking mainline" or "breaking devo".
But sometimes I'd join a team who let this "local stack" hygiene slip. So when I'd join, I'd tidy this up and get it running locally. But then over time, I admit sometimes I'd let this slip on my own service, or it'd get too complicated to have a perfect simulated version of every dependency.
But now, having great local running versions is absolutely critical. Agents move so fast and push changes so often that the cost of a regression in mainline is a huge problem that stalls progress. If an agent has to wait an hour for their tests to run on every iteration, it'll still take forever to get anything done (even though you can get a bunch done in parallel).
Fortunately, it's also way easier than before to build high-quality simulators of your dependencies and wire it all up locally. It's a relatively simple task for a coding agent. Sometimes it requires a little code refactoring to make it happen, but it's easy. So do it!
I spend a lot of my time making sure my coding agents have a setup where they can run high-fidelity tests locally (or at least pre-merge in a dedicated stack). And for tests that can't happen there, I focus on high-quality automated testing in a post-merge, pre-production environment, and building the machinery to automatically address failures.
The more things change…
In the world of frontier engineering, I'm not really doing new things that I wasn't doing before; I just spend a ton of my time on things that I used to do only occasionally. The economics have changed; instead of focusing on new team member onboarding docs every few months, I update them daily. Instead of begrudgingly accepting style checkers to gate changes in my codebases, I embrace them. Instead of fixing flaky tests begrudgingly, I do so urgently (ok but still a little begrudgingly).
Some of these tasks aren't as glamorous or endorphin-generating as coding was. They don't always activate the same part of the brain as before. Some of these were tasks that languished on backlogs not only because "we weren't given enough time to do them" (I'll leave a rant about ownership for another day), but also because they were boring. Sure today I'm still doing that inventing, designing, researching, and exploring, but proportionally more of my time is spent on tasks that feel more like eating my vegetables.
I already had this "part 2" sketched out when I posted part 1, so that didn't leave me a lot of wiggle room to bring in the ideas that folks shared with me after part 1. So that means I now have to sign up for a part 3 to actually do that. There have been some really insightful things that folks have told me about, from things that worry them about the future, to some angles to what we can learn from cognitive science and psychology in Cat Hicks' What helps an identity crisis?, to thinking about how job role definition can help or hurt as things change, with Annie Vella's talk called The Specialization Trap. I don't know what shape part 3 will take yet, but it'll probably talk about this kind of thing.

What am I missing here? What do you do now that you weren't doing before? What are you still doing, but where you've found a new way to do it more easily or more effectively? Or what hasn't changed at all? What parts aren't solved yet that need something to change? And in lieu of a Dune-universe-like revolution against 'thinking machines', what should we do about it?