In 2020, a page like my portfolio homepage went through three people in a row. Someone in product decided what to say and to whom. Someone in design decided how it would be understood. Someone in engineering built it and made sure it didn't break. Each one owned a part, and the next person caught the previous one's mistakes.
This week I redesigned that homepage with AI agents. Claude Code designed and wrote the code; Codex reviewed it. It moved fast, and along the way I learned something I didn't expect: agents do a lot of the work of all three crafts, but they don't bring the judgment of any of them. If you don't bring it, nobody does.
Product
- In 2020
- One person decided what to say and to whom.
- Today the agent
- Keeps producing versions, even when the request is wrong.
- Still your job
- Deciding what must land in 3 seconds.
Design
- In 2020
- Another decided how it would be understood.
- Today the agent
- Takes adjectives literally and swings to extremes.
- Still your job
- Format, word limit and who it is for.
Engineering
- In 2020
- Another built it and made sure it held up.
- Today the agent
- Writes the code and calls it done.
- Still your job
- Having another model review it, and asking for proof.
What follows are four lessons from that project, grouped by craft. Engineering gets two, because that's where I made the most mistakes: reviewing and testing. Each has what happened, a piece you can play with, what I'd do differently and a template to copy. At the end you can build your own request using the three roles.
Product
1. The hard part isn't making. It's deciding
The third animation on the homepage went through eight versions in one morning. One with lots of text. One made of abstract dots. One with icons. Another with speed bars. They all moved nicely, and none of them explained what I wanted to show.
The agent wasn't the problem. It did exactly what I asked, and what I asked was badly framed. In 2020, someone in product would have asked the question before anything got drawn: what does a visitor need to understand in three seconds? After the eighth version I stopped and wrote that sentence down. I picked three seconds myself, thinking of an audience that skims. The next version used a Gantt chart to show how I went from working in a line to working in parallel, and it finally said what I'd been trying to explain.
An agent rarely asks why unless you tell it to, which is why the template below does. And when a new version costs a minute, it's tempting to ask for another one instead of figuring out what went wrong.
What I'd do differently
- Before asking for anything, write in one sentence what the person needs to understand.
- Start with three options. If none of them gets the idea across, fix the request before asking for more.
- Judge each option against that sentence, not against how I feel that minute.
Context: [what it is and who it's for].Whoever sees this must understand in 3 seconds that [the sentence].Before proposing anything, tell me whether that sentence is clear or what it's missing.Then propose 3 different options, say in one line whether each meets the sentence, and recommend one.
Design
2. AI agents go to extremes
It happened to me again and again with the animations. When I asked for something "simpler", the agent stripped so much that the drawing stopped meaning anything. When I asked for "clearer", it filled the screen with text. It swung from one end to the other and never landed in the middle.
What was missing was design judgment: adjectives aren't instructions. What helped was swapping them for concrete decisions. I picked formats people recognize without explanation, like a dashboard, a counter or a screen with a cursor. I capped the text at three words per animation. And I told the agent who it was for: someone who comes to the site to evaluate my work and won't stop to read.
Too abstract: nobody gets what it is
What I'd do differently
- Don't ask for "simpler" or "clearer". Ask for a specific format: dashboard, list, counter, before and after.
- Put a word limit in writing.
- When the answer swings to one extreme, name the other one as the limit: "simpler, without making it abstract".
Show [the idea] using a familiar format: [dashboard / counter / list / before and after].[number] words max inside the graphic.Audience: [who]. They need to get it without reading a paragraph.Don't make it abstract and don't fill it with text.
Engineering
3. Whoever writes it doesn't review it
This is the biggest change in how I work. I use two agents from different companies: Claude Code, from Anthropic, writes the code and the copy; Codex, from OpenAI, reviews it with one instruction: try to reject it. Nothing ships without its approval.
In my experience, two different models get different things wrong. The homepage chat had two bugs that Claude Code had called done. If someone left the page just as they unlocked a secret, the chat got stuck on "typing". If the connection dropped at the wrong moment, the retry button disappeared. Codex found both.
It wasn't the only time. For this article alone, Codex flagged 25 Spanish sentences that sounded AI-written. In the code review it asked for six changes; five were real, and it only approved once they were fixed. That's the job a demanding engineering teammate used to do.
What I'd do differently
- Always review with a different model from the one that wrote the work, and give it the cases I want checked.
- Ask it to look for problems, not to say whether it's fine.
- If it flags a possible bug, ask how to reproduce it. If there's no way, ask what it saw.
git diff | codex exec "Review this change as if you were going to reject it. You didn't write it.Look for: edge cases, what happens if the user leaves halfway, what happens if the connection fails, copy that doesn't make sense.Reply with a list ordered by severity, with file and line. End with APPROVE or CHANGES-REQUESTED."
4. If I didn't test it, it isn't done
The other half of engineering is proof. Before publishing, an automated browser checked the homepage at desktop and phone sizes, in the site's three languages, and returned screenshots and measurements of the main pieces.
Those measurements caught two mistakes. I'd been told a change was halfway done, and measuring showed it had never been started. Another time, on the phone, a fixed control showed up on top of the chat. Neither one appeared in the agent's summary. The screenshots let me compare what the agent said with what was actually on screen.
What I'd do differently
- Write down what "done" means before starting: which screens, which languages, which cases.
- Ask for proof you can look at: screenshots, measurements, logs.
- Look at those screenshots myself, not at the summary.
Before saying it's ready, test on desktop (1440 px) and phone (375 px), in [languages].For each combination, give me a screenshot and one line on what you measured.If something fails, don't call it done: tell me what failed and what you'll change.
What didn't work
With the sound, I was missing design judgment, and I didn't see it in time. I asked an agent to add sound effects to a short video. It handed back the file with volume and sync measurements and called it done. When I played it, the synthetic beeps sounded cheap.
I had accepted the measurements before listening. For the next version I used music and sound effects made by people, with each effect timed to a movement on screen. Now I listen to, read and look at everything that depends on taste myself. No measurement tells me whether it sounds like me.
Build your request with the three roles
This request gathers the decisions of all three roles before you start. Fill in the fields and paste the result into Claude, Codex or whatever AI you use. Then run what you get back through the three reviews below.
Fill in the fields and press "Build my request".Review it as product
You're a product lead. Read this and tell me in 3 lines: who it's for, what that person understands in 3 seconds, and what they still need in order to decide. If the answer isn't clear, say so.Review it as design
You're a senior designer. Check whether this can be understood without reading: format, hierarchy and amount of text. Flag what's extra and what's too abstract. Propose one change, the most important one.Review it as engineering
You're an engineer and your job is to break this. List what can fail: edge cases, phone, slow connection, languages. Order by severity and explain how to reproduce each problem.What I'm taking with me
Agents don't replace product, design or engineering. They do much of the work, but each craft's judgment has to live in someone. On this project, that someone was me, with a second model reviewing everything.
That's why understanding all three crafts matters to me today as much as mastering my own. A designer who knows what product would ask and what engineering would break works better with agents, and with people too.
I'm still learning to work this way. If you work with agents too, I'd like to know how you check what they give you.
Want to see how I applied this?
Nova, the AI on my homepage, knows the project. Ask it anything, or look at the cases.