Content & AI

AI can generate 100 interfaces. It still cannot decide which one should exist.

When generation is free, teams stop being limited by what they can make and start being limited by what they can evaluate.

Six small generated screens beside one larger screen under review, its top row highlighted as the decision somebody made

More options were never the bottleneck

A design critique used to contain a scarcity problem. There were three directions because the team had time to make three directions.

Now there can be thirty before lunch. You can ask for a dashboard at three densities, a mobile flow with five navigation patterns, a landing page in ten styles, a dozen empty states and forty headlines. The marginal cost of another concept approaches zero.

That sounds like creative freedom. It can also become a new form of indecision. When generation is cheap, teams stop being limited by what they can make and start being limited by what they can evaluate. The bottleneck moves from production capacity to judgment capacity.

The jagged frontier is a product design problem

The BCG and Harvard field experiment on GPT-4 is useful because it avoided the usual framing. Consultants performed substantially better and faster on tasks inside the model's capability frontier. On a task outside that frontier, people using AI were more likely to produce incorrect answers.

The important word is jagged. Capability does not rise smoothly from easy to hard.

Product design has the same shape. AI can be very good at generating layout variants, applying known interaction patterns, summarising interviews, proposing copy, finding consistency problems, turning tokens into code and drafting edge-case lists.

Then it fails on questions that look simpler. Which user should this screen optimise for? Is this a usability problem or a positioning problem? Is the stakeholder request evidence or an opinion? Should we add another control or remove the capability from this surface? Should an AI action happen automatically, require approval, or not exist at all?

Those are not harder because they need more pixels. They are harder because they need context, consequences and a model of the business.

Generation, selection, validation, responsibility

Product design needs cleaner vocabulary for AI-era work, and I would separate it into four questions.

Four questions in AI-era design work — generation where AI is excellent, selection which needs the constraints, validation which needs a method, and responsibility for when it fails in production — with responsibility highlighted as the layer generated concepts hide
  1. Generation. What are plausible ways this could look or work? AI is already excellent here. 2. Selection. Which option best satisfies the constraints we actually have? This requires knowing the constraints, which is where teams often discover they do not agree. 3. Validation. What evidence would tell us the selected option is better? This requires method rather than taste, and it is why sizing a study correctly still matters. 4. Responsibility. What happens when this decision fails in production?

The fourth is the layer generated concepts routinely hide. A beautiful agent screen is easy. Deciding who can reverse an automated action, how the audit trail works, what confidence threshold triggers human review, and what happens to downstream records when the model is wrong is product design.

Seniority is mostly the quality of rejection

Junior designers are evaluated by what they produce. Senior designers should increasingly be evaluated by what they prevent. That does not mean conservatism. It means understanding the cost surface.

A senior designer can look at an elegant generated concept and say that the interaction creates a second source of truth, that the AI suggestion is impossible to audit, that the filter will be unusable at fifty clients, that the navigation works in the prototype because the data is clean, that the chart is strong visually and answers no operational question, or that the automation needs approval because recovery costs too much.

Those comments rarely appear in a portfolio. They are often the reason the shipped product survives contact with real users.

Taste is trained selection, not vibes

Taste is becoming a popular explanation for what humans retain, and it is vague enough to mean nothing. For product design I would define it as trained selection under constraints.

It contains aesthetics, and also knowing what to leave out, recognising when a pattern is overused, sensing when hierarchy is fighting the task, understanding how brand and interaction reinforce each other, and knowing which details deserve craft rather than standardisation.

Taste improves with exposure, but exposure alone is not enough. It requires feedback. A designer who has produced beautiful work for years without seeing what happened after launch has style, not judgment.

The trap: mistaking plausibility for evidence

Generated interfaces arrive with an unfair advantage. They look finished.

A polished prototype is psychologically persuasive. Stakeholders can imagine using it, engineers can imagine building it, and the team starts discussing icon choice and animation timing before anyone has shown that the workflow matches user behaviour.

That is not new, but AI compresses the time between idea and convincing artefact so aggressively that the error happens before the team notices the assumption. The defence is procedural.

  1. Write the user problem and the decision criteria before generating concepts. 2. Mark assumptions explicitly. 3. Generate structurally different approaches, not five skins of one layout. 4. Evaluate against criteria that existed before the concepts did. 5. Test the highest-risk assumption at the lowest fidelity that can answer it. 6. Only then invest in the polished direction.

If the model generates the problem statement, the criteria, the concepts and the critique in one uninterrupted loop, you have not created independent evaluation. You have created a very coherent echo chamber.

A hiring test for the AI era

If I were hiring a senior product designer now, I would spend less time on whether the final screens look polished, because AI is making polish easy to obtain. I would ask them to walk through one decision.

What did you initially think the problem was? What changed your mind? Which alternative looked better but lost, and what constraint killed it? What evidence did you trust, and what did you reject? What did you deliberately leave manual? What would make you reverse the decision after launch?

A designer who can answer those is showing the part AI cannot render for them.

The point

Making interfaces was never the whole profession. It was the part easiest to see, and AI is exposing that distinction by making visible output abundant.

When one person can generate a hundred plausible screens, the valuable question stops being whether you can design this. It becomes whether this should exist, and what you are willing to trade for it.

That is a better question for the profession anyway.

Frequently asked questions

Can AI replace a product designer?
It replaces a growing share of generation: layouts, variants, copy options, scaffolding and documentation drafts. It does not replace selection or responsibility. Deciding which option satisfies real constraints, what evidence is sufficient, and who is accountable when the decision fails in production requires context about the business that a model generating plausible screens does not have.
What is the jagged technological frontier?
It is the finding, from a field experiment with BCG consultants, that AI performs extremely well on some tasks and poorly on nearby tasks that appear equally difficult. Performance does not rise smoothly with task difficulty. In design work this means a model can produce excellent layout variants and then fail at deciding which user a screen should serve.
How should design teams evaluate AI-generated concepts?
Write the problem statement and decision criteria before generating anything, so the criteria cannot be shaped by the output. Generate structurally different approaches rather than visual variations of one layout. Then evaluate against the pre-existing criteria and test the riskiest assumption at the lowest fidelity that can answer it, before investing in polish.
What does design taste mean in practice?
Trained selection under constraints. It includes aesthetics, and also knowing what to leave out, recognising an overused pattern, sensing when hierarchy fights the task, and knowing which details deserve craft rather than standardisation. It develops through exposure combined with feedback about what happened after launch, which is why style and judgment are not the same thing.

Sources

This thinking, applied

design judgment on a team a few days a week

Share this

Keep reading

All posts