The artifact just became the cheapest thing in hiring. The room test is the live follow-up you can’t generate, and it’s where interviews actually get decided in 2026.

Seven or eight years ago I made it to the final round at a large international company. Four interview stages, maybe five. Portfolio review, case studies, team conversations, the whole gauntlet. I passed every one. The last step was a conversation with the VP of design, and everyone around the process treated it as the easy one. The handshake round.
She asked me how design thinking had changed the way I work.
The label threw me. Design thinking was just becoming a term of its own back then, a thing with a capital D that people put on slides and certifications. I practiced it every day: framing the problem before touching solutions, testing assumptions on real users, iterating in the open. But where I worked, nobody called it design thinking. We called it solving the problem. So I asked her, honestly, what do you mean by design thinking? The room went silent.
She explained, briefly and politely, what she meant. And while she talked, I watched her make up her mind in real time. I could see the verdict assembling behind her eyes before I said another word. I spent the rest of the interview demonstrating that I did everything the term described, that I had been doing it for years without the label. It did not matter. Every answer after my question arrived at a door that had already closed.
For years I filed that story under bad luck. Wrong vocabulary, wrong year, wrong VP. It took sitting on the other side of the table, running these conversations weekly as a manager, to understand what actually happened. I did not fail a knowledge test. The work was real and the practice behind it was real. I failed something else: the live, unprepared moment where you translate your own judgment for a stranger, under pressure, with no artifact to hide behind.
I call that moment the room test. And I think it is about to become the entire interview.
The room I run now
Not long ago I sat on the other side of that silence. I have blurred the details; the two questions are not.

The strongest case study I had seen in months. A banking onboarding flow, which I can judge closely because I run one. Clean problem framing, research artifacts, a decision log, before and after numbers. The candidate presented it fluently. If the interview had ended at the presentation, we would have been talking about an offer.
Then a colleague asked the first question: you sequenced identity verification before showing the user any value, most teams do the opposite, what made you choose that order? A pause. The candidate walked us through the screen we were already looking at, again.
I asked the second question: what almost shipped instead? Nothing came back. Not a weak answer. No answer. There was no other version. There was no almost.
I want to be careful with this story, because it is not about catching anyone, and this article will not become one of those. I do not know how that case study was made, and I did not need to know. What the two questions established was narrower: the judgment we were hiring for was not in the room that day. Whether it existed somewhere else, in a different form, on a better afternoon, the room could not say. Rooms are not truth machines. They are pressure tests. Mine failed one years ago while carrying eight years of real practice.
But the two stories share one structure, mine and the candidate’s. The artifact survived and the conversation did not. And the reason this matters more now than it did in 2019 is simple: the artifact used to be expensive to make, it is not anymore.
The inversion
For twenty years the design career ran on a quiet division of labor. The artifact carried the proof: the portfolio, the case study, the deck. The interview existed to verify the artifact, to confirm that the person in front of you made the thing behind them. Expensive object, cheap conversation.

AI inverted that economics in roughly two years. A polished case study, complete with research narrative, decision log, and before and after metrics, can now be assembled in an afternoon by someone who was never near the project. So can the deck. So can the take home exercise, which is why hiring managers are quietly abandoning it. The visual signal that used to mean I can produce work at this level now means I have a subscription.
Pawel Klasa, a product designer writing in Bootcamp, made the case this spring for walking away from the portfolio entirely: “AI can fake a portfolio in an afternoon.” His answer is a public catalogue of continuous work, writing, talks, shipped experiments, a body of evidence too long and too messy to counterfeit. It is a good answer, and it is also a decade of homework. Most designers do not have a decade.
What everyone has, whether they want it or not, is the conversation. The artifact is now the cheapest thing in the hiring pipeline, and the conversation is the expensive one. The proof did not disappear, it moved.
One thing you can do with that this week: inventory your own proof. List everything you would show someone to establish your level, then mark each item as artifact or live. If everything on the list can be attached to an email, the list is weaker than it was two years ago, through no fault of yours.
Writers felt it first
Design is not the first craft to hit this wall. Writers got there a year earlier, because text was the first thing AI made free.
Matt Lillywhite, an essayist who publishes in The Daily Draft, wrote a piece in May about the folklore of spotting AI writing. He spends most of it dismantling the popular tells: the punctuation theories, the suspicion of similes, the idea that a qualifier is a confession. Then he lands somewhere more interesting: the giveaways worth taking seriously live outside the text entirely, in the author. Publish something you barely understand, and eventually a human being asks you to expand on a point, or challenge a statistic you never verified, and the gap between the writing and the writer opens up in front of a live audience, “while they gradually realize you don’t fully understand your own argument.”
Replace author with designer, article with case study, live event with portfolio review. Nothing else in the sentence needs to change.
Design hiring is already moving in this direction. Tom Scott, a recruiter who runs design and design leadership searches and writes the Verified Insider newsletter, described the shift in February: “Live, in-progress work can outperform a polished portfolio.” Interviews built around messy Figma files that are days old, concepts that have not been validated, decisions that have not been defended before. He reports seeing these formats far more often now, and his read on who struggles is precise: candidates whose articulation outruns their execution once the conversation leaves the script.
Now the honest caveat, before this argument gets too pleased with itself. The live conversation is not unfakeable either. Shraddha Sunil and Mudit Saraf, writing in Harvard Business Review in June after interviewing 120 talent acquisition leaders and analyzing more than 6,000 screening sessions, found that candidates can perform convincingly in remote interviews with real time AI assistance: “the ability to perform well in interviews is becoming infinitely scalable and practically free.” One asterisk you should hold while reading that: both authors cofounded an interview screening company, so treat the finding as directional evidence from people with a stake, not neutral research.
But notice what real time assistance can and cannot carry. It can carry the first question, the one every candidate saw coming. It struggles with the second one: the follow up that departs from every script, that asks for the almost, the tradeoff, the regret, the specific Tuesday when the decision got made. Recall can be delegated to a hidden window. The texture of having been there cannot. Which is why rooms are decided at the second question, and why the useful response to all of this is not paranoia about other people.
It is asking whether your own defense of your own work is script deep. Before your next interview, volunteer the messy file. Walk someone through work that is two weeks old and unresolved. If that idea frightens you, that fear is information.
What the room actually measures
“The hand can never consistently produce better than the eye can discern.” — Julie Zhuo
Seniority is two independent measurements, not one ladder. Level is the scope you own, granted by the organization. Stage is your maturity as a practitioner, accumulated rep by rep, and no reorg can hand it to you or take it away.

Level shows up in artifacts. Look at what someone owns and shipped, and you can read their scope off the deliverables. But artifacts are exactly what became generatable. Stage never lived in artifacts. Stage shows up in real time, under questioning, in the space between a hard question and your first honest sentence.
The room test is the Stage axis made visible. Specifically, a room reads you on four axes.
Whether a hostile question lands as information or as threat. The mature practitioner hears an attack on the work as data about the work. The early one hears it as data about themselves, and you can watch the difference in the first three seconds.
Whether your patterns transfer under pressure. Anyone can apply their process to the project they prepared. The room asks you to apply it to a variant you have never seen, live, which is the only condition under which a pattern proves it is a pattern and not a memory.
Can you defend a tradeoff you made six months ago, including the part where you would now decide differently? The best answer to what almost shipped is never nothing. It is a specific option, a specific reason, and usually a small regret.
Whether you know where your own work is weak before the room finds it. The hiring side sees the same pattern from the opposite chair: the candidates who name their own gaps before being asked are the ones who read as credible.
Notice what is missing from that list. Polish, which is now generatable. Recall, which is now delegable. And confidence, which was always performable. Confidence theater can survive a presentation. It rarely survives a follow up, because theater has no almosts.
One more thing the list refuses, and I refuse it deliberately. Failing a room does not make you a fraud. I failed one carrying nearly a decade of real practice, because one specific skill was missing: translating private judgment into public defense, in someone else’s vocabulary, at conversation speed. The old model treats that skill as theater, as interviewing well, as something beneath the real work. That framing is wrong on the facts, not merely the feelings. When the judgment is real, the defense is the last mile of the work.
My own data points the same direction, with the caveats it always carries. The assessment I built asks scenario questions instead of collecting self ratings, which makes it, in effect, a private room: nobody is watching, but you cannot pattern match your way to the senior sounding answer. Around one hundred designers have taken it, a small and self selected sample, so read these as instrument observations, not population statistics. When I rewrote the question bank so that scope signals and maturity signals could separate, four in ten profiles came off the diagonal: their Level and their Stage told different stories. And nearly half of the people who start it call themselves senior, while the assessed results center a full level lower, a comparison of unmatched groups, so a shape, not a verdict. The narrow claim both numbers support: most of us have never had our self assessment tested, and the first test moves it.
Reps, not talent
Here is where this article leaves the AI panic genre, because the panic genre ends at naming the problem and this is where the useful part starts. The room test has a property that polish never had: it is trainable, in public, for free, starting Monday. I watch it being trained every week, because I run these rooms from the manager’s chair: calibrations, promotion cases, portfolio reviews, stakeholder challenges. The designers who pass them are not the most talented ones I manage. They are the ones with reps. Four reps, specifically.

Narrate your decisions before anyone forces you to. In crits, in standups, in the Figma comment, say the why out loud while the decision is fresh. This was exactly the rep missing from my VP interview: eight years of practice and zero repetitions of explaining that practice in anyone else’s vocabulary. The judgment existed. The translation did not, because I had never once rehearsed it.
Close the loop on every project. Go back after shipping and name what worked, what did not, and what you would change. The question what almost shipped instead has an answer only if you kept your almosts. Most designers throw them away the day the final version wins.
Invite the hostile question early, in rooms where it costs nothing. Ask a peer to attack your strongest case study for ten minutes. Ask your lead to play the skeptical stakeholder. The first time someone questions your work face to face should not be the time it matters.
“A designer who delegates strategy cannot ask for the title of excellent.” — Soleio Cuervo
Sit in on reviews you are not required to attend and watch how someone a stage beyond you takes a hard question: how long they pause, what they concede, when they push back. Judgment under pressure is learnable by observation long before it is learnable by experience.
None of these require talent, a budget, or permission. They require a tolerance for small early embarrassment, which is precisely the price that most people optimizing their artifacts are trying to avoid paying. Pay it in the cheap rooms.
Stop optimizing the artifact
I will not pretend the room deserves the weight it is about to carry.
Rooms misfire. Mine did, seven years ago, rejecting real judgment over a missing label, and I have no doubt rooms I run have misfired since, in ways I did not catch. Rooms can be gamed, and the HBR data says the gaming is industrializing. And remote work spent five years shrinking the number of rooms we sit in at all, which means fewer cheap chances to practice and higher stakes when a room finally arrives. The industry has not solved any of that, and I cannot solve it here.
The room gets the weight anyway. Not because it is fair, but because it is the last assessment standing that you must take in person, as yourself, with no artifact in front of you. Everything else in the pipeline, the portfolio, the case study, the take home, the polished answer to the expected question, can now be produced without you. The room is the only place left where your presence is the evidence.
So the practical conclusion is a ratio. Most designers spend something like twenty hours polishing the artifact for every hour rehearsing its defense. That ratio was rational when the artifact was the proof. The proof moved. Stop optimizing the artifact. Start rehearsing the defense.
The VP asked me one question, and I lost the room. It took me years to understand that the question was never really about design thinking. It was about whether I could stand next to my own work without the work speaking for me. I could not, then. I can now, and not because I became smarter or more senior, because I got reps.
A case study can be generated. A witness cannot. In every room that matters from here on, you are not there to present the work.
You are there to prove you were present when it happened.
I turned the two axis model into a free self assessment: scenario based, about ten minutes, and the output is a profile, not a grade. Treat it as a private room. You can take it at designlevel.io/assessment. These articles are also becoming a book later this year, chapter by chapter.
If you think I am wrong about any of this, tell me. The strongest corrections to this argument have come from readers, not from me.
Resources cited:
- Matt Lillywhite, The Obvious Ways To Spot Someone Secretly Writing With AI (The Daily Draft)
- The same essay, open version on Lillywhite’s Substack
- Pawel Klasa, Your design portfolio is performance. Your public work is evidence. (Bootcamp)
- Tom Scott, FAQ: Design Hiring in 2026 (Verified Insider)
- Shraddha Sunil and Mudit Saraf, AI Has Broken Hiring. Here’s How to Fix It. (Harvard Business Review)
- Julie Zhuo, The Death of Product Development as We Know It
- Soleio Cuervo, quoted in Julie Zhuo, How to Spot a World-Class Designer
- Part 4, The promotion that didn’t take
- The Design Level assessment
AI can fake your portfolio. It can’t fake the second question. was originally published in UX Collective on Medium, where people are continuing the conversation by highlighting and responding to this story.
This post first appeared on Read More