Google is testing a “Buy” button in Gemini on a few products first. The question that tests conviction with evidence.
The detail worth dwelling on is the gap between the announcement and the test. Earlier this year, according to TechCrunch, Google introduced the Universal Commerce Protocol and said shoppers would be able to buy eligible products through Gemini and AI Mode using a Google-hosted checkout. The flow in this test appears different, bringing up a Flipkart-branded checkout instead. The reasons aren't public, but the shape will be familiar to anyone who has run a product organization. A direction gets stated with confidence. Then the evidence for it is gathered in a small, bounded experiment before anyone commits to scale, while most users keep seeing the ordinary experience. A firm view tested with a narrow proof isn't a question of commerce strategy. It's a behavior, and it's the one this interview question is built to find.
Why they ask it
On paper, two kinds of manager look alike. One has strong opinions. The other has strong opinions that are usually right, and can show why. This question separates them by asking for both halves at once: the knowing and the backing.
At senior levels, both failures are expensive. A manager who can't hold a position gives way to whoever argues loudest, and the organization loses the benefit of their judgment. A manager who holds positions without evidence wins arguments they should have lost, and the organization pays for it later. The interviewer wants to see how the candidate turned private certainty into something colleagues could check. The certainty itself proves nothing. What counts is how it was converted.
The trap
The most common failure is the vindication story. The candidate saw the answer, everyone else was wrong, the candidate pushed, and events proved them right. It sounds strong, and it fails for three reasons.
First, it has an outcome and no evidence. Being proved right later shows the panel the candidate was right. It doesn't show they were rigorous, and a good interviewer can't tell it apart from luck. Second, it casts colleagues as obstacles, which suggests poor judgment about people even when the technical judgment was sound. Third, the claim was never sized. Nothing in the story says what the candidate would have accepted as proof that they were wrong, so the conviction can't be told apart from stubbornness.
The opposite failure is quieter. The candidate gathered data, laid out options, and let the group decide. That's good collaboration, but it answers a different question. This one asks about a time the candidate knew, and a story with no conviction in it has skipped the premise.
Applying STAR-T
Situation. Pick a moment where the stakes were real and the disagreement was reasonable. The opposing view should have had merit. If it was obviously wrong, there was nothing to back up. Our team was preparing to rebuild the billing service. I believed the failures came from the retry logic, not the architecture. The senior engineers disagreed, and on the evidence we had at the time, their case was sound.
Task. State what the candidate was actually responsible for. It wasn't being right. It was making the claim testable before a costly decision locked in. My job was to either stop the rebuild with evidence or get out of its way quickly.
Action. This is most of the answer, and it should cover four things in order. First, where the conviction came from: a pattern seen before, a signal in the data, a specific anomaly. Second, the smallest test that could settle it. Third, the result that would have changed the candidate's mind, written down in advance. Fourth, how the evidence was presented to the people who disagreed. I proposed patching the retry logic for the worst-affected customer segment only, and agreed with the engineering lead beforehand that if error rates didn't fall clearly, I'd back the rebuild. I shared the raw logs, not a summary of them. Committing to a threshold in advance is the detail that makes the story credible. It shows the conviction came with its own way to be proved wrong.
Result. Give the outcome in concrete terms: the metric that moved, what was decided, what it cost. Then give the second result, which is what happened to the working relationship. The rebuild was deferred, and the engineering lead co-presented the findings. A result where the doubters became co-owners of the decision is worth more than one where they were simply overruled.
Trade-off. This question rewards naming a cost, so include one. Narrow tests are slow, and they can mislead. A fix that works for one segment may not hold for all of them. We lost several weeks on the rebuild timeline, and I had to own that if the patch had failed. A candidate who names the cost shows they weighed the risk before acting, not after.
The follow-up that breaks weak answers
The follow-up is almost always a version of this: What would have told you that you were wrong, and did you decide that before or after you saw the data?
Weak answers improvise here. The candidate names a threshold on the spot, and because it fits the result a little too neatly, it reads as having been set afterwards. Strong answers already have it. The candidate names the threshold, says who agreed to it, and explains why that number and not a looser one.
A second probe often follows: When were you this confident and wrong? A candidate who has a ready answer, and who can say what they now check that they didn't check then, has shown the calibration the original question was testing for. A claim sized to its evidence holds up under follow-up questions. An unsized one tends to fall apart.
Score your answer against the director’s bar
Q: Tell me about a time you knew what to do and were prepared to back up your claims.
Bank the one story where your conviction came with its own evidence, and rehearse it until the threshold you set in advance comes out without prompting. Try it free →
