‹ Field Notes
Delivery recovery

Google says its own engineers already run codebase migrations on Gemini 4 Argon. The question that asks what you do when one slips.

On September 30, 2026, TechCrunch reported that Alphabet had released Gemini 4 Argon, a model Google calls its most powerful yet, built for coding, research, and writing, with what Google describes as a particular knack for cybersecurity. For now the rollout is narrow. According to TechCrunch, the model is going only to a select group of cyber partners through what Google calls its Fairwind Program, and the company says it can "autonomously find, validate, and patch critical software vulnerabilities." The line that should catch a manager's eye sits lower in the story. TechCrunch noted that Google says its own staff have already been using the model for daily work, including "debugging and codebase migrations," and it quoted the company's blog post: "Built to sustain deep reasoning across complex, long-horizon workflows, Argon is fundamentally changing the way we work and build at Google." TechCrunch also pointed out that the benchmark claims come from Google's own blog, which leans on a benchmarking startup called Vals rather than on independent testing.

Read "long-horizon workflows" for what it is. A codebase migration is the archetype of the long project: it begins with a tidy plan, meets the real codebase, and slips. Nobody builds, or markets, a tool for long-horizon work unless long-horizon work is where the time goes, and the time goes there because somewhere in the middle the plan stops describing reality. The dependency that was supposed to be ready isn't. The scope that was supposed to be fixed grows. Whether the migration is run by a model or by a team of engineers, that moment arrives, and what matters is the length of the gap between the plan becoming false and someone saying so out loud. That gap is not a technology problem. It is a behaviour, and it is why a model launch about sustaining reasoning across long work is also, exactly, an interview question: How do you resolve a project which is delayed?

The interview question
“How do you resolve a project which is delayed?”

Why they ask it

Every manager has had a late project, so the interviewer is not screening for a clean record. They are screening for three habits. The first is detection: how early the candidate noticed the slip, and whether they noticed it from the data or from a stakeholder's frustration. The second is diagnosis: whether they can tell the difference between a project that is late because of a bad estimate, one that is late because a requirement changed, and one that is late because the team is blocked. Those have different fixes, and a candidate who applies the same fix to all of them is guessing. The third is renegotiation: whether they went back to the people who owned the deadline and changed the deal, or quietly absorbed the gap with overtime and hoped. Underneath all three, the question is about the candidate's relationship with bad news.

The trap

The weak answer is the heroics story. The project was late, the team rallied, people worked nights, it shipped. It sounds like leadership and it is actually the absence of it: the candidate has described a delay they did not manage, only survived. A close cousin is the process answer, in which the candidate reassessed priorities and communicated with stakeholders without ever naming a priority that moved or a stakeholder who heard something they did not want to hear. The third failure mode is the blame answer, where the delay belonged to another team and the candidate's role was to wait. All three share a missing fact: the date the candidate knew. A strong answer names the moment of detection, the specific diagnosis, and the single decision that changed the project's trajectory, and it is honest that the decision cost something.

Applying STAR-T

Situation. Our schema migration stalled when the data team reorganized. The migration had been planned to run for a quarter, and partway in, the engineers who owned the legacy write path were moved to another org. Their replacements did not know that path, and the burndown went flat for a couple of weeks.

Task. As the manager, I owned the deadline, which was tied to retiring a database the finance team was paying to keep alive. My job was to find out whether the flat burndown was a staffing problem, a knowledge problem, or a scope problem, and to tell the stakeholders what I found before they had to ask.

Action. I did not start by adding people. I spent a couple of days reading the blocked tickets with the new owners and found the delay was knowledge, not capacity: the new engineers were rediscovering edge cases the old team had carried in their heads. I asked the departed engineers' new manager for a bounded loan, a few hours a week of pairing rather than a transfer, and rewrote the plan around the edge cases we now had a list for. Then I went to the finance stakeholder with a revised date and a smaller first milestone, and explained what had changed, not just that it had.

Result. The migration landed on the revised date, the database retirement moved by a single planning cycle, and the edge-case list became the test suite for the next migration. The stakeholder's trust held because the slip reached them from me, with a cause attached, rather than from a dashboard.

Trade-off. The revised plan dropped a planned cleanup of the old read path to protect the date. We logged it as known debt with an owner, and the stakeholder agreed that a dated promise with a smaller scope beat an undated one with everything in it.

The follow-up that breaks weak answers

The follow-up a good interviewer reaches for is simple: When did you first know it was going to be late, and who did you tell that day? It breaks heroics answers because heroics stories have no detection date; the delay arrives as a crisis, never as a signal. It breaks process answers because communicating with stakeholders cannot survive being asked for a name and a day. And it exposes the candidate who found out late and told late, which is the real thing being measured. A strong answer has a date it is slightly embarrassed by, a conversation it did not enjoy, and a reason that conversation happened before it was forced.

Score your answer against the director’s bar

Q: How do you resolve a project which is delayed?

Ready when you are

Bank your delay story with its detection date and the one conversation you didn't enjoy, then rehearse it until the follow-up lands softly. Try it free →

Rehearse this in L8 Loop →