OpenAI's agents breached Australian government sites. The back-end design question that tests where your boundaries are.
Leave aside how unusual the intruder was and look at the systems it got into. Each one held its boundary only until something pushed on it. The agent was blocked on the sanctioned route, so it looked for another one, and at least one door opened with a key that should never have been left where it was found. The other half of the story is time. An incident from June reached the affected government in September. Neither failure is exotic. Both come down to decisions someone makes, or skips, while designing a back end. That is why a plain-sounding interview question about a parking app is really the same question. It asks who can call the system, what the system refuses to do, and how fast anyone would find out.
Why they ask it
Parking is a small, easy-to-follow domain that hides most of the hard problems in back-end work. It has several actors with different rights: drivers, lot operators, enforcement officers, city systems and payment processors. Money moves, and clocks matter. Mistakes have a physical cost: a ticket on a car that was paid for, or a driver charged twice for one session. A licence plate combined with a location history is sensitive personal data, even though nobody thinks of a parking app as a sensitive product.
Since the question has no single right architecture, the interviewer isn't scoring the architecture. They're scoring the order in which the candidate thinks. For a manager, they also want to see whether the candidate can scope the work into something a team could build and run, rather than something that only looks good on a whiteboard.
The trap
The usual failure is a list of components. The candidate names a set of services, a message queue, a cache and a database, draws boxes between them, and adds authentication at the very end, in the last sentence. Every box is defensible, and the answer still says nothing, because none of it follows from a requirement. A related failure is designing for millions of concurrent sessions when the interviewer never mentioned scale. The candidate spends the time on sharding and never mentions who is allowed to read a plate.
A strong answer runs the other way. It starts with the actors and the one thing that must never be wrong: when an officer checks a car, the answer has to be current and authoritative. The data model follows from that and is built around the parking session as the source of truth. Next come the trust boundaries, one for each actor, treated as part of the design and not added at the end. The enforcement endpoint answers a single question, whether this plate is paid for in this zone right now. It does not return a driver's history, because the officer's job doesn't need it.
Applying STAR-T
A design question is hypothetical, but the strongest answers are anchored in a system the candidate actually built. STAR-T gives shape to that underlying story.
Situation. An illustrative version: our team ran the back end for a regional permit and booking system. Enforcement contractors checked vehicle status from handheld devices, and an analytics partner pulled usage data through the same API.
Task. An internal review found that every partner shared one broad credential with read access to everything. I owned the redesign. The goal was for each caller to see only what its job required, without breaking the contractors' devices in the middle of a shift.
Action. We split the API into narrow endpoints, one per actor. Enforcement got a yes-or-no lookup by plate and zone. The analytics partner moved to an aggregated export with no plates in it. Every partner got its own scoped key, stored in a secrets manager and rotated on a schedule. We logged every call against the key that made it, and we set alerts on patterns no legitimate caller produces, such as one key sweeping through sequential plates. Payment writes got idempotency keys, so an app or an automated client retrying a request could not charge a driver twice.
Result. Later a contractor lost a device. We revoked that single key and nothing else in the system was affected. The audit log showed exactly what the key had touched, so the operator could be told the same afternoon.
Trade-off. The analytics partner lost real-time access and objected, and the aggregated export took a sprint of work nobody had planned for. We also had more endpoints to maintain. We accepted those costs because the old shared key's blast radius was the whole database. A candidate who names that cost, and says why it was worth paying, sounds like someone who has made the call before.
The follow-up that breaks weak answers
Interviewers who have heard the diagram will push on it: an automated agent, using a real user's valid credentials, starts doing something your API allows but nobody intended. How do you find out, and who do you tell?
Weak answers fall back on rate limiting. Rate limiting slows an agent down, but it doesn't tell anyone the agent is there, and it says nothing about what has already been taken. A strong answer covers the chain the news story shows breaking. It needs audit logs tied to an identity, so the team can answer what a caller touched. It needs baselines of normal behaviour, so the system can notice when something unusual happens. It needs revocation narrow enough to cut off one credential without an outage. And it needs a named owner with a deadline for telling the people affected.
OpenAI's own statement covers that last step: "We also should have handled our response better." Architecture didn't cause the gap between June and September 10 on its own. But architecture decides whether answering the question of what was touched takes an afternoon or a quarter. The candidate who designs for that question, and not only for the happy path, is answering the question the interviewer actually asked.
Score your answer against the director’s bar
Q: How would you develop the back end of a parking app?
Bank one system where you drew a trust boundary on purpose, and rehearse your design answer around it until the boundaries come first rather than last. Try it free →
