In short
Interview an agentic AI engineer on failure, not on building. The demo is the easy part and a model helps anyone produce it. What separates a strong candidate is how they reason about an agent that stopped halfway, spent too much, or took an action it should not have. Four conversations, no take-home.
We have written about what an agentic AI engineer is and where to find them. This is the part in between: once you have a candidate in front of you, how you tell a real one from someone who has wired a framework to a model.
The difficulty is specific to this role. Technical assessment has got harder across the board: a 2026 Karat survey of 400 engineering leaders found 71 per cent think skills are harder to assess than they were, with take-home signal degrading fastest. For agentic work it is sharper still, because the visible output of an agent project is a weekend of work and a model will help produce it. Everything that distinguishes a strong engineer is in the parts that never show up in a demo.
The loop: four conversations
None of these needs the candidate to write code on the spot, and all of them are hard to fake.
1. A system they shipped (45 minutes)
Ask them to walk you through an agent they built that ran in production. Not a prototype. Go deep on four things: what tools it could call, how it held state between steps, how they knew whether a run had gone well, and the worst thing it ever did.
Listen for: specifics about failure, cost and limits. People who have run agents unattended have long, precise and slightly weary answers. People who have not describe the architecture diagram.
2. The failure drill (60 minutes)
This is the most useful hour in the loop. Give them a scenario and ask them to reason through it out loud. For example: an agent was meant to complete six steps for a customer. It finished three, then timed out. Step two created a record in a billing system. Token spend on this run was ten times the usual. The customer is asking what happened.
Ask what they would check first, what state the system is in, whether the run can be resumed or has to be unwound, and what they would change so it cannot happen the same way again.
Listen for: idempotency without being prompted. A strong candidate asks immediately whether re-running step two would create a second billing record. They separate "resume" from "retry". They treat the cost spike as a symptom of the agent being confused rather than a separate problem.
The best signal in the whole loop is a candidate who asks, unprompted, what happens if step two runs twice.
3. Designing the tool surface (45 minutes)
Describe an agent you actually need and ask them to design the tools it can call. Then ask the question that matters: what would you deliberately not let it do, and where would you put the human?
Listen for: restraint. Strong candidates talk about the smallest set of tools that does the job, about which actions need approval, and about who in the business should own that boundary. A candidate who gives the agent broad access to everything because it is more capable has not been burned yet.
4. How they know it works (30 minutes)
Ask how they evaluated an agent before trusting it with real work. Evaluating a sequence of actions is harder than checking a single answer against an expected output, because there is often no single right path.
Listen for: an evaluation harness they built and maintained, how they caught regressions when a model or prompt changed, and honesty about what they could not measure.
What to leave out
- The take-home. "Build an agent in four hours" tests the part a model does well and hides the part you are hiring for. We have written about what take-homes cost you more generally.
- Framework trivia. Which orchestration library someone has used tells you little. The frameworks change every few months; the judgement about failure does not.
- Model knowledge for its own sake. Useful context, not the job. The hard problems here are distributed systems problems.
Keep the loop short
The four conversations above fit in one long day or two short ones. Strong agentic engineers are scarce and have options, so a loop that drags over three weeks loses them to the company that decided faster. Our own data says briefs with the loop agreed at intake close in the low teens of days; briefs that reopen the level partway through run past 35.
Every week the search runs has a cost. At the AUD 180-220k band for senior AI and ML engineers in Sydney, which is what these searches price against, the cost of vacancy calculator will show you what each week of delay is costing. Before the loop starts, the JD grader will check that the advert is attracting the right people in the first place.
FAQ
How do you interview an agentic AI engineer?
Focus on failure rather than building. Four conversations work well: a deep walk-through of an agent they shipped to production, a failure drill where they reason through a run that stopped partway with side effects, a tool-surface design exercise that asks what they would deliberately not let the agent do, and a discussion of how they evaluated agents before trusting them. None requires a take-home.
What questions should I ask an agentic AI engineer?
Ask about the worst thing an agent they built ever did, what happens if a step with side effects runs twice, where they put the human approval boundary and who decided, what tools they deliberately withheld from an agent, and how they caught regressions when a model or prompt changed. Strong candidates give specific, practical answers about failure and cost.
Should I give an agentic AI engineer a take-home test?
No. The visible output of an agent project is quick to produce and a model will help, so a take-home tests the easy part and hides the skills you are hiring for: handling partial failure, safe retries, cost control and evaluation. A structured conversation about a real failure gives far better signal.
What is the most important skill in an agentic AI engineer?
Judgement about failure. The strongest candidates think immediately about idempotency, whether a run can be resumed or must be unwound, and what the system should never be allowed to do without a human. These are distributed systems skills applied to a component that is non-deterministic by design.