Applied AI Engineer
Job Description
About the AI Software Factory
The AI Software Factory is a new software-engineering function inside R&D.; We build the systems that let AI agents do real engineering work on IFS software: design, implementation, testing, review.
We are after an order-of-magnitude improvement in delivery speed. That is a hypothesis we intend to prove or disprove in production, not a slogan. Faster only counts if the software still works, stays secure, and can be supported afterwards.
Our output is software. Orchestration, evaluation, codebase analysis, verification, and the interfaces that connect all of that to the way engineers actually work. Product teams are our users and our pilot partners. The team is new, so the first pilots and much of the technical architecture are still open questions. You would be joining to turn the proposition into something that runs, rather than to inherit a finished system.
One of our core deliverables is a reusable framework for parallel software engineering at IFS. It has to define how a feature gets broken into work several agents can do at once, how dependencies and shared state limit that, what context, tools and guardrails each agent receives, where an engineer reviews or decides, and how separately produced changes come back together as one releasable result. It must cover the full lifecycle: plan, design, build, test, review, document, operate. Parallel code generation on its own is not the goal.
Why this role?
Pointing one agent at one ticket is increasingly common. Building a repeatable framework where many agents plan, build, test and review parts of one outcome at the same time and still produce coherent software is not, and you would be helping define how it works.
You get a meaningful slice of that capability. Clear responsibility for working components and the evidence behind one pilot, with senior architecture support behind you, instead of a queue of disconnected tickets. What you build is production engineering infrastructure: evaluation, regression and codebase-analysis systems used to decide whether an agentic workflow is safe enough to expand. Those results inform whether a pilot proceeds, where human controls stay necessary, and which practices get adopted more widely across IFS.
Our interview process uses realistic work samples from the Factory's problem space, such as agent-generated changes and evaluation evidence, so both sides can look at actual work instead of talking around it.
The problem this role exists to solve
An agent opens a pull request and CI goes green. That proves less than it looks. The agent may have written the tests itself, missed an indirect dependency, or made a change that works alone and collides with what another agent is doing three files away.
So verification splits in two:
Is this agent's output correct? Eval suites that mean something, deterministic checks, human validation where judgement cannot be avoided, and regression tests that keep known failures fixed when a model or prompt changes.
Do many agents' outputs compose? A codebase is a graph of dependencies. Running agents at the same time means understanding blast radius, spotting overlapping work, and verifying the merged result rather than approving a collection of individually green pull requests.
You help build that framework and prove it on real pilots. Principles become working software: work decomposition and dependency models, agent orchestration and isolation, coordination and reintegration, and the checks that show the composed result holds. You own substantial components and the evidence they produce, with architecture direction from a senior engineer. You contribute to the Factory-wide architecture without being expected to define it alone.
What you'll do
You help define and implement the framework for parallel software engineering across planning, design, build, test, review, documentation and operation. For each phase that means making it explicit what agents execute, what engineers review, and what stays human-owned.
Much of the work is mechanism. Turning a feature or engineering objective into a dependency-aware work graph: which tasks can run concurrently, which have to be sequenced, what context each agent needs, and where results have to synchronise. Then the controls that make running them at once safe, which means isolation of concurrent changes, dependency and blast-radius analysis, detection of overlapping work, failure and retry handling, and verification that a batch holds together as a whole.
You build and maintain an eval suite for a real agentic engineering workflow, covering planning, implementation, testing and review. You wire a pilot codebase and its CI pipeline into the Factory's verification harness: ground truth, smoke tests, full-suite execution, reporting someone can actually read. Every agent failure you observe becomes a durable regression test, so it stays fixed across model, prompt and tooling changes.
Some of it is judgement rather than code. Reviewing held-out agent outputs through a structured human-validation process, then working out when automated judging agrees with independent human reviewers closely enough to be worth trusting. Investigating runs that failed or came out strange, and making a defensible call: what failed, why it matters, which check backs the conclusion.
You also work with pilot teams on real delivery. Set the delegate/review/own boundary, watch where parallel work succeeds or falls apart, and turn what you learn into reusable capability instead of one-off pilot fixes.
Your centre of gravity is building and proving the parallel software-engineering framework through one pilot. The reusable capability is the product. The pilot is where its assumptions meet a real codebase, a real delivery workflow and a real engineering team.
In your first six months: take one pilot workflow from an engineering objective through dependency-aware decomposition, parallel agent execution, human review, reintegration and end-to-end evaluation against its real codebase and CI. Deliver reusable framework components, a documented set of known failure modes, and evidence showing where parallelism is safe, where work has to sequence, and why.
Requirements
Department: Finance
Function: Engineering
Experience Level: Mid-Senior Level