We operate what we build, so everyone here carries a pager for something they wrote. That is the job: systems built to still be running a year after launch.
Most AI engineering roles end at the demo. Someone else takes it to production, someone else again gets the alert when it drifts, and the person who designed the retrieval layer never finds out which of their assumptions was wrong.
Here the same person does all three. You scope the work with the client, build it in two-week increments, write the evaluation that decides whether it ships, and then sit on the rota that keeps it alive. It is a longer loop than most engineering jobs, and it is the only way we know to build systems that are still working a year on.
If you would not want to be woken up for it, you should not have shipped it.
That is the whole deal. Everything below is a consequence of it.
The demo
Engineer A
Production
Team B
The 2am alert
On-call C
We hire slowly and against a signed engagement rather than a forecast, so every one of these gets weighed properly.
Not just shipped it. We want to hear about the thing that broke at 2am, what the graph looked like, and what you changed so it could not happen the same way twice.
Every engagement here starts with telling a client which of their ideas are not worth building. That is impossible if you cannot hold a boundary on your own certainty first.
You will be in the room from the first sprint, explaining a trade-off to someone senior and non-technical. This is not a role with an account manager in between.
The whole company argues that an AI system you cannot evaluate is a liability. We hire people who already reach for the eval before the demo.
What we do not screen on
Each one costs us candidates other firms would filter out.
The process is short because a long one mostly measures who can afford a long one.
Forty-five minutes with an engineer who does the work, about what you have built and what you would have done differently. No take-home before we have met.
Written down so you can hold us to it, and so you know exactly what you are agreeing to.
A real change to a real system, reviewed and merged. Onboarding here is a pull request, not a slide deck, and the first one is deliberately small.
One evaluation suite becomes yours: what it measures, what it gates, and the conversation with the client when it fails a release it should have failed.
You take the pager for a system you helped build, alongside someone who has carried it longer. Nobody goes on call for code they have never touched.
Six things that are true of this job whether or not they suit you. Better you find out now than in month three.
You carry the pager
For what you write, on a rota that includes the people who run this company.
Client-facing from sprint one
You explain your own trade-offs. There is no account manager translating for you.
Two-week cycles
Something runnable at the end of each. No quarter-long stretches with nothing to show.
We say no to work
You will help decide which engagements we turn down, and you will be asked to defend it.
No timesheets
We do not bill by the hour, so there is nothing to account for in six-minute blocks.
Compensation, openly
Discussed in the first conversation with a real number, not at the end of the process.
We hire in small numbers, usually against a signed engagement rather than a forecast, so there are no open roles listed today. These are the kinds of work we hire for. We read everything that arrives, and we reply to most.
Not on the list? Send it anyway. We do not need you to have used our exact stack.
What you would work on
A link is enough. We read every one, and we reply to most.
Bring one process you think an agent could run.
We'll tell you straight whether it's worth building — and what it would cost if it is.