About
I build AI systems for teams that have to trust the output, and I check that the output is actually trustworthy. Those two halves rarely live in the same person, and the gap between them is where most AI adoption stalls.
Before any of this, I spent more than twenty years in K-12 education, the last stretch as a vice-principal. When I moved into AI work, I did K-8 curriculum and assessment QC at an AI-education company: writing items, auditing rubrics, watching where a model’s confidence and a real person’s reality came apart. That is where I learned what a broken workflow looks like before it shows up in front of a customer, and what shipping the wrong answer costs.
So the work I do now is horizontal. The domain was education. The skill is getting AI to do real work and proving it holds up.
What I do
- Workflow audits. I find the steps in a team’s workflow actually worth handing to AI, ranked by payoff and risk, so effort goes where it pays. Here’s how I run one.
- Agentic systems. I build workflows that check their own output before a human sees it. Here’s one that reviews itself.
- LLM quality control. Runnable harnesses that catch the failures an accuracy dashboard misses, including a judge that waves through broken work. Here’s the QC framework and the judge-trust proof of concept.
How I work
- I instrument, I don’t assert. A guardrail you haven’t watched fail closed is a guess.
- If it isn’t reproducible, it didn’t happen. Every result links to runnable code.
- Adoption is the goal. A clever system nobody trusts enough to use still isn’t finished.
Working together
I’m open to remote roles in applied AI enablement, full-time, fractional, or focused consulting. If you’re putting AI into a real workflow and you need to know it works before a customer does, that’s my lane.
- Email: josh.durey@gmail.com
- GitHub: github.com/jdurey