Product vision, evaluation frameworks, and the guardrails that keep AI systems reliable once they're live — plus the real, working code behind it. Ask an AI grounded in my actual resume and GitHub projects, not a summary written to impress you.
I've spent recent years taking LLM and RAG systems from initial vision to production — setting product direction, debating architecture, and building the evaluation frameworks and guardrails that keep them reliable once real users depend on them.
My background spans product, content, and customer-experience leadership, most recently as Global Manager of Digital Client Experience at OANDA, owning the AI/automation portfolio end to end — chatbots, autonomous agents, and the human-in-the-loop review systems that keep them accountable.
I apply the same rigor in my own time: open-source experiments comparing chunking strategies, evaluation frameworks, and multi-agent orchestration — real, runnable code, not slideware.
Multi-tenant document classification, extraction, and routing via Gemini Flash, with RAG for long documents and an eval gate before any config change ships.
View project →Five versioned pipelines comparing chunking and retrieval strategies for a RAG audit system — measured trade-offs, not guesses.
View project →Scores a "loose" prompt strategy against a guardrailed one using LLM-as-a-judge metrics — quantifying hallucination risk instead of eyeballing it.
View project →A router agent classifies and dispatches to a knowledge-base agent (hand-built vector search) or an escalation agent — with live latency/token tracing.
View project →The tool linked above — grounded document Q&A on Cloudflare Workers, with a deterministic pre-filter, citation validation, and a kill switch.
Try it →K-Means/DBSCAN clustering and predictive modeling for marketing segmentation — a 5-person team project; my role was preprocessing, pipeline automation, and K-means clustering.
View project →