←  Back to home Deep Dives

Case Studies

Long-form system and platform designs for production AI. Each one states its assumptions up front, argues the trade-offs, and commits to a decision, because the reasoning is the deliverable, not just the conclusion.

01
Platform Design · Regulated Gaming

Enterprise Agentic Automation Platform

One multi-use-case agent platform, the seven-phase lifecycle that governs it, and the argument for which use case to build first. Buy the runtime, build the governance, and keep a human on every decision that touches money, KYC or responsible gaming.

Google ADK Vertex AI MCP A2A Human-in-the-loop Evals Governance
Scope
Architecture · lifecycle · RACI · 6-month delivery plan
Length
~14 min read
Read case study →
02
System Design · Streaming

Conversational AI Assistant for Streaming Search

Grounded Czech intent-to-catalog mapping as strict structured output, at 1.2M conversations a day inside a 5–10 s p99 budget. A bounded deterministic pipeline instead of an agent, where the schema, not the prompt, is what makes hallucination impossible.

Hybrid RAG pgvector Structured Outputs FastAPI Redis Golden Datasets Model Tiering
Scope
Pipeline · grounding · tech choices · MVP cut · risks
Length
~10 min read
Read case study →
03
Data Trust · B2B SaaS

Where AI Was Considered and Rejected

A pre-quote validation service that gates a contract price with TRUSTED, REVIEW or BLOCKED. The deterministic engine produces 100% of the number, no agent runs anywhere, exactly one model call is allowed to read English and never arithmetic — and the most useful verdict hands back no number at all.

Rules Engine Agent vs Workflow Databricks Salesforce Prompt Evals Data Governance Human-in-the-loop
Scope
Scoping · design · working prototype · evals · ADR log
Length
~15 min read
Read case study →
04
People Analytics · Fairness

What the Rating Actually Measures

Evidence-grounded support for performance conversations. It gathers a person's evidence across five systems, scores it against a rubric, flags where a manager's rating is unsupported, abstains on the quarter of the roster it cannot score fairly, and never proposes a rating. The strongest predictor of the rating turned out to be how much a person writes about themselves.

Identity Resolution Redaction Gate Rubric Design Claim Extraction Bias Audit Abstention Human-in-the-loop
Scope
Analysis · working proof of concept · E1–E7 evals · ops blueprint
Length
~16 min read
Read case study →
// More coming

More case studies land here as they are written, evaluation frameworks, agent tooling, and the things that only show up once a system has been in production for a while.

Let’s connect

Build reliable AI.
Start with a conversation.

Production AI, platform architecture and technical leadership.