Skip to content
← All work
2025–2026LiveClient work

Enterprise AI Learning Platform

It began with a client’s question: how do we use AI with our product? Fifteen months later it’s an enterprise learning platform whose AI guide knows each learner and never invents what it doesn’t know. I’ve owned it end to end since taking it over, and three enterprises have signed on, with a fourth close.

Technical owner, end to end

Client work under confidentiality. Names, code and screens are withheld; this describes the problem and my approach in general terms.

Ownership
11 months, end to end
Today’s code
~73% mine · ~1,170 commits
Enterprises
3 signed, 1 close
Platform
195 APIs · 98 screens
AI guide
18 tools, memory, RAG
Security
123 RLS policies
An architecture diagram in three zones. Enterprises: a central authoring hub publishing read-only content with the same IDs to Enterprise A, B and C (signed) and D (close), each with its own database and deployment. The guide: the learner’s context flowing into an AI guide built in three layers (voice, conversation flow, methodology), with memory, retrieval, 18 tools and guardrails. The quality loop: a synthetic cohort, a simulated working day, a method validator, a model-swap bench (24/30 vs 19/30 at about 38% of the cost) and a daily health read.

01 / problem

The problem

It started in July 2025 with a high-level question from a client in leadership and workplace learning: how do we use AI with our product? The request arrived as a wish list: a model trained on their material, a voice clone of the author, a video avatar, and an assistant embedded everywhere.

I turned the question into a plan before writing code. A written response set out the approach in phases, with a cost model at three levels of use. Real-time video went to the backlog, because the technology couldn't do it live, and the voice clone was scheduled behind the core. Then a six-week proof of concept answered the questions that mattered first: retrieval over their material on Bedrock, the same question run across Claude and Nova models with and without retrieval side by side, and a "three doors" demo across three parts of the methodology. Feedback from real users grew into a test set of 500 questions.

One design choice from that prototype is still the spine of the guide today: separating how the AI talks from what it knows. A single conversational-style layer, shared by every domain and tuned against that feedback, sits on top of each domain's knowledge.

In November 2025 the client handed me their early-stage learning platform to bring the guide into the product. Since then I've owned it end to end: architecture, AI, infrastructure, design and interface, and the conversations about where it goes next. It has grown from one pilot to three signed enterprises.

02 / build

What I built

  • A guide designed like a character, built like a system. Three prompt layers (voice, conversation flow and methodology) are assembled on every turn with the learner's progress, plan and journey. A long-term memory is consolidated twice a day: observations with confidence, open threads that stop being raised once they've been asked enough, and a growth arc. Learners can read it for themselves. Retrieval runs over licensed source material, truncated before it ever leaves the server, plus semantic search over the learner's own past conversations. Eighteen tools run in a streamed loop with a stall watchdog. The guardrails mean it never invents methods, never speaks as the human author, never claims memory it doesn't have, and checks the full conversation before saying it can't see something.
  • The experience, orchestrated end to end. I directed and built the interface, the interactions and the admin tools. A design system of 181 tokens, with a living admin screen that audits stray colors against the brand. Motion built into the interface, including a guide character that reacts as replies arrive and respects reduced motion. A video course player with transcripts and cleaned subtitles, a drag-and-drop journey studio, a guided onboarding tour, printable generated reports, and an installable mobile app.
  • One codebase, many enterprises. Every enterprise gets its own isolated database and deployment. Content is authored once in a central hub and published to every tenant read-only, keeping the same IDs so nothing downstream breaks, and a new tenant is provisioned by script. On top of that: org trees, seats and invite flows, operator dashboards, and team sentiment that is never shown for fewer than four people. 123 row-level security policies across 61 tables.
  • Tested by agents, not just by me. A synthetic cohort of personas living out storylines on a schedule, graded like report cards. A simulated working day across many sites, with a replay monitor. A validator that runs every practice method through the real interface. And a model-swap bench that runs the real app with only the model swapped, proven identical down to the input tokens.
  • Run like a product. Nine scheduled jobs, web push and email nudges that respect preferences, Stripe plans and access tiers, per-call AI cost logging, and a daily health read: metrics gathered read-only, read by a model, delivered as a morning email.

03 / signal

What it shows

How I take an open question like "how do we use AI with our product?": plan it in phases, price it, prove the core in weeks, and keep the parts that work. Then taking that from early stage to enterprise-ready, without losing the craft. The AI is treated like a teammate with a job description: layered instructions, memory it earns and shows, tools it can only use when the learner can, and evaluation before any model change. The bench paid for itself on its first sweep. A smaller model scored 24 of 30 on the client's own rubric, against 19 for the model in production, at about 38% of the cost.

It also shows where the care goes: privacy designed in, with memory learners can see and sentiment that can't single anyone out. Environments that can't be confused, with test tooling that never touches the wrong database. And failure handling in the places people actually feel it.

Gallery 1 / 2

Where it started: the six-week proof of concept. The shared conversational style, kept apart from each domain’s knowledge and tuned by feedback, is still the spine of the guide.