Build, test, break, learn.
Not every experiment needs to become a product. Some answer a question. Some test an architecture. Some look at how a model behaves. Some become parts of larger systems.
The point is to learn by building.
Run it yourself.
Working demos that show what happens underneath: which steps stay deterministic, where a model earns its place, and where a person decides.
Every demo runs in your browser on scripted, invented data. There are no live model calls, and every person, company and record is fictional.
Equity Mining Alert
A Dealer Flywheel service-lane alert: flag customers with equity as they check in, text a salesperson, escalate if nobody claims it.
AI Recruiting System
An assistant runs first contact and screening. Hard requirements stay rules, and it escalates instead of guessing.
Knowledge Intelligence
A retrieval system that shows the chunks it used, cites them, and says when the evidence isn't enough.
Operational Intelligence
Find the records that need attention between systems, with the evidence for each exception.
Intelligent Workflow
An inbound lead workflow that marks which steps are deterministic, which use AI, and which need a person.
Open questions.
Each starts as a concept and becomes an experiment only once something has actually been built.
Agent State & Memory
How should an agent keep useful state across a multi-step interaction without old history crowding out the current decision?
Structured Outputs
When should an LLM return constrained, machine-readable output instead of natural language?
Retrieval & Grounding
How do chunking, metadata, ranking and context affect whether an answer is correct and traceable?
Human-in-the-Loop
Where should automatic execution stop and human judgment begin?
Tool Calling
How should an agent decide when to reason, look something up, call software, ask for clarification or escalate?
Evaluation
How do we measure AI systems where “it ran” and “it was right” aren't the same thing?
Failure Handling
What should a system do when information is missing, a tool fails, or confidence is low?